I let an agent edit 34 phone clips into a 31 s reel. The text is drawn by code, not by an image model

Wait 5 sec.

I built a skill plus a small Python engine so a coding agent (Claude Code, Codex, Gemini CLI, Cursor, Antigravity) can edit a short reel from raw clips. The example in the repo: 34 phone clips from a client, cut into a 31 second reel with no voiceover. The cuts sit on a 115 BPM grid, and it spent zero credits on generated video. The design decision that mattered: the model never draws text. Images are generated without text, and the engine composes subtitles, captions and title cards itself, with Pillow and ffmpeg, when it renders. The text comes out sharp and timed to the frame, and changing the brand or the language is one line instead of regenerating every shot. What it doesn't do: it doesn't replace a video editor, and I have only tested it on Windows. macOS and Linux are untested. Question for people who edit video: what would you try to make with it, and what is missing before you could use it on your own footage? I built this. It's free (MIT): https://github.com/Octonove/invokard-studio (the example is in examples/schippers).   submitted by   /u/AntonioJBer [link]   [comments]