
“AI solved motion graphic,” says X. The repo behind the Opus 5.5 music video shows five revisions and a client
X says AI solved motion graphics in one prompt. The Opus 5.5 video's code shows five revisions, a brief for scene agents and a colour bug found 4 days in.
In this article5
A 2:36 lyric video for the song “I'm Upping My P(doom)” is rendered entirely from code: according to the README of the mexicat/pdoom-video repository, the concept, the word-level alignment of lyrics to audio, the three.js renderer and all 17 scenes were made by Claude Opus 5.5 in Claude Code. This is not generative video but a deterministic program: 56 TypeScript files, 22,290 lines, exported through headless Chrome and ffmpeg, with the 4K version taking about 2.5 hours to render. On X the clip spread under the caption “AI solved motion graphic”, next to stories of videos made “with one prompt”. The repository tells a different story: five revisions of the treatment, a human acting as the client, separate instructions for the agents that wrote the scenes, and manual fixes to the word timings. And four days after release, an outside PR found that every render had come out with a colour shift.
“I made this with one prompt using Opus 5.5.” That is how Donald Jewkes captioned his video for the same song on 23 September, adding that he talked to his computer for five minutes, Claude worked for 12 hours, and he woke up to the result. The post reached 3.4 million views. The next day developer Giacomo Magnanini (mexicat on X) showed his own version, drawn in code with three.js, and on 27 September Min Choi reposted it with the caption “AI solved motion graphic”.
We checked both “one prompt” and “solved” against the primary source. Magnanini published the Claude Opus 5.5 video as an open MIT repository, together with the treatment, a guide for scene authors and the full commit history. Those files show how the “Claude Code writes code, code renders video” setup actually works, and that picture is more useful than the viral caption.
How one P(doom) song became a benchmark for Opus 5.5
The song belongs to none of the video makers. The lyrics were written by osmarks and MusicPerson, the original was generated in Udio in November 2024, and on 9 September 2026 a user called deckard posted a pop version made in Suno under the title Claude-Pop. A track about AGI, the paperclip maximiser and RLHF turned into a shared task: everyone works from the same 2:36 audio, only the visuals differ.
On 22 September the user other__reality posted a video for it, calling it the best visual design of all the models he had tested. In his prompt Donald Jewkes names the sources of that video: the JohnHeibel/PDoomVideo repository, where frames are painted with the p5.brush library. A relay followed: Jewkes, then pleometric following Donald's workflow, then mexicat in a different stylistic direction.
| JohnHeibel/PDoomVideo | Donald Jewkes | mexicat/pdoom-video | |
|---|---|---|---|
| Drawn with | p5.js and p5.brush, brush strokes | Seedance 2.5 and fal images as a base, JS animation on top | three.js, GLSL shaders, TypeScript |
| What the human said | The Clawd character and interesting visuals for every line | A prompt of about 1,700 words | Several rounds of tweaks, five treatment revisions |
| Instructions for subagents | ANIMATION_GUIDE.md, written by Opus | Not published | ENGINE.md for scene authors |
| Code | Open | Closed, only the prompt is public | Open, MIT |
Jewkes's “one prompt” is worth reading in full. It is roughly 1,700 words of production spec: ElevenLabs and fal keys (dictated, so fal comes out as “foul”), about two thousand dollars of generation credits, an instruction to use up the whole limit of the Max plan, characters, JS animation layered over generated video in a rotoscope style, and a closing “make no mistakes”. It is one prompt in the honest sense that there was one message. But it is a full production brief, not a line saying “make a music video”.
What is inside mexicat/pdoom-video
The renderer's key property sits in the first line of the README: every frame is a deterministic function of song time. That is why the browser preview and the offline export in 1080p60 or 4K60 produce the same frame. Not a single pixel in the video comes from a generative video model. Everything you see (the TikZ unicorn, the oscilloscope with a loss curve, the form stamped SAFE ENOUGH, the field of paperclips) is drawn with shaders and Canvas.
A three.js music video with word-by-word karaoke typography. Renderer in TypeScript (bun + Vite), export through headless Chrome and ffmpeg, audio analysis in Python. The code is MIT; the song and lyrics are not covered by the licence.
The pipeline has two halves: the audio is analysed once, and the video is recomputed from code as often as you like.
- 1
Analyse the audio
Demucs separates the vocals from the music, two acoustic CTC engines align every word to the vocal, and Whisper cross-checks the result. Tempo (132.007 BPM), beats, bars and loudness envelopes are computed separately.
- 2
Store the timings
The output goes into data/lyrics.json (start and end of every word) and data/audio.json (beats, sections, drum hits). The renderer needs nothing else.
- 3
Split the video into scenes
17 modules, one per “tableau”: open, loss, prompt, hook, room, shoggoth and on to outro. Scene windows are tied to lyric lines and snapped to the beat grid.
- 4
Preview in the browser
A Vite server plays the video in real time, with frame stepping and looping of a single scene.
- 5
Export
A script drives headless Chrome frame by frame, pipes raw frames to ffmpeg and encodes with x264. Motion blur is built from 4 to 324 subframes per frame, depending on how fast the image is moving.
Scenes find a lyric line by its content, not by seconds: lyrics.get('sudden drop'), then the start of the word they need. Fix the alignment and every scene moves with it. That is how Claude Code can rework parts of the video without breaking the sync of the rest.
Brief, director and scene authors: how Claude Code ran the team
The most interesting parts of the repository are two documents next to the code. docs/TREATMENT.md is the treatment and style bible: the idea in one paragraph, a palette with HEX codes, four font families and their roles, karaoke rules and a description of every scene. It also spells out what must not appear in the video: purple-and-cyan neon, glowing brains, Matrix rain, and, separately, “nothing that looks AI-generated”.
Revision 2 retired the realistic engraved human eye (opening and "What did Ilya see?"): the client found it uncanny.
The client here is a human, and the treatment keeps a record of their decisions revision by revision, up to the fifth. The realistic eye was removed because the client found it uncanny. The blue accent was removed because it read as creepy. The captions in the frame corners were removed. Magnanini himself puts it more modestly in his post: the result is impressive “after a few rounds of tweaking”.
The second document, docs/ENGINE.md, is addressed to “scene authors”. In the scene table every scene has an owner: A1–A8, B1 and lead. The rules for authors read like a contractor handbook: do not touch files outside your scene, do not edit the timeline, and for engine changes “ask the lead”. We read this as a setup with one director agent and several scene agents. The JohnHeibel repository says it outright: Opus wrote ANIMATION_GUIDE.md to brief the subagents it ran in parallel.
One more line in ENGINE.md explains why the model can make video at all. The main way to check a scene is to render stills, “then LOOK at the PNGs with the Read tool”. Claude Code does not watch video, but it can read images. So quality control works like an animator with a storyboard: pull frames at the right seconds, look, fix. For motion there are contact sheets and short clips, and the performance budget is strict: a target of under 25 ms per frame.
$ cd app && bun scripts/render.ts stills --t 12.5,13.0,14.2 --only open --out ../out/wip/open$ bun scripts/render.ts sheet --from 1.5 --to 9 --n 16 --cols 4 --only open --out ../out/wip/open/sheet.png# then the agent opens the PNGs with Read and compares them to the treatment
What neither the model nor the author noticed
On 28 September a PR from an outside developer arrived in the repository, also written together with Claude Opus 5.5. The gist: frames were piped to ffmpeg as sRGB but converted to yuv420p with the default BT.601 matrix and no colour tags. Players, browsers and YouTube decode untagged HD video as BT.709. The PR description sums it up: “every render came out shifted”, with pure green #20C040 turning into (20, 169, 62) in Chrome. The faulty encoding had been in the repository since the first commit on 24 September.
This shows exactly where checking by stills stops working. The PNGs the agent looks at through Read are taken before encoding, so their colours are correct. The bug lived in the last step, where the model never looked, and a hue shift in a finished video is impossible to spot by eye without something to compare it to. An outside developer found it by comparing pixel values.
There are other places where automation fell short. analysis/align.py contains manual word-boundary fixes set from spectrograms and QA plots. In the final chorus the lead vocal is buried under the backing vocals, CTC alignment finds nothing there, and the words are placed by hand with a confidence of 0.35–0.45. The video on YouTube, as the README admits, is an old render with four subframes per frame, where fast motion breaks into steps. The best version only exists if you render it yourself, which takes about 2.5 hours on an M5 Pro and produces a 13 GB file in 4K with default settings.
If you render video from a browser
Check which colour matrix ffmpeg uses to convert frames to yuv420p and whether it writes colour tags. Without an explicit BT.709 you will get a hue shift in any similar pipeline, and you will not see it in stills taken before encoding.
What to take away if you make video with Claude Code
Our opinion, which you are free to dispute: what AI has mastered here is not motion graphics but the director's job. The model keeps 17 scenes in its head, writes a brief for them, hands them out to workers and checks the result frame by frame. Aesthetic calls, such as dropping the uncanny eye or the blue accent, are still made by a human, and the repository records them as revisions. If the next video like this ships without a single client revision and with an open history, we will have been wrong.
Three techniques from the repository are worth taking, and they apply to any video made from code, with the music video just one case:
- Timings separate from the picture. Audio and lyrics are analysed once into JSON, and scenes look words up by content. Edits do not break the sync.
- The treatment as a contract. Palette, fonts, bans and a description of every scene live in one file, and every revision is appended to it. The agent writing the fifth scene sees the same rules as the first one.
- Checking by frames. Stills and contact sheets the agent inspects itself, plus the one step it cannot see: the final file after encoding.
If you want to build a pipeline like this without three.js and shaders, we cover the same “describe the scene, Claude writes the code, the code renders an MP4” approach in our guide to motion video with Remotion, and the full path from script to finished video in the course on motion video with AI. For animation on a website rather than in a video, see the ready-made animated components of React Bits. We wrote about how several agents work under one lead in our breakdown of the Paperclip agent orchestrator, and how much a task costs on Opus 5.5 compared with GPT-6 Sol in a separate article.
Repository versions, view counts and descriptions as of 28 September 2026; check the current ones.
Sources10expand
- Giacomo Magnanini, “mexicat/pdoom-video”, README, checked 28 September 2026 — https://github.com/mexicat/pdoom-video
- Giacomo Magnanini, “I'm Upping My P(doom): treatment & style bible”, docs/TREATMENT.md, 25 September 2026 — https://github.com/mexicat/pdoom-video/blob/main/docs/TREATMENT.md
- Giacomo Magnanini, “Engine guide (for scene authors)”, docs/ENGINE.md, 26 September 2026 — https://github.com/mexicat/pdoom-video/blob/main/docs/ENGINE.md
- HEOJUNFO, “render: encode with BT.709 and tag the stream (#3)”, 28 September 2026 — https://github.com/mexicat/pdoom-video/pull/3
- mexicat, post with the video, 24 September 2026 — https://x.com/_mexicat/status/2103108369569726802
- Donald Jewkes, “I made this with one prompt using Opus 5.5”, 23 September 2026 — https://x.com/donaldjewkes/status/2102801274173587569
- Donald Jewkes, full prompt, 23 September 2026 — https://x.com/donaldjewkes/status/2102801469976248500
- John Heibel, “JohnHeibel/PDoomVideo”, README, checked 28 September 2026 — https://github.com/JohnHeibel/PDoomVideo
- deckard, “Claude-Pop - I'm Upping My P(Doom)”, 9 September 2026 — https://x.com/slimer48484/status/2097752569212756134
- Min Choi, “Opus 5.5 is pretty wild”, 27 September 2026 — https://x.com/minchoi/status/2104227722377757031
Read next
React Bits: the best visual component library for React in 2026September 2, 2026
Agents now get a boss, a task queue and a budget. Do you need one if you already work in Claude Code?September 26, 2026
GPT-6 Sol is half the price of Opus 5.5 per token. So why is a task only 20% cheaper?September 26, 2026
Comments