CodeariaAcademy
Article cover: an object from the world of film and music built out of code, with the caption “A music video written in code”
September 28, 202613 min readAI ContentClaude CodeAI Agents

“AI solved motion graphic,” says X. The repo behind the Opus 5.5 music video shows five revisions and a client

X says AI solved motion graphics in one prompt. The Opus 5.5 video's code shows five revisions, a brief for scene agents and a colour bug found 4 days in.

In this article5
In short

A 2:36 lyric video for the song “I'm Upping My P(doom)” is rendered entirely from code: according to the README of the mexicat/pdoom-video repository, the concept, the word-level alignment of lyrics to audio, the three.js renderer and all 17 scenes were made by Claude Opus 5.5 in Claude Code. This is not generative video but a deterministic program: 56 TypeScript files, 22,290 lines, exported through headless Chrome and ffmpeg, with the 4K version taking about 2.5 hours to render. On X the clip spread under the caption “AI solved motion graphic”, next to stories of videos made “with one prompt”. The repository tells a different story: five revisions of the treatment, a human acting as the client, separate instructions for the agents that wrote the scenes, and manual fixes to the word timings. And four days after release, an outside PR found that every render had come out with a colour shift.

“I made this with one prompt using Opus 5.5.” That is how Donald Jewkes captioned his video for the same song on 23 September, adding that he talked to his computer for five minutes, Claude worked for 12 hours, and he woke up to the result. The post reached 3.4 million views. The next day developer Giacomo Magnanini (mexicat on X) showed his own version, drawn in code with three.js, and on 27 September Min Choi reposted it with the caption “AI solved motion graphic”.

We checked both “one prompt” and “solved” against the primary source. Magnanini published the Claude Opus 5.5 video as an open MIT repository, together with the treatment, a guide for scene authors and the full commit history. Those files show how the “Claude Code writes code, code renders video” setup actually works, and that picture is more useful than the viral caption.

How one P(doom) song became a benchmark for Opus 5.5

The song belongs to none of the video makers. The lyrics were written by osmarks and MusicPerson, the original was generated in Udio in November 2024, and on 9 September 2026 a user called deckard posted a pop version made in Suno under the title Claude-Pop. A track about AGI, the paperclip maximiser and RLHF turned into a shared task: everyone works from the same 2:36 audio, only the visuals differ.

On 22 September the user other__reality posted a video for it, calling it the best visual design of all the models he had tested. In his prompt Donald Jewkes names the sources of that video: the JohnHeibel/PDoomVideo repository, where frames are painted with the p5.brush library. A relay followed: Jewkes, then pleometric following Donald's workflow, then mexicat in a different stylistic direction.

JohnHeibel/PDoomVideoDonald Jewkesmexicat/pdoom-video
Drawn withp5.js and p5.brush, brush strokesSeedance 2.5 and fal images as a base, JS animation on topthree.js, GLSL shaders, TypeScript
What the human saidThe Clawd character and interesting visuals for every lineA prompt of about 1,700 wordsSeveral rounds of tweaks, five treatment revisions
Instructions for subagentsANIMATION_GUIDE.md, written by OpusNot publishedENGINE.md for scene authors
CodeOpenClosed, only the prompt is publicOpen, MIT

Jewkes's “one prompt” is worth reading in full. It is roughly 1,700 words of production spec: ElevenLabs and fal keys (dictated, so fal comes out as “foul”), about two thousand dollars of generation credits, an instruction to use up the whole limit of the Max plan, characters, JS animation layered over generated video in a rotoscope style, and a closing “make no mistakes”. It is one prompt in the honest sense that there was one message. But it is a full production brief, not a line saying “make a music video”.

What is inside mexicat/pdoom-video

The renderer's key property sits in the first line of the README: every frame is a deterministic function of song time. That is why the browser preview and the offline export in 1080p60 or 4K60 produce the same frame. Not a single pixel in the video comes from a generative video model. Everything you see (the TikZ unicorn, the oscilloscope with a loss curve, the form stamped SAFE ENOUGH, the field of paperclips) is drawn with shaders and Canvas.

A three.js music video with word-by-word karaoke typography. Renderer in TypeScript (bun + Vite), export through headless Chrome and ffmpeg, audio analysis in Python. The code is MIT; the song and lyrics are not covered by the licence.

commit this monthchecked 28 September 2026

The pipeline has two halves: the audio is analysed once, and the video is recomputed from code as often as you like.

  1. 1

    Analyse the audio

    Demucs separates the vocals from the music, two acoustic CTC engines align every word to the vocal, and Whisper cross-checks the result. Tempo (132.007 BPM), beats, bars and loudness envelopes are computed separately.

  2. 2

    Store the timings

    The output goes into data/lyrics.json (start and end of every word) and data/audio.json (beats, sections, drum hits). The renderer needs nothing else.

  3. 3

    Split the video into scenes

    17 modules, one per “tableau”: open, loss, prompt, hook, room, shoggoth and on to outro. Scene windows are tied to lyric lines and snapped to the beat grid.

  4. 4

    Preview in the browser

    A Vite server plays the video in real time, with frame stepping and looping of a single scene.

  5. 5

    Export

    A script drives headless Chrome frame by frame, pipes raw frames to ffmpeg and encodes with x264. Motion blur is built from 4 to 324 subframes per frame, depending on how fast the image is moving.

Scenes find a lyric line by its content, not by seconds: lyrics.get('sudden drop'), then the start of the word they need. Fix the alignment and every scene moves with it. That is how Claude Code can rework parts of the video without breaking the sync of the rest.

Brief, director and scene authors: how Claude Code ran the team

The most interesting parts of the repository are two documents next to the code. docs/TREATMENT.md is the treatment and style bible: the idea in one paragraph, a palette with HEX codes, four font families and their roles, karaoke rules and a description of every scene. It also spells out what must not appear in the video: purple-and-cyan neon, glowing brains, Matrix rain, and, separately, “nothing that looks AI-generated”.

Revision 2 retired the realistic engraved human eye (opening and "What did Ilya see?"): the client found it uncanny.
docs/TREATMENT.md in mexicat/pdoom-video · I'm Upping My P(doom): treatment & style bible, 25 September 2026

The client here is a human, and the treatment keeps a record of their decisions revision by revision, up to the fifth. The realistic eye was removed because the client found it uncanny. The blue accent was removed because it read as creepy. The captions in the frame corners were removed. Magnanini himself puts it more modestly in his post: the result is impressive “after a few rounds of tweaking”.

The second document, docs/ENGINE.md, is addressed to “scene authors”. In the scene table every scene has an owner: A1–A8, B1 and lead. The rules for authors read like a contractor handbook: do not touch files outside your scene, do not edit the timeline, and for engine changes “ask the lead”. We read this as a setup with one director agent and several scene agents. The JohnHeibel repository says it outright: Opus wrote ANIMATION_GUIDE.md to brief the subagents it ran in parallel.

One more line in ENGINE.md explains why the model can make video at all. The main way to check a scene is to render stills, “then LOOK at the PNGs with the Read tool”. Claude Code does not watch video, but it can read images. So quality control works like an animator with a storyboard: pull frames at the right seconds, look, fix. For motion there are contact sheets and short clips, and the performance budget is strict: a target of under 25 ms per frame.

How a scene is checked, per ENGINE.md
$ cd app && bun scripts/render.ts stills --t 12.5,13.0,14.2 --only open --out ../out/wip/open
$ bun scripts/render.ts sheet --from 1.5 --to 9 --n 16 --cols 4 --only open --out ../out/wip/open/sheet.png
# then the agent opens the PNGs with Read and compares them to the treatment

What neither the model nor the author noticed

On 28 September a PR from an outside developer arrived in the repository, also written together with Claude Opus 5.5. The gist: frames were piped to ffmpeg as sRGB but converted to yuv420p with the default BT.601 matrix and no colour tags. Players, browsers and YouTube decode untagged HD video as BT.709. The PR description sums it up: “every render came out shifted”, with pure green #20C040 turning into (20, 169, 62) in Chrome. The faulty encoding had been in the repository since the first commit on 24 September.

This shows exactly where checking by stills stops working. The PNGs the agent looks at through Read are taken before encoding, so their colours are correct. The bug lived in the last step, where the model never looked, and a hue shift in a finished video is impossible to spot by eye without something to compare it to. An outside developer found it by comparing pixel values.

There are other places where automation fell short. analysis/align.py contains manual word-boundary fixes set from spectrograms and QA plots. In the final chorus the lead vocal is buried under the backing vocals, CTC alignment finds nothing there, and the words are placed by hand with a confidence of 0.35–0.45. The video on YouTube, as the README admits, is an old render with four subframes per frame, where fast motion breaks into steps. The best version only exists if you render it yourself, which takes about 2.5 hours on an M5 Pro and produces a 13 GB file in 4K with default settings.

If you render video from a browser

Check which colour matrix ffmpeg uses to convert frames to yuv420p and whether it writes colour tags. Without an explicit BT.709 you will get a hue shift in any similar pipeline, and you will not see it in stills taken before encoding.

What to take away if you make video with Claude Code

Our opinion, which you are free to dispute: what AI has mastered here is not motion graphics but the director's job. The model keeps 17 scenes in its head, writes a brief for them, hands them out to workers and checks the result frame by frame. Aesthetic calls, such as dropping the uncanny eye or the blue accent, are still made by a human, and the repository records them as revisions. If the next video like this ships without a single client revision and with an open history, we will have been wrong.

Three techniques from the repository are worth taking, and they apply to any video made from code, with the music video just one case:

  1. Timings separate from the picture. Audio and lyrics are analysed once into JSON, and scenes look words up by content. Edits do not break the sync.
  2. The treatment as a contract. Palette, fonts, bans and a description of every scene live in one file, and every revision is appended to it. The agent writing the fifth scene sees the same rules as the first one.
  3. Checking by frames. Stills and contact sheets the agent inspects itself, plus the one step it cannot see: the final file after encoding.

If you want to build a pipeline like this without three.js and shaders, we cover the same “describe the scene, Claude writes the code, the code renders an MP4” approach in our guide to motion video with Remotion, and the full path from script to finished video in the course on motion video with AI. For animation on a website rather than in a video, see the ready-made animated components of React Bits. We wrote about how several agents work under one lead in our breakdown of the Paperclip agent orchestrator, and how much a task costs on Opus 5.5 compared with GPT-6 Sol in a separate article.

Repository versions, view counts and descriptions as of 28 September 2026; check the current ones.

Sources10expand
  1. Giacomo Magnanini, “mexicat/pdoom-video”, README, checked 28 September 2026 — https://github.com/mexicat/pdoom-video
  2. Giacomo Magnanini, “I'm Upping My P(doom): treatment & style bible”, docs/TREATMENT.md, 25 September 2026 — https://github.com/mexicat/pdoom-video/blob/main/docs/TREATMENT.md
  3. Giacomo Magnanini, “Engine guide (for scene authors)”, docs/ENGINE.md, 26 September 2026 — https://github.com/mexicat/pdoom-video/blob/main/docs/ENGINE.md
  4. HEOJUNFO, “render: encode with BT.709 and tag the stream (#3)”, 28 September 2026 — https://github.com/mexicat/pdoom-video/pull/3
  5. mexicat, post with the video, 24 September 2026 — https://x.com/_mexicat/status/2103108369569726802
  6. Donald Jewkes, “I made this with one prompt using Opus 5.5”, 23 September 2026 — https://x.com/donaldjewkes/status/2102801274173587569
  7. Donald Jewkes, full prompt, 23 September 2026 — https://x.com/donaldjewkes/status/2102801469976248500
  8. John Heibel, “JohnHeibel/PDoomVideo”, README, checked 28 September 2026 — https://github.com/JohnHeibel/PDoomVideo
  9. deckard, “Claude-Pop - I'm Upping My P(Doom)”, 9 September 2026 — https://x.com/slimer48484/status/2097752569212756134
  10. Min Choi, “Opus 5.5 is pretty wild”, 27 September 2026 — https://x.com/minchoi/status/2104227722377757031

Comments