a skill by Dancan254, brought here by kt
voiceover video
paste this link into your ai. it will know what to do.
Turn a voice recording (voice note, narration, podcast snippet) into a fully animated, brand-styled video — local word-level transcription, a timed shot list, kinetic typography, camera moves, real photos and licensed video clips of the people and products it mentions, synthesized sound design and a voice-ducked music bed, rendered as a 1080x1920 Short (or 1920x1080). Also takes a to-camera phone recording: the opening line and sign-off stay on camera as face shots and everything between is animated. With no recording at all, writes an explainer script, voices each cartoon character offline, a
Voiceover Video Skill
Takes an audio file and builds a complete edit around it: every shot is an HTML/GSAP scene timed to the exact spoken word, rendered frame by frame in headless Chromium, then muxed with a mixed soundtrack (voice + synthesized SFX + ducked music). The judgement — what each line should look like — is yours. The scripts handle transcription, rendering, audio, and encoding.
Given a to-camera video instead of audio, the opening line and the sign-off stay on camera as face shots; everything between them is animated over the voice from the same take.
Given a to-camera video shot against a green screen (presenter mode), the speaker is keyed out and stays on screen in the brand's background for the whole video, moving between full-frame, split and close-up layouts while graphics build beside them, with full-screen cutaways for diagrams and chapter cards. Steps 2a and 2b trim and key the take; the Presenter blocks in scene-blocks.md lay it out.
Given only a topic or a script (script mode), there is no recording: you write the script, Step 1b voices it with one synthetic voice per character, and original cartoon hosts act it out. Everything after Step 1b runs on the generated <work>/voice.wav as if it were a recording.
SKILL_DIR = the directory containing this SKILL.md.
Always load SKILL_DIR/references/scene-blocks.md before writing the shot list. It holds the block catalogue, the pacing rules, and the sound-cue vocabulary.
Step 0 — Gather inputs
| Field | Required | Example |
|---|---|---|
audio, video or topic | Yes | ~/Downloads/voice-note.m4a · a to-camera recording ~/Movies/take-1.mp4 · "explain Kafka consumer groups" (script mode) |
script | No | the script, plain or with [FACE] / [VOICE] sections; captions are checked against it. Sections start on their own line with exactly [FACE] or [VOICE]. In script mode, Name: line lines, and it is the input |
cast | No | script mode: who's in it and which voice, e.g. teacher Mama Log, sidekick Pip (squeaky) |
format | No | vertical 1080x1920 (default) · landscape 1920x1080 |
style | No | for a video: bookends (default: face shots open and close) · presenter (green screen: the speaker stays on screen throughout) |
kit | No | brand kit folder for this video, e.g. a client's ~/clients/acme-kit (default: the user's own kit) |
template | No | visual theme id from templates/templates.json (default kinetic) |
music | No | synth (default) · path to a royalty-free track · none |
model | No | small (default) · base for clean audio, ~2x faster |
vocab | No | names and terms the speaker uses: "Kubernetes, Kafka, Jane Doe" |
slug | No | inferred from the topic, e.g. java-origin |
Brand kit: kit if given, then ./brand.json, then ~/.config/voiceover-video/brand.json, then the bundled SKILL_DIR/brand.example.json. A kit is a folder holding brand.json (version 2) and the fonts and logos it names; SKILL_DIR/examples/kits/ has two to copy. Output goes to <kit output.dir>/<slug>/ (default ~/voiceover-videos); work files go to <output dir>/<slug>/work/. Create it now and set <work> to that path for the rest of the workflow:
mkdir -p <output dir>/<slug>/work
work="<output dir>/<slug>/work"Pick the defaults and proceed. Don't interrogate.
A version-1 brand file (no "version": 2; scripts say so): convert it once. The original is kept as brand.v1.json:
python3 SKILL_DIR/scripts/init_kit.py --from ~/.config/voiceover-video/brand.jsonFirst run only — no kit anywhere: the example kit ships someone else's handle and colours, so ask before rendering. Four questions, each with its default, answered in one message:
| Ask | Default |
|---|---|
| Name and handle shown on the video | required, no default |
Look: midnight-pink, carbon-cyan, ink-amber, violet-signal, or their own colours | midnight-pink |
| Fonts: display and code — a Google font name, or a font file they have | Archivo / Geist Mono |
| Where finished videos go | ~/voiceover-videos |
python3 SKILL_DIR/scripts/init_kit.py --name "Their Name" --handle @theirhandle [--preset carbon-cyan] \
[--primary '#ff6600' --bg '#0d1117'] [--display Archivo --mono 'Geist Mono'] [--output-dir ~/voiceover-videos]A client's brand (a company video): build a kit from what the client supplies — never fetch a company's logo or font from the web. Ask for the brand colours, the font files or Google names, and the logo files (a mark, and a wordmark for dark and for light backgrounds), then:
python3 SKILL_DIR/scripts/init_kit.py --name Acme --handle acme.com --primary '#e50914' --bg '#0a0a0a' \
--display ~/client/fonts/AcmeSans.woff2 --mono 'JetBrains Mono' --mark ~/client/logo-mark.svg \
--wordmark-on-dark ~/client/wordmark-white.svg --wordmark-on-light ~/client/wordmark.svg \
--background glow --out ~/clients/acme-kitA light brand also passes --bg-alt and --text-alt: dark themes use them. --background is theme (the theme's own canvas), solid, glow, gradient or grid; --background-image <file> uses a picture.
In script mode there is no file to check; go to Step 1b after setup. Otherwise, if the audio or video path is missing or not a file, say so and stop. With a video, pass the video file wherever a step below takes <audio>; ffmpeg reads the voice from its audio track.
Step 1 — Setup (first run, and whenever the kit changes)
bash SKILL_DIR/scripts/setup.sh [<kit>]Checks ffmpeg, node, faster_whisper and numpy; installs playwright-core and its matching Chromium build; downloads GSAP and the themes' fonts into SKILL_DIR/assets/, and installs the kit's fonts and logos into SKILL_DIR/assets/kits/<id>/, so several clients' kits live side by side. Prints ready or names the missing piece. Re-running is a no-op for an unchanged kit.
Step 1a — Approve the look (first run, and every new kit)
python3 SKILL_DIR/scripts/fill_template.py "$work" 6 --format landscape --brand <kit> --template <template-id> --board
node SKILL_DIR/scripts/render.js board "$work"/board.html "$work"/board.png
python3 SKILL_DIR/scripts/brand_kit.py check <kit>The board shows the kit inside the theme: headline and highlight, a panel, a lower third, captions, a pill, a stamp, the end card with the logo, and the colour swatches. render.js board fails if a kit font fell back to a default face. brand_kit.py check prints the contrast report (exit 1 below WCAG AA) and warns when a logo has no variant for the canvas. Show the board and wait for a yes; if you cannot view images, give the user board.png and the contrast report. A wrong colour costs seconds here and a full render later. Pick the theme in Step 4 first if the user hasn't — the board renders one theme.
Step 1b — Script mode: write and voice the script
Only when there is no recording. Load SKILL_DIR/references/explainers.md first. It has the explainer shape, how to pick the analogy, the cast roles and the script format.
- Write the analogy mapping, then the script as
Name: linelines, to<work>/script.txt. Show both to the user and wait for a yes; the script is the cheapest thing to change. - Voice it (first run:
bash SKILL_DIR/scripts/setup.sh --voicesfetches the voice and alignment models, ~500 MB):
python3 SKILL_DIR/scripts/speak.py "$work"/script.txt "$work" [--cast "$work"/cast.json]Writes voice.wav, speech.json/speech.js (who speaks when; the hosts read it), and words.json + transcript.txt. speak.py runs faster-whisper over the voice track and maps what it hears onto each line, so every script word carries the time it is spoken at. A line recognition can't read (a short line with a hard name, a very squeaky voice) keeps estimated times; the result line names each one (… 14/15 lines aligned (1 estimated: pip line 7)) — tell the user. --no-align estimates every line, for a fast draft only. Tell the user which voice each character got and ask them to listen to voice.wav; you cannot. Re-voicing a line costs seconds, re-timing thirty shots does not.
From here, <audio> in every later step is "$work"/voice.wav. Skip Step 2: words.json is already timed and spelled as the script. In Step 3, script.txt is the reference and fixes are rarely needed.
Step 2 — Transcribe
Tell the user the estimate first: roughly 2–3x realtime on CPU for small.
python3 SKILL_DIR/scripts/transcribe.py <audio> --outdir "$work" --model small --vocab "<vocab>"Writes words.json (every word with start/end) and transcript.txt — one phrase per line as [start] word@time word@time …. Read transcript.txt; the per-word times are what you cut to.
Run it in the background for anything over two minutes. Never wait silently. The first run downloads the model, so it may sit at low CPU for a few minutes.
Long steps outlast shell time limits. Transcribing, keying, rendering frames and encoding a take of several minutes can each run longer than an agent's command timeout, and a killed step loses its work. Start them detached and poll the log: nohup <command> > "$work"/<step>.log 2>&1 &, then read the log until it prints its Next: line. key_greenscreen.py resumes where it stopped. A killed render-frames.sh doesn't: re-run it with a frame range starting at the first frame missing from frames/.
Step 2a — Trim the take (audio, and presenter mode)
python3 SKILL_DIR/scripts/trim_take.py <audio-or-video> "$work" [--in auto] [--out auto]Cuts the dead air before the first word (keeps 0.5s) and after the last (keeps 2.5s for the outro), writes a lossless voice.wav, shifts words.json and transcript.txt so the first kept moment is 0, and records the cut in take.json. Pass --in/--out in seconds to cut by hand (a false start, a reach for the camera); cut points inside the speech make an excerpt, such as a Short from a long talk, and drop the words outside them. From here on, <audio> is "$work"/voice.wav and <duration> is the length it prints. Skip it for bookends video: face shots are cut from the untrimmed recording.
Step 2b — Key the speaker (presenter mode)
python3 SKILL_DIR/scripts/key_greenscreen.py "$work" --preview 30 [--mask x,y,w,h …]
python3 SKILL_DIR/scripts/key_greenscreen.py "$work" [--mask x,y,w,h …] [--quality master] # detached for long takesIt measures the screen from the take and prints its colour and margin; a take with no usable screen fails with the reason, so fall back to bookends. Look at presenter-preview.png first: hair, glasses and shoulders should have clean edges. A name banner or logo burned into the recording needs a --mask over it (source pixels); a mask that cuts a shoulder leaves a notch, which presenterStyle({fade:"left"}) hides. The full run writes presenter/fNNNNN.webp, one per edit frame (~4 frames/s on 6 cores, so a 6-minute take is about 45 minutes): tell the user, run it detached, and re-run the same command if it stops; it keys only what is missing. --quality master keeps full colour for the sharpest edges at about ten times the disk.
Step 3 — Proofread the captions
Captions are burned in. Scan transcript.txt for misheard words — proper nouns and technical terms break first (observed: San Micro Systems → Sun Microsystems, Ok → Oak, CNC++ → C/C++, Right once → Write once). Write <work>/fixes.json mapping the raw token to its fix; an empty string drops the token:
{ "San": "Sun", "Ok.": "Oak.", "CNC++,": "C/C++,", "alias,": "", "that@50.22": "data" }A plain key fixes the token everywhere. For a common word that was misheard once, key it as token@time, with the start time printed in transcript.txt, and only that word changes.
Optionally write <work>/keywords.json, an array of words that stay in the accent colour once spoken: ["Kafka", "Java@3.00"]. A plain word flags every occurrence; @time flags one. Keep it to names and the few nouns the video is about, one per phrase at most.
python3 SKILL_DIR/scripts/build_captions.py "$work"Writes <work>/words.js. Never change timestamps. Tell the user about any token you could not resolve instead of guessing.
With a script, it is the reference for spelling: a token that differs from it goes in fixes.json. Captions follow what was said, so an ad-libbed line stays; tell the user where the take left the script.
Step 4 — Shot list (confirm before building)
Load references/scene-blocks.md. Write a shot table — one row per shot, cut on word times:
# in-out line (spoken) block sound
01 0.00-2.70 Did you know Java wasn't… slam + strike hit@0.44 hit@2.12
02 2.70-6.10 a completely different problem highlight-box whoosh, hit@4.60
03 6.10-8.30 Back in the early 1990s vhs + counter tick, hit@7.20Rules: a new shot every 1–4 seconds; a hit only on a word that deserves it; captions hidden whenever the spoken word is the visual. Find photos and clips before the table is final (Step 5). See SKILL_DIR/examples/kafka-vs-rabbitmq/ for a complete worked shot list.
Pick the template (theme) now. Read templates/templates.json and choose the id whose mood matches the topic. Each theme changes type, colour, captions and motion (default shot entry, shake, flash), so it is a real choice, not a palette swap:
| id | Pick it for |
|---|---|
kinetic | fast, dark, neon — tech explainers, launches, listicles (default) |
documentary | founder stories, company origins, biographies — serif, film grain, slow fades |
newsroom | announcements, funding, outages, "this week in tech" — condensed type, ticker, wipes |
blueprint | system design, architecture, "how X works under the hood" — grid paper, dashed outlines |
brutalist | hot takes, myth busting — light background, hard black borders, hard cuts, big shake |
aurora | AI and SaaS launches, dev tools — gradients, frosted glass, soft entries |
minimal | thought leadership, deep dives — light, editorial, calm |
retro | history of tech, CLI demos, hacking stories — CRT amber, glitch cuts |
If the user asked for a specific look, use that. Pass it to fill_template.py with --template <id>.
In script mode, the hosts carry the video: cut shots on line(who, n) times from speech.json, give the sidekick a bubble for their lines (say()), name each term with a .pill the first time it's spoken, and use lanes and tokens for anything that queues, flows or is numbered. The Host, Speech bubble, Term pill, Lane and tokens and Failure and recovery blocks in scene-blocks.md cover it.
In presenter mode, every row names a layout: full (the speaker centred, a tag or headline beside them), split (speaker left, a panel building the point on the right), close (punch in for a slam), or cut (a full-screen graphic, the speaker stepped out), plus chapter cards for long talks. A layout may hold 3–8s in a talk of several minutes as long as something in it moves; a Short still cuts every 1–4s.
With a video in bookends style, the first and last rows are face shots (Face hook, Face sign-off). The hook runs from 0 to the last word of the script's first [FACE] section (no script: the first sentence). The sign-off runs from the first word of the last [FACE] section to the end. A series badge goes on shot 02.
Show the table and the media list (every photo and clip, with its source and licence) to the user. Wait for a yes — this is the expensive part to change later.
Step 5 — Source photos and clips
Whenever the audio names a person, company, product or event, show it: the founder on stage, the product launch, the person saying the line. find_media.py searches licensed sources (Wikimedia Commons, Openverse, Internet Archive, and Pexels with PEXELS_API_KEY) alongside the open web (Bing images) and YouTube:
python3 SKILL_DIR/scripts/find_media.py search "$work" "Yang Zhilin portrait" # photos
python3 SKILL_DIR/scripts/find_media.py search "$work" "Yang Zhilin keynote" --kind video # talks
python3 SKILL_DIR/scripts/find_media.py fetch "$work" m6 --name yang-launch
python3 SKILL_DIR/scripts/find_media.py fetch "$work" m8 --name yang-gtc --section 120-150
python3 SKILL_DIR/scripts/find_media.py fetch "$work" --url "https://youtu.be/…" --name demo --section 30-45
python3 SKILL_DIR/scripts/find_media.py search "$work" "facepalm" --kind gif # reactions
python3 SKILL_DIR/scripts/find_media.py fetch "$work" m11 --name facepalmGifs come from Commons, and from GIPHY when GIPHY_API_KEY is set (GIPHY results are marked ⚠). A fetched gif becomes an mp4 in clips/src/; cut it like any clip, adding --loop so a 2-second gif fills a longer shot. Never put a .gif in an <img>: the browser plays it on its own clock, so every render would differ. One or two reaction beats per video, on a punchline, never on the point itself.
Results without a licence are marked ⚠. When a licensed result is as good a shot, take it; otherwise use the best shot. Prefer the subject's own channel (a company's official YouTube) over re-uploads. Search queries that work: the name plus the company, then name + event ("keynote", "launch", "interview"). Watermarked stock sites are filtered out.
A YouTube or page fetch downloads the whole video once into clips/src/.cache/ (720p for anything over 15 minutes), then cuts --section (≤120s) from it, so later sections of the same talk are instant. Tell the user the first fetch of a long talk takes a few minutes. To find a quote inside a section, run transcribe.py on the fetched clip and use the word times as --from in Step 6.
Look at every photo and every preview sheet before using it. Search results lie: the wrong person with the same name, a news site's logo or banner burned into the image, a thumbnail with text on it. Crop around a banner with object-position or pick another result. If you cannot view images, list each file with its source and what you expect it to show, and ask the user to confirm.
Logos: https://cdn.jsdelivr.net/npm/simple-icons@13/icons/<slug>.svg. Every fetch is recorded in <work>/credits.json; log hand-sourced files yourself. The report lists every ⚠ file so the user knows which footage belongs to someone else before posting.
Step 6 — Author the composition
python3 SKILL_DIR/scripts/fill_template.py "$work" <duration> --format vertical --template <template-id> [--brand <kit>]
# or --format landscape for 1920x1080
# omit --template to use the default kinetic themeColours in shots come from the kit: write var(--brand-primary), var(--brand-primary-ink) (the brand colour as readable text), var(--brand-text), var(--brand-muted), var(--brand-surface2), var(--brand-success), var(--brand-error), never a hex value, so the same shots render in any client's colours. Logos are in BRAND.logos (mark, wordmark.onDark, wordmark.onLight); see Logo end card and Corner mark in scene-blocks.md.
<duration> = last word end + ~2.5s for the outro; with a video, last word end + 0.5s, and never past the recording's length. Get the recording length with:
ffprobe -v error -show_entries format=duration -of csv=p=0 <video>Writes <work>/index.html from the template with brand, geometry and duration filled, and links the vendored assets as <work>/vendor.
In presenter mode, the speaker layer is already in the template. Set its look once, then give every shot the layout from the shot list; presenterOff() before a full-screen cutaway, and chapter() builds a chapter card and steps the speaker out for it:
presenterStyle({halo:true, fade:"left"}); // push:.03 adds a slow zoom, off by default: it softens footage
presenter("full", 0, 2.8); shot("s01", 0, 2.8, "cut");
presenter("split", 2.8, 6.4); shot("s02", 2.8, 6.4); fromRight("#s02p", 2.85);
presenterOff(6.4); shot("s03", 6.4, 10.3, "zoom");
chapter("s12", 58.0, 61.1, "01", "The patch was never<br>the bottleneck");The Presenter blocks in scene-blocks.md show the panel, checklist, diagram and stamp pieces these layouts use.
With a video in bookends style, extract the camera frames for both face shots, using the shot list's in/out times:
bash SKILL_DIR/scripts/extract_face.sh <video> <work> vertical <hook-in> <hook-out> <signoff-in> <signoff-out>For every clip shot, cut the fetched clip to the shot's in/out. The size is the box it fills — vertical or landscape for full-bleed, or the .pip box size like 900x620. --from is the second inside the clip to start at. Add --audio when the clip's own sound should play — a founder's line, a crowd:
bash SKILL_DIR/scripts/extract_clip.sh "$work"/clips/src/torvalds-talk.mp4 "$work" 900x620 torvalds 8.9 12.6 --from 20 --audio
bash SKILL_DIR/scripts/extract_clip.sh "$work"/clips/src/facepalm.mp4 "$work" 900x620 facepalm 31.2 33.4 --loopClip audio ducks under the narration automatically, so it only really plays over a pause in the voice. For the speaker's line to land, place the clip over a gap in the recording (a scripted [CLIP] beat) or leave it silent and put the quote on screen with a Portrait quote.
B-roll is the same mechanism. A full-bleed or picture-in-picture clip playing under narration is just a clip() (or the alias broll()). Use the B-roll under narration block in scene-blocks.md for the full-bleed + lower-third pattern.
Replace the demo shots between BEGIN SHOTS / END SHOTS (markup) and BEGIN TIMELINE / END TIMELINE (GSAP) with your shot list, using the helpers the template already defines. Keep the outer #world and #cam containers intact: shot(), slam(), hit(), rise(), pop(), stagger(), drift(), kenBurns(), lowerThird(), ticker(), typer(), counter(), terminal(), faceCam(), clip(), broll(), the NOCAP ranges, and the CAP_STYLE constant. Every helper that makes noise pushes its own sound cue. See SKILL_DIR/examples/kafka-vs-rabbitmq/ for a complete worked composition.
The display style (.xl) is uppercase and width-expanded; captions are condensed. Size headlines for the expanded width — a vertical frame fits ~6 characters at 250px.
Step 7 — QA stills (loop until clean)
Pick one timestamp per shot, at the moment of densest content:
node SKILL_DIR/scripts/render.js stills "$work"/index.html "$work"/stills 0.6,2.4,4.9,…
bash SKILL_DIR/scripts/contact-sheet.sh "$work"/stills "$work"/contact.jpgPass timestamps as a comma-separated list with no spaces; spaces parse as NaN.
Measure first — this needs no eyes and runs in seconds:
node SKILL_DIR/scripts/render.js check "$work"/index.html ["$work"/check.json]It seeks to each shot at 25%, 50% and 85% of its window and reports text past the frame edge, text sitting under the caption box, shots that render nothing at all three moments, and PAGE ERROR lines. Exit 1 means findings. Fix and re-run until it is clean; it catches clipped headlines that a full render would waste minutes on.
Then, if you can view images, read contact.jpg for what measurement cannot judge: crops that cut off a face or subject, an image that does not match the line, and layouts that fit the frame but read badly. Fix, re-render only the affected stills, and look again.
If you cannot view images, say so in the report and ask the user to look at contact.jpg before Step 9. Do not start Step 9 with a known defect — a full render costs minutes.
Step 8 — Sound
node SKILL_DIR/scripts/render.js cues "$work"/index.html "$work"/cues.json
python3 SKILL_DIR/scripts/synth_audio.py "$work"/cues.json <duration> "$work" --template <template-id> --drop <time-of-final-slam>Writes sfx.wav and music.wav. Pass the same <template-id> as Step 6: the music follows the theme's tempo, key and layers, and varies per video (the slug folder). Omit --drop if the video has no final slam; otherwise use the time of the last big hit. Optional pacing flags:
--drums-from <t>— bring the drums in at<t>seconds.--quiet <a>:<b>— duck the music betweenaandbseconds (repeatable).
With music set to a file, pass --no-music and give that file to Step 10. With none, pass --no-music and nothing else.
You cannot hear the result. Say so in the report and ask the user to listen.
Step 9 — Render frames
bash SKILL_DIR/scripts/render-frames.sh [--blur 4] "$work"/index.html "$work"/frames <duration> [workers] [from_frame to_frame]Frames render supersampled (2x) and are saved as lossless PNG: crisp edges and exact brand colours, about 1.2 MB a frame (~4 GB for 108s) and 2–3x slower than a draft. Run it in the background. For a quick preview cut, prefix VV_QUALITY=draft (1x JPEG); render the final with the default. The optional frame range re-renders a single shot after a fix; each bound is its own argument, so it is safe under zsh. Re-render a range with the same quality as the rest, or Step 10 refuses to mix them.
Pass --blur 4 (first) for the final render only: it adds motion blur and makes the render about 8× slower than the default. Render drafts and single-shot fixes without it, unless re-rendering a range of a blurred final.
Step 10 — Mix and encode
out="<brand.output.dir>/<slug>/<slug>.mp4"
bash SKILL_DIR/scripts/mix-encode.sh "$work" <audio> <duration> "$out" [music-file]Normalises the voice, ducks the music under it, lays the SFX on top, pads everything to the full duration, and targets −14 LUFS. Video is H.264 at CRF 16 (slow preset, capped at 16 Mbps so film grain can't balloon the file), converted and tagged as BT.709 so phones show the brand colours as designed. A 108s vertical lands around 200 MB: high enough to survive the platform's own re-encode.
Step 11 — Verify, then report
Before reporting, extract a frame from the encoded file at a shot you changed, read it, and check the duration and size:
ffmpeg -v error -y -ss <t> -i <out.mp4> -frames:v 1 -vf scale=360:-1 <work>/verify.jpg
ffprobe -v error -show_entries format=duration,size -of compact <out.mp4>A still rendered from the HTML is not proof the video contains the fix.
<slug>.mp4 · 1080x1920 · 108.2s · 112 MB · -14.7 LUFS
shots 30 · sound cues 117 · images 4 · clips 2 · captions fixed 6
credits James Gosling 2008 (Wikimedia, CC BY-SA) · Solitary oak (geograph, CC BY-SA)
unverified music taste — listen before posting · timeline years are stylistic
Next: preview it on a phone, then post it with the credits in the descriptionmix-encode.sh prints raw ffprobe output; reformat it into the line above for the report. Get the credits with find_media.py credits "$work" and give them to the user ready to paste into the description. Name any unlicensed clip on its own line.
Non-negotiables
- Shot list confirmed before authoring. Rendering is cheap; redesigning thirty shots is not.
- Every image and clip is looked at before it is used. Never ship media you have not seen.
- Every file is credited. Licensed or not, each photo and clip goes in the description credits, and
the report names every ⚠ file.
- Every fix is verified in the encoded file, not just in a still.
- Transcription is local. Never upload the audio. Voices are synthesized locally too.
- Characters are original. Never draw, name or imitate an existing cartoon, film or game character.
- Synthetic voices are named in the report, so the user can disclose them when they post.
- Colours come from the kit's
--brand-*tokens, never hex values in shots. One brand accent;--brand-successonly for success states,--brand-erroronly for errors. - A client's logos and fonts come from the client. Never fetch a company's logo or font from the web, and never put a competitor's logo in a video unless the client asks for it.
- Timelines are deterministic: no
Math.random(), noDate.now(). The grain uses a seeded PRNG. - Captions never cover the element the viewer is meant to read; hide them via
NOCAPinstead. - Report every image credit and every invented detail (dates, labels) that isn't in the audio.
- Never commit rendered files or
work/. faceCam()in/out times match theextract_face.shranges exactly; never re-time a face shot alone.
The same holds for clip() and extract_clip.sh.
Failure states
| Symptom | Cause | Fix |
|---|---|---|
Executable doesn't exist at …ms-playwright… | playwright-core updated, its Chromium build isn't installed | re-run setup.sh |
| Fix visible in stills but not in the video | frames never re-rendered — a range passed as one quoted string renders zero frames | use render-frames.sh with the range as two separate args; check frame mtimes |
| Video shorter than the audio | an audio filter trimmed the stream | mix-encode.sh pads and trims to the duration; don't hand-roll the mix |
| Output file is hundreds of MB | film grain defeats compression at constant CRF | keep the -maxrate cap in mix-encode.sh |
frames/ mixes high-quality PNG and draft JPEG frames | a range was re-rendered with a different VV_QUALITY | re-render the whole video with one setting |
| Accent colour looks orange or washed out on a phone | an encode without the BT.709 conversion and tags | use mix-encode.sh; don't hand-roll the encode |
| Frame render fills the disk | high-quality PNG frames are ~1.2 MB each | free space, or preview with VV_QUALITY=draft and render the final once |
| Wrong or fallback font in stills | fonts not downloaded for this brand | re-run setup.sh with the brand file |
PAGE ERROR in render output | a script error in the timeline | fix it; GSAP only warns on missing selectors, so also check each shot visually |
| Whisper sits at low CPU for minutes | model download on first run | expected once; the model is cached afterwards |
| Shot renders blank | shot() start ≥ end, or the shot id is misspelled | check the shot row's in/out times |
missing frame face/… or clips/… stops the render | a faceCam() / clip() range is wider than the extracted one | re-run extract_face.sh / extract_clip.sh with that shot's in/out |
find_media.py search returns nothing | query too specific, or a source is down (its error prints on stderr) | the name plus the company; then --kind video; then a logo or an era look |
| YouTube fetch crawls or fails with a challenge warning | no JS runtime, or yt-dlp is out of date | needs node or deno on PATH; pipx upgrade yt-dlp |
| Clip shows the wrong moment | --from is in the fetched clip's seconds, not the original's | subtract the --section start |
| Old clip sound still in the mix | a stale clips/<name>.wav from an earlier cut | re-run extract_clip.sh for that clip; it removes the old .wav |
kokoro-onnx is not installed / Voice model missing / Cannot load the 'base' alignment model | script mode set up without voices, or offline before the models were fetched | setup.sh --voices (online, once) |
Unknown voice from speak.py | a cast file names a voice that doesn't exist | pick one from the list it prints |
| A host never moves its mouth | its who doesn't match the speaker name in the script | use the lowercase name from speech.json |
speech.js has no line N for … in PAGE ERROR | line()/say() asks for a line the script doesn't have | count that speaker's lines in speech.json from 0 |
| Face shots look grey and washed out | HDR (HLG) phone recording, tone-mapped without metadata | record in SDR (iPhone: Settings › Camera › Formats, HDR Video off) |
… is a version-1 brand file | a brand.json from before kits | init_kit.py --from <that file> (keeps the original as brand.v1.json) |
Brand kit … is not installed, or changed since it was | the kit is new or was edited after setup | setup.sh <kit> |
font … did not load from render.js board | a kit font file is missing or broken, or a Google family name is misspelled | fix the kit's fonts, re-run setup.sh <kit> |
| Logo invisible on the end card | a dark logo on a dark canvas, or an SVG with no size | add wordmark.onDark; give the <img> a height (see Logo end card) |
keep it where your ai can reach it.
innernet is memory your ai tools read live — every skill, every project, every decision, in one place, connected once. save this skill to yours, or publish one of your own as a link like this.