--- name: video-review description: Run TJ's post-edit QA pass on a rendered cut BEFORE he watches it — the internal review that catches the things that waste his time. Use this skill when TJ says "video-review", "/video-review", "review the cut", "QA the cut", "check the edit", "screen this cut", "did I miss anything", or after ANY render in /edit or /finalcut (it is the mandatory final gate of both). It scaffolds the review automatically — a per-segment contact sheet from the actual render, a voice-vs-music stem-balance table, and a premature-cutoff report against the word-level transcript — then the agent LOOKS and proposes fixes. Catches: black/blown-out/pointed-at-the-ground footage, lines cut mid-sentence, music louder than the voice, vertical text mashing/unsafe captions, sub-1s flashes. Proposes fixes; doesn't silently auto-fix. --- You are running the **internal review** on a rendered cut, before TJ ever watches it. This is the verification gate that saves him time — born 2026-06-27 on Koh Samui_2, after cutting the dead pointed-at-the-ground footage and a mid-word audio clip turned a weak first version into a strong one with two targeted fixes. This is **both** a standalone command (run it on any rendered cut, anytime) **and the mandatory last step inside `/edit` and `/finalcut`** — the same checklist, one file, two entry points, so it is never skipped before a cut reaches TJ. ## Governing principle — RE-ROUTE, don't DISCARD (don't be rigid) This is the master rule; it outranks every check below. A flag identifies a defect and proposes the **lightest fix that PRESERVES what's good** — the review never throws material away by reflex. **Judge PICTURE and AUDIO separately. Never let a verdict on one kill the other.** - **Bad picture + worth-keeping audio** (camera pointed down during a great riff — the data-net / dating-app clip, 2026-06-27) → KEEP the audio, float it over B-roll with `asrc`. Do NOT drop the segment. This is THE save. - **Good picture + junk audio** (wind, whisper hallucination, a mumble) → keep the shot as `mute`/`nat` B-roll. Don't lose a beautiful frame because the sound is dead. - **Drop a segment ONLY when BOTH the picture is unusable AND the audio isn't worth keeping.** Before proposing to cut or replace anything for a *picture* problem, **read that clip's transcript (the helper prints each spoken segment's gist) and LISTEN — decide if the line earns its place.** A flag is a prompt to judge, not an order to delete. Rules here are heuristics, not laws; the rule's *purpose* (a cut that's beautiful AND keeps TJ's voice) is the master, not its letter — when unsure, ASK TJ. (`feedback_creative_rules_are_heuristics`, `feedback_dialogue_is_central`) Keep this as ONE principle, not a pile of edge cases — it governs every picture/audio trade-off so the core checklist stays simple. ## The boundary — be honest about it The command **automates the scaffolding** (frame extraction, stem measurement, transcript-vs-cut comparison, flag generation). It does **NOT** replace judgment: "is this frame actually crap?" is still **you looking at the contact sheet**. What this buys is making that looking **systematic, fast, and non-optional** — not a step you shortcut because the render looked fine in motion. `feedback_verify_visuals_not_metadata` is the law here: a clean duration/spec is a hint to LOOK, never a conclusion. ## Run it ``` python3 .claude/skills/video-review/review_cut.py /render/.mp4 ``` It infers the EDL (`/edl/_edl.json`), the stems (`/render/stems/_{voice,music}.wav`), and the transcripts (`/transcript/out/`). It writes `/render/_review/CONTACT_.jpg` and prints the stem table, the dialogue out-points, and a FLAG list. (Pass the EDL explicitly as the 2nd arg if it lives elsewhere; `--transcript DIR` to override.) Then **Read the contact sheet** and walk the checklist below. The render-and-stems must already exist (this reviews a finished render — it does not re-cut). ## The checklist 1. **Junk-frame scan (YOUR EYES).** Open `CONTACT_.jpg` — one frame per segment, sampled from the actual render, labeled `idx clip mode / off +dur`. Hunt for: - **black / heavily backlit** (a dark subject against a bright window — the C1557 gong) → swap clip/in-point. - **blown-out / overexposed** sky or highlights → pick a graded in-point or another clip (never re-grade his footage). - **dead "pointed-at-the-ground/ceiling" footage** — camera not actually on anything (the C1565 beach conversation shot at the tiles). Per the governing principle: check the segment's gist FIRST. If the audio is worth keeping → **float it over beauty with `asrc`** (set the segment's video to a beautiful B-roll clip, keep the line via `asrc:(origclip,time)`) — this was THE key save. Only if the audio is also throwaway do you drop it. - **tilted / obviously off** frames. Shake can't be seen in a still — flag the segment for a human playback pass, don't claim it's smooth. When a frame is bad, pull the source clip at that time to confirm, then propose a specific replacement clip + in-point, or an `asrc` float over named B-roll. 2. **Premature-cutoff check.** The script compares each spoken segment's out-point to the word-level transcript and flags `CUTOFF` (the out splits a word) and `CUT?` (the sentence continues right after). Every flag: decide if the line **completes** — never end a cut mid-sentence or before the punchline (`feedback_dialogue_is_central`). Fix by extending `dur` to the natural end (check the next word's timestamp) or moving the out. A trailing complete thought ("…yeah", a finished question) is fine — use judgment, the flag is a prompt to check, not an order. 3. **Mix-balance check.** The stem table shows voice vs music max-dB at every dialogue cluster and a few visual beats. Confirm: under speech music sits **~8–16 dB below the voice peaks**; on visual beats music **swells up**. `MIX` flag = music within 3 dB of the voice (it will mask the line — usually a *quiet* source the music drowns, so **raise that beat's `vg`** before ducking music deeper, per `feedback_video_music_direction` / the Koh Samui mix learnings). `CLIP` flag = voice hitting the limiter (max ≥ −0.3 dB) → pull that beat's `vg` down. You can't hear it — tune by levels, then TJ confirms by ear (`TJ is the ears`). 4. **Vertical / text safe-area check** (9:16 reels + any burned captions). On the portrait contact sheet: is the **subject still framed** (heads not cropped — adjust `px` bias)? Do any captions/burned text **collide, stack, or sit outside the safe margins** (TJ's "vertical mashing of text")? Reels open on the subject in the first 1–2s (`feedback_social_hook_subject_first`). 5. **Cut-sanity.** `FLASH` flag = a sub-1s segment (subliminal) unless intended. `SILENT-FLOAT` flag = a seg with `asrc` but `mode:"mute"` — the engine zeroes gain BEFORE reading asrc, so the floated line is silent; floats must be `voice` mode (caught on Bangkok 2026-07-11: all 15 floats — the film's heart-lines — rendered silent, and the cluster analysis masked it because mute segs read as visual beats. TJ heard '30s of no dialogue' and thought it was pacing). Also watch for **hard audio slams** (a loud nat/engine/reaction beat cutting to silence with no L-cut trail → ears get hit) and **same-clip back-to-back jump-cuts** that can double a spoken line via the 0.30s trail bleed (set `{"trail":0.0}` on the golden seg + its same-clip neighbors — the Koh Samui "famous in Samui" gotcha). 6. **Censor check (when a clean/advertiser-friendly upload is claimed).** If the cut censors profanity, verify the ARTIFACT — but **MEASURE the word-window dB, do NOT rely on transcribing.** Hard-won twice (2026-06-27): - **The mute must actually be silence.** `ffmpeg volume=0:enable='between(t,…)'` only attenuated ~6 dB (left an audible −16 dB remnant) — the `areverse` chain breaks its timeline. The reliable mute is **hard sample-zeroing the wav in Python** (`_censor_wav` in `koh-samui_2/assemble_film.py`: read int16 wav, `data[L*sr:R*sr]=0`, write back) → true −91 dB silence. - **Transcribing to verify gives FALSE POSITIVES.** mlx_whisper *hallucinates* the censored word (or `****`) over the silence from sentence context — it'll "report" profanity that isn't audible. Confirm by `volumedetect` on the **tight word window** of `render/stems/_voice.wav` (and the `.m4a` FCP plays): want ≈ −91 dB. Get the word's exact span from the word-level transcript, measure THAT span (a window even slightly wider catches the loud neighbours and misreads as residual). Selective MUTE beats a bleep (less corny, fits TJ's aesthetic); if a word can't be muted without gutting the line (e.g. "fuck off, boy"), drop the beat instead. - **Censor DEPTH is TJ's creative + distribution call — do NOT default to "maximally clean."** A bulletproof −91 dB silence is advertiser-logic; a LIGHT / barely-muted curse is often funnier (the half-heard bleep IS the joke), and TJ will knowingly trade some reach for the laugh, especially on a small video. PRESENT the choice (full-clean vs comedic-partial) and let him pick; don't bleach it unilaterally. And don't tidy away the funny TAIL chasing runtime — e.g. cutting "...dumb as hell" off the censored line dropped the exact beat TJ wanted (Koh Samui 2, 2026-06-27; the over-edit trap, [[feedback_creative_rules_are_heuristics]]). 7. **Report back.** Summarize as PASS or a short FLAG list: `timestamp → symptom → likely cause → proposed fix`. **Propose, don't silently rewrite** the cut. Apply fixes only on TJ's go (or obvious mechanical swaps you call out). Re-render with `--seg A B` for fast section iteration instead of a full render. ## The dialogue pass — MANDATORY on any talking-head or VO cut (added 2026-09-05, after TJ caught four of these by hand in one evening) Run **before** the contact-sheet pass, on the cut list and the word-level transcript: ``` python3 .claude/skills/video-review/dialogue_audit.py --parts [--fix-out cuts_fixed.json] ``` It checks, per cut: (1) the OUT clips the last word; (2) the IN lands inside a word; (3) a **mid-sentence pause** >1.5 s between word STARTS (a stumble or a think left in the take — TJ reads it as "a bad take" and even as "out of sync"); (4) a **blooper at the head** — the first two words recur within 4 s, i.e. a restart, so the cut should start on the second attempt; (5) with `--parts`, it transcribes the first and last 2.5 s of every rendered part and confirms the expected first and last word are actually there. - **Cut from WORD-level timestamps, never sentence/segment ones.** Segment ends put 24 of 28 tutorial cuts at or before the last word's end ("…for a Mac" clipped). - **Word END times from whisper are unreliable across pauses** (they stretch to the next word). Measure gaps start-to-start; when in doubt, re-transcribe the 6–8 s around the join with `whisper-large-v3-mlx` at word level, or read the 50 ms energy envelope — that is what exposed "another great / another great". - **A restart in the take is not a gap to trim, it is a first attempt to drop.** Start on the second attempt. - Pauses and bloopers are jump cuts on a static talking head; cut them, TJ accepts the jump. Cover with the cutaway if one is already scheduled there. - **Never report "fixed" without the head/tail transcription of the rendered part.** The log records what finished; only the audio says what is there. ## Where it sits - **Inside `/edit`:** the render step ends by running this before the cut is shown to TJ. A render is not "done" until it has passed the review. - **Inside `/finalcut`:** the cut should have passed this before conforming; re-run on the derived reel (the vertical crop + any captions are exactly what check 4 exists for). - **Standalone:** run anytime TJ asks, or after any manual tweak to a cut. - **Not this skill:** creative cut decisions (`/edit`), music/structure *variants* (a separate concern — don't muddy review with exploration), copy/thumbnails (`/distribute`). ## Housekeeping - Review artifacts go in `/render/_review/` (contact sheet + extracted seg frames). Safe to overwrite per render; not a deliverable. - **Process hygiene (2026-09-05):** every `-loop 1` still gets `-t` and the output `-shortest`; a background render that goes quiet gets `ps aux | grep ffmpeg` before any status is reported (30 min lost to an unbounded still). A VO longer than its segment is silently clipped by `-t` — concatenate the picture ranges first, then lay the VO. - LIVING DOC: when the review misses something that later bit us, or catches a new class of defect, add the check here so it's never relearned.