--- name: edit description: Edit a dump of video clips into whatever TJ actually wants — a cinematic story/film, a montage, a highlight reel, a talking-head + B-roll piece, or a simple assembly. Use this skill when TJ says "edit", "/edit", "edit these clips", "edit this footage", "cut this footage", "let's edit", "make a reel/montage/video", or drops a folder of clips and wants something made. DO NOT assume the output type — the first move is to look at the clips and ASK TJ what this is and what he wants, then build it together. Drives the Vision Editing Layer pipeline (ingest → analyze → editorial → render → FCPXML) with review gates. Mechanical steps automated; the creative cut is interactive. --- You are editing TJ's video footage. **Do not assume he wants a montage.** Your FIRST job is to understand the footage and the intent, then build it with him, holding the review gates. Layout (restructured 2026-06-22; was `vision-experiment/`). The macro `Entrepreneurship Projects/YouTube TJSMind/` holds the engine and the projects as SIBLINGS: - **`pipeline/`** — SYSTEM engine only (set-and-forget): `analyze.py`, `beatmap.py`, `gen_fcpxml.py`, `emit_film_fcpxml.py`, `audio/` (beatmap cache). - **`pipeline/audio/`** — the SHARED MUSIC LIBRARY now lives IN THE REPO here (moved off the external drive 2026-06-22): mood folders `Afrobeats/ Ambient mild documentary/ Aerials/ Urban Hip Hop Boom Bap/ Deepdive/ …` + `beatmap_*.json` cache. Scripts set `AUDDIR` to this path. Music is system/pipeline infrastructure, so it lives with the engine and no longer needs the drive mounted. - **`projects//`** — one folder per video, each self-contained: its OWN `assemble_film.py` (+ `assemble.py` for the older beat-montage) + `proxy/ frames/ transcript/ edl/ render/ signals/`. Existing: `huacachina/` (sand-dunes, shipped), `7lakes/` (current reference), `7 Immigrant Books/` (in-flight; folder name not yet slug-normalized). - Rationale docs live loose in the macro: `vision-editing-layer-plan.md`, `video-editing-pipeline.md`, `POSTMORTEMS.md` (2026-06-21 audio entry). **Source FOOTAGE — external drive (4K).** The drive now holds ONLY per-project footage (audio moved to the repo): ``` /Volumes//TJSMind Working Folder/ ← macro (on-disk name is "TJSMind Working Folder", NOT "TJ's Mind Working Folder") ├── Huacachina/ ├── 🇵🇪7 Lakes Ausangate Trek/footage/ + distribution/ ← (distribution/ = tour-listing screenshots when a video is review-framed) └── /footage/ ``` Mirror this in the repo: each project gets its own self-contained workspace at `projects//` (its own `assemble_film.py` copy with that project's `FF`=footage path, shared `AUDDIR`, `CLIP_DUR`, and `SEGS`/`SEGS_REEL`; plus `proxy/ frames/ transcript/ edl/ render/`). **Starting point for a NEW project:** for a **dialogue-driven doc** copy `projects/koh-samui/assemble_film.py` (TJ's gold standard — duck-DOMINANT `music_envelope`, per-source `vg` balance, `.mov` support, looping music bed, `asrc` monologue-over-B-roll); for a **music-led cinematic-review** copy `projects/7lakes/assemble_film.py`. Set `PROJ`, `FF`, `EXT`, `CLIP_DUR`, `TRACKS`, `SEGS`/`SEGS_REEL` (if the footage is MIXED `.mov`/`.mp4`, resolve extension per-clip with a `cext()`/`cpath()` helper instead of one global `EXT` — e.g. Koh Phangan, where `C1623`–`C1647` are `.mp4` and the rest `.mov`). `huacachina/` is the older beat-montage reference. Both share the shared-audio + `MACRO`-folder layout. ## TJ's taste — bake these in (memories, but listed here so they're not missed) - **Anti-corny.** Reach past the conventional move; unique/thoughtful/understated over slick/upbeat/formulaic. He picked boom-bap over "upbeat travel" sound for this reason. Use as a tie-breaker. (`feedback_tj_creative_aesthetic`) - **Dialogue is the differentiator.** His voice/personality is what makes it his, not a template. Mine footage for reactions, reflections, riffs and weave them through EVERY cut — even short Reels. Don't understate it; bring more over time. (`feedback_dialogue_is_central`) - **Hook on the subject (him) in the first 1-2s** for social — never open on scenery; an audio/voice hook is strongest. Long-form too: open cold on the most intriguing line and cut the orientation to one breath — a long setup is where people skip (Koh Phangan, 2026-06-28). (`feedback_social_hook_subject_first`) - **L-cuts / J-cuts + SMOOTH AUDIO TRANSITIONS** (`feedback_l_cuts_j_cuts`): never hard-cut loud audio (engine/reaction/dialogue) to the next scene — it slams the ears. `assemble_film.py` now trails every audible segment's audio ~0.30s past its picture-cut and overlap-mixes it into the next (a built-in L-cut crossfade). `asrc` audio-decouple floats a line over a better visual. Carry smooth transitions through every edit. - **BUT the 0.30s trail BACKFIRES on consecutive jump-cuts within the SAME clip** (e.g. three C1458 chunks in a row): it bleeds one spoken line over the next = a doubled/muddy voice that wrecks a golden moment (Koh Samui "Buddy, I'm famous in Samui", 2026-06-25). Per-seg opt **`{"trail":0.0}`** disables the bleed to keep a line CLEAN — set it on the golden seg AND its same-clip neighbors. (Gotcha: any new per-seg opt must be threaded through the `main()` timeline-dict (`"trail":opts.get("trail")`) AND read as `s["trail"] if s.get("trail") is not None else TRAIL` — else it silently no-ops because 0.0 is falsy / the key is absent.) - **"Audio low" in a dialogue section is usually RELATIVE masking, not absolute level** — measure the voice vs music STEMS before boosting, then pick the right lever for the case: if music is ALREADY deep-ducked (~−31) and voice is still buried, punch the VOICE `vg` up (a loud source already limits at peaks, so a vg bump lifts body/presence; don't overdrive into the limiter); if "hard to hear when I speak" and the music ISN'T deeply ducked under speech yet, deepen the global `DUCK` instead (e.g. 0.06 on Koh Phangan) rather than nudging vg. Measure which case you're in first. - **Change of pace.** Alternate fast action with slow reflection; the signature move is dropping the MUSIC OUT so a reflective line breathes, then swelling it back. End reflective beats on stillness, not a hype drop. - **DON'T OVER-EDIT — keep it simple, let it breathe.** TJ's gold-standard dialogue-doc magic is **voice floating over BEAUTY** (monologue over bright scenery, cutting to the speaker's face only on gut-punches), NOT a wall of talking heads. Chasing a stricter matched-voice rule on Koh Samui over-cut it and *lost* the special sauce; the fix was FEWER cuts + more breathing scenery + continuity (2026-06-26). "Artificial AI chop over-edits" is the thing to avoid. Editorial heuristic (a DEFAULT, not a law): a voice plays over the person saying it, or over scenery with no people — not floated over a different non-talking person; but cutting a landscape that includes someone is often fine — just **don't overdo the float-over-mismatch** (corollary: when a line NAMES someone — "subscribe to this man, Hamid" — that person must be ON screen, not just TJ). This same voice-over-beauty move also **hides soft-focus footage**: where TJ shot himself out of focus (background-focused), keep the soft talking-head ~1-3s to ground him, then float the rest of the line over beautiful B-roll rather than holding on the soft face (Koh Phangan, 2026-06-28). (`feedback_creative_rules_are_heuristics`, `feedback_edit_complete_lines_ground_then_float`) - **Start from the version that WORKS; apply targeted fixes — don't rebuild.** When a cut is fundamentally right, iterate on it (swap bad clips, fix audio, trim whole redundant beats) rather than recutting from scratch to "bring the good across." (`project_koh_samui_gold_standard`) - **Let every line COMPLETE — never cut dialogue mid-sentence** (the thing that broke Koh Phangan: "so much of the voice is cut off before the sentence ends, unsatisfying"). Split-and-float makes it easy to cut mid-word. Procedure: get WORD-level timing for every spine clip (`mlx_whisper --output-format json --word-timestamps True`) and set each line's audio window to end on its **last word + ~0.35s breath**; the float (last seg of a line) must run far enough to finish the sentence — don't trust round-number durations. TRAP: whisper's SENTENCE-level SRT boundaries are coarse, so verify the START timestamp holds the intended words (a cut once pulled "very confusing, right" from the wrong second instead of the line meant). (`feedback_edit_complete_lines_ground_then_float`, `feedback_dialogue_is_central`) - **Afrobeats is TJ's SOUND SIGNATURE** (decided 2026-06-21). Default to it for TJSMind edits — honest (he's Nigerian), unique, anti-corny, chill-yet-energetic. Growing library now lives in the SHARED drive folder `/Volumes//TJSMind Working Folder/audio/Afrobeats/` (from UpBeats, free for creators) — shared across all projects. Use the `afrobeats` track; `MUSIC_OFFSET` starts short cuts in the groove. - **Music direction** (`feedback_video_music_direction`): long-form = ONE cohesive track throughout (splicing two songs reads as a corny seam); match the track's energy ARC to the film's; the track must be long enough to avoid a loop seam (only the longest tracks cover 2-3 min); short and long can be DIFFERENT tracks/vibes (Huacachina: Urban for the short, Ambient for the full). Energy-arc the candidates with `beatmap.py` and pick by arc-fit + length + vibe, and say why. - Keep his profanity/rawness (it's for his circle, not IG-influencer reach) unless he says otherwise. ## Step 0 — Look, then ASK (never skip) 1. **Survey the footage yourself first**: list clips + durations; build a thumbnail **contact sheet** (ffmpeg thumbnails → `tile=` filter, keep sheets <2000px) so you can see everything at once; **transcribe the long clips** (mlx-whisper, word timestamps) to find the dialogue/emotional arc. Note "feeling/couldn't"-type loops = whisper hallucinating on wind = silent B-roll. 2. **Then ASK** what it is, what he wants out of it (film/story, montage, highlight, talking-head, selects), who it's for + where (YouTube 16:9 / Reels 9:16 / both), vibe + length, music vs voice. Bring a recommendation; research current best-practice if it's a big creative call. 3. State the plan back as a **review gate** before cutting. ## Workflows - **Cinematic film (`assemble_film.py`)** — the dialogue-driven story. Hand-timed segments, each with an AUDIO MODE (`mute`/`nat`/`react`/`hook`/`voice`/`resolve`); music has a volume envelope that DROPS under the reflection and swells back. Arc: hook → build → drop (action) → reflection (music out, voice clear) → resolve (stillness, fade to black). Renders ` [--render]`. Optional per-seg opts `{px: crop-bias, asrc: (clip,sin)}` (see Capabilities). This is the flagship path for trip/story footage. - **Beat-synced montage (`assemble.py`)** — pure energy, cuts snap to beats. `<30|60> --render`; `--natbed` for engine bed. - **Highlight / talking-head / simple assembly** — lighter treatments; transcript-driven spine for talking-head. Deliver a **matched set** when asked: flagship 16:9 → derive vertical Reel → long-form. A/B different music tracks (group the library into moods, render the same cut against 2) — the track changes everything. ## Audio architecture — the WORKING pattern (hard-won; don't deviate without measuring) `assemble_film.py` renders in passes for reliability: 1. **Pre-extract each segment to a local temp clip** — one accurate **output seek** (`-i clip -ss sin -t dur`, NOT `-ss` before+after `-i`, which mis-windows audio). Avoids 30+ simultaneous deep seeks deadlocking the mixer. 2. Per-seg audio: reset PTS (`asetpts=PTS-STARTPTS`), gain/lowpass per mode, fades via **`areverse,afade=t=in,areverse`** — NEVER `afade=t=out` (it silences the whole clip from ~0.5s in ffmpeg 8.x when combined with a seek; this silently ate all the dialogue once). 3. Concat-copy temps → assembled cut. **Extract the assembled audio to a clean WAV** before mixing (mixing from the concat mp4's audio makes amix swallow the music at boundaries). 4. **Pre-bake music+envelope to a WAV** (the `volume=eval=frame` envelope lags ~1s if applied inside the final amix), then sum the two WAVs. 5. Keep music OUT of the voice hook (arrival clips have their OWN background music → overlap = double beat); swell it in over the TAIL of the hook line (no dead air, no clash). ### DUCKING + MIX BALANCE — the gold-standard pattern (Koh Samui 2026-06-25; TJ's "best work, the music mix is brilliant") This is the single biggest quality lever. Bake it into every cut. 6. **Duck-DOMINANT envelope** (`music_envelope(clusters, total)` in `koh-samui/assemble_film.py`): music = `MUSIC_BASE` on visual beats, ducks HARD to `DUCK` across every spoken cluster (fast duck-in ~0.12s after speech starts, gentle ~0.55s swell-back after). The OLD bug: a "swell to BASE at hook_end" breakpoint landed *inside* the opening speech cluster and pushed music back up under the voice — music ended up LOUDER than the voice. The fix is to make ducking dominate: never emit a BASE breakpoint inside a cluster. Levels that worked: `MUSIC_BASE=0.62`, `DUCK=0.13`. 7. **`voice_clusters` = `voice`/`react` ONLY** (NOT `resolve`). Resolve beats are musical/visual → music must stay UP for the ending swell. Including resolve ducks the music exactly when you want it to soar. 8. **Per-source `vg` balancing — measure the VOICE STEM, not vibes.** Sources differ wildly: quiet/`asrc`/distant lines need `vg≈1.8`; loud/excited/close lines (a shouted "kaaaah!", a close talking-head) CLIP at `vg≈1.4`. After a render, `volumedetect` the **voice stem** per dialogue beat — any beat hitting `max_volume: 0.0 dB` is clipping the limiter → pull its `vg` down to ~1.0. Don't use one global vg. 9. **Mix target (measure both stems separately — they're saved):** during speech, music should sit **~8–10 dB below the voice PEAKS** (audible bed, not overpowering); on visual beats music **swells ~+13 dB**. Koh Samui open went from broken (music −17 / voice −28, i.e. music *louder*) to right (voice −8 peak / music −17 peak; music +13 dB on the vista). TJ's note verbatim: "lower the dB of music during speech, increase during the visual vibes — audible but not overpowering." **Debugging audio = measure the rendered file** with `ffmpeg ... -af volumedetect` in 0.25s windows, and **bisect the filter chain**. A perceptual symptom ("cuts to music" / "music overpowers voice") can have a different cause ("voice was never there" / "envelope swelled mid-speech"). The voice + music stems are preserved in `render/stems/` — measure them independently. See the 2026-06-21 postmortem. ## Capabilities in assemble_film.py - `px` (portrait crop bias -1..+1): shift the 9:16 center-crop to keep an off-center subject framed when cropping 16:9→9:16. - `asrc: (clip, sin)`: source a segment's AUDIO from a different clip/time than its video — float a line over a better visual (used to put "…in a desert" over the subject sitting, not an empty pan). The L/J-cut foundation. **An asrc float's mode must be `voice` (or another gain>0 mode), NEVER `mute`** — the engine zeroes gain before reading asrc, silencing the line (Bangkok 2026-07-11). - Music drop-out envelope, nat-sound layering, accurate frame-aligned output. ## Internal review — MANDATORY before TJ sees ANY rendered cut (`/video-review`) A render is NOT done until it has passed `/video-review`. After every render, run it and act on the flags BEFORE showing TJ: ``` python3 .claude/skills/video-review/review_cut.py /render/.mp4 ``` It scaffolds the review automatically — a per-segment contact sheet from the actual render, a voice-vs-music stem table, and a premature-cutoff report vs the word-level transcript — then **you LOOK at the contact sheet** and propose fixes. It catches the exact things that waste TJ's time: black/blown-out/**pointed-at-the-ground** footage (float the audio over beauty with `asrc`), **lines cut mid-sentence / before the punchline**, music louder than the voice, vertical text mashing, sub-1s flashes. This gate was born 2026-06-27 on Koh Samui_2 — cutting dead ground-footage + a mid-word clip turned a weak v1 into a strong cut with two targeted fixes. See the `video-review` skill for the full checklist. (It is ALSO a standalone command TJ can invoke anytime.) ## FCPXML / conform to Final Cut (after TJ approves the cut — hold the gate) **Hand off to `/finalcut`** — that's the dedicated command for conforming an approved cut into Final Cut (emits frame-aligned FCPXML, relinks 4K by path, mutes scratch, rides the rendered mix as a connected lane; handles landscape AND the vertical reel crop). `/edit` ends at the approved render; `/finalcut` takes it into FCP for finishing. (Legacy beat-montage path: `python3 gen_fcpxml.py [music.mp3]`.) ## Non-negotiables - **Intent before action.** Look + ask first. - **Rules are heuristics, not laws — know when to deviate; when unsure, ASK.** Don't over-learn one correction into a blanket policy (the over-edit trap). The rule's *purpose* is the master, not its letter. (`feedback_creative_rules_are_heuristics`) - **When TJ says it's off, believe him and verify against the ARTIFACT — don't insist.** Look at the frame / transcribe the audio to settle it (twice the intro "stutter" was a real stale-stem bug while I claimed it was fixed). Pin the EXACT beat before re-cutting; use `--seg A B` fast section renders to iterate instead of full renders. - **Frame taste overrides the numbers.** View frames; a high-sharpness frame can be blown-out/rotated/junk. Re-check any window you picked by in-point. - **Review gates:** planned timeline before render; **`/video-review` before TJ sees the render**; cut before FCPXML. - **TJ is the ears.** You can't hear the mix — tune by levels, then have him confirm voice/reaction/music balance; iterate fast. - **Be honest about automated vs. judgment.** Plumbing is automated; selection/hook/arc/pacing/taste is you + TJ. ## Outputs & housekeeping - Keep `render/` clean: only the CURRENT deliverables at top level (they keep the same names on re-render), everything exploratory in `render/_archive/` with a README. Re-renders overwrite in place. - **Version meaningfully-different CUTS as distinct NAMED sets (render + EDL + stems + FCPXML), non-destructively.** Use a different cut tag so filenames don't collide (Koh Samui kept `flagship_ambient` / `flagship_coco` / `hybrid_coco` side by side). The original 8.7-min cut was fully restorable WITHOUT re-rendering ONLY because its EDL + stems had distinct names — nothing overwrote. Never overwrite a version TJ might want back. (`project_koh_samui_gold_standard`) - Naming: `{cut}_{track}_{orient}.mp4`; EDLs in `edl/`; dialogue word-timestamps in `transcript/`. ## Known limits - Few distinct "action" clips → fast sections need reuse; flag it, offer the shorter cut. - Stills can't show pan smoothness/shake — lean on flow signals + a human playback pass. - Portrait center-crop loses wide-shot sweep; use `px` bias, or offer a blurred-pad fill. - True L/J cuts (audio leading/trailing ACROSS a cut) not built yet — only same-length audio-decouple. ## Dialogue cuts — hard rules (2026-09-05, Claude Code tutorial talking head) - Cut on **word-level** timestamps; out = last word end + 0.25 s (capped 0.06 s before the next word), in = first word start − 0.10 s. - Before any render reaches TJ, run `.claude/skills/video-review/dialogue_audit.py` on the cut list: it catches clipped words, in-points inside words, mid-sentence pauses (>1.5 s start-to-start) and bloopers (a restarted opening). Then transcribe the head and tail of every rendered part. - A restart in the take ("another great… another great part") means START on the second attempt, not trim the gap. Whisper merges restarts; re-transcribe the join at word level with the large model to see them.