1. Home
  2. How it works

The pipeline

From one idea to a published Short.

Shotcoda is twelve scripted stages that pass files to each other. An LLM director makes the creative calls, open models generate every asset on your GPU, and deterministic code does the timing, editing, mastering and checks.

Design principles

Everything generated locally

Pictures, motion, voice, music, sound effects, captions and graphics come from open-weight models on your own hardware. There is no stock footage and no paid generation API.

The narration is the clock

The voice is generated first. Every cut, caption and graphic is timed from the spoken words, so nothing drifts and every caption is exact.

Built for agents

Creative intent lives in two JSON files. Any stage can be re-run on its own, so an AI agent can operate the whole studio the same way a person would.

00
Finds what to make

Research

Scores candidate topics by search demand, trend and how comparable Shorts perform. The winner becomes a brief with a verified fact sheet, where every number has a source.

Engine
Google Keyword Planner + YouTube Data API + primary sources
Runs on
CPU · free APIs
Typical time
minutes
Writes
brief.md
01
Writes the film

Script & shot list

The director writes a hook-first narration and a shot list: image prompt, one camera move, overlay and sound cue per shot. Everything lands in one editable JSON plan.

Engine
Local Qwen3.8-27B, or the Claude / ChatGPT account you already have
Runs on
GPU 0 + 1
Typical time
≈ 1 min
Writes
plan.json
02
Designs a voice and speaks it

Narration

A narrator is designed from a text description, so no real person is cloned. Each sentence is spoken in several takes and transcribed back, and the cleanest take wins. Word-level timings come out too.

Engine
Qwen3-TTS 1.7B VoiceDesign + Base, faster-whisper
Runs on
GPU 1 + CPU
Typical time
≈ 4 min
Writes
vo.wav · words.json
03
Paints every shot

Keyframes

Vertical 1088×1920 keyframes are painted in one consistent visual world. A contact sheet goes to the director, human or agent, who picks or regenerates frames before any motion is spent.

Engine
Z-Image Turbo 6B via ComfyUI
Runs on
GPU 0
Typical time
≈ 26 s / image
Writes
keyframes/ · picks.json
04
Scores it

Music

An original instrumental score is generated to the narration length. It is made locally and registered in no Content ID pool.

Engine
ACE-Step 1.5 XL turbo
Runs on
GPU 1
Typical time
24–37 s / track
Writes
music.flac
05
Times cuts to speech

Timeline

The narration is the clock. Every cut lands 0.12 s before its first word, and captions are exact to the spoken word.

Engine
Word-level alignment (difflib + num2words)
Runs on
CPU
Typical time
< 5 s
Writes
timeline.json
06
Animates the frames

Motion

Image-to-video with the high-noise expert on one card and the low-noise expert on the other. Only a small latent file crosses between them, so two 16 GB cards do what would otherwise need one very large one.

Engine
Wan 2.2 I2V A14B + lightx2v 4-step
Runs on
GPU 0 → GPU 1
Typical time
≈ 3.5 s / frame
Writes
clips/*.mp4
07
Smooths and upscales

Finish

The 16 fps clips become smooth 32 fps footage at the full 1080×1920 resolution.

Engine
RIFE v4.6 frame interpolation + Lanczos upscale
Runs on
iGPU + CPU
Typical time
12–28 s / clip
Writes
clips_hd/*.mp4
08
Cuts, captions, motion graphics

Edit & graphics

The whole edit is generated as code: cut effects, film grain, animated captions where the spoken word turns amber, count-up numbers, charts, maps and exact shapes.

Engine
HyperFrames (HTML + GSAP), headless Chrome render
Runs on
CPU
Typical time
≈ 2.5 min render
Writes
picture.mp4
09
Mixes to platform loudness

Sound & master

Whooshes, impacts and risers are synthesised in code, so there are no sample licences. The music ducks under the voice, and the mix is mastered to −14 LUFS and −1 dBTP.

Engine
Synthesised SFX, sidechain ducking, FFmpeg loudnorm
Runs on
CPU
Typical time
< 1 min
Writes
final.mp4
10
Checks its own work

Quality gate

Seven deterministic checks: resolution, frame rate, duration, length, loudness, true peak and unplanned black frames. Then a local vision model scores one frame per shot and flags garbled text or wrong subjects.

Engine
FFprobe / EBU R128 checks + local vision-model judge
Runs on
CPU + GPU
Typical time
≈ 1–3 min
Writes
qc/report.json
11
Uploads it for you

Publish

Title, description, tags and AI disclosure are written from the fact sheet. The video uploads on a schedule, or waits for your OK. You decide.

Engine
Official YouTube, Facebook and Instagram APIs
Runs on
CPU
Typical time
seconds
Writes
live Short

The creative plan

One JSON file holds the whole film.

The director writes beats of narration, and each beat gets shots. Change a line, a prompt or an overlay and re-run just the stage it affects.

  • Clip, still or graphic. The fast profile animates 12–14 shots with AI and uses Ken Burns stills and full-screen motion graphics for the rest. It still cuts to a new picture every 2–3 seconds.
  • Overlays as data. Hook titles, year stamps, count-up numbers, bar charts, route maps and exact polygons are all driven by the plan.
  • Reproducible. Seeds come from shot IDs, so re-running a stage gives the same result until you change the prompt.
// saturn10/plan.json (excerpt, real episode)
{
  "title": "Saturn's Hexagon Has a New Twin",
  "voice_design": "Male documentary narrator in his forties.
     Deep, warm baritone… measured, suspenseful pacing",
  "music_prompt": "cinematic ambient space underscore…
     72 bpm, instrumental, no vocals",
  "beats": [{
    "vo": "Above its north pole, a six-sided jet stream never
           stops spinning. It's 30,000 kilometers across.",
    "shots": [{
      "id": "s03", "kind": "clip",
      "image_prompt": "a gas giant's polar vortex seen from
         above at an oblique angle, swirling golden clouds…",
      "motion_prompt": "the cloud bands rotate around the
         center, camera holds steady"
    }, {
      "id": "s04", "kind": "graphic", "bg_shot": "s03",
      "overlay": "big_number",
      "overlay_value": "30,000",
      "overlay_label": "KILOMETERS ACROSS",
      "sfx": "impact"
    }]
  }]
}

Quality control

It checks its own work before anyone sees it.

A contact sheet of every shot, plus a three-tier gate: machine checks, a local vision-model judge, and a final review by a person or an agent.

Contact sheet of all shots from the Shotcoda Short about the 52 Hertz whale
TIER 1

Deterministic checks

1080×1920, 30 fps, duration within 0.25 s of the timeline, under 3 minutes, −14 ± 1 LUFS, true peak ≤ −1 dBTP, no unplanned black frames.

TIER 2

Vision-model judge

A local multimodal model scores one frame per shot against its prompt. It hard-fails garbled pseudo-text, deformed anatomy, overlapping or cut-off text, and pictures that contradict the story.

TIER 3

Final watch

Facts on screen match the fact sheet, names are pronounced right, music never masks the voice, and the ending loops cleanly into the hook.

Details

Pipeline questions

Is Shotcoda fully automatic, or does a person have to steer it?

It can run end to end without a person. Two checkpoints are built in, the shot plan and the keyframe picks, and an AI agent or a human can handle both. You can keep a final approval before publishing, or let it post on a schedule.

Which LLM directs the pipeline?

Any LLM that can write JSON. Shotcoda ships with a local Qwen3.8-27B served by llama.cpp, which also runs the vision quality check. For the strongest scripts you can plug in the Claude or ChatGPT account you already have. No other subscription is involved.

Why generate keyframes first instead of text-to-video?

Keyframes keep every shot in one visual world, they are cheap to review, and they anchor the video model. A bad frame is caught in 26 seconds instead of wasting minutes of motion generation.

How are exact numbers, shapes and maps drawn?

Not by the image model. Image models cannot count sides or spell reliably, so every number, chart, map and exact shape is drawn as a motion graphic in code (HTML + GSAP) on top of the generated picture.

What does the edit look like under the hood?

One generated HTML file per episode, rendered frame by frame in headless Chrome by HyperFrames. Graphics cost only code, rendering is deterministic and a linter catches layout mistakes before the render.

Build my video pipeline

Want this pipeline on your machine?

Tell me what you make and where you publish. I'll check your hardware, build the pipeline around your format and brand, and set it up on your machine. No pricing tables: every setup is different, so you get a straight answer.

What do you need?

Your details are used only to reply. Privacy