PIPE_VIDEO_2026-07-29.md

pipelines/PIPE_VIDEO_2026-07-29.md

PIPE_VIDEO — the cinematic/video-generation lane dossier

Research date 2026-07-29. Every model claim below is either VERIFIED (primary source fetched this

session, URL inline) or UNVERIFIED-EXCLUDED (named, evidence given, not adopted). No name is

carried on an aggregator's word alone.

0. THE HONEST FRAME — this lane ships no frames

The game's cinematics are UE-rendered, and that is already ruled, not my call to reopen:

docs/translation/T99_Translation_Cinematic.md §2.3/§2.4 builds every scene as a Sequencer

Level Sequence assembled by NeoStack + Python/Remote Control off the seven-column scene row, delivered

via MRQ; a grep of that doc for AI-video generation returns nothing but a *deferral* row (line 479,

animation-content generation out of scope). docs/MARKETING_WEB_PROGRAM.md lane 1 puts marketing motion

on Movie Render Graph EXR masters + OBS/NVENC AV1 playthroughs — real captures, gated on P2/P3

surfaces existing. So this lane is support: internal mood reels, shot-idea previs reference, and

motion-tier concept exploration. Its outputs are **previz-tagged, never shipped frames, never public

frames**, unless Josh rules otherwise.

The licence tiering falls straight out of the ruled posture (5090_SETUP_RUNBOOK.md:1150-1153,

quoting REALM_ANALYSIS_ART_PIPELINE_2026-07-27.md §2.3 as binding): unrestricted-by-default in any

load-bearing path; anything MAU-capped, territory-excluded, revenue-ceilinged or non-commercial is

previz-only, never load-bearing. Applied here, exactly two models clear the bar for anything that

could ever touch a public pixel — Wan 2.2 and SeedVR2, both Apache-2.0. Everything else in this

dossier, including the best-quality option, is internal-only by that existing rule.

---

1. THE VERIFIED STACK

JobPRIMARYFALLBACKWhy
Image→video mood reel from a Lane-P/R concept plateLTX-2.3 (22B)Wan 2.2LTX is the quality/AV leader; Wan is the only licence-clean one
Anything that may ever face the publicWan 2.2 — no other optionnoneApache-2.0 is the only clean tier
Fast iteration / co-resident windowHunyuanVideo-1.5 (8.3B)Wan 2.2 fp8the only model with a verified 14GB floor
Upscale 720p → deliverySeedVR2 (3B/7B)re-generate at resApache-2.0, native in ComfyUI since v0.28.0
Motion/animation referenceNOT THIS LANE → ARDY/Kimodosee §1.3

1.1 Adopted — verified

Open-sourced 2026-01-06 (GlobeNewswire release).

Current checkpoint ltx-2.3-22b-dev + ltx-2.3-22b-distilled-1.1, fp8/fp4 casts, spatial/temporal

x2 upscalers (HF Lightricks/LTX-2.3). ComfyUI: built-in

LTXVideo nodes via Manager (model card, verbatim); ComfyUI v0.8.0 changelog: *"Added LTXV 2 model

support"* (docs.comfy.org/changelog).

LICENCE — read it before the first GPU hour. "LTX-2 Community License Agreement"

(raw LICENSE). Three clauses that bind us:

(a) *"Entities with annual revenues of at least $10,000,000 … are required to obtain a paid commercial

use license"* — a revenue ceiling, which the §2.3 posture classifies previz-only;

(b) *"Licensor claims no rights in the Output you generate"* — output ownership is clean;

(c) Attachment A §5 forbids disseminating output *"without expressly and intelligibly disclaiming

that the information and/or content is machine generated."* **(c) is the operative reason no LTX frame

goes on the marketing site without an AI-generated disclosure** — a Josh business call, not a tech one.

plain Apache 2.0, no MAU cap, no territory exclusion, no output restriction**

(raw LICENSE.txt). Repo last

updated 2026-03-17, 16.9k stars (github.com/Wan-Video). ComfyUI native.

This is the only model in the lane that may touch a load-bearing or public surface.

offloading enabled)"*; 480p/720p, 121 frames (~5s @24fps), 1080p via super-resolution; released

2025-11-20 (HF tencent/HunyuanVideo-1.5). ComfyUI

supported with an official guide.

LICENCE — the aggregators are wrong. Multiple 2026 blogs call it "Apache 2.0." Direct fetch of the

LICENSE says Tencent Hunyuan Community License Agreement, with a 100M MAU request-a-licence

clause and, verbatim: *"THIS LICENSE AGREEMENT DOES NOT APPLY IN THE EUROPEAN UNION, UNITED KINGDOM AND

SOUTH KOREA"* (raw LICENSE).

Territory-excluded → internal previz only, permanently, by the ruled posture. Same trap class as

Hunyuan3D-2.1 already carries in LOCAL_3D_ASSET_GEN.md.

Apache-2.0 (ByteDance-Seed/SeedVR,

HF ByteDance-Seed/SeedVR2-3B). **Native in ComfyUI

v0.28.0 (2026-07-15): "Native SeedVR2 image and video upscaling (#14424)"** — verified in the changelog,

so the generate-at-720p / upscale-for-delivery path needs no third-party node pack.

1.2 Cosmos — verified real, EXCLUDED from the primary stack on three independent grounds

docs/pipeline_review/tech_research/COSMOS3_DEEP_DIVE.md is still accurate and I re-verified its load-bearing

facts rather than trusting it. Cosmos 3 Nano (16B): licence OpenMDW-1.1, *"ready for commercial and

non-commercial use"* — the cleanest licence of anything in this dossier, no MAU, no territory, no output

claim (HF nvidia/Cosmos3-Nano, re-fetched today). Output caps

verified: 256p/480p/720p, 5-400 frames, default 189. Excluded anyway because:

1. VRAM — the transformer alone is >29GB in BF16 on a 32GB 5090, needing enable_sequential_cpu_offload

just to boot (COSMOS3_DEEP_DIVE §1, hands-on report). Against the runbook's binding 29-of-32GB arithmetic

(5090_SETUP_RUNBOOK.md:854), that is an exclusive-machine model for a support lane.

2. No ComfyUI support. Comfy-Org/ComfyUI issue #14228,

"Support for NVIDIA Cosmos 3 model family?" — **opened 2026-06-02, still OPEN, zero maintainer comments,

no branches or PRs**, verified today. ComfyUI's actual Cosmos support stops at the original Cosmos and

Cosmos-Predict2 (2B/14B) — and Predict2.5 is in limited-maintenance mode per NVIDIA's own repo.

3. Wrong training objective — physical-AI curated, model card's own intended use is robotics/AV; it is

not benchmarked as a creative/aesthetic video model and its own tech report flags artifacting.

Verdict: WATCH, re-check at benchmark day. If ComfyUI lands Cosmos 3 and a quantized NIM checkpoint runs

on consumer Blackwell, its licence is the best in the field and it would jump straight to a public-surface-

eligible tier alongside Wan.

1.3 The leads — confirmed or killed, each with a URL

LeadVerdictEvidence
NVIDIA Cosmos 3CONFIRMED REAL, EXCLUDED§1.2 above; 3 independent grounds
Wan 2.5 / Wan 2.7 open weightsUNVERIFIED-EXCLUDED — killedgithub.com/Wan-Video lists only Wan2.2, Wan2.1, Wan-Dancer, Wan-skills; huggingface.co/Wan-AI newest is Wan-Dancer-14B then Wan2.2-Animate. No Wan2.5 or Wan2.7 weights exist. Artificial Analysis states Wan 2.5 "breaks from tradition by not being open weights." The wan27.org / oakgen.ai / spheron "Wan 2.7 open source" pages are SEO content contradicted by both primary orgs.
"ARDY"CONFIRMED REAL — WRONG LANEnv-tlabs/ardy: autoregressive-diffusion human MOTION generation (SIGGRAPH 2026), released 2026-07-10, code Apache-2.0, weights NVIDIA Open Model License, tested RTX 4090, ~14GB for the text encoder. Not video. Route to the animation/motion lane — and it is a genuinely strong lead there.
"Kimodo"CONFIRMED REAL — WRONG LANEnv-tlabs/kimodo, kinematic motion diffusion, open-sourced 2026-03-16, five variants (SOMA/G1/SMPL-X). ARDY is its real-time successor. Same routing.
ACE-Step 1.5 / "XL SFT" / "XL Turbo" / "Excel Bass"NOT THIS LANE — not verified hereMusic/SFX generation. Deliberately not fake-verified in a video dossier; belongs to the audio pipeline's own licence-read. Flagged so it is not silently dropped.
Wan-Dancer-14BCONFIRMED REAL — watch itemHF Wan-AI/Wan-Dancer-14B: Apache 2.0, released 2026-07-13, music-synced long-form dance video. ComfyUI support is on the TODO, not implemented. Niche (festival/dance scenes); revisit when ComfyUI lands.
Lyra 2.0CONFIRMED, BLOCKEDWeights are NVIDIA Internal Scientific Research license — *"may not use … in a production environment or for the purpose of generating works for sale."* Already recorded in COSMOS3_DEEP_DIVE §3. Excluded.
"Gemma 4 generates video"UNVERIFIED — not adoptedAggregator claim only; no primary source. Not carried.

1.4 VRAM budget + scheduling class (the concurrency law)

ModelConfigVRAMClass
HunyuanVideo-1.58.3B, offload on14GB (VERIFIED, model card)windowed — the only member that can plausibly share the card
Wan 2.2A14B, fp8~16-22GB (ESTIMATE — measure)windowed/overnight
SeedVR2-3Bupscale, tiled~8-14GB (ESTIMATE — tiling-dependent)overnight
LTX-2.3 distilled22B, fp8~32GB claimed → assume EXCLUSIVEovernight, exclusive
Cosmos 3 Nano16B, BF16>29GB + offloadnot installed

Scheduling class for the whole lane: OVERNIGHT / EXCLUSIVE by default, never always-on. This is the

lowest-priority GPU consumer in the house — it yields to the UE editor, to the 3D-gen benchmark, and to the

14B runtime reference. The runbook's binding arithmetic (:854-858) says heavy lanes run **sequentially,

not concurrently**, and that an OOM from our own co-residency sloppiness would be a *false* hardware-upgrade

trigger. This lane inherits that verbatim: it takes the card only when nothing else wants it, releases it on

demand, and never launches from the same worklist pass as a 3D-gen or Hunyuan3D run.

---

2. INSTALL PLAN

Environment lane (D-1): option (a) — WSL2 Ubuntu 24.04 + the Puget Docker App Pack comfy_ui flavor.

Already evidence-ridered in the runbook (:1169-1187): the pack's ComfyUI image is

nvidia/cuda:12.8.0-runtime-ubuntu24.04 + torch from the cu128 index — the known-good sm_120 stack — and its

install-time model selector already offers LTX-Video, so this lane's host is vendor-pre-validated rather

than hand-built. Carry the recorded caveats: fork the pack and pin torch==2.9.0+cu128 (upstream is

unpinned), hand-finish the driver step under WSL2 (host-driver passthrough), skip team_llm. Keep

ComfyUI-on-Windows (option b) as the fast-iteration lane for prompt work that needs no exotic wheels —

this lane is unusual in that ALL FOUR primaries have first-class ComfyUI paths, so option (b) is genuinely

viable here even though the 3D lane needs (a).

Disk placement (RULED layout). Everything in this lane lands on the Gen4 4TB (models / ComfyUI /

projects cache): weights under models/{diffusion_models,vae,text_encoders,upscale_models}, the WSL2 vhdx,

ComfyUI output/ and input/. Nothing from this lane touches the Gen5 4TB — that drive is UE engine

hot path + DDC, and a 40GB checkpoint pull competing with shader-compile I/O is exactly the interference the

layout ruling exists to prevent. Long-term reel archive → 8TB SATA, monthly sweep.

Install ORDER (each step verifies before the next):

1. WSL2 Ubuntu 24.04 distro exists + nvidia-smi passes inside it (Stage 2 prerequisite).

2. Forked Puget pack setup.shcomfy_ui flavor, torch pinned, model volumes host-mapped to the Gen4 path.

3. Launch ComfyUI, confirm port 8188 is in the Stage-6F firewall pre-authorization (already listed).

4. Wan 2.2 first — it is the licence-clean baseline; prove one 720p I2V generation, previz-tagged.

5. SeedVR2 (native nodes, v0.28.0+) — prove one 720p→1080p upscale on that clip.

6. HunyuanVideo-1.5 — prove the 14GB-offload claim on our card; this is the co-residency probe.

7. LTX-2.3 distilled fp8 last — it is the one that may not fit; failing here costs nothing because 4-6

already gave a working lane.

Thursday-night download list. Sizes below are ESTIMATES from parameter counts (bf16 ≈ 2 bytes/param,

fp8 ≈ 1) except where a card states otherwise — verify at pull time, do not budget the night on them:

PullEst. sizePriority
Wan 2.2 (T2V + I2V checkpoints, fp8 where offered)~30-60 GB1 — must
SeedVR2-3B (+7B optional)~7 GB (+15)2 — must
HunyuanVideo-1.5 (8.3B + super-res)~18-22 GB3
LTX-2.3-22b-distilled fp8 (+ x2 upscalers)~22-26 GB4
Shared text encoders / VAEs (ComfyUI repackaged)~10-20 GBrides 1
Total~90-150 GBfits the Gen4 with room

Anonymous HF pulls; none of the four primaries is a gated repo (verify at pull — if one has gated, stop

and surface it rather than creating an account, per §5).

---

3. THE INTEGRATION CONTRACT

Video generation is NOT a DR-2 asset class, and I am not inventing a twelfth one.

docs/ASSET_DROPIN_CONTRACT.md §2 enumerates exactly eleven classes; none is video. The only adjacent row is

Cinematic scene: T0_Scene_Spec_Registry.scene_id/scene_anchor_idscene_spec_ref, final asset a

UE Level Sequence at /Game/Cinematics/Scenes/<scene_id>, swap by manifest re-point. This lane feeds the

prompt/reference side of that row. It never writes ue_asset_path, never advances generation_status.

The reference-tier addressing rule — reuses DR-2's discipline without touching its schema:

id they reference** (§1's invariant, applied to reference material) — scene_id, realm_id, boss_id.

(apache|ceilinged|territory_excluded), prompt, generation_prompt_hash, generation_tier, decay_stage,

previz: true}`. The generating model's licence is named in every artifact — that is the licence law,

mechanized: a file whose sidecar says territory_excluded can be filtered out of any public-surface

candidate set by a script, not by someone remembering.

("live is never a stored claim") extends here: previz existing is not an asset existing.

generation_tier='realm' and, for the fairy realm, decay_stage ∈ {height, present} — so the mood-reel

pair for the fairy realm is generated as a pair by construction, matching D1/D3.

paintover plate (never a lossy intermediate into a paintover); EXR only if a plate needs float. Never

cooked, never in /Game/, never in an always-cook directory.

The QA gate that judges outputs (three teeth, all existing rubric families — zero new gates):

1. Reveal discipline — the same ladder MARKETING_WEB_PROGRAM.md §"Binding disciplines" binds on

marketing surfaces: no House-of-Velheim pattern before its in-game reveal point, no late-arc reveals.

Runs BEFORE any frame leaves the repo, not before it is shown publicly.

2. §17.1 living-heritage — the ruled Lane-R/Lane-P rule extends from stills to motion: **real living-

tradition architecture and iconography are never raw-generated.** For those subjects the lane runs I2V

off a Lane-R/Lane-P plate (inheriting the reference's proportions — the "superset by process" property),

never T2V off a prompt.

3. Image/feel critic read — the existing critic loop reads the reel; for realm subjects the

RB-RLM-PROVENANCE sibling rubric's before-state line applies (shown a fairy-realm frame, the critic must

say what the place looked like BEFORE, per DR-2 §6-D).

---

4. THE AUTONOMY CONTRACT

Dispatch — from registry rows, script-emitted, never hand-typed (the standing worklist rule). Three

generators:

(canonical_set_piece, architect_appearance, boss_encounter_*, convergence_node,

prologue/epilogue/chapter_bridging_cinematic) → previs reference boards, feeding the human-director

Sequencer pass described in Cinematic §2.4. Scene-driven-track rows get nothing from this lane — they

assemble without a director pass by design, and previz there would be waste.

decay-paired where decay_bearing is set.

pre-capture shot-idea boards to the person framing the shot, and stops there.

Promote / demote gate. A model is PROMOTED to primary for a job only after (a) benchmark-day A/B on the

fixed shot set against the same critic rubric, and (b) a licence re-read at that sitting — the GPU-free

pre-flight REALM_ANALYSIS_ART_PIPELINE §8.5 already schedules. DEMOTE on any of: a licence change, an

upstream repo going quiet or closing weights (the Wan 2.5/2.7 pattern is live proof this happens), a critic

pass-rate drop below its band, or a measured VRAM regression that breaks serialization. Demotion is

mechanical and does not need a sitting.

Unattended vs attended.

run; previz-tagged writes only; zero registry writes; zero /Game/ writes; the run holds the GPU

exclusively and exits by a hard wall-clock so the nightly soak and the morning UE window are never

contested.

Josh; any subject drawn from a living tradition (tooth 2 above); any run that would coexist with a UE

session. Two named agent-failure classes from memory apply directly — agents backgrounding the slow final

verify, and agents describing a frame the artifact contradicts. Mitigation: the lane's completion is

the sidecar files on disk plus a director re-read of the newest output, never an agent's summary.

---

5. JOSH-MINIMUM

Genuinely short — this is one of the cleanest lanes for hands-off setup.

1. Nothing. All four primaries are anonymous public HF/GitHub pulls. No account, no card, no token, no

EULA click. If a pull turns out to be gated at download time, stop and surface it rather than creating

an account (account creation is not mine to do).

2. One real decision, and it is a business call, not a technical one: whether any public marketing

surface may ever carry a generated frame. If yes, LTX-2's Attachment A §5 compels an intelligible

machine-generated disclaimer on that surface, and the §2.3 revenue-ceiling posture makes it previz-tier

anyway. RECOMMENDATION: no — keep every public frame on real captures, which is already the ruled

Capture lane and costs this lane nothing. Surfaced as a brief, not a bare question.

3. The Puget burn-in/benchmark sheet ask (already on record in the runbook as a Josh QC action) — this

lane's overnight exclusive runs are the machine's longest sustained GPU load, so that thermal/power

baseline is the one this lane most depends on.

---

6. BENCHMARK-DAY QUESTIONS

1. Measured peak VRAM and wall-clock per model at a fixed 720p/121-frame I2V job. Every ESTIMATE in §1.4

is replaced with an observation, and the runbook's own "measure them on day one" correction applies here too.

2. Does ltx-2.3-22b-distilled fp8 actually fit in 32GB headless? If no, LTX leaves the lane entirely and

Wan 2.2 becomes both primary and fallback — a materially simpler, fully-Apache lane.

3. Co-residency probe: can HunyuanVideo-1.5 at its 14GB offload floor complete a job while the UE editor

holds the project open, without OOM? This is the only question that could earn this lane a *windowed*

class instead of *overnight-exclusive*.

4. Licence re-read of all four, folded into the existing GPU-free pre-flight — plus a re-check of the

Cosmos 3 ComfyUI issue (#14228) and of whether Wan has released anything past 2.2.

5. Critic pass-rate per subject class (realm / creature / architecture / crowd) on a fixed 6-shot set —

which model wins where, since "one primary for everything" is unlikely to survive contact.

6. I2V-from-a-Lane-P-plate vs T2V-from-prompt: does the paintover's proportion-inheritance property

survive into motion? If it does, tooth 2 stops being a restriction and becomes the default method.

7. SeedVR2 720p→1080p/4K vs re-generating at higher resolution — does the upscale hold identity, and is it

cheaper than the native-res run?

8. ComfyUI native nodes vs custom packs for each model on the day — which is on the maintained path, given

this lane's whole install argument rests on native support.

Generated by harness/site/structure_site.py — the URL path is the repo path. review root