pipelines/PIPE_VIDEO_2026-07-29.md
Research date 2026-07-29. Every model claim below is either VERIFIED (primary source fetched this
session, URL inline) or UNVERIFIED-EXCLUDED (named, evidence given, not adopted). No name is
carried on an aggregator's word alone.
The game's cinematics are UE-rendered, and that is already ruled, not my call to reopen:
docs/translation/T99_Translation_Cinematic.md §2.3/§2.4 builds every scene as a Sequencer
Level Sequence assembled by NeoStack + Python/Remote Control off the seven-column scene row, delivered
via MRQ; a grep of that doc for AI-video generation returns nothing but a *deferral* row (line 479,
animation-content generation out of scope). docs/MARKETING_WEB_PROGRAM.md lane 1 puts marketing motion
on Movie Render Graph EXR masters + OBS/NVENC AV1 playthroughs — real captures, gated on P2/P3
surfaces existing. So this lane is support: internal mood reels, shot-idea previs reference, and
motion-tier concept exploration. Its outputs are **previz-tagged, never shipped frames, never public
frames**, unless Josh rules otherwise.
The licence tiering falls straight out of the ruled posture (5090_SETUP_RUNBOOK.md:1150-1153,
quoting REALM_ANALYSIS_ART_PIPELINE_2026-07-27.md §2.3 as binding): unrestricted-by-default in any
load-bearing path; anything MAU-capped, territory-excluded, revenue-ceilinged or non-commercial is
previz-only, never load-bearing. Applied here, exactly two models clear the bar for anything that
could ever touch a public pixel — Wan 2.2 and SeedVR2, both Apache-2.0. Everything else in this
dossier, including the best-quality option, is internal-only by that existing rule.
---
| Job | PRIMARY | FALLBACK | Why |
|---|---|---|---|
| Image→video mood reel from a Lane-P/R concept plate | LTX-2.3 (22B) | Wan 2.2 | LTX is the quality/AV leader; Wan is the only licence-clean one |
| Anything that may ever face the public | Wan 2.2 — no other option | none | Apache-2.0 is the only clean tier |
| Fast iteration / co-resident window | HunyuanVideo-1.5 (8.3B) | Wan 2.2 fp8 | the only model with a verified 14GB floor |
| Upscale 720p → delivery | SeedVR2 (3B/7B) | re-generate at res | Apache-2.0, native in ComfyUI since v0.28.0 |
| Motion/animation reference | NOT THIS LANE → ARDY/Kimodo | — | see §1.3 |
Open-sourced 2026-01-06 (GlobeNewswire release).
Current checkpoint ltx-2.3-22b-dev + ltx-2.3-22b-distilled-1.1, fp8/fp4 casts, spatial/temporal
x2 upscalers (HF Lightricks/LTX-2.3). ComfyUI: built-in
LTXVideo nodes via Manager (model card, verbatim); ComfyUI v0.8.0 changelog: *"Added LTXV 2 model
support"* (docs.comfy.org/changelog).
LICENCE — read it before the first GPU hour. "LTX-2 Community License Agreement"
(raw LICENSE). Three clauses that bind us:
(a) *"Entities with annual revenues of at least $10,000,000 … are required to obtain a paid commercial
use license"* — a revenue ceiling, which the §2.3 posture classifies previz-only;
(b) *"Licensor claims no rights in the Output you generate"* — output ownership is clean;
(c) Attachment A §5 forbids disseminating output *"without expressly and intelligibly disclaiming
that the information and/or content is machine generated."* **(c) is the operative reason no LTX frame
goes on the marketing site without an AI-generated disclosure** — a Josh business call, not a tech one.
Wan-Video/Wan2.2) — **Apache-2.0, verified by direct fetch of the LICENSE:plain Apache 2.0, no MAU cap, no territory exclusion, no output restriction**
(raw LICENSE.txt). Repo last
updated 2026-03-17, 16.9k stars (github.com/Wan-Video). ComfyUI native.
This is the only model in the lane that may touch a load-bearing or public surface.
offloading enabled)"*; 480p/720p, 121 frames (~5s @24fps), 1080p via super-resolution; released
2025-11-20 (HF tencent/HunyuanVideo-1.5). ComfyUI
supported with an official guide.
LICENCE — the aggregators are wrong. Multiple 2026 blogs call it "Apache 2.0." Direct fetch of the
LICENSE says Tencent Hunyuan Community License Agreement, with a 100M MAU request-a-licence
clause and, verbatim: *"THIS LICENSE AGREEMENT DOES NOT APPLY IN THE EUROPEAN UNION, UNITED KINGDOM AND
SOUTH KOREA"* (raw LICENSE).
Territory-excluded → internal previz only, permanently, by the ruled posture. Same trap class as
Hunyuan3D-2.1 already carries in LOCAL_3D_ASSET_GEN.md.
Apache-2.0 (ByteDance-Seed/SeedVR,
HF ByteDance-Seed/SeedVR2-3B). **Native in ComfyUI
v0.28.0 (2026-07-15): "Native SeedVR2 image and video upscaling (#14424)"** — verified in the changelog,
so the generate-at-720p / upscale-for-delivery path needs no third-party node pack.
docs/pipeline_review/tech_research/COSMOS3_DEEP_DIVE.md is still accurate and I re-verified its load-bearing
facts rather than trusting it. Cosmos 3 Nano (16B): licence OpenMDW-1.1, *"ready for commercial and
non-commercial use"* — the cleanest licence of anything in this dossier, no MAU, no territory, no output
claim (HF nvidia/Cosmos3-Nano, re-fetched today). Output caps
verified: 256p/480p/720p, 5-400 frames, default 189. Excluded anyway because:
1. VRAM — the transformer alone is >29GB in BF16 on a 32GB 5090, needing enable_sequential_cpu_offload
just to boot (COSMOS3_DEEP_DIVE §1, hands-on report). Against the runbook's binding 29-of-32GB arithmetic
(5090_SETUP_RUNBOOK.md:854), that is an exclusive-machine model for a support lane.
2. No ComfyUI support. Comfy-Org/ComfyUI issue #14228,
"Support for NVIDIA Cosmos 3 model family?" — **opened 2026-06-02, still OPEN, zero maintainer comments,
no branches or PRs**, verified today. ComfyUI's actual Cosmos support stops at the original Cosmos and
Cosmos-Predict2 (2B/14B) — and Predict2.5 is in limited-maintenance mode per NVIDIA's own repo.
3. Wrong training objective — physical-AI curated, model card's own intended use is robotics/AV; it is
not benchmarked as a creative/aesthetic video model and its own tech report flags artifacting.
Verdict: WATCH, re-check at benchmark day. If ComfyUI lands Cosmos 3 and a quantized NIM checkpoint runs
on consumer Blackwell, its licence is the best in the field and it would jump straight to a public-surface-
eligible tier alongside Wan.
| Lead | Verdict | Evidence |
|---|---|---|
| NVIDIA Cosmos 3 | CONFIRMED REAL, EXCLUDED | §1.2 above; 3 independent grounds |
| Wan 2.5 / Wan 2.7 open weights | UNVERIFIED-EXCLUDED — killed | github.com/Wan-Video lists only Wan2.2, Wan2.1, Wan-Dancer, Wan-skills; huggingface.co/Wan-AI newest is Wan-Dancer-14B then Wan2.2-Animate. No Wan2.5 or Wan2.7 weights exist. Artificial Analysis states Wan 2.5 "breaks from tradition by not being open weights." The wan27.org / oakgen.ai / spheron "Wan 2.7 open source" pages are SEO content contradicted by both primary orgs. |
| "ARDY" | CONFIRMED REAL — WRONG LANE | nv-tlabs/ardy: autoregressive-diffusion human MOTION generation (SIGGRAPH 2026), released 2026-07-10, code Apache-2.0, weights NVIDIA Open Model License, tested RTX 4090, ~14GB for the text encoder. Not video. Route to the animation/motion lane — and it is a genuinely strong lead there. |
| "Kimodo" | CONFIRMED REAL — WRONG LANE | nv-tlabs/kimodo, kinematic motion diffusion, open-sourced 2026-03-16, five variants (SOMA/G1/SMPL-X). ARDY is its real-time successor. Same routing. |
| ACE-Step 1.5 / "XL SFT" / "XL Turbo" / "Excel Bass" | NOT THIS LANE — not verified here | Music/SFX generation. Deliberately not fake-verified in a video dossier; belongs to the audio pipeline's own licence-read. Flagged so it is not silently dropped. |
| Wan-Dancer-14B | CONFIRMED REAL — watch item | HF Wan-AI/Wan-Dancer-14B: Apache 2.0, released 2026-07-13, music-synced long-form dance video. ComfyUI support is on the TODO, not implemented. Niche (festival/dance scenes); revisit when ComfyUI lands. |
| Lyra 2.0 | CONFIRMED, BLOCKED | Weights are NVIDIA Internal Scientific Research license — *"may not use … in a production environment or for the purpose of generating works for sale."* Already recorded in COSMOS3_DEEP_DIVE §3. Excluded. |
| "Gemma 4 generates video" | UNVERIFIED — not adopted | Aggregator claim only; no primary source. Not carried. |
| Model | Config | VRAM | Class |
|---|---|---|---|
| HunyuanVideo-1.5 | 8.3B, offload on | 14GB (VERIFIED, model card) | windowed — the only member that can plausibly share the card |
| Wan 2.2 | A14B, fp8 | ~16-22GB (ESTIMATE — measure) | windowed/overnight |
| SeedVR2-3B | upscale, tiled | ~8-14GB (ESTIMATE — tiling-dependent) | overnight |
| LTX-2.3 distilled | 22B, fp8 | ~32GB claimed → assume EXCLUSIVE | overnight, exclusive |
| Cosmos 3 Nano | 16B, BF16 | >29GB + offload | not installed |
Scheduling class for the whole lane: OVERNIGHT / EXCLUSIVE by default, never always-on. This is the
lowest-priority GPU consumer in the house — it yields to the UE editor, to the 3D-gen benchmark, and to the
14B runtime reference. The runbook's binding arithmetic (:854-858) says heavy lanes run **sequentially,
not concurrently**, and that an OOM from our own co-residency sloppiness would be a *false* hardware-upgrade
trigger. This lane inherits that verbatim: it takes the card only when nothing else wants it, releases it on
demand, and never launches from the same worklist pass as a 3D-gen or Hunyuan3D run.
---
Environment lane (D-1): option (a) — WSL2 Ubuntu 24.04 + the Puget Docker App Pack comfy_ui flavor.
Already evidence-ridered in the runbook (:1169-1187): the pack's ComfyUI image is
nvidia/cuda:12.8.0-runtime-ubuntu24.04 + torch from the cu128 index — the known-good sm_120 stack — and its
install-time model selector already offers LTX-Video, so this lane's host is vendor-pre-validated rather
than hand-built. Carry the recorded caveats: fork the pack and pin torch==2.9.0+cu128 (upstream is
unpinned), hand-finish the driver step under WSL2 (host-driver passthrough), skip team_llm. Keep
ComfyUI-on-Windows (option b) as the fast-iteration lane for prompt work that needs no exotic wheels —
this lane is unusual in that ALL FOUR primaries have first-class ComfyUI paths, so option (b) is genuinely
viable here even though the 3D lane needs (a).
Disk placement (RULED layout). Everything in this lane lands on the Gen4 4TB (models / ComfyUI /
projects cache): weights under models/{diffusion_models,vae,text_encoders,upscale_models}, the WSL2 vhdx,
ComfyUI output/ and input/. Nothing from this lane touches the Gen5 4TB — that drive is UE engine
hot path + DDC, and a 40GB checkpoint pull competing with shader-compile I/O is exactly the interference the
layout ruling exists to prevent. Long-term reel archive → 8TB SATA, monthly sweep.
Install ORDER (each step verifies before the next):
1. WSL2 Ubuntu 24.04 distro exists + nvidia-smi passes inside it (Stage 2 prerequisite).
2. Forked Puget pack setup.sh → comfy_ui flavor, torch pinned, model volumes host-mapped to the Gen4 path.
3. Launch ComfyUI, confirm port 8188 is in the Stage-6F firewall pre-authorization (already listed).
4. Wan 2.2 first — it is the licence-clean baseline; prove one 720p I2V generation, previz-tagged.
5. SeedVR2 (native nodes, v0.28.0+) — prove one 720p→1080p upscale on that clip.
6. HunyuanVideo-1.5 — prove the 14GB-offload claim on our card; this is the co-residency probe.
7. LTX-2.3 distilled fp8 last — it is the one that may not fit; failing here costs nothing because 4-6
already gave a working lane.
Thursday-night download list. Sizes below are ESTIMATES from parameter counts (bf16 ≈ 2 bytes/param,
fp8 ≈ 1) except where a card states otherwise — verify at pull time, do not budget the night on them:
| Pull | Est. size | Priority |
|---|---|---|
| Wan 2.2 (T2V + I2V checkpoints, fp8 where offered) | ~30-60 GB | 1 — must |
| SeedVR2-3B (+7B optional) | ~7 GB (+15) | 2 — must |
| HunyuanVideo-1.5 (8.3B + super-res) | ~18-22 GB | 3 |
| LTX-2.3-22b-distilled fp8 (+ x2 upscalers) | ~22-26 GB | 4 |
| Shared text encoders / VAEs (ComfyUI repackaged) | ~10-20 GB | rides 1 |
| Total | ~90-150 GB | fits the Gen4 with room |
Anonymous HF pulls; none of the four primaries is a gated repo (verify at pull — if one has gated, stop
and surface it rather than creating an account, per §5).
---
Video generation is NOT a DR-2 asset class, and I am not inventing a twelfth one.
docs/ASSET_DROPIN_CONTRACT.md §2 enumerates exactly eleven classes; none is video. The only adjacent row is
Cinematic scene: T0_Scene_Spec_Registry.scene_id/scene_anchor_id → scene_spec_ref, final asset a
UE Level Sequence at /Game/Cinematics/Scenes/<scene_id>, swap by manifest re-point. This lane feeds the
prompt/reference side of that row. It never writes ue_asset_path, never advances generation_status.
The reference-tier addressing rule — reuses DR-2's discipline without touching its schema:
previz/<canon_id>/<generation_prompt_hash>/ on the Gen4 drive, **addressed by the canon id they reference** (§1's invariant, applied to reference material) — scene_id, realm_id, boss_id.
(apache|ceilinged|territory_excluded), prompt, generation_prompt_hash, generation_tier, decay_stage,
previz: true}`. The generating model's licence is named in every artifact — that is the licence law,
mechanized: a file whose sidecar says territory_excluded can be filtered out of any public-surface
candidate set by a script, not by someone remembering.
complete. The referenced row stays pending. DR-2 §3's honesty("live is never a stored claim") extends here: previz existing is not an asset existing.
generation_tier='realm' and, for the fairy realm, decay_stage ∈ {height, present} — so the mood-reel
pair for the fairy realm is generated as a pair by construction, matching D1/D3.
paintover plate (never a lossy intermediate into a paintover); EXR only if a plate needs float. Never
cooked, never in /Game/, never in an always-cook directory.
The QA gate that judges outputs (three teeth, all existing rubric families — zero new gates):
1. Reveal discipline — the same ladder MARKETING_WEB_PROGRAM.md §"Binding disciplines" binds on
marketing surfaces: no House-of-Velheim pattern before its in-game reveal point, no late-arc reveals.
Runs BEFORE any frame leaves the repo, not before it is shown publicly.
2. §17.1 living-heritage — the ruled Lane-R/Lane-P rule extends from stills to motion: **real living-
tradition architecture and iconography are never raw-generated.** For those subjects the lane runs I2V
off a Lane-R/Lane-P plate (inheriting the reference's proportions — the "superset by process" property),
never T2V off a prompt.
3. Image/feel critic read — the existing critic loop reads the reel; for realm subjects the
RB-RLM-PROVENANCE sibling rubric's before-state line applies (shown a fairy-realm frame, the critic must
say what the place looked like BEFORE, per DR-2 §6-D).
---
Dispatch — from registry rows, script-emitted, never hand-typed (the standing worklist rule). Three
generators:
T0_Scene_Spec_Registry rows whose scene_type routes to the director-in-loop track (canonical_set_piece, architect_appearance, boss_encounter_*, convergence_node,
prologue/epilogue/chapter_bridging_cinematic) → previs reference boards, feeding the human-director
Sequencer pass described in Cinematic §2.4. Scene-driven-track rows get nothing from this lane — they
assemble without a director pass by design, and previz there would be waste.
generation_tier='realm' (the ten ruled realm rows) → realm mood reels, decay-paired where decay_bearing is set.
pre-capture shot-idea boards to the person framing the shot, and stops there.
Promote / demote gate. A model is PROMOTED to primary for a job only after (a) benchmark-day A/B on the
fixed shot set against the same critic rubric, and (b) a licence re-read at that sitting — the GPU-free
pre-flight REALM_ANALYSIS_ART_PIPELINE §8.5 already schedules. DEMOTE on any of: a licence change, an
upstream repo going quiet or closing weights (the Wan 2.5/2.7 pattern is live proof this happens), a critic
pass-rate drop below its band, or a measured VRAM regression that breaks serialization. Demotion is
mechanical and does not need a sitting.
Unattended vs attended.
run; previz-tagged writes only; zero registry writes; zero /Game/ writes; the run holds the GPU
exclusively and exits by a hard wall-clock so the nightly soak and the morning UE window are never
contested.
Josh; any subject drawn from a living tradition (tooth 2 above); any run that would coexist with a UE
session. Two named agent-failure classes from memory apply directly — agents backgrounding the slow final
verify, and agents describing a frame the artifact contradicts. Mitigation: the lane's completion is
the sidecar files on disk plus a director re-read of the newest output, never an agent's summary.
---
Genuinely short — this is one of the cleanest lanes for hands-off setup.
1. Nothing. All four primaries are anonymous public HF/GitHub pulls. No account, no card, no token, no
EULA click. If a pull turns out to be gated at download time, stop and surface it rather than creating
an account (account creation is not mine to do).
2. One real decision, and it is a business call, not a technical one: whether any public marketing
surface may ever carry a generated frame. If yes, LTX-2's Attachment A §5 compels an intelligible
machine-generated disclaimer on that surface, and the §2.3 revenue-ceiling posture makes it previz-tier
anyway. RECOMMENDATION: no — keep every public frame on real captures, which is already the ruled
Capture lane and costs this lane nothing. Surfaced as a brief, not a bare question.
3. The Puget burn-in/benchmark sheet ask (already on record in the runbook as a Josh QC action) — this
lane's overnight exclusive runs are the machine's longest sustained GPU load, so that thermal/power
baseline is the one this lane most depends on.
---
1. Measured peak VRAM and wall-clock per model at a fixed 720p/121-frame I2V job. Every ESTIMATE in §1.4
is replaced with an observation, and the runbook's own "measure them on day one" correction applies here too.
2. Does ltx-2.3-22b-distilled fp8 actually fit in 32GB headless? If no, LTX leaves the lane entirely and
Wan 2.2 becomes both primary and fallback — a materially simpler, fully-Apache lane.
3. Co-residency probe: can HunyuanVideo-1.5 at its 14GB offload floor complete a job while the UE editor
holds the project open, without OOM? This is the only question that could earn this lane a *windowed*
class instead of *overnight-exclusive*.
4. Licence re-read of all four, folded into the existing GPU-free pre-flight — plus a re-check of the
Cosmos 3 ComfyUI issue (#14228) and of whether Wan has released anything past 2.2.
5. Critic pass-rate per subject class (realm / creature / architecture / crowd) on a fixed 6-shot set —
which model wins where, since "one primary for everything" is unlikely to survive contact.
6. I2V-from-a-Lane-P-plate vs T2V-from-prompt: does the paintover's proportion-inheritance property
survive into motion? If it does, tooth 2 stops being a restriction and becomes the default method.
7. SeedVR2 720p→1080p/4K vs re-generating at higher resolution — does the upscale hold identity, and is it
cheaper than the native-res run?
8. ComfyUI native nodes vs custom packs for each model on the day — which is on the maintained path, given
this lane's whole install argument rests on native support.