MOTION_TOOL_SURVEY_2026-08-07.md

pipelines/MOTION_TOOL_SURVEY_2026-08-07.md

THE EXEMPLAR-TO-MOTION TOOL SURVEY — 2026-08-07

What this is. The chartered survey answering docs/spine/DECISIONS_PENDING_JOSH.md:4027-4031:
*"TOOLING: an exemplar-to-motion tool survey is chartered (image-to-video / consistent-character
motion: the NVIDIA stack and the open field)… Pilot = Rock Throw itself (A_E_003), the named case,
end to end: exemplar -> judged -> motion strip -> judged."* It is a survey plus a proof plan.
It is not a purchase, not a run, and it adopts nothing.
Authority. PROPOSAL-TIER, subordinate to canon. It says HOW, never WHAT. Canon binding here:
the exemplar-first motion law (docs/spine/DECISIONS_PENDING_JOSH.md:4017-4031), the light law and
four-beat grammar (docs/D1-27_ELEMENT_VISUAL_LAW.md), and the A_E_003 authoring contract
(docs/ABILITY_EXEMPLAR_DOSSIER_A_E_003.md). Where this document and canon disagree, canon wins
and this document is the defect.
Honest register. The count of acceptable ability visuals is 0
(docs/spine/DECISIONS_PENDING_JOSH.md:4037-4038) and this document moves it nowhere. Every claim
in §1's requirement tables, §2's candidate tables and §4's proof plan is marked
[MEASURED-ON-BOX] (verified on this machine today), [PRIMARY] (vendor licence / model card
/ API fetched today, URL given), [VENDOR CLAIM] (the vendor's own number, not reproduced by
us), or [THIN EVIDENCE] (named, not ranked). §3's shortlist prose and §5's findings are
SUMMARIES of those tagged rows — they inherit the tag of the row they summarize and are not
re-tagged sentence by sentence, so a claim in §3 is checkable only by reading back to its §2 row.
No VRAM figure in this document is a measurement of *our* card — every one of them is an estimate
and is labelled so, which is precisely what §4 exists to fix.
PROOF-BEFORE-SPEND is law here (docs/proposals/music/MEASURED_LOOP_PROOF.md:15, Josh
verbatim: *"Give me a few sample tracks so I can know you can actually do this and improve before I
go spend $1000 more."*). §4 is the measured sample that must exist before any dollar moves.

---

§0 THE HEADLINE, IN FIVE LINES

1. No new tool needs installing; only weights and a measurement. The A2 ledger row

(docs/PIPELINE_LEDGER.md:91) describes a survey gap — "run the chartered exemplar-to-motion

tool survey" — and this document is the artifact that fills it. (A2 was never flagged BLOCKED;

A3 is, "BLOCKED behind A2".) The survey's own answer is that every conditioning mechanism this

pipeline needs is already installed on this box as a native ComfyUI workflow template —

first/last-frame anchoring, reference-image identity, trajectory control, depth/pose control,

video upscale, frame interpolation and frame extraction. What is missing is **weights on disk and

one measured run**, not a tool. [MEASURED-ON-BOX]

2. The dossier's own strip shape rules out plain image-to-video. The exemplar sits *mid-strip* —

three cast frames come before it (docs/ABILITY_EXEMPLAR_DOSSIER_A_E_003.md:153-156). A

first-frame-only I2V model cannot produce those three frames from the exemplar. The route must

anchor the exemplar as a last frame as well as a first. That single structural fact selects

the shortlist.

3. The whole recommended route is Apache-2.0 and costs $0. Wan 2.2 / 2.1 weights are anonymous

public pulls onto a drive with 2936.29 GiB free (2.87 TiB; 3.15 TB decimal). [MEASURED-ON-BOX]

4. The hosted tier is a real fallback and is cheap — the API nodes are already installed, so it

is credits-only — but it carries two blockers a spend gate cannot buy away: the face lane's

standing exclusion on sending Josh's floor sheet to a third party, and the fact that this pilot's

caster is a seven-year-old, against a Runway minors clause whose scope is

product-line-limited and is stated exactly in §2.3.

5. **Two aggregator claims that would have cost us are false, and both were killed by primary

sources fetched today**: LTX-2.3 is *not* Apache-2.0, and no Wan model past 2.2/Animate-2 has open

weights.

---

§1 THE REQUIREMENT — derived from the dossier, not from the tools

A survey that starts from what the models do produces a wish list. This one starts from what the

A_E_003 dossier already specifies and asks each tool whether it can consume and emit it.

1.1 What the tool must CONSUME

#InputCanon source
I1ONE graded exemplar frame — full-figure (face AND body AND dress), the signature-read pause: the stone hanging at hip height, the socket open at the feet, the throwing arm uncommitteddocs/ABILITY_EXEMPLAR_DOSSIER_A_E_003.md:66-78; whole-entity floor docs/spine/DECISIONS_PENDING_JOSH.md:4068-4075
I2The motion plan — the four beats with authored per-beat content: CAST (earth cracks, grain rises, stone follows, arrives half a beat late), BODY (a seven-year-old's honest overhand throw), FLIGHT (flat, fast, heavy, shedding grit), IMPACT (surface dents/spalls first, stone usually splits, target loses footing), AFTERMATH (fragments scatter, loose earth slides back into the still-open socket)docs/ABILITY_EXEMPLAR_DOSSIER_A_E_003.md:131-149; the grammar itself docs/D1-27_ELEMENT_VISUAL_LAW.md:26-50; the Earth dispatch beats docs/D1-27_ELEMENT_VISUAL_LAW.md:307-310
I3No per-frame text description. Identity, scene, light and palette are carried by the exemplar and are never re-described per frame. Independent per-tile text prompts are BANNED. This does not mean the graph has no text input — see C5, which discloses the one every shortlisted template carries and rules the discipline on itdocs/spine/DECISIONS_PENDING_JOSH.md:4017-4022; restated as the working rule at docs/ABILITY_EXEMPLAR_DOSSIER_A_E_003.md:157-161

1.2 What the tool must EMIT

#Output requirementCanon source
O18–12 frames, allocated 3 cast / 1–2 flight / 2 impact / 2 aftermath, with the exemplar itself sitting at the signature pausedocs/ABILITY_EXEMPLAR_DOSSIER_A_E_003.md:153-156
O2Frames, not a clip. The deliverable is the boards-law reviewable surface: phone-viewable, one vertical scroll. A video file is an intermediate, not the artifactdocs/ABILITY_EXEMPLAR_DOSSIER_A_E_003.md:155; viewability floor docs/spine/DECISIONS_PENDING_JOSH.md:4223-4230
O3Identity holds on every frame against the graded keystone — G1 re-applied to the strip as G7docs/ABILITY_EXEMPLAR_DOSSIER_A_E_003.md:178-179, 192-194
O4The light law holds on every frame. Earth never emits; it reads by occlusion, by the dust it puts into existing daylight, and by silhouette. The seven banned classes are armed against the output: no aura/halo/corona/rim-light without a source, no ring or disc of light in the air, no arcane geometry, no glowing eyes, no self-lit stone, no particle sparkle, no magic light shaftdocs/D1-27_ELEMENT_VISUAL_LAW.md:78-84 (the seven), :278 (Earth never emits), :111 (the arming, and its asymmetry — see §4 instrument 2); gate G2 docs/ABILITY_EXEMPLAR_DOSSIER_A_E_003.md:180-181
O5Source persistence — the open socket stays visible and legible across the whole strip, and the downrange scatter appears and persists. *"A frame where the stone appears at the hand with no socket has failed regardless of quality"*gate G3 docs/ABILITY_EXEMPLAR_DOSSIER_A_E_003.md:182-184; the universal cast rule docs/D1-27_ELEMENT_VISUAL_LAW.md:32
O6Period + world hold — no modern surface anywhere in frame; the setting stays Ch 3 Manggaraigate G4 docs/ABILITY_EXEMPLAR_DOSSIER_A_E_003.md:185-186

1.3 The five derived constraints that actually select the tool and discipline its inputs

time. A first-frame-conditioned I2V model can only run forward from the exemplar, so it can emit

flight/impact/aftermath and structurally cannot emit the cast beat without either (a) anchoring

the exemplar as the last frame of a cast segment, or (b) treating the exemplar as a

*reference* rather than a *frame*. **Plain I2V is therefore not a complete answer to this

pipeline, at any quality.** This is the sharpest requirement in the survey and it was not visible

from the tool side.

particle sparkle, rim light and energy rings when asked for a magic effect. O4 forbids all seven.

The mitigations are structural, not prompt-side: prefer conditioning that is geometric

(first/last frame, reference image, drawn trajectory, depth/pose control) over conditioning that is

semantic (text describing the effect), and run the armed banned-class scan on the OUTPUT frames

rather than trusting the prompt.

to the UE editor, and never co-launches with a 3D-gen run

(docs/pipeline_review/tech_research/PIPE_VIDEO_2026-07-29.md:118-123). A route that needs more

than ~29 GB in one graph is out.

review surface, which would let it sit in the previz tier

(docs/pipeline_review/tech_research/PIPE_VIDEO_2026-07-29.md:19-24). Recommend against. The

dossier makes the graded frames *"the look-target the real VFX pass is judged against"*

(docs/ABILITY_EXEMPLAR_DOSSIER_A_E_003.md:149), and a look-target you cannot legally re-derive

once revenue crosses a ceiling is a dependency, not a reference. Ratify only unrestricted-licence

routes; keep ceilinged routes named and available for previz.

Disclosed rather than buried, because it looks at first like a conflict with I3: all three

shortlisted templates load CLIPLoader + CLIPTextEncodevideo_wan2_2_14B_flf2v.json (4

loaders / 8 encodes), video_wan_vace_14B_ref2v.json (4 / 4), video_wan_ati.json (2 / 4) — and

every one of them loads umt5_xxl_fp8_e4m3fn_scaled. **Wan requires a umt5 text conditioning

input**; there is no Wan graph without one, so "no text at all" is not an available configuration

and claiming it would be false. [MEASURED-ON-BOX] **THE RULED DISCIPLINE, which keeps the L4022

ban fully intact: the positive conditioning is ONE shared segment-level motion-plan string per

segment**, derived from the dossier's own four beats

(docs/ABILITY_EXEMPLAR_DOSSIER_A_E_003.md:131-149) — the arm coming round, the stone's flat

heavy line, the surface denting, the fragments and the loose earth sliding back. It is **never

per-frame, and it never re-describes identity, scene, light or palette** — those are carried

by the anchor/reference pixels and by the exemplar alone, which is precisely what L4022 bans

re-describing. Anything a text string could add about the *look* belongs in the negative slot as

the banned-class vocabulary, never in the positive slot as an instruction. The string is authored

once, recorded with the run, and is the same string across both arms so the blind comparison in §4

is a comparison of models rather than of prompts.

1.4 The dependencies, stated honestly — G6 is the canon one, and it is not the only one

The strip cannot be generated until G6 passes — Josh's grade on the exemplar frame, *"the

only gate that moves the acceptable-visuals count 0 → 1. Until it passes, everything stays at concept

tier and no derivation spends GPU"* (docs/ABILITY_EXEMPLAR_DOSSIER_A_E_003.md:190-191). §4 respects

that: the capability measurement runs on a throwaway frame and never touches the exemplar.

G6 is the only CANON dependency; it is not the only dependency. Two production ones sit beside

it and are named here so §4's rung headers are not read as "one gate away": **Arm A additionally

needs the derived, separately-gated pre-cast still** (§3.1's stated cost — it must pass G1-G5 in its

own right before it may anchor a segment), and **Arm B additionally needs a measured, provenance-

recorded VACE quant** that fits 32 GB, since the shipped fp16 does not (§4 Rung 0b). Both are $0 and

neither is blocked by G6, which is exactly why they are scheduled before it.

---

§2 THE CANDIDATE TABLE

Licence posture key: CLEAN = unrestricted (Apache/MIT/OpenMDW), eligible for a load-bearing

look-target · CEILINGED = commercial-free below a revenue threshold, previz-only per the ruled

posture · TERRITORY = markets excluded, previz-only permanently · HOSTED = vendor ToS governs.

"On box" means the ComfyUI workflow template and its nodes are installed today at

D:/comfyui/.venv-art/Lib/site-packages/comfyui_workflow_templates_json/templates/, on ComfyUI

0.29.0 (D:/comfyui/ComfyUI/comfyui_version.py:3). Video weights are not yet pulled — the

only files in models/diffusion_models/ are flux-2-klein-base-4b, qwen_image_edit_2511_fp8mixed,

z_image_turbo_bf16 and seedvr2_3b_int8_convrot. The scope of that sentence is that one directory:

models/checkpoints/ separately holds ace_step_1.5_turbo_aio.safetensors (the music lane's

all-in-one, not a video model), and no claim is made here about the other model directories.

[MEASURED-ON-BOX]

2.1 LOCAL / OPEN — the primary field

ToolConditioning mechanism (what holds identity)LicenceOn box?VRAM fit @ 32 GBMeets C1?Evidence
Wan 2.2 14B FLF2VWanFirstLastFrameToVideo on the I2V checkpoints (wan2.2_i2v_high/low_noise_14B_fp8_scaled + umt5_xxl_fp8 + wan_2.1_vae, with the lightx2v_4steps LoRAs)First AND last frame pixel anchoring. Identity is carried by the anchor frames themselves — no adapter, no textCLEAN — Apache-2.0YESvideo_wan2_2_14B_flf2v.json~16–22 GB [ESTIMATE, unmeasured]YEStemplate contents [MEASURED-ON-BOX]; licence [PRIMARY] https://raw.githubusercontent.com/Wan-Video/Wan2.2/main/LICENSE.txt (verified in PIPE_VIDEO_2026-07-29.md:54-58); node documented [PRIMARY] https://docs.comfy.org/tutorials/video/wan/wan2_2
Wan 2.1 VACE-14B ref2v / flf2vWanVaceToVideo; reference images + masks + optional control videoTrue reference-image conditioning (context adapter). The exemplar is an identity anchor the model honours throughout, not just an endpoint frameCLEAN — Apache-2.0YESvideo_wan_vace_14B_ref2v.json, video_wan_vace_flf2v.json, _v2v, _inpainting, _outpaintingtemplate ships fp16 (wan2.1_vace_14B_fp16) ≈ 28 GB — does not fit with the encoder. GGUF/fp8 community builds exist and are the routeYEStemplate [MEASURED-ON-BOX]; quant build verified to exist today [PRIMARY] HF API for QuantStack/Wan2.1_14B_VACE-GGUF — repo present, licence tag apache-2.0, siblings include Wan2.1_14B_VACE-Q4_0.gguf, -Q5_K_M.gguf, -BF16.gguf (a community re-quantization, so its provenance is recorded before use per §4 Rung 0b); mechanism family named in the CVPR-2026 literature review of R2V/VACE-class adapters
Wan 2.1 ATIWanTrackToVideo (Wan2_1-I2V-ATI-14B_fp8_e4m3fn + clip_vision_h)Drawn trajectory control. The motion plan literally becomes the input: the stone's ballistic line and the arm's arc are tracks on the exemplarCLEAN — Wan 2.1 Apache-2.0 lineageYESvideo_wan_ati.jsonfp8 14B, ~16–20 GB [ESTIMATE]forward onlytemplate contents [MEASURED-ON-BOX]
Wan 2.2 14B Fun Control (wan2.2_fun_control_high/low_noise_14B_fp8_scaled)Control-video conditioning — pose / depth / canny driving a 2.2-quality generationCLEAN — Apache-2.0YESvideo_wan2_2_14B_fun_control.json (+ fun_camera, fun_inpaint, 5B variants)fp8, ~16–22 GB [ESTIMATE]with a rendered driver[MEASURED-ON-BOX]
Wan2.2-Animate-2-14BReference image + driving video → identity-preserved animation; *"high-fidelity motion generation and strong identity preservation by eliminating intermediate motion extractors"*CLEAN — Apache-2.0not templated*"default settings… tuned for 8× A800… 480P on 2× A800"* [VENDOR CLAIM] — does not fit as shippedneeds a driver video[PRIMARY] https://huggingface.co/Wan-AI/Wan2.2-Animate-2-14B ; created 2026-07-14, Diffusers + Distilled variants 2026-08-06 [PRIMARY] HF API
Wan 2.1 SCAIL-2 character replacement (wan2.1_14B_SCAIL_2_fp16 + sam3.1_multiplex)Masked character replacement inside an existing videoCLEAN lineageYESvideo_wan21_scail2_character_replacement.json (+ int8)int8 variant is the fitn/a (later lane)[MEASURED-ON-BOX]
LTX-2.3 22Bvideo_ltx2_3_flf2v.json, _i2v, _ia2v, _ic_lora, _id_lora; IC-LoRA union-control and depth/canny-to-video on the LTX-2 19B lineFLF2V plus an ID-LoRA (ltx-2.3-id-lora-talkvid-3k) plus IC-LoRA structural controlCEILINGED — "LTX-2 Community License Agreement": *"Entities with annual revenues of at least $10,000,000… are required to obtain a paid commercial use license"*; output ownership clean; mandatory machine-generated disclosure on disseminationYES (6 LTX-2.3 + 7 LTX-2 templates)ltx-2.3-22b-distilled-fp8 (~22 GB) plus a gemma_3_12B_it_fp4 text encoder → assume exclusive, may not fitYEStemplates [MEASURED-ON-BOX]; licence [PRIMARY] https://raw.githubusercontent.com/Lightricks/LTX-2/main/LICENSE — and the HF card's own tag reads ltx-2-community-license-agreement, https://huggingface.co/Lightricks/LTX-2.3. The ID-LoRA is trained on talking-head video, which is the wrong domain for a full-figure action strip [THIN EVIDENCE on transfer]
HunyuanVideo-1.5 720p I2VFirst-frame I2VTERRITORY — Tencent Community; *"DOES NOT APPLY IN THE EUROPEAN UNION, UNITED KINGDOM AND SOUTH KOREA"* → previz-only permanentlyYESvideo_hunyuan_video_1.5_720p_i2v.json14 GB with offload [VENDOR CLAIM]noPIPE_VIDEO_2026-07-29.md:59-68 [PRIMARY, re-read not repeated]
Kandinsky-5 Lite I2V · Capybara v0.1 I2V · HuMo 17B · causal-forcing I2V · WanMove · Bernini-Rvarious I2V / human-centric / streamingUNVERIFIEDYES (all templated)unmeasuredunknown[THIN EVIDENCE] — named so they are a decision rather than a rediscovery; not ranked, no licence read performed
SeedVR2 3B/7B int8Video/image restoration + upscale; conditioning is the source latent, no text encoderCLEAN — Apache-2.0YES — templates *and already wired in this repo* (harness/asset_factory/seedvr2_refine.py)3B int8 on disk todayn/a — finisher[MEASURED-ON-BOX]; harness/asset_factory/seedvr2_refine.py:1-45
FILM frame interpolation (film_net_fp16) · SDPose wholebody video→pose · Depth-Anything-3 video depthstrip densification; driver extractionCLEAN/unverifiedYESutility_video_frame_interpolation.json, utility_sdpose_ood_video_to_pose_map.json, utility_depth_anything3_video_depth_estimation.jsonsmallutilities[MEASURED-ON-BOX]

2.2 THE NVIDIA STACK

ToolMechanismLicenceOn box?VRAMVerdict
Cosmos 3nvidia/Cosmos3-Nano (16B), Cosmos3-Super (65B), Cosmos3-Edge, and dedicated Cosmos3-Super-Image2Video + -Image2Video-4Step configurationsimage-to-video world-foundation modelOpenMDW-1.1 — the cleanest licence in the field, commercial use permitted, no MAU, no territory, no output claimNONano BF16 >29 GB + offload; but community INT4-AWQ, NVFP4-AWQ, FP8 and NF4 builds of both Nano and Super now exist on HF — the checkable ids, all confirmed present in today's HF model listing: Reza2kn/Cosmos3-Nano-INT4-AWQ, Reza2kn/Cosmos3-Nano-NVFP4-AWQ, Reza2kn/Cosmos3-Nano-FP8, SanDiegoDude/Cosmos3-Nano-nf4, Intel/Cosmos3-Super-int4-AutoRound, plus first-party nvidia/Cosmos3-Super-Image2Video and nvidia/Cosmos3-Super-Image2Video-4StepWATCH — and the VRAM objection is now weaker than PIPE_VIDEO recorded. The remaining blocker is the ComfyUI gap: issue #14228 "Support for NVIDIA Cosmos 3 model family?" is still OPEN, opened 2026-06-02, no maintainer comments, no branches or PRs — re-verified today [PRIMARY] https://github.com/Comfy-Org/ComfyUI/issues/14228 ; model list [PRIMARY] HF API
Cosmos-Predict2 (2B/14B) I2Vimage-to-videoNVIDIA Open Model Licence, commercial permittednative ComfyUI support exists for the Predict2 line2B is the fitEVALUATE at benchmark day only. Training objective is physical-AI/robotics, not aesthetic; PIPE_VIDEO_2026-07-29.md:88-90 records the mismatch and upstream limited-maintenance
GEM-X · Kimodo · ARDYvideo→SOMA mocap; text→motion; real-time motionApache-2.0 code + NVIDIA Open Model weightsWRONG LANE, and correctly so. These emit skeletal animation data, not pixels. They serve E4/G2, not the strip. docs/MODEL_STACK_AUDIT_AUG2026.md:161-162; PIPE_VIDEO_2026-07-29.md:101-102. Named here so this survey is not re-run against them
NVIDIA Lyra 2.0Internal Scientific Research: *"may not… generate works for sale"*EXCLUDED, standing (PIPE_VIDEO_2026-07-29.md:105)

2.3 HOSTED — the fallback tier

The material finding: this tier needs no installation. The ComfyUI API-node templates are already

on the box, so the hosted fallback is credits only, with zero setup cost. [MEASURED-ON-BOX]

ServiceInstalled templateConditioning shapeIndicative pricePosture
Vidu Q3api_vidu_start_end_to_video.json, api_vidu_reference_to_video.json, api_vidu_q3_image_to_video.jsonstart+end frame AND reference-to-video — the exact two grammars C1 requiresnot quoted herethe best-shaped hosted candidate for this pipeline
Wan 2.6 / 2.7 (API)api_wan2_6_i2v.json, api_wan2_7_i2v.json, api_wan2_7_r2v.json, api_wan2_7_video_edit.jsonI2V + reference-to-videonot quoted herethe newest Wan quality with the same conditioning grammar as the local route — the natural hosted escalation. Independently confirms 2.6/2.7 are API-only
Kling v3api_kling_v3_video.json, api_kling_o3_video_edit.jsonI2V / edit~$0.112/s at 1080p [VENDOR-AGGREGATOR CLAIM]ToS read owed
Runway Gen-4 Turbo / Gen-4.5 / Aleph 2api_runway_gen4_turo_image_to_video.json, api_runway_aleph2_video_edit.jsonI2V / video edit~$0.15/s [AGGREGATOR CLAIM]see the minor clause below
ByteDance Seedance 1.5api_bytedance_seedance1_5_image_to_video.jsonI2VToS read owed
Luma Ray3.2 · Grok reference-to-video · Hailuo/MiniMax · Google Gemini Omni Flash video edittemplatedvariousnamed, not ranked

Two blockers on this tier that money does not buy away.

1. The standing exclusion. The face lane's own survey already excluded *"every API node class in

/object_info… They fail the box requirement, and using them would send Josh's own floor sheet to a

third party, which is not this lane's call"*

(build/3d/model_survey/face_2026_08_06/EXCLUDED.json). The exemplar frame is that floor

sheet's derivative. This is a director-tier call, not a lane's.

2. The subject is a seven-year-old. The pilot's caster is CHAR_0001 at the age-7 keystone

(docs/ABILITY_EXEMPLAR_DOSSIER_A_E_003.md:73-78). Runway's usage policy restricts *"Characters

based on the face or voice of a person under the age of 18"* [PRIMARY]

https://runway.com/safety/usage-policy — **and the clause's SCOPE has to be stated with it,

because scope is half the finding.** It does not sit in the blanket platform prohibitions; it sits

under the heading "Additional Character & Game Worlds Policies", which the policy itself scopes

to the Runway Characters & Game Worlds product line. It is therefore **not, on its face, a rule

about Gen-4 Turbo image-to-video**, which is the endpoint this survey would actually call

(api_runway_gen4_turo_image_to_video.json). Re-verified today [PRIMARY]. **On top of that

scope question, the honest read is that the clause targets real minors, not a synthetic character,

so what remains is a genuine ambiguity rather than a categorical ban.**

Stating it either way would be a defect: overclaiming it is the care-inflation the standing

calibration forbids; ignoring it risks an account action on the account that holds the pipeline.

The correct handling is §4's recommendation — **measure the hosted capability on a non-protagonist

stand-in subject**, so the question never has to be answered to get the number.

---

§3 THE RANKED SHORTLIST

1. Wan 2.2 14B FLF2V — WanFirstLastFrameToVideoTHE ROUTE, with a cost stated up front

The only candidate that is simultaneously CLEAN-licensed, already templated on this box, and

structurally able to satisfy C1: run the cast beat with the exemplar as the *last* frame, and the

flight→impact→aftermath run with it as the *first*, so every frame in the strip is anchored to graded

pixels at at least one end. It reuses the I2V checkpoints (no separate FLF2V weight set), and the

4-step lightx2v LoRAs make iteration cheap.

THE COST, NAMED RATHER THAN BURIED — FLF2V NEEDS A SECOND ANCHOR IT DOES NOT HAVE. A first/last

node takes two frames. The cast segment supplies the exemplar as the *last* one; **there is no

canon artifact for the first one.** A neutral pre-cast still has to be MADE, and the only lawful way

to make it is to derive it from the graded exemplar by image-editing / conditioning — the

qwen_image_edit_2511_fp8mixed route already on disk is the obvious instrument — because

text-prompting a fresh still from scratch is exactly the per-tile-text architecture L4022 bans, and

a second independently-generated still would put a second, ungraded identity into the strip's own

anchor set. **That derived still is itself an artifact that must pass G1-G5 before it may anchor

anything** (identity, light law, source-in-frame, period, legibility), which is a whole extra gated

sub-deliverable and a second place for identity to drift.

**This weakens the FLF2V-first ranking, and the honest statement of it is that VACE needs no second

anchor at all** — a reference-conditioned run takes the exemplar and the beat targets and nothing

else, so Arm B has zero derived-artifact surface where Arm A has one gated one. That is a

stronger argument for VACE than §3.2 previously made, and it is now made here.

FLF2V still leads, for three reasons that survive the cost, and the ranking is held rather than

flipped: (a) its weights are first-party Comfy-Org repackages of Apache-2.0 Wan 2.2, where the

only VACE build that fits 32 GB is a community re-quantization whose provenance is an open item

— an unmeasured third-party artifact is a real cost too, and it sits on Arm B's *only* viable

configuration rather than on an accessory; (b) FLF2V gives pixel-exact endpoint anchoring, which

is a stronger identity guarantee per frame than reference conditioning, and the derived pre-cast

still inherits its pixels from the graded exemplar rather than from a new generation, so the

identity provenance chain is unbroken even though the artifact count is higher; (c) Wan 2.2 is a

generation ahead of Wan 2.1 VACE on quality. **The ranking is explicitly provisional and Rung 0/1

can overturn it:** if the derived pre-cast still fails its own G1-G5, or if the VACE quant measures

clean and its ref2v cast frames hold identity better, Arm B becomes the route — which is why §4

runs both blind rather than running the winner.

2. Wan 2.1 VACE-14B ref2v — THE IDENTITY ARM, and the one with no second-anchor problem

When pixel anchoring is not enough — and on a 3-frame cast run away from the anchor it may not be —

VACE conditions on the exemplar as a reference with mask control, which is the mechanism class

built for exactly this problem. Same Apache family, one node, one checkpoint swap. **Its structural

advantage over Arm A is that it needs no second anchor**: the exemplar alone conditions the run, so

the cast beat costs no derived, separately-gated still. Its cost is the mirror of that: the shipped

template is fp16 (≈28 GB, does not fit beside the encoder) and a GGUF/fp8 build has to be sourced

from a community re-quantizer and its provenance recorded — which is why §4 Rung 0b measures it as

its own sub-rung rather than assuming it.

3. Wan 2.1 ATI — WanTrackToVideoTHE MOTION-PLAN ARM

The closest thing in the open field to *"the motion plan is the input"*: the stone's flat, fast,

heavy ballistic line and the shoulder coming round are drawn as trajectories on the exemplar

rather than described in text — which is the single best mitigation for C2, because a drawn arc

cannot ask for a glow. Third rather than first only because it is Wan 2.1-quality and forward-only,

so it complements the FLF2V route rather than replacing it.

Named behind the shortlist, deliberately: LTX-2.3 FLF2V is the quality rung and it is fully

templated here, but it is CEILINGED (C4) and may not fit beside a 12B text encoder (C3) — keep it as

the previz/quality comparator, not the ratified route. Cosmos 3 has the best licence in the

entire field and now has quantized builds; it is one ComfyUI PR away from jumping to the top and

should be re-checked at every benchmark day. HunyuanVideo-1.5 stays territory-excluded and

previz-only. Wan2.2-Animate-2 is the strongest identity mechanism published this cycle and is

Apache-2.0, but it needs a driving video and does not fit as shipped — its moment comes when the A1

wave has an in-engine or Blender-rendered throw to drive it with, and at that point it is the best

route in this document.

---

§4 THE PROOF-BEFORE-SPEND PLAN

The law, applied: a capability is demonstrated on a measured sample before Josh's money moves.

The music lane is the exemplar — three playable samples and a per-dimension diff before the

$1,000 (docs/proposals/music/MEASURED_LOOP_PROOF.md:15, 215-232). Here the analogue is: an 8–12

frame strip with an identity number and a banned-class scan on every frame, produced entirely on

owned hardware with Apache-2.0 weights, before any hosted credit is bought.

RUNG 0a — THE FLF2V CAPABILITY MEASUREMENT · $0 · unblocked today · does not touch the exemplar

the two lightx2v_4steps LoRAs, umt5_xxl_fp8_e4m3fn_scaled, wan_2.1_vae. Land on the Gen4 path

per the ruled disk layout (PIPE_VIDEO_2026-07-29.md:139-143); 2936.29 GiB free (2.87 TiB /

3.15 TB decimal) [MEASURED-ON-BOX].

any existing landed plate, explicitly not the exemplar, which does not exist yet (G6).

resolution, whether the two-expert high/low-noise pair serialises inside 32 GB. Land them as a

generation record beside the run, in the shape the 3D lane already uses.

video_wan2_2_5B_ti2v.json (*"the 5B fits 8 GB"* is the vendor's figure, [VENDOR CLAIM], and

this rung is where it stops being one) and record the quality cost honestly.

RUNG 0b — THE VACE CAPABILITY MEASUREMENT · $0 · unblocked today · its own sub-rung, because Arm B does not ship in a form that fits

Rung 0a covers one arm. Arm B cannot be run from what the box ships: video_wan_vace_14B_ref2v.json

loads wan2.1_vace_14B_fp16 at ≈28 GB, which does not fit in 32 GB beside the umt5 encoder and

the VAE. Arm B therefore has a pull-and-measure step of its own, and it is a third-party one,

which changes what has to be recorded.

tag apache-2.0, siblings include Wan2.1_14B_VACE-Q4_0.gguf, -Q5_K_M.gguf, -BF16.gguf

[PRIMARY, HF API fetched today]. Plus umt5_xxl_fp8_e4m3fn_scaled and wan_2.1_vae if 0a has not

already landed them.

not a first-party Comfy-Org repackage like Arm A's weights. Record, beside the run: the exact repo

id and revision SHA, the file and its hash, the upstream weights it claims to derive from

(Wan-AI/Wan2.1-VACE-14B), the licence tag as fetched, and the fetch date. A look-target derived

from an unattributed binary is the C4 dependency problem in a different costume.

on the same throwaway plate 0a used, and emit the same numbers: peak VRAM, wall clock, frames,

resolution.

fits 32 GB, Arm B is out on a measured axis and the shortlist's §3.2 is corrected in place

rather than carried.

RUNG 1 — THE MEASURED SAMPLE · $0 · blocked on G6, on Rung 0a/0b, and (Arm A only) on the derived pre-cast still

The sample: the A_E_003 Rock Throw exemplar frame → the 8–12 frame strip, exactly the shape

docs/ABILITY_EXEMPLAR_DOSSIER_A_E_003.md:153-156 specifies. Two arms, run blind against the same

frames:

the 3 cast frames). Segment 2: exemplar as FIRST frame → flight, impact, aftermath.

**THE PRE-CAST STILL IS ITS OWN DELIVERABLE, and its derivation route is named rather than

assumed: it is DERIVED FROM THE GRADED EXEMPLAR by image-editing / conditioning** — the

qwen_image_edit_2511_fp8mixed route already on disk, or the same FLF2V/VACE stack run backwards

a beat — and it is never text-prompted from scratch, which is the L4022 ban and would also put

a second ungraded identity into the strip's own anchor set. **It takes its own gate pass — G1

through G5 — before it may anchor anything**, on the same instruments as the strip frames below.

If it cannot pass them, Arm A has failed on a measured axis and Arm B is the route (§3.1).

second anchor and no derived artifact** — this is the structural advantage §3.2 names.

so the blind comparison is between models and not between prompts.

THE ASSEMBLY STEP, corrected. Extract the beat frames deterministically, then assemble them into

the single phone-viewable vertical sheet O2 requires. **harness/asset_factory/contact_sheet.py is

NOT the tool for this** — it takes a .glb per candidate and shells out to Blender to render four

views of each, so it cannot assemble video frames at all. The right pattern is the **2D PIL grid this

repo already uses in two places**: Image.new + ImageDraw with labelled cells, exactly as

harness/asset_factory/age_keystones.py:1235, 1283, 1316 and

harness/asset_factory/objref_pick_sheet.py:56-57 build their sheets, and as

build/3d/characters/body_specs/a18_face_refine.py:156-159 builds its labelled comparison strip.

This is a small piece of NEW TOOLING — a frame-strip assembler of perhaps forty lines against an

established in-repo pattern — and it is declared as such rather than hidden, correcting this plan's

earlier claim that nothing here is new tooling.

**The instruments — two exist in this repo, one is an agent-time critic read, and one required axis

is UNMET:**

1. Identity distance — exemplar_match, and its real implementation and real scope. The metric

is a face copy index measured on HEAD CROPS: build/3d/characters/body_specs/a18_face_refine.py:139

records "exemplar_match": ident["match"] from merge_route_a_v8.measure_head, which returns

RACE.find_face(...)["copy_index"] — a match of the approved inner face over scale and position.

harness/asset_factory/seedvr2_refine.py:14-20 narrates the resulting numbers (0.91889

unrefined, drifting to 0.90621 / 0.88914 under the refine ladder, the drift Josh ruled against)

but does not implement the metric. Report it per frame, not per strip. **Its scope is the

face**, which is stated so it is not over-claimed.

2. The banned-class scan — and D1-27's arming is ASYMMETRIC, so the rule has two branches. The

vocabulary is already armed (docs/D1-27_ELEMENT_VISUAL_LAW.md:111) and the seven classes are at

:78-84. The law's own arming: *"On the dispatch grammar any hit fails, because that surface

already forbids negation and a banned token there can only be an instruction to draw the banned

thing. On a tier read — design prose written for a person, where negation carries its meaning — a

hit fails only when its own sentence is affirmative."* Applied here:

graph): any hit fails, full stop. Negation does not survive the graph, so a banned token in

a prompt is an instruction to draw the banned thing.

for a person — a hit fails only when its own sentence is AFFIRMATIVE. A critic writing *"no

rim-light without a source; the dust reads by occlusion"* has scored a PASS and must not be

counted as a hit; a critic writing *"a faint rim-light sits along the shoulder"* has scored a

FAIL. Scoring the negated sentence as a hit is instrument-misapplication, and it is called out

here because it is the easy error and it would fail correct frames.

3. Socket persistence (O5/G3) — counted by hand across the strip and reported as N-of-N. A strip

that loses the socket has failed regardless of quality.

4. UNMET — the body/dress axis has no instrument. Instrument 1 measures the face and nothing

else, but the governing floor is the whole entity: *face, body, hands, dress all judged*

(G1, docs/ABILITY_EXEMPLAR_DOSSIER_A_E_003.md:178-179; the whole-entity beauty floor,

docs/spine/DECISIONS_PENDING_JOSH.md:4068-4075). **No numeric instrument in this repo measures

body proportion, hand count or garment consistency across a strip.** This is not folded into

instrument 1 and it is not quietly dropped: for this pilot it is carried by the same hand-authored

per-frame critic read as instrument 2, reported as a named prose verdict per frame, and it is

boarded as a real tooling gap for whoever builds the strip lane at volume.

THE COST OF THE CRITIC READ, in the honest unit. Instruments 2, 3 and 4 are **hand-authored

per-frame reads**, on the precedent this repo already set: harness/ability_vfx/judge.py is a

four-axis judge in which *"every one of the 96 rows below was judged by OPENING its four candidates

side by side. 384 of 384 candidates were opened; zero rows are carried as unread."* At 8-12 frames ×

2 arms — plus the derived pre-cast still on Arm A — that is **roughly 17-25 frames opened, read

against four axes, and written up. That is real work, and its currency is AGENT TIME, not

dollars.** The rung is $0 in spend and it is not $0 in effort; a plan that reported only the dollar

figure would be under-reporting what it costs to run.

RUNG 2 — THE GRADE

Arms A and B go to Josh as the strip, per G7

(docs/ABILITY_EXEMPLAR_DOSSIER_A_E_003.md:192-194) and only after G6. **Only his grade moves the

count.** Everything above is evidence for his eye, never a substitute for it.

THE SPEND GATE

**No dollars until Rungs 0a, 0b and 1 are measured AND both local arms have failed on a named,

measured axis.** "It looked worse" is not a failure record; "identity fell to X on the cast frames

and the banned-class scan hit affirmatively on Y of Z frames" is.

AND THAT FAILURE RECORD IS NOT SELF-CERTIFIED BY THIS LANE. A lane that both writes the failure

record and spends against it is grading its own homework, and the whole point of PROOF-BEFORE-SPEND

is that someone else looks first. So: **the failure record routes to the DIRECTOR and to Josh, with

the samples attached, and no dollar moves without that sign-off.** This is the MEASURED_LOOP_PROOF

precedent applied exactly as Josh stated it — *"Give me a few sample tracks so I can know you can

actually do this and improve before I go spend $1000 more"*

(docs/proposals/music/MEASURED_LOOP_PROOF.md:15). The samples are seen BEFORE the money moves,

not summarized after. What routes up is the strips themselves, both arms, plus the per-frame numbers

and the named axis of failure — not a verdict.

The $0 floor is STRUCTURAL, not a promise this lane is making. Verified on this box today: there

is no ComfyUI API key fileD:/comfyui/ComfyUI/user/default/comfy.settings.json,

D:/comfyui/ComfyUI/.env and D:/comfyui/ComfyUI/comfy_api_key.json are all absent, and a

recursive search for those filenames anywhere under D:/comfyui/ComfyUI returns nothing — and **no

vendor API environment variable is provisioned**: no RUNWAY*, KLING*, VIDU* or DASHSCOPE*

variable exists in the environment. [MEASURED-ON-BOX] The installed api_* templates are therefore

graph shapes with no credential behind them: **the hosted tier cannot spend a cent until someone

deliberately provisions a key**, which is itself a visible, auditable act. The spend gate below is a

policy on top of a floor that already holds mechanically.

If both arms fail that way and the director and Josh sign off, the hosted probe is bounded at $50

for the entire round. At the

indicative $0.11–0.15/s and a 5-second clip, $50 buys roughly 60–90 clips across the four

best-shaped hosted candidates (Vidu Q3 start-end + reference-to-video, Wan 2.7 R2V, Kling v3, Runway

Gen-4 Turbo) — an order of magnitude more than the 8–12 frames the strip needs, so the bound is

generous rather than tight. Anything past $50 is a new brief, not a continuation.

Two hard preconditions on that spend, both from §2.3:

or a caster-free prop cast. The question the probe answers is *"can a hosted model hold a graded

frame's identity, light law and socket across the four beats?"*, and that question does not need

the child plate to leave the box. It measures the capability at $50 and settles the ToS and

floor-sheet questions to zero exposure. Only if the stand-in probe *succeeds* does the real

question — send the protagonist plate or not — reach the director, and it arrives with evidence

attached instead of as a bare fork.

measures the wrong subject. The pipeline's hardest identity problem is specifically a seven-year-old

child's face and body at full figure, and an adult stand-in that passes proves less than it appears

to. Counter: the four axes the probe is actually testing — bidirectional anchoring, light-law

discipline, socket persistence and beat legibility — are subject-independent, and identity hold is

the one axis where the local route already has the structural advantage (the plate never

leaves). If a hosted model wins on the other three, that is exactly when the protagonist question

is worth putting to Josh.

THE $0-FIRST PATH, STATED PLAINLY

ComfyUI 0.29.0 is installed and running (harness/asset_factory/keystone_direct.py:106-107

HOST = "http://127.0.0.1:8189", COMFY_IN = "D:/comfyui/ComfyUI/input"). Every template this

survey recommends is already on disk. The weights are anonymous public Apache-2.0 pulls (Arm B's is a

community re-quantization and carries the Rung 0b provenance record). The disk has **2936.29 GiB

free** (2.87 TiB / 3.15 TB decimal). The identity instrument exists; the frame-strip assembler is

forty lines of new code against an in-repo pattern; the light-law, socket and body/dress axes are

agent-time critic reads. And the hosted floor is structural — no API key file and no vendor API

environment variable is provisioned on this box. **The entire recommended route through Rung 2 costs

zero dollars and consumes only GPU hours and agent time we already own** — which is why the spend

gate above should be hard rather than nominal.

---

§5 FINDINGS FOR THE DIRECTOR

1. A2's next step should be SPLIT — no new tool needs installing; only weights and a measurement.

Stated precisely, because the earlier framing over-reached: the A2 row

(docs/PIPELINE_LEDGER.md:91) describes a survey gap — "run the chartered exemplar-to-motion

tool survey (NVIDIA stack + open field)" — and this document is the artifact that fills it. **A2

was never flagged BLOCKED; the row that carries that flag is A3** ("BLOCKED behind A2"). So

the correction is not that the ledger mis-framed a blocker; it is that A2's single NEXT STEP cell

fuses two steps of very different readiness. Recommend splitting it into (a) measure the route

(unblocked, tonight, Rungs 0a/0b) and (b) run the pilot (blocked on G6, and on the Arm A

pre-cast still). Fused, the cell hides an hour of work that could already be done, and it holds A3

behind it.

2. **docs/MODEL_STACK_AUDIT_AUG2026.md:156 is CONFIRMED CORRECT and its evidence can be

strengthened. "Wan 2.2 14B FLF2V" is real — but there is no separate Wan2.2-FLF2V checkpoint**

on HF (the org's newest FLF2V weights are Wan2.1-FLF2V-14B-720P, 2025-04-17). The 2.2 FLF2V

route is the WanFirstLastFrameToVideo node on the 2.2 I2V checkpoints, which the

installed template proves. Worth writing down before someone hunts for a checkpoint that does not

exist.

3. The Wan-closed-weights finding now has two independent controls. The HF org listing fetched

today returns 27 models ending at Wan2.2-Animate-2-14B-Distilled-Diffusers (2026-08-06) — nothing

past 2.2/Animate-2 — and the box's own installed API templates (api_wan2_6_i2v.json,

api_wan2_7_i2v.json, api_wan2_7_r2v.json) show 2.6/2.7 existing only as remote services. The

"Wan 3.0 open weights, April 2026" claim still circulating on aggregator sites is false on both

controls. The aggregators are named so the counter-claim is checkable rather than a gesture at

"SEO content": they are the same wan27.org / oakgen.ai / spheron pages

PIPE_VIDEO_2026-07-29.md:100 already identified and killed for the Wan-2.5/2.7 version of this

claim — one family of pages, re-badged a version number later, and re-killed on the same two

primary controls.

4. LTX-2.3's licence is being widely mis-reported as Apache-2.0 right now — including on pages

that read as first-party-adjacent. The raw LICENSE fetched today is the **LTX-2 Community License

Agreement** with the $10M ceiling and the machine-generated-disclosure clause, and the HF card's

own metadata tag reads ltx-2-community-license-agreement. PIPE_VIDEO_2026-07-29.md:47-53 was

right; hold that line, and treat any future "LTX is Apache" claim as an aggregator error.

5. Cosmos 3's exclusion should be narrowed, not overturned. PIPE_VIDEO_2026-07-29.md:82-84

excluded it partly on VRAM; community INT4-AWQ / NVFP4 / FP8 / NF4 builds of Nano *and* Super now

exist, and NVIDIA ships dedicated Cosmos3-Super-Image2Video and -4Step configurations. The

ComfyUI ground (#14228, still open, zero maintainer comments, no PRs — re-verified today) is

the one that still holds. Same shape as the TRELLIS.2 narrowing in

MODEL_STACK_AUDIT_AUG2026.md:326-348: a real exclusion that was treated as broader than it is.

6. A route worth boarding rather than building now. Wan2.2-Animate-2 (Apache-2.0, reference image

A1 wave can render a greybox throw to drive it with — the somatic form is already authored

(docs/ABILITY_EXEMPLAR_DOSSIER_A_E_003.md:137-140), and a rendered proxy is a legitimate driver

that costs no licence surface. Boarded, not adopted, and named so it is a decision.

7. Owed and not done by this lane, per the wave override: the docs/DOC_MAP.md registration row

for this document, and its PIPE_VIDEO_2026-07-29.md in-place supersession

(MODEL_STACK_AUDIT_AUG2026.md:487 routes video/motion-strip verdicts to that dossier). Both are

director-landed.

8. Nothing here was surfaced as a bare question. The one genuine fork — whether the protagonist's

graded plate may go to a hosted service — is answered with a recommendation (stand-in probe first,

$50 bound) and its strongest objection, in §4.

---

§6 DERIVATION

DERIVED FROM:

this document answers verbatim; the ban on per-tile text prompts (L4022); the identity anchor

(L4023-4024); the named pilot A_E_003 (L4030-4031); the honest register, acceptable = 0 (L4037-4038).

and every derived frame are judged full-figure, not as a crop.

phone-viewable frame sheet and not a video file (O2).

§4 (L131-149) — the four-beat VFX breakdown → I2. §5 (L151-170) — the motion plan, the strip's

frame budget and the exemplar's mid-strip position → O1 and the C1 derivation. §6 (L172-199) — the

gate order G1-G8 → O3-O6 and §4's grading hook. §7 (L201-207) — the queued inputs → §4's rungs.

and the source-in-frame universal rule → O5. § the seven banned classes (L78-84) and their

arming (L111), **including the arming's stated ASYMMETRY — any hit fails on dispatch grammar, an

affirmative-only hit fails on a tier read** → O4, the C2 constraint, and the two-branch rule in

§4's instrument 2. § Earth (L278, L307-310) — Earth never emits; the per-beat dispatch strings the

strip must satisfy.

same table carrying the actual BLOCKED flag ("BLOCKED behind A2") → the finding in §5.1.

→ C4; the VRAM/scheduling law (L108-123) → C3; the Hunyuan territory read (L59-68); the Cosmos

exclusion (L75-93) narrowed in §5.5; the disk layout (L139-143); the ARDY/Kimodo wrong-lane routing

(L101-102).

Wan-family recommendation and its stated reasoning (L167-172); queue item 5 (L270-275) naming this

survey; the routing table (L487).

applied as §4's shape.

class and its stated reason → §2.3 blocker 1 and §4's stand-in recommendation.

instrument 4. build/3d/characters/body_specs/a18_face_refine.py:139 is where the metric is

actually recorded ("exemplar_match": ident["match"]); build/3d/characters/body_specs/merge_route_a_v8.py:386-408

is measure_head, which returns RACE.find_face(...)["copy_index"] — a **face copy index on a

head crop**, which is why the body/dress axis is named as unmet rather than folded in.

harness/asset_factory/seedvr2_refine.py:1-45 supplies the ComfyUI-on-box method (template

transcribed node-for-node, re-verified against the running server) and narrates the baseline

numbers (0.91889 / 0.90621 / 0.88914) in its docstring; it does not implement the metric.

opened, zero rows carried unread → the precedent for §4's per-frame critic read and the basis for

costing it in agent time.

and build/3d/characters/body_specs/a18_face_refine.py:156-159 — the established in-repo

Image.new / ImageDraw labelled-grid pattern → the frame-strip assembler §4 declares as new

tooling. harness/asset_factory/contact_sheet.py is explicitly NOT the assembler — it consumes

a .glb per candidate and shells to Blender, so it cannot assemble video frames.

D:/comfyui/ComfyUI/input) → the $0-first path.

the installed template set at

D:/comfyui/.venv-art/Lib/site-packages/comfyui_workflow_templates_json/templates/; the model

filenames and node classes inside video_wan2_2_14B_flf2v.json, video_wan_vace_14B_ref2v.json,

video_wan_ati.json, video_ltx2_3_flf2v.json, video_ltx2_3_id_lora.json,

video_wan2_2_14B_fun_control.json, video_ltx2_depth_to_video.json,

utility_seedvr2_3b_int8_upscale_video.json, utility_video_frame_interpolation.json,

utility_sdpose_ood_video_to_pose_map.json, and the api_* hosted templates; the

CLIPLoader / CLIPTextEncode / umt5_xxl_fp8_e4m3fn_scaled counts inside the three shortlisted

templates → C5; the contents of D:/comfyui/ComfyUI/models/diffusion_models (and, separately,

models/checkpoints/ace_step_1.5_turbo_aio.safetensors); 2936.29 GiB free on D: (2.87 TiB /

3.15 TB decimal); and the absence of user/default/comfy.settings.json, .env and

comfy_api_key.json anywhere under D:/comfyui/ComfyUI, plus the absence of any RUNWAY* /

KLING* / VIDU* / DASHSCOPE* environment variable → the structural $0 floor in §4's spend gate.

Wan2.2-Animate-2-14B-Distilled-Diffusers 2026-08-06); the HF model search for Cosmos3, which

returns all seven ids §2.2 now pastes (Reza2kn/Cosmos3-Nano-INT4-AWQ,

Reza2kn/Cosmos3-Nano-NVFP4-AWQ, Reza2kn/Cosmos3-Nano-FP8, SanDiegoDude/Cosmos3-Nano-nf4,

Intel/Cosmos3-Super-int4-AutoRound, nvidia/Cosmos3-Super-Image2Video,

nvidia/Cosmos3-Super-Image2Video-4Step); the HF model API for

QuantStack/Wan2.1_14B_VACE-GGUF (repo present, apache-2.0, GGUF siblings) → Rung 0b;

https://huggingface.co/Wan-AI/Wan2.2-Animate-2-14B ;

https://raw.githubusercontent.com/Lightricks/LTX-2/main/LICENSE ;

https://huggingface.co/Lightricks/LTX-2.3 ; https://docs.comfy.org/tutorials/video/wan/wan2_2 ;

https://github.com/Comfy-Org/ComfyUI/issues/14228 ; https://runway.com/safety/usage-policy —

from which the minors clause was read together with the section heading it sits under,

"Additional Character & Game Worlds Policies", scoped by the policy to the Characters & Game

Worlds product line → §2.3 blocker 2.

NOT DERIVED — authored judgment, and why it had no canon home:

and the exemplar's position within it; nothing in canon says what that implies for a tool's

conditioning grammar, because canon does not name tools. The inference — exemplar-at-index-4 means

first-frame-only I2V cannot produce the cast beat — is mine, and it is the load-bearing one in the

document. If it is wrong, the shortlist is wrong.

a ceilinged model for an internal surface. I recommend the stricter bar because the dossier makes

the strip a durable look-target. This is a judgment against the letter of the existing posture and

is named as such so it can be overruled.

arithmetically from published per-second pricing against the strip's actual frame need, and it is

deliberately generous so that the bound never becomes the reason a measurement was not taken.

subject; the idea of measuring hosted capability on a non-protagonist subject to avoid the question

entirely is mine.

animator with a greybox render from the A1 wave. Both halves are canon-grounded; the pairing is not.

quantization, not a measurement, and §4 Rungs 0a/0b exist specifically to replace them.

not weigh a derived-but-gated anchor artifact against an unmeasured third-party quant. The three

reasons given — first-party weights, pixel-exact anchoring with unbroken identity provenance, and a

generation of quality — are my weighting, the ranking is declared PROVISIONAL against Rungs 0/1,

and the counter-argument (VACE needs no second anchor at all) is stated in the same paragraph so it

can be picked up rather than having to be discovered.

what to do with a model that structurally requires a text encoder. One shared segment-level

motion-plan string, never per-frame and never re-describing identity/scene/light, is my reading of

the narrowest thing that satisfies both the ban and the graph. It is named so it can be tightened

to an empty positive prompt if the director prefers.

frames into a review sheet; the nearest are 3D contact sheets and 2D keystone grids. Declaring a

new forty-line assembler rather than stretching an existing tool is a judgment, and it corrects

this plan's earlier and wrong claim that nothing here is new tooling.

---

*Authored 2026-08-07, combat-schema wave Lane D. PROPOSAL-TIER: it rules nothing, adopts nothing and

spends nothing. Adoption routes through docs/pipeline_review/tech_research/PIPE_VIDEO_2026-07-29.md

as an in-place supersession, per docs/MODEL_STACK_AUDIT_AUG2026.md:479-487. Owed at landing: a

docs/DOC_MAP.md row with a named consumer (ledger row A2).*

Generated by harness/site/structure_site.py — the URL path is the repo path. review root