pipelines/MOTION_TOOL_SURVEY_2026-08-07.md
What this is. The chartered survey answering docs/spine/DECISIONS_PENDING_JOSH.md:4027-4031:
*"TOOLING: an exemplar-to-motion tool survey is chartered (image-to-video / consistent-character
motion: the NVIDIA stack and the open field)… Pilot = Rock Throw itself (A_E_003), the named case,
end to end: exemplar -> judged -> motion strip -> judged."* It is a survey plus a proof plan.
It is not a purchase, not a run, and it adopts nothing.
Authority. PROPOSAL-TIER, subordinate to canon. It says HOW, never WHAT. Canon binding here:
the exemplar-first motion law (docs/spine/DECISIONS_PENDING_JOSH.md:4017-4031), the light law and
four-beat grammar (docs/D1-27_ELEMENT_VISUAL_LAW.md), and the A_E_003 authoring contract
(docs/ABILITY_EXEMPLAR_DOSSIER_A_E_003.md). Where this document and canon disagree, canon wins
and this document is the defect.
Honest register. The count of acceptable ability visuals is 0
(docs/spine/DECISIONS_PENDING_JOSH.md:4037-4038) and this document moves it nowhere. Every claim
in §1's requirement tables, §2's candidate tables and §4's proof plan is marked
[MEASURED-ON-BOX] (verified on this machine today), [PRIMARY] (vendor licence / model card
/ API fetched today, URL given), [VENDOR CLAIM] (the vendor's own number, not reproduced by
us), or [THIN EVIDENCE] (named, not ranked). §3's shortlist prose and §5's findings are
SUMMARIES of those tagged rows — they inherit the tag of the row they summarize and are not
re-tagged sentence by sentence, so a claim in §3 is checkable only by reading back to its §2 row.
No VRAM figure in this document is a measurement of *our* card — every one of them is an estimate
and is labelled so, which is precisely what §4 exists to fix.
PROOF-BEFORE-SPEND is law here (docs/proposals/music/MEASURED_LOOP_PROOF.md:15, Josh
verbatim: *"Give me a few sample tracks so I can know you can actually do this and improve before I
go spend $1000 more."*). §4 is the measured sample that must exist before any dollar moves.
---
1. No new tool needs installing; only weights and a measurement. The A2 ledger row
(docs/PIPELINE_LEDGER.md:91) describes a survey gap — "run the chartered exemplar-to-motion
tool survey" — and this document is the artifact that fills it. (A2 was never flagged BLOCKED;
A3 is, "BLOCKED behind A2".) The survey's own answer is that every conditioning mechanism this
pipeline needs is already installed on this box as a native ComfyUI workflow template —
first/last-frame anchoring, reference-image identity, trajectory control, depth/pose control,
video upscale, frame interpolation and frame extraction. What is missing is **weights on disk and
one measured run**, not a tool. [MEASURED-ON-BOX]
2. The dossier's own strip shape rules out plain image-to-video. The exemplar sits *mid-strip* —
three cast frames come before it (docs/ABILITY_EXEMPLAR_DOSSIER_A_E_003.md:153-156). A
first-frame-only I2V model cannot produce those three frames from the exemplar. The route must
anchor the exemplar as a last frame as well as a first. That single structural fact selects
the shortlist.
3. The whole recommended route is Apache-2.0 and costs $0. Wan 2.2 / 2.1 weights are anonymous
public pulls onto a drive with 2936.29 GiB free (2.87 TiB; 3.15 TB decimal). [MEASURED-ON-BOX]
4. The hosted tier is a real fallback and is cheap — the API nodes are already installed, so it
is credits-only — but it carries two blockers a spend gate cannot buy away: the face lane's
standing exclusion on sending Josh's floor sheet to a third party, and the fact that this pilot's
caster is a seven-year-old, against a Runway minors clause whose scope is
product-line-limited and is stated exactly in §2.3.
5. **Two aggregator claims that would have cost us are false, and both were killed by primary
sources fetched today**: LTX-2.3 is *not* Apache-2.0, and no Wan model past 2.2/Animate-2 has open
weights.
---
A survey that starts from what the models do produces a wish list. This one starts from what the
A_E_003 dossier already specifies and asks each tool whether it can consume and emit it.
| # | Input | Canon source |
|---|---|---|
| I1 | ONE graded exemplar frame — full-figure (face AND body AND dress), the signature-read pause: the stone hanging at hip height, the socket open at the feet, the throwing arm uncommitted | docs/ABILITY_EXEMPLAR_DOSSIER_A_E_003.md:66-78; whole-entity floor docs/spine/DECISIONS_PENDING_JOSH.md:4068-4075 |
| I2 | The motion plan — the four beats with authored per-beat content: CAST (earth cracks, grain rises, stone follows, arrives half a beat late), BODY (a seven-year-old's honest overhand throw), FLIGHT (flat, fast, heavy, shedding grit), IMPACT (surface dents/spalls first, stone usually splits, target loses footing), AFTERMATH (fragments scatter, loose earth slides back into the still-open socket) | docs/ABILITY_EXEMPLAR_DOSSIER_A_E_003.md:131-149; the grammar itself docs/D1-27_ELEMENT_VISUAL_LAW.md:26-50; the Earth dispatch beats docs/D1-27_ELEMENT_VISUAL_LAW.md:307-310 |
| I3 | No per-frame text description. Identity, scene, light and palette are carried by the exemplar and are never re-described per frame. Independent per-tile text prompts are BANNED. This does not mean the graph has no text input — see C5, which discloses the one every shortlisted template carries and rules the discipline on it | docs/spine/DECISIONS_PENDING_JOSH.md:4017-4022; restated as the working rule at docs/ABILITY_EXEMPLAR_DOSSIER_A_E_003.md:157-161 |
| # | Output requirement | Canon source |
|---|---|---|
| O1 | 8–12 frames, allocated 3 cast / 1–2 flight / 2 impact / 2 aftermath, with the exemplar itself sitting at the signature pause | docs/ABILITY_EXEMPLAR_DOSSIER_A_E_003.md:153-156 |
| O2 | Frames, not a clip. The deliverable is the boards-law reviewable surface: phone-viewable, one vertical scroll. A video file is an intermediate, not the artifact | docs/ABILITY_EXEMPLAR_DOSSIER_A_E_003.md:155; viewability floor docs/spine/DECISIONS_PENDING_JOSH.md:4223-4230 |
| O3 | Identity holds on every frame against the graded keystone — G1 re-applied to the strip as G7 | docs/ABILITY_EXEMPLAR_DOSSIER_A_E_003.md:178-179, 192-194 |
| O4 | The light law holds on every frame. Earth never emits; it reads by occlusion, by the dust it puts into existing daylight, and by silhouette. The seven banned classes are armed against the output: no aura/halo/corona/rim-light without a source, no ring or disc of light in the air, no arcane geometry, no glowing eyes, no self-lit stone, no particle sparkle, no magic light shaft | docs/D1-27_ELEMENT_VISUAL_LAW.md:78-84 (the seven), :278 (Earth never emits), :111 (the arming, and its asymmetry — see §4 instrument 2); gate G2 docs/ABILITY_EXEMPLAR_DOSSIER_A_E_003.md:180-181 |
| O5 | Source persistence — the open socket stays visible and legible across the whole strip, and the downrange scatter appears and persists. *"A frame where the stone appears at the hand with no socket has failed regardless of quality"* | gate G3 docs/ABILITY_EXEMPLAR_DOSSIER_A_E_003.md:182-184; the universal cast rule docs/D1-27_ELEMENT_VISUAL_LAW.md:32 |
| O6 | Period + world hold — no modern surface anywhere in frame; the setting stays Ch 3 Manggarai | gate G4 docs/ABILITY_EXEMPLAR_DOSSIER_A_E_003.md:185-186 |
time. A first-frame-conditioned I2V model can only run forward from the exemplar, so it can emit
flight/impact/aftermath and structurally cannot emit the cast beat without either (a) anchoring
the exemplar as the last frame of a cast segment, or (b) treating the exemplar as a
*reference* rather than a *frame*. **Plain I2V is therefore not a complete answer to this
pipeline, at any quality.** This is the sharpest requirement in the survey and it was not visible
from the tool side.
particle sparkle, rim light and energy rings when asked for a magic effect. O4 forbids all seven.
The mitigations are structural, not prompt-side: prefer conditioning that is geometric
(first/last frame, reference image, drawn trajectory, depth/pose control) over conditioning that is
semantic (text describing the effect), and run the armed banned-class scan on the OUTPUT frames
rather than trusting the prompt.
to the UE editor, and never co-launches with a 3D-gen run
(docs/pipeline_review/tech_research/PIPE_VIDEO_2026-07-29.md:118-123). A route that needs more
than ~29 GB in one graph is out.
review surface, which would let it sit in the previz tier
(docs/pipeline_review/tech_research/PIPE_VIDEO_2026-07-29.md:19-24). Recommend against. The
dossier makes the graded frames *"the look-target the real VFX pass is judged against"*
(docs/ABILITY_EXEMPLAR_DOSSIER_A_E_003.md:149), and a look-target you cannot legally re-derive
once revenue crosses a ceiling is a dependency, not a reference. Ratify only unrestricted-licence
routes; keep ceilinged routes named and available for previz.
Disclosed rather than buried, because it looks at first like a conflict with I3: all three
shortlisted templates load CLIPLoader + CLIPTextEncode — video_wan2_2_14B_flf2v.json (4
loaders / 8 encodes), video_wan_vace_14B_ref2v.json (4 / 4), video_wan_ati.json (2 / 4) — and
every one of them loads umt5_xxl_fp8_e4m3fn_scaled. **Wan requires a umt5 text conditioning
input**; there is no Wan graph without one, so "no text at all" is not an available configuration
and claiming it would be false. [MEASURED-ON-BOX] **THE RULED DISCIPLINE, which keeps the L4022
ban fully intact: the positive conditioning is ONE shared segment-level motion-plan string per
segment**, derived from the dossier's own four beats
(docs/ABILITY_EXEMPLAR_DOSSIER_A_E_003.md:131-149) — the arm coming round, the stone's flat
heavy line, the surface denting, the fragments and the loose earth sliding back. It is **never
per-frame, and it never re-describes identity, scene, light or palette** — those are carried
by the anchor/reference pixels and by the exemplar alone, which is precisely what L4022 bans
re-describing. Anything a text string could add about the *look* belongs in the negative slot as
the banned-class vocabulary, never in the positive slot as an instruction. The string is authored
once, recorded with the run, and is the same string across both arms so the blind comparison in §4
is a comparison of models rather than of prompts.
The strip cannot be generated until G6 passes — Josh's grade on the exemplar frame, *"the
only gate that moves the acceptable-visuals count 0 → 1. Until it passes, everything stays at concept
tier and no derivation spends GPU"* (docs/ABILITY_EXEMPLAR_DOSSIER_A_E_003.md:190-191). §4 respects
that: the capability measurement runs on a throwaway frame and never touches the exemplar.
G6 is the only CANON dependency; it is not the only dependency. Two production ones sit beside
it and are named here so §4's rung headers are not read as "one gate away": **Arm A additionally
needs the derived, separately-gated pre-cast still** (§3.1's stated cost — it must pass G1-G5 in its
own right before it may anchor a segment), and **Arm B additionally needs a measured, provenance-
recorded VACE quant** that fits 32 GB, since the shipped fp16 does not (§4 Rung 0b). Both are $0 and
neither is blocked by G6, which is exactly why they are scheduled before it.
---
Licence posture key: CLEAN = unrestricted (Apache/MIT/OpenMDW), eligible for a load-bearing
look-target · CEILINGED = commercial-free below a revenue threshold, previz-only per the ruled
posture · TERRITORY = markets excluded, previz-only permanently · HOSTED = vendor ToS governs.
"On box" means the ComfyUI workflow template and its nodes are installed today at
D:/comfyui/.venv-art/Lib/site-packages/comfyui_workflow_templates_json/templates/, on ComfyUI
0.29.0 (D:/comfyui/ComfyUI/comfyui_version.py:3). Video weights are not yet pulled — the
only files in models/diffusion_models/ are flux-2-klein-base-4b, qwen_image_edit_2511_fp8mixed,
z_image_turbo_bf16 and seedvr2_3b_int8_convrot. The scope of that sentence is that one directory:
models/checkpoints/ separately holds ace_step_1.5_turbo_aio.safetensors (the music lane's
all-in-one, not a video model), and no claim is made here about the other model directories.
[MEASURED-ON-BOX]
| Tool | Conditioning mechanism (what holds identity) | Licence | On box? | VRAM fit @ 32 GB | Meets C1? | Evidence |
|---|---|---|---|---|---|---|
Wan 2.2 14B FLF2V — WanFirstLastFrameToVideo on the I2V checkpoints (wan2.2_i2v_high/low_noise_14B_fp8_scaled + umt5_xxl_fp8 + wan_2.1_vae, with the lightx2v_4steps LoRAs) | First AND last frame pixel anchoring. Identity is carried by the anchor frames themselves — no adapter, no text | CLEAN — Apache-2.0 | YES — video_wan2_2_14B_flf2v.json | ~16–22 GB [ESTIMATE, unmeasured] | YES | template contents [MEASURED-ON-BOX]; licence [PRIMARY] https://raw.githubusercontent.com/Wan-Video/Wan2.2/main/LICENSE.txt (verified in PIPE_VIDEO_2026-07-29.md:54-58); node documented [PRIMARY] https://docs.comfy.org/tutorials/video/wan/wan2_2 |
Wan 2.1 VACE-14B ref2v / flf2v — WanVaceToVideo; reference images + masks + optional control video | True reference-image conditioning (context adapter). The exemplar is an identity anchor the model honours throughout, not just an endpoint frame | CLEAN — Apache-2.0 | YES — video_wan_vace_14B_ref2v.json, video_wan_vace_flf2v.json, _v2v, _inpainting, _outpainting | template ships fp16 (wan2.1_vace_14B_fp16) ≈ 28 GB — does not fit with the encoder. GGUF/fp8 community builds exist and are the route | YES | template [MEASURED-ON-BOX]; quant build verified to exist today [PRIMARY] HF API for QuantStack/Wan2.1_14B_VACE-GGUF — repo present, licence tag apache-2.0, siblings include Wan2.1_14B_VACE-Q4_0.gguf, -Q5_K_M.gguf, -BF16.gguf (a community re-quantization, so its provenance is recorded before use per §4 Rung 0b); mechanism family named in the CVPR-2026 literature review of R2V/VACE-class adapters |
Wan 2.1 ATI — WanTrackToVideo (Wan2_1-I2V-ATI-14B_fp8_e4m3fn + clip_vision_h) | Drawn trajectory control. The motion plan literally becomes the input: the stone's ballistic line and the arm's arc are tracks on the exemplar | CLEAN — Wan 2.1 Apache-2.0 lineage | YES — video_wan_ati.json | fp8 14B, ~16–20 GB [ESTIMATE] | forward only | template contents [MEASURED-ON-BOX] |
Wan 2.2 14B Fun Control (wan2.2_fun_control_high/low_noise_14B_fp8_scaled) | Control-video conditioning — pose / depth / canny driving a 2.2-quality generation | CLEAN — Apache-2.0 | YES — video_wan2_2_14B_fun_control.json (+ fun_camera, fun_inpaint, 5B variants) | fp8, ~16–22 GB [ESTIMATE] | with a rendered driver | [MEASURED-ON-BOX] |
| Wan2.2-Animate-2-14B | Reference image + driving video → identity-preserved animation; *"high-fidelity motion generation and strong identity preservation by eliminating intermediate motion extractors"* | CLEAN — Apache-2.0 | not templated | *"default settings… tuned for 8× A800… 480P on 2× A800"* [VENDOR CLAIM] — does not fit as shipped | needs a driver video | [PRIMARY] https://huggingface.co/Wan-AI/Wan2.2-Animate-2-14B ; created 2026-07-14, Diffusers + Distilled variants 2026-08-06 [PRIMARY] HF API |
Wan 2.1 SCAIL-2 character replacement (wan2.1_14B_SCAIL_2_fp16 + sam3.1_multiplex) | Masked character replacement inside an existing video | CLEAN lineage | YES — video_wan21_scail2_character_replacement.json (+ int8) | int8 variant is the fit | n/a (later lane) | [MEASURED-ON-BOX] |
LTX-2.3 22B — video_ltx2_3_flf2v.json, _i2v, _ia2v, _ic_lora, _id_lora; IC-LoRA union-control and depth/canny-to-video on the LTX-2 19B line | FLF2V plus an ID-LoRA (ltx-2.3-id-lora-talkvid-3k) plus IC-LoRA structural control | CEILINGED — "LTX-2 Community License Agreement": *"Entities with annual revenues of at least $10,000,000… are required to obtain a paid commercial use license"*; output ownership clean; mandatory machine-generated disclosure on dissemination | YES (6 LTX-2.3 + 7 LTX-2 templates) | ltx-2.3-22b-distilled-fp8 (~22 GB) plus a gemma_3_12B_it_fp4 text encoder → assume exclusive, may not fit | YES | templates [MEASURED-ON-BOX]; licence [PRIMARY] https://raw.githubusercontent.com/Lightricks/LTX-2/main/LICENSE — and the HF card's own tag reads ltx-2-community-license-agreement, https://huggingface.co/Lightricks/LTX-2.3. The ID-LoRA is trained on talking-head video, which is the wrong domain for a full-figure action strip [THIN EVIDENCE on transfer] |
| HunyuanVideo-1.5 720p I2V | First-frame I2V | TERRITORY — Tencent Community; *"DOES NOT APPLY IN THE EUROPEAN UNION, UNITED KINGDOM AND SOUTH KOREA"* → previz-only permanently | YES — video_hunyuan_video_1.5_720p_i2v.json | 14 GB with offload [VENDOR CLAIM] | no | PIPE_VIDEO_2026-07-29.md:59-68 [PRIMARY, re-read not repeated] |
| Kandinsky-5 Lite I2V · Capybara v0.1 I2V · HuMo 17B · causal-forcing I2V · WanMove · Bernini-R | various I2V / human-centric / streaming | UNVERIFIED | YES (all templated) | unmeasured | unknown | [THIN EVIDENCE] — named so they are a decision rather than a rediscovery; not ranked, no licence read performed |
| SeedVR2 3B/7B int8 | Video/image restoration + upscale; conditioning is the source latent, no text encoder | CLEAN — Apache-2.0 | YES — templates *and already wired in this repo* (harness/asset_factory/seedvr2_refine.py) | 3B int8 on disk today | n/a — finisher | [MEASURED-ON-BOX]; harness/asset_factory/seedvr2_refine.py:1-45 |
FILM frame interpolation (film_net_fp16) · SDPose wholebody video→pose · Depth-Anything-3 video depth | strip densification; driver extraction | CLEAN/unverified | YES — utility_video_frame_interpolation.json, utility_sdpose_ood_video_to_pose_map.json, utility_depth_anything3_video_depth_estimation.json | small | utilities | [MEASURED-ON-BOX] |
| Tool | Mechanism | Licence | On box? | VRAM | Verdict |
|---|---|---|---|---|---|
Cosmos 3 — nvidia/Cosmos3-Nano (16B), Cosmos3-Super (65B), Cosmos3-Edge, and dedicated Cosmos3-Super-Image2Video + -Image2Video-4Step configurations | image-to-video world-foundation model | OpenMDW-1.1 — the cleanest licence in the field, commercial use permitted, no MAU, no territory, no output claim | NO | Nano BF16 >29 GB + offload; but community INT4-AWQ, NVFP4-AWQ, FP8 and NF4 builds of both Nano and Super now exist on HF — the checkable ids, all confirmed present in today's HF model listing: Reza2kn/Cosmos3-Nano-INT4-AWQ, Reza2kn/Cosmos3-Nano-NVFP4-AWQ, Reza2kn/Cosmos3-Nano-FP8, SanDiegoDude/Cosmos3-Nano-nf4, Intel/Cosmos3-Super-int4-AutoRound, plus first-party nvidia/Cosmos3-Super-Image2Video and nvidia/Cosmos3-Super-Image2Video-4Step | WATCH — and the VRAM objection is now weaker than PIPE_VIDEO recorded. The remaining blocker is the ComfyUI gap: issue #14228 "Support for NVIDIA Cosmos 3 model family?" is still OPEN, opened 2026-06-02, no maintainer comments, no branches or PRs — re-verified today [PRIMARY] https://github.com/Comfy-Org/ComfyUI/issues/14228 ; model list [PRIMARY] HF API |
| Cosmos-Predict2 (2B/14B) I2V | image-to-video | NVIDIA Open Model Licence, commercial permitted | native ComfyUI support exists for the Predict2 line | 2B is the fit | EVALUATE at benchmark day only. Training objective is physical-AI/robotics, not aesthetic; PIPE_VIDEO_2026-07-29.md:88-90 records the mismatch and upstream limited-maintenance |
| GEM-X · Kimodo · ARDY | video→SOMA mocap; text→motion; real-time motion | Apache-2.0 code + NVIDIA Open Model weights | — | — | WRONG LANE, and correctly so. These emit skeletal animation data, not pixels. They serve E4/G2, not the strip. docs/MODEL_STACK_AUDIT_AUG2026.md:161-162; PIPE_VIDEO_2026-07-29.md:101-102. Named here so this survey is not re-run against them |
| NVIDIA Lyra 2.0 | — | Internal Scientific Research: *"may not… generate works for sale"* | — | — | EXCLUDED, standing (PIPE_VIDEO_2026-07-29.md:105) |
The material finding: this tier needs no installation. The ComfyUI API-node templates are already
on the box, so the hosted fallback is credits only, with zero setup cost. [MEASURED-ON-BOX]
| Service | Installed template | Conditioning shape | Indicative price | Posture |
|---|---|---|---|---|
| Vidu Q3 | api_vidu_start_end_to_video.json, api_vidu_reference_to_video.json, api_vidu_q3_image_to_video.json | start+end frame AND reference-to-video — the exact two grammars C1 requires | not quoted here | the best-shaped hosted candidate for this pipeline |
| Wan 2.6 / 2.7 (API) | api_wan2_6_i2v.json, api_wan2_7_i2v.json, api_wan2_7_r2v.json, api_wan2_7_video_edit.json | I2V + reference-to-video | not quoted here | the newest Wan quality with the same conditioning grammar as the local route — the natural hosted escalation. Independently confirms 2.6/2.7 are API-only |
| Kling v3 | api_kling_v3_video.json, api_kling_o3_video_edit.json | I2V / edit | ~$0.112/s at 1080p [VENDOR-AGGREGATOR CLAIM] | ToS read owed |
| Runway Gen-4 Turbo / Gen-4.5 / Aleph 2 | api_runway_gen4_turo_image_to_video.json, api_runway_aleph2_video_edit.json | I2V / video edit | ~$0.15/s [AGGREGATOR CLAIM] | see the minor clause below |
| ByteDance Seedance 1.5 | api_bytedance_seedance1_5_image_to_video.json | I2V | — | ToS read owed |
| Luma Ray3.2 · Grok reference-to-video · Hailuo/MiniMax · Google Gemini Omni Flash video edit | templated | various | — | named, not ranked |
Two blockers on this tier that money does not buy away.
1. The standing exclusion. The face lane's own survey already excluded *"every API node class in
/object_info… They fail the box requirement, and using them would send Josh's own floor sheet to a
third party, which is not this lane's call"*
(build/3d/model_survey/face_2026_08_06/EXCLUDED.json). The exemplar frame is that floor
sheet's derivative. This is a director-tier call, not a lane's.
2. The subject is a seven-year-old. The pilot's caster is CHAR_0001 at the age-7 keystone
(docs/ABILITY_EXEMPLAR_DOSSIER_A_E_003.md:73-78). Runway's usage policy restricts *"Characters
based on the face or voice of a person under the age of 18"* [PRIMARY]
https://runway.com/safety/usage-policy — **and the clause's SCOPE has to be stated with it,
because scope is half the finding.** It does not sit in the blanket platform prohibitions; it sits
under the heading "Additional Character & Game Worlds Policies", which the policy itself scopes
to the Runway Characters & Game Worlds product line. It is therefore **not, on its face, a rule
about Gen-4 Turbo image-to-video**, which is the endpoint this survey would actually call
(api_runway_gen4_turo_image_to_video.json). Re-verified today [PRIMARY]. **On top of that
scope question, the honest read is that the clause targets real minors, not a synthetic character,
so what remains is a genuine ambiguity rather than a categorical ban.**
Stating it either way would be a defect: overclaiming it is the care-inflation the standing
calibration forbids; ignoring it risks an account action on the account that holds the pipeline.
The correct handling is §4's recommendation — **measure the hosted capability on a non-protagonist
stand-in subject**, so the question never has to be answered to get the number.
---
WanFirstLastFrameToVideo — THE ROUTE, with a cost stated up frontThe only candidate that is simultaneously CLEAN-licensed, already templated on this box, and
structurally able to satisfy C1: run the cast beat with the exemplar as the *last* frame, and the
flight→impact→aftermath run with it as the *first*, so every frame in the strip is anchored to graded
pixels at at least one end. It reuses the I2V checkpoints (no separate FLF2V weight set), and the
4-step lightx2v LoRAs make iteration cheap.
THE COST, NAMED RATHER THAN BURIED — FLF2V NEEDS A SECOND ANCHOR IT DOES NOT HAVE. A first/last
node takes two frames. The cast segment supplies the exemplar as the *last* one; **there is no
canon artifact for the first one.** A neutral pre-cast still has to be MADE, and the only lawful way
to make it is to derive it from the graded exemplar by image-editing / conditioning — the
qwen_image_edit_2511_fp8mixed route already on disk is the obvious instrument — because
text-prompting a fresh still from scratch is exactly the per-tile-text architecture L4022 bans, and
a second independently-generated still would put a second, ungraded identity into the strip's own
anchor set. **That derived still is itself an artifact that must pass G1-G5 before it may anchor
anything** (identity, light law, source-in-frame, period, legibility), which is a whole extra gated
sub-deliverable and a second place for identity to drift.
**This weakens the FLF2V-first ranking, and the honest statement of it is that VACE needs no second
anchor at all** — a reference-conditioned run takes the exemplar and the beat targets and nothing
else, so Arm B has zero derived-artifact surface where Arm A has one gated one. That is a
stronger argument for VACE than §3.2 previously made, and it is now made here.
FLF2V still leads, for three reasons that survive the cost, and the ranking is held rather than
flipped: (a) its weights are first-party Comfy-Org repackages of Apache-2.0 Wan 2.2, where the
only VACE build that fits 32 GB is a community re-quantization whose provenance is an open item
— an unmeasured third-party artifact is a real cost too, and it sits on Arm B's *only* viable
configuration rather than on an accessory; (b) FLF2V gives pixel-exact endpoint anchoring, which
is a stronger identity guarantee per frame than reference conditioning, and the derived pre-cast
still inherits its pixels from the graded exemplar rather than from a new generation, so the
identity provenance chain is unbroken even though the artifact count is higher; (c) Wan 2.2 is a
generation ahead of Wan 2.1 VACE on quality. **The ranking is explicitly provisional and Rung 0/1
can overturn it:** if the derived pre-cast still fails its own G1-G5, or if the VACE quant measures
clean and its ref2v cast frames hold identity better, Arm B becomes the route — which is why §4
runs both blind rather than running the winner.
When pixel anchoring is not enough — and on a 3-frame cast run away from the anchor it may not be —
VACE conditions on the exemplar as a reference with mask control, which is the mechanism class
built for exactly this problem. Same Apache family, one node, one checkpoint swap. **Its structural
advantage over Arm A is that it needs no second anchor**: the exemplar alone conditions the run, so
the cast beat costs no derived, separately-gated still. Its cost is the mirror of that: the shipped
template is fp16 (≈28 GB, does not fit beside the encoder) and a GGUF/fp8 build has to be sourced
from a community re-quantizer and its provenance recorded — which is why §4 Rung 0b measures it as
its own sub-rung rather than assuming it.
WanTrackToVideo — THE MOTION-PLAN ARMThe closest thing in the open field to *"the motion plan is the input"*: the stone's flat, fast,
heavy ballistic line and the shoulder coming round are drawn as trajectories on the exemplar
rather than described in text — which is the single best mitigation for C2, because a drawn arc
cannot ask for a glow. Third rather than first only because it is Wan 2.1-quality and forward-only,
so it complements the FLF2V route rather than replacing it.
Named behind the shortlist, deliberately: LTX-2.3 FLF2V is the quality rung and it is fully
templated here, but it is CEILINGED (C4) and may not fit beside a 12B text encoder (C3) — keep it as
the previz/quality comparator, not the ratified route. Cosmos 3 has the best licence in the
entire field and now has quantized builds; it is one ComfyUI PR away from jumping to the top and
should be re-checked at every benchmark day. HunyuanVideo-1.5 stays territory-excluded and
previz-only. Wan2.2-Animate-2 is the strongest identity mechanism published this cycle and is
Apache-2.0, but it needs a driving video and does not fit as shipped — its moment comes when the A1
wave has an in-engine or Blender-rendered throw to drive it with, and at that point it is the best
route in this document.
---
The law, applied: a capability is demonstrated on a measured sample before Josh's money moves.
The music lane is the exemplar — three playable samples and a per-dimension diff before the
$1,000 (docs/proposals/music/MEASURED_LOOP_PROOF.md:15, 215-232). Here the analogue is: an 8–12
frame strip with an identity number and a banned-class scan on every frame, produced entirely on
owned hardware with Apache-2.0 weights, before any hosted credit is bought.
wan2.2_i2v_high_noise_14B_fp8_scaled, wan2.2_i2v_low_noise_14B_fp8_scaled, the two lightx2v_4steps LoRAs, umt5_xxl_fp8_e4m3fn_scaled, wan_2.1_vae. Land on the Gen4 path
per the ruled disk layout (PIPE_VIDEO_2026-07-29.md:139-143); 2936.29 GiB free (2.87 TiB /
3.15 TB decimal) [MEASURED-ON-BOX].
video_wan2_2_14B_flf2v.json template on a throwaway conditioning pair —any existing landed plate, explicitly not the exemplar, which does not exist yet (G6).
resolution, whether the two-expert high/low-noise pair serialises inside 32 GB. Land them as a
generation record beside the run, in the shape the 3D lane already uses.
video_wan2_2_5B_ti2v.json (*"the 5B fits 8 GB"* is the vendor's figure, [VENDOR CLAIM], and
this rung is where it stops being one) and record the quality cost honestly.
Rung 0a covers one arm. Arm B cannot be run from what the box ships: video_wan_vace_14B_ref2v.json
loads wan2.1_vace_14B_fp16 at ≈28 GB, which does not fit in 32 GB beside the umt5 encoder and
the VAE. Arm B therefore has a pull-and-measure step of its own, and it is a third-party one,
which changes what has to be recorded.
QuantStack/Wan2.1_14B_VACE-GGUF — verified present today, licence tag apache-2.0, siblings include Wan2.1_14B_VACE-Q4_0.gguf, -Q5_K_M.gguf, -BF16.gguf
[PRIMARY, HF API fetched today]. Plus umt5_xxl_fp8_e4m3fn_scaled and wan_2.1_vae if 0a has not
already landed them.
not a first-party Comfy-Org repackage like Arm A's weights. Record, beside the run: the exact repo
id and revision SHA, the file and its hash, the upstream weights it claims to derive from
(Wan-AI/Wan2.1-VACE-14B), the licence tag as fetched, and the fetch date. A look-target derived
from an unattributed binary is the C4 dependency problem in a different costume.
video_wan_vace_14B_ref2v.json with the GGUF loader substituted for the fp16 diffusion loader,on the same throwaway plate 0a used, and emit the same numbers: peak VRAM, wall clock, frames,
resolution.
fits 32 GB, Arm B is out on a measured axis and the shortlist's §3.2 is corrected in place
rather than carried.
The sample: the A_E_003 Rock Throw exemplar frame → the 8–12 frame strip, exactly the shape
docs/ABILITY_EXEMPLAR_DOSSIER_A_E_003.md:153-156 specifies. Two arms, run blind against the same
frames:
the 3 cast frames). Segment 2: exemplar as FIRST frame → flight, impact, aftermath.
**THE PRE-CAST STILL IS ITS OWN DELIVERABLE, and its derivation route is named rather than
assumed: it is DERIVED FROM THE GRADED EXEMPLAR by image-editing / conditioning** — the
qwen_image_edit_2511_fp8mixed route already on disk, or the same FLF2V/VACE stack run backwards
a beat — and it is never text-prompted from scratch, which is the L4022 ban and would also put
a second ungraded identity into the strip's own anchor set. **It takes its own gate pass — G1
through G5 — before it may anchor anything**, on the same instruments as the strip frames below.
If it cannot pass them, Arm A has failed on a measured axis and Arm B is the route (§3.1).
second anchor and no derived artifact** — this is the structural advantage §3.2 names.
so the blind comparison is between models and not between prompts.
THE ASSEMBLY STEP, corrected. Extract the beat frames deterministically, then assemble them into
the single phone-viewable vertical sheet O2 requires. **harness/asset_factory/contact_sheet.py is
NOT the tool for this** — it takes a .glb per candidate and shells out to Blender to render four
views of each, so it cannot assemble video frames at all. The right pattern is the **2D PIL grid this
repo already uses in two places**: Image.new + ImageDraw with labelled cells, exactly as
harness/asset_factory/age_keystones.py:1235, 1283, 1316 and
harness/asset_factory/objref_pick_sheet.py:56-57 build their sheets, and as
build/3d/characters/body_specs/a18_face_refine.py:156-159 builds its labelled comparison strip.
This is a small piece of NEW TOOLING — a frame-strip assembler of perhaps forty lines against an
established in-repo pattern — and it is declared as such rather than hidden, correcting this plan's
earlier claim that nothing here is new tooling.
**The instruments — two exist in this repo, one is an agent-time critic read, and one required axis
is UNMET:**
1. Identity distance — exemplar_match, and its real implementation and real scope. The metric
is a face copy index measured on HEAD CROPS: build/3d/characters/body_specs/a18_face_refine.py:139
records "exemplar_match": ident["match"] from merge_route_a_v8.measure_head, which returns
RACE.find_face(...)["copy_index"] — a match of the approved inner face over scale and position.
harness/asset_factory/seedvr2_refine.py:14-20 narrates the resulting numbers (0.91889
unrefined, drifting to 0.90621 / 0.88914 under the refine ladder, the drift Josh ruled against)
but does not implement the metric. Report it per frame, not per strip. **Its scope is the
face**, which is stated so it is not over-claimed.
2. The banned-class scan — and D1-27's arming is ASYMMETRIC, so the rule has two branches. The
vocabulary is already armed (docs/D1-27_ELEMENT_VISUAL_LAW.md:111) and the seven classes are at
:78-84. The law's own arming: *"On the dispatch grammar any hit fails, because that surface
already forbids negation and a banned token there can only be an instruction to draw the banned
thing. On a tier read — design prose written for a person, where negation carries its meaning — a
hit fails only when its own sentence is affirmative."* Applied here:
graph): any hit fails, full stop. Negation does not survive the graph, so a banned token in
a prompt is an instruction to draw the banned thing.
for a person — a hit fails only when its own sentence is AFFIRMATIVE. A critic writing *"no
rim-light without a source; the dust reads by occlusion"* has scored a PASS and must not be
counted as a hit; a critic writing *"a faint rim-light sits along the shoulder"* has scored a
FAIL. Scoring the negated sentence as a hit is instrument-misapplication, and it is called out
here because it is the easy error and it would fail correct frames.
3. Socket persistence (O5/G3) — counted by hand across the strip and reported as N-of-N. A strip
that loses the socket has failed regardless of quality.
4. UNMET — the body/dress axis has no instrument. Instrument 1 measures the face and nothing
else, but the governing floor is the whole entity: *face, body, hands, dress all judged*
(G1, docs/ABILITY_EXEMPLAR_DOSSIER_A_E_003.md:178-179; the whole-entity beauty floor,
docs/spine/DECISIONS_PENDING_JOSH.md:4068-4075). **No numeric instrument in this repo measures
body proportion, hand count or garment consistency across a strip.** This is not folded into
instrument 1 and it is not quietly dropped: for this pilot it is carried by the same hand-authored
per-frame critic read as instrument 2, reported as a named prose verdict per frame, and it is
boarded as a real tooling gap for whoever builds the strip lane at volume.
THE COST OF THE CRITIC READ, in the honest unit. Instruments 2, 3 and 4 are **hand-authored
per-frame reads**, on the precedent this repo already set: harness/ability_vfx/judge.py is a
four-axis judge in which *"every one of the 96 rows below was judged by OPENING its four candidates
side by side. 384 of 384 candidates were opened; zero rows are carried as unread."* At 8-12 frames ×
2 arms — plus the derived pre-cast still on Arm A — that is **roughly 17-25 frames opened, read
against four axes, and written up. That is real work, and its currency is AGENT TIME, not
dollars.** The rung is $0 in spend and it is not $0 in effort; a plan that reported only the dollar
figure would be under-reporting what it costs to run.
Arms A and B go to Josh as the strip, per G7
(docs/ABILITY_EXEMPLAR_DOSSIER_A_E_003.md:192-194) and only after G6. **Only his grade moves the
count.** Everything above is evidence for his eye, never a substitute for it.
**No dollars until Rungs 0a, 0b and 1 are measured AND both local arms have failed on a named,
measured axis.** "It looked worse" is not a failure record; "identity fell to X on the cast frames
and the banned-class scan hit affirmatively on Y of Z frames" is.
AND THAT FAILURE RECORD IS NOT SELF-CERTIFIED BY THIS LANE. A lane that both writes the failure
record and spends against it is grading its own homework, and the whole point of PROOF-BEFORE-SPEND
is that someone else looks first. So: **the failure record routes to the DIRECTOR and to Josh, with
the samples attached, and no dollar moves without that sign-off.** This is the MEASURED_LOOP_PROOF
precedent applied exactly as Josh stated it — *"Give me a few sample tracks so I can know you can
actually do this and improve before I go spend $1000 more"*
(docs/proposals/music/MEASURED_LOOP_PROOF.md:15). The samples are seen BEFORE the money moves,
not summarized after. What routes up is the strips themselves, both arms, plus the per-frame numbers
and the named axis of failure — not a verdict.
The $0 floor is STRUCTURAL, not a promise this lane is making. Verified on this box today: there
is no ComfyUI API key file — D:/comfyui/ComfyUI/user/default/comfy.settings.json,
D:/comfyui/ComfyUI/.env and D:/comfyui/ComfyUI/comfy_api_key.json are all absent, and a
recursive search for those filenames anywhere under D:/comfyui/ComfyUI returns nothing — and **no
vendor API environment variable is provisioned**: no RUNWAY*, KLING*, VIDU* or DASHSCOPE*
variable exists in the environment. [MEASURED-ON-BOX] The installed api_* templates are therefore
graph shapes with no credential behind them: **the hosted tier cannot spend a cent until someone
deliberately provisions a key**, which is itself a visible, auditable act. The spend gate below is a
policy on top of a floor that already holds mechanically.
If both arms fail that way and the director and Josh sign off, the hosted probe is bounded at $50
for the entire round. At the
indicative $0.11–0.15/s and a 5-second clip, $50 buys roughly 60–90 clips across the four
best-shaped hosted candidates (Vidu Q3 start-end + reference-to-video, Wan 2.7 R2V, Kling v3, Runway
Gen-4 Turbo) — an order of magnitude more than the 8–12 frames the strip needs, so the bound is
generous rather than tight. Anything past $50 is a new brief, not a continuation.
Two hard preconditions on that spend, both from §2.3:
or a caster-free prop cast. The question the probe answers is *"can a hosted model hold a graded
frame's identity, light law and socket across the four beats?"*, and that question does not need
the child plate to leave the box. It measures the capability at $50 and settles the ToS and
floor-sheet questions to zero exposure. Only if the stand-in probe *succeeds* does the real
question — send the protagonist plate or not — reach the director, and it arrives with evidence
attached instead of as a bare fork.
measures the wrong subject. The pipeline's hardest identity problem is specifically a seven-year-old
child's face and body at full figure, and an adult stand-in that passes proves less than it appears
to. Counter: the four axes the probe is actually testing — bidirectional anchoring, light-law
discipline, socket persistence and beat legibility — are subject-independent, and identity hold is
the one axis where the local route already has the structural advantage (the plate never
leaves). If a hosted model wins on the other three, that is exactly when the protagonist question
is worth putting to Josh.
ComfyUI 0.29.0 is installed and running (harness/asset_factory/keystone_direct.py:106-107 —
HOST = "http://127.0.0.1:8189", COMFY_IN = "D:/comfyui/ComfyUI/input"). Every template this
survey recommends is already on disk. The weights are anonymous public Apache-2.0 pulls (Arm B's is a
community re-quantization and carries the Rung 0b provenance record). The disk has **2936.29 GiB
free** (2.87 TiB / 3.15 TB decimal). The identity instrument exists; the frame-strip assembler is
forty lines of new code against an in-repo pattern; the light-law, socket and body/dress axes are
agent-time critic reads. And the hosted floor is structural — no API key file and no vendor API
environment variable is provisioned on this box. **The entire recommended route through Rung 2 costs
zero dollars and consumes only GPU hours and agent time we already own** — which is why the spend
gate above should be hard rather than nominal.
---
1. A2's next step should be SPLIT — no new tool needs installing; only weights and a measurement.
Stated precisely, because the earlier framing over-reached: the A2 row
(docs/PIPELINE_LEDGER.md:91) describes a survey gap — "run the chartered exemplar-to-motion
tool survey (NVIDIA stack + open field)" — and this document is the artifact that fills it. **A2
was never flagged BLOCKED; the row that carries that flag is A3** ("BLOCKED behind A2"). So
the correction is not that the ledger mis-framed a blocker; it is that A2's single NEXT STEP cell
fuses two steps of very different readiness. Recommend splitting it into (a) measure the route
(unblocked, tonight, Rungs 0a/0b) and (b) run the pilot (blocked on G6, and on the Arm A
pre-cast still). Fused, the cell hides an hour of work that could already be done, and it holds A3
behind it.
2. **docs/MODEL_STACK_AUDIT_AUG2026.md:156 is CONFIRMED CORRECT and its evidence can be
strengthened. "Wan 2.2 14B FLF2V" is real — but there is no separate Wan2.2-FLF2V checkpoint**
on HF (the org's newest FLF2V weights are Wan2.1-FLF2V-14B-720P, 2025-04-17). The 2.2 FLF2V
route is the WanFirstLastFrameToVideo node on the 2.2 I2V checkpoints, which the
installed template proves. Worth writing down before someone hunts for a checkpoint that does not
exist.
3. The Wan-closed-weights finding now has two independent controls. The HF org listing fetched
today returns 27 models ending at Wan2.2-Animate-2-14B-Distilled-Diffusers (2026-08-06) — nothing
past 2.2/Animate-2 — and the box's own installed API templates (api_wan2_6_i2v.json,
api_wan2_7_i2v.json, api_wan2_7_r2v.json) show 2.6/2.7 existing only as remote services. The
"Wan 3.0 open weights, April 2026" claim still circulating on aggregator sites is false on both
controls. The aggregators are named so the counter-claim is checkable rather than a gesture at
"SEO content": they are the same wan27.org / oakgen.ai / spheron pages
PIPE_VIDEO_2026-07-29.md:100 already identified and killed for the Wan-2.5/2.7 version of this
claim — one family of pages, re-badged a version number later, and re-killed on the same two
primary controls.
4. LTX-2.3's licence is being widely mis-reported as Apache-2.0 right now — including on pages
that read as first-party-adjacent. The raw LICENSE fetched today is the **LTX-2 Community License
Agreement** with the $10M ceiling and the machine-generated-disclosure clause, and the HF card's
own metadata tag reads ltx-2-community-license-agreement. PIPE_VIDEO_2026-07-29.md:47-53 was
right; hold that line, and treat any future "LTX is Apache" claim as an aggregator error.
5. Cosmos 3's exclusion should be narrowed, not overturned. PIPE_VIDEO_2026-07-29.md:82-84
excluded it partly on VRAM; community INT4-AWQ / NVFP4 / FP8 / NF4 builds of Nano *and* Super now
exist, and NVIDIA ships dedicated Cosmos3-Super-Image2Video and -4Step configurations. The
ComfyUI ground (#14228, still open, zero maintainer comments, no PRs — re-verified today) is
the one that still holds. Same shape as the TRELLIS.2 narrowing in
MODEL_STACK_AUDIT_AUG2026.md:326-348: a real exclusion that was treated as broader than it is.
6. A route worth boarding rather than building now. Wan2.2-Animate-2 (Apache-2.0, reference image
A1 wave can render a greybox throw to drive it with — the somatic form is already authored
(docs/ABILITY_EXEMPLAR_DOSSIER_A_E_003.md:137-140), and a rendered proxy is a legitimate driver
that costs no licence surface. Boarded, not adopted, and named so it is a decision.
7. Owed and not done by this lane, per the wave override: the docs/DOC_MAP.md registration row
for this document, and its PIPE_VIDEO_2026-07-29.md in-place supersession
(MODEL_STACK_AUDIT_AUG2026.md:487 routes video/motion-strip verdicts to that dossier). Both are
director-landed.
8. Nothing here was surfaced as a bare question. The one genuine fork — whether the protagonist's
graded plate may go to a hosted service — is answered with a recommendation (stand-in probe first,
$50 bound) and its strongest objection, in §4.
---
DERIVED FROM:
docs/spine/DECISIONS_PENDING_JOSH.md § the exemplar-first motion law (L4017-4031) — the charterthis document answers verbatim; the ban on per-tile text prompts (L4022); the identity anchor
(L4023-4024); the named pilot A_E_003 (L4030-4031); the honest register, acceptable = 0 (L4037-4038).
docs/spine/DECISIONS_PENDING_JOSH.md § whole-entity beauty floor (L4068-4075) — why the exemplarand every derived frame are judged full-figure, not as a crop.
docs/spine/DECISIONS_PENDING_JOSH.md § viewability floor (L4223-4230) — why the deliverable is aphone-viewable frame sheet and not a video file (O2).
docs/ABILITY_EXEMPLAR_DOSSIER_A_E_003.md §2 (L64-97) — the exemplar frame spec → I1.§4 (L131-149) — the four-beat VFX breakdown → I2. §5 (L151-170) — the motion plan, the strip's
frame budget and the exemplar's mid-strip position → O1 and the C1 derivation. §6 (L172-199) — the
gate order G1-G8 → O3-O6 and §4's grading hook. §7 (L201-207) — the queued inputs → §4's rungs.
docs/D1-27_ELEMENT_VISUAL_LAW.md § the four-beat grammar (L26-50) — CAST/FLIGHT/IMPACT/AFTERMATHand the source-in-frame universal rule → O5. § the seven banned classes (L78-84) and their
arming (L111), **including the arming's stated ASYMMETRY — any hit fails on dispatch grammar, an
affirmative-only hit fails on a tier read** → O4, the C2 constraint, and the two-branch rule in
§4's instrument 2. § Earth (L278, L307-310) — Earth never emits; the per-beat dispatch strings the
strip must satisfy.
docs/PIPELINE_LEDGER.md A2 row (L91) — the survey gap this document fills, and the A3 row on thesame table carrying the actual BLOCKED flag ("BLOCKED behind A2") → the finding in §5.1.
docs/pipeline_review/tech_research/PIPE_VIDEO_2026-07-29.md — the ruled licence tiering (L19-24)→ C4; the VRAM/scheduling law (L108-123) → C3; the Hunyuan territory read (L59-68); the Cosmos
exclusion (L75-93) narrowed in §5.5; the disk layout (L139-143); the ARDY/Kimodo wrong-lane routing
(L101-102).
docs/MODEL_STACK_AUDIT_AUG2026.md §Category 5b (L149-172) — the two-motion-programs split, theWan-family recommendation and its stated reasoning (L167-172); queue item 5 (L270-275) naming this
survey; the routing table (L487).
docs/proposals/music/MEASURED_LOOP_PROOF.md (L15, L215-232) — the proof-before-spend exemplar,applied as §4's shape.
build/3d/model_survey/face_2026_08_06/EXCLUDED.json — the standing exclusion of every API nodeclass and its stated reason → §2.3 blocker 1 and §4's stand-in recommendation.
exemplar_match, its implementation and its scope → §4's instrument 1 and the UNMET instrument 4. build/3d/characters/body_specs/a18_face_refine.py:139 is where the metric is
actually recorded ("exemplar_match": ident["match"]); build/3d/characters/body_specs/merge_route_a_v8.py:386-408
is measure_head, which returns RACE.find_face(...)["copy_index"] — a **face copy index on a
head crop**, which is why the body/dress axis is named as unmet rather than folded in.
harness/asset_factory/seedvr2_refine.py:1-45 supplies the ComfyUI-on-box method (template
transcribed node-for-node, re-verified against the running server) and narrates the baseline
numbers (0.91889 / 0.90621 / 0.88914) in its docstring; it does not implement the metric.
harness/ability_vfx/judge.py:1-27 — the four-axis hand-authored judge, 384 of 384 candidatesopened, zero rows carried unread → the precedent for §4's per-frame critic read and the basis for
costing it in agent time.
harness/asset_factory/age_keystones.py:1235, 1283, 1316, harness/asset_factory/objref_pick_sheet.py:56-57 and build/3d/characters/body_specs/a18_face_refine.py:156-159 — the established in-repo
Image.new / ImageDraw labelled-grid pattern → the frame-strip assembler §4 declares as new
tooling. harness/asset_factory/contact_sheet.py is explicitly NOT the assembler — it consumes
a .glb per candidate and shells to Blender, so it cannot assemble video frames.
harness/asset_factory/keystone_direct.py:106-107 — the live ComfyUI client seam (port 8189, D:/comfyui/ComfyUI/input) → the $0-first path.
D:/comfyui/ComfyUI/comfyui_version.py:3 (0.29.0);the installed template set at
D:/comfyui/.venv-art/Lib/site-packages/comfyui_workflow_templates_json/templates/; the model
filenames and node classes inside video_wan2_2_14B_flf2v.json, video_wan_vace_14B_ref2v.json,
video_wan_ati.json, video_ltx2_3_flf2v.json, video_ltx2_3_id_lora.json,
video_wan2_2_14B_fun_control.json, video_ltx2_depth_to_video.json,
utility_seedvr2_3b_int8_upscale_video.json, utility_video_frame_interpolation.json,
utility_sdpose_ood_video_to_pose_map.json, and the api_* hosted templates; the
CLIPLoader / CLIPTextEncode / umt5_xxl_fp8_e4m3fn_scaled counts inside the three shortlisted
templates → C5; the contents of D:/comfyui/ComfyUI/models/diffusion_models (and, separately,
models/checkpoints/ace_step_1.5_turbo_aio.safetensors); 2936.29 GiB free on D: (2.87 TiB /
3.15 TB decimal); and the absence of user/default/comfy.settings.json, .env and
comfy_api_key.json anywhere under D:/comfyui/ComfyUI, plus the absence of any RUNWAY* /
KLING* / VIDU* / DASHSCOPE* environment variable → the structural $0 floor in §4's spend gate.
Wan-AI (27 models, newest Wan2.2-Animate-2-14B-Distilled-Diffusers 2026-08-06); the HF model search for Cosmos3, which
returns all seven ids §2.2 now pastes (Reza2kn/Cosmos3-Nano-INT4-AWQ,
Reza2kn/Cosmos3-Nano-NVFP4-AWQ, Reza2kn/Cosmos3-Nano-FP8, SanDiegoDude/Cosmos3-Nano-nf4,
Intel/Cosmos3-Super-int4-AutoRound, nvidia/Cosmos3-Super-Image2Video,
nvidia/Cosmos3-Super-Image2Video-4Step); the HF model API for
QuantStack/Wan2.1_14B_VACE-GGUF (repo present, apache-2.0, GGUF siblings) → Rung 0b;
https://huggingface.co/Wan-AI/Wan2.2-Animate-2-14B ;
https://raw.githubusercontent.com/Lightricks/LTX-2/main/LICENSE ;
https://huggingface.co/Lightricks/LTX-2.3 ; https://docs.comfy.org/tutorials/video/wan/wan2_2 ;
https://github.com/Comfy-Org/ComfyUI/issues/14228 ; https://runway.com/safety/usage-policy —
from which the minors clause was read together with the section heading it sits under,
"Additional Character & Game Worlds Policies", scoped by the policy to the Characters & Game
Worlds product line → §2.3 blocker 2.
NOT DERIVED — authored judgment, and why it had no canon home:
and the exemplar's position within it; nothing in canon says what that implies for a tool's
conditioning grammar, because canon does not name tools. The inference — exemplar-at-index-4 means
first-frame-only I2V cannot produce the cast beat — is mine, and it is the load-bearing one in the
document. If it is wrong, the shortlist is wrong.
a ceilinged model for an internal surface. I recommend the stricter bar because the dossier makes
the strip a durable look-target. This is a judgment against the letter of the existing posture and
is named as such so it can be overruled.
arithmetically from published per-second pricing against the strip's actual frame need, and it is
deliberately generous so that the bound never becomes the reason a measurement was not taken.
subject; the idea of measuring hosted capability on a non-protagonist subject to avoid the question
entirely is mine.
animator with a greybox render from the A1 wave. Both halves are canon-grounded; the pairing is not.
quantization, not a measurement, and §4 Rungs 0a/0b exist specifically to replace them.
not weigh a derived-but-gated anchor artifact against an unmeasured third-party quant. The three
reasons given — first-party weights, pixel-exact anchoring with unbroken identity provenance, and a
generation of quality — are my weighting, the ranking is declared PROVISIONAL against Rungs 0/1,
and the counter-argument (VACE needs no second anchor at all) is stated in the same paragraph so it
can be picked up rather than having to be discovered.
what to do with a model that structurally requires a text encoder. One shared segment-level
motion-plan string, never per-frame and never re-describing identity/scene/light, is my reading of
the narrowest thing that satisfies both the ban and the graph. It is named so it can be tightened
to an empty positive prompt if the director prefers.
frames into a review sheet; the nearest are 3D contact sheets and 2D keystone grids. Declaring a
new forty-line assembler rather than stretching an existing tool is a judgment, and it corrects
this plan's earlier and wrong claim that nothing here is new tooling.
---
*Authored 2026-08-07, combat-schema wave Lane D. PROPOSAL-TIER: it rules nothing, adopts nothing and
spends nothing. Adoption routes through docs/pipeline_review/tech_research/PIPE_VIDEO_2026-07-29.md
as an in-place supersession, per docs/MODEL_STACK_AUDIT_AUG2026.md:479-487. Owed at landing: a
docs/DOC_MAP.md row with a named consumer (ledger row A2).*