PIPE_VOICE_2026-07-29.md

pipelines/PIPE_VOICE_2026-07-29.md

PIPE_VOICE — the TTS / VO lane dossier (5090 build)

Scope: partial-VO (RULED, D-VO-POLICY 2026-07-24) — generated hero/canonical lines only; bulk and

ambient stay subtitle; high-care cultural roles subtitle-or-synthetic. BAKE-TIME only: there is no

runtime TTS lane in canon and this dossier does not create one (ENGINE_OPTIMIZATION_DOCTRINE.md

§2.7.2, U-39). Every model below was checked against a live primary source on 2026-07-29.

---

1. THE VERIFIED STACK

**Verdict up front: the repo's Kokoro-only local ruling is CORRECT on licence and CANNOT serve the

13 cultural-register rows.** Kokoro v1.0 ships 54 voices across 8 languages — American and British

English plus es/fr/hi/it/ja/pt-br/zh — and *zero* non-Western English accents. Our registers are

English-primary with culturally-registered speech (Manggarai, Balinese, Khmer/Cham, Sinhalese/Vedda,

Tamil, Swahili, Ge'ez, San, Baka/Mbuti/Fang, Yoruba/Dahomey, Shona, fairy-realm otherworld). Kokoro

would voice all thirteen in the same two accents — the flattening defect, and the registry itself

forbids it ("distinct from every neighbouring island register", REGISTER_MANGGARAI_FLORES). The new

finding below closes that hole without touching the cloning-consent posture.

CORRECTED 2026-07-30 to the RULED VO policy (GAP-133; docs/spine/DECISIONS_PENDING_JOSH.md

twenty-sixth sitting item 4, Josh verbatim "Row 4 yes c"). The stack table below assigns a model

to a content class, and its "Bulk / ambient / crowd / UI → Kokoro-82M" row contradicted this

dossier's own scope line — which already said bulk and ambient stay subtitle under D-VO-POLICY.

The ruling resolves it as DATA rather than as a vendor split: T0_Dialogue_Line.vo_tier carries

voiced_performance / synthesis_permitted / subtitle_only, and synthesis_permitted is

confined to NAMED NON-STORY classes — barks, ambient crowd, vendor one-liners — never story

dialogue and never a hard-line beat. Read every model row below as the lane that serves a tier,

not as a licence to synthesize a class: a row is voiced because its vo_tier says so.

JobPickLicenceVRAMSource
Hero / canonical recurring (15 nodes)ElevenLabs (cloud), behind the §4.4.1 clone gateSubscription; commercial licence from Starter $6/mon/aelevenlabs.io/pricing — re-verified: "Everything in free, plus – Commercial License" begins at Starter
Cultural-register named NPC — NEW PRIMARYQwen3-TTS-12Hz-1.7B-VoiceDesignApache 2.0~5-6 GBHF card · GitHub — released 2026-01-22
Register variant / dialect timbresQwen3-TTS-12Hz-1.7B-CustomVoice (9 premium timbres, gender/age/language/dialect)Apache 2.0~5-6 GBsame GitHub
Bulk / ambient / crowd / UI — only where vo_tier = synthesis_permitted (named non-story)Kokoro-82MApache 2.0~0-2 GB (CPU-runnable)HF card — v1.0 2025-01-27, 82M, 54 voices / 8 langs
Zero-GPU scratch + temp trackPiperMIT0 (CPU)licence corroborated across sources; MIT
Expressive / cloning-capable (GATED, never default)Chatterbox-Turbo 350M (EN) · Chatterbox Multilingual (23 langs incl. Swahili, Hindi, Malay)MIT2.3 GB / 3.5 GB (NVIDIA-published)HF · Turbo · NVIDIA ACE plugin docs
Bench-only challengersHiggs Audio V2 3B (Apache 2.0, 2025-07) · IndexTTS-2 (Apache 2.0) · CosyVoice 3 (Apache 2.0) · StyleTTS2 (MIT)all permissive3-8 GBsee §6

LICENCE KILLS — verified, do not install into any shippable path:

no commercial licence is purchasable from anyone. Hard kill; the MPL-2.0 *code* licence is the

trap that makes people think it is clear. (discussion)

deployment/product embedding requires separate commercial licensing. **This is the sharpest trap in

the field**: its predecessor Higgs Audio V2 3B *is* Apache 2.0, so "Higgs is Apache" is now false.

(HF card)

(CC-BY-NC-SA) · DramaBox (LTX-2 Community) — all non-commercial output. Excluded.

VibeVoice in commercial or real-world applications… intended for research and development purposes

only,"* and the TTS code was pulled from the repo once after misuse. Licence permits, vendor posture

and supply-chain volatility exclude. EXCLUDED-BY-POSTURE, revisit only if the wording changes.

(github.com/microsoft/VibeVoice)

Watermark note (honest, not a blocker): every Chatterbox output carries Resemble's Perth neural

watermark by default. Imperceptible, survives MP3 — it makes shipped VO detectably Chatterbox-generated.

MIT permits removal; keeping it is arguably a provenance asset. Ship-time call, flagged in §5.

THE USER'S SEARCH-AI LEADS — adjudicated:

(github.com/ace-step/ACE-Step-1.5). **Music lane, not

this one** — hand to PIPE_MUSIC. The repo's benchmark-gated status stands.

models, not audio. ARDY = autoregressive diffusion interactive human motion, SIGGRAPH 2026, inference

code Apache 2.0 / weights NVIDIA Open Model License (commercial permitted)

(nv-tlabs/ardy); Kimodo = kinematic motion diffusion, 700h

commercially-friendly mocap (nv-tlabs/kimodo). Route to the

animation lane — they are genuinely valuable there.

or ComfyUI node by any of these names exists in any source found. Two plausible mis-hearings, neither

adoptable as named: *SDXL Base / SDXL Turbo* (image lane) or the generic Bass/Treble utility nodes in

niknah/audio-general-ComfyUI. Do not adopt an unverifiable name.

---

2. INSTALL PLAN

Environment lane (per runbook D-1): this lane rides option (b), ComfyUI-on-Windows — no WSL2

needed, all four picks are pure-PyTorch Windows-native. The integration vehicle is

diodiogod/TTS-Audio-Suite v5.6.0 (MIT, ComfyUI custom nodes, 19 engines incl. Chatterbox

classic+multilingual, Qwen3-TTS, IndexTTS-2, Higgs 2/3, CosyVoice3; explicitly "Windows: no additional

system dependencies") — one node graph, one venv, the same ComfyUI the image lane already installs at

Stage 6b. Kokoro + Piper get their own tiny CPU-only venv (.venv-tts-cpu) so bulk/ambient batches

never touch the GPU and can run concurrently with UE.

Disk placement (per the RULED layout): Gen4 4TB — D:\ComfyUI\models\TTS\ (all weights),

D:\tts\out\ (working WAVs). Gen5 stays UE hot path + DDC, no model weights. 8TB SATA —

E:\archive\vo_masters\ (24-bit masters) and E:\archive\vo_reference\ (any consented reference clips,

which never leave that drive).

Install ORDER (audio lane rides its own scheduled window — Stage 6c defers it off day one, correctly):

1. .venv-tts-cpupip install kokoro soundfile piper-tts — smoke one line, CPU, zero GPU. (10 min)

2. ComfyUI Manager → install TTS-Audio-Suite → run its installer script (Python 3.12+). (20 min)

3. Qwen3-TTS weights + tokenizer → smoke VoiceDesign on one register description. (20 min)

4. Chatterbox Turbo + Multilingual → smoke, gate-locked: config flag clone_enabled=false by

default so no unattended job can reach cloning. (15 min)

5. Bench challengers only on benchmark day, never at install.

Thursday-night download list (~23 GB total):

ItemGB
Kokoro-82M (weights ~327 MB + voices)0.4
Piper voices ×5 (en_US, en_GB)0.3
Qwen3-TTS-12Hz-1.7B-VoiceDesign (BF16)~4.5
Qwen3-TTS-12Hz-1.7B-CustomVoice~4.5
Qwen3-TTS-12Hz-0.6B-Base (fast baseline)~2.5
Qwen3-TTS-Tokenizer-12Hz~0.2
Chatterbox-Turbo (350M) + Multilingual (500M)~1.8
Higgs Audio V2 3B (bench only)~6.5
IndexTTS-2 (bench only)~2.5

CONCURRENCY LAW — this lane's budget and class:

Sub-laneVRAMRAMClass
Kokoro + Piper (bulk/ambient/UI/scratch)0 GB (CPU)~4 GBALWAYS-ON — safe beside UE and every other lane
Qwen3-TTS 1.7B (cultural registers)~5-6 GB~8 GBWINDOWED — yields to UE cook / 3D-gen
Chatterbox Turbo / Multilingual2.3 / 3.5 GB (NVIDIA-published)~6 GBWINDOWED, and attended (clone-capable)
Full VO foundry batch (queue drain)≤8 GB~16 GBOVERNIGHT

The whole lane fits under 8 GB VRAM. It is the *cheapest* GPU tenant on the machine and should be

scheduled to fill gaps around the 3D/image lanes, never to contend with them.

---

3. THE INTEGRATION CONTRACT

Formats. Masters: 48 kHz / 24-bit mono WAV → E:\archive\vo_masters\. UE import: 48 kHz /

16-bit mono PCM WAV (matches the 48 kHz SFX standard already in the audio stack). ElevenLabs returns

44.1 kHz — resample to 48 kHz at ingest, never at import.

DR-2 addressing (already contracted). T0_Voice_Registry.voice_id → ue_sound_asset_path, final

drop /Game/Audio/Voice/<cat>/V_<voice_id>, swap = path takeover, tooth = cue-fired telemetry +

QT-9 §17.12 gender-lock. Five-state machine: pending → in_progress (prompt hash set) → complete

(path + timestamp) with rejected / regeneration_required branches. VO regeneration_trigger set:

line-text change, voice_substrate change, bound-register change, or model swap.

Two schema asks this lane needs (both ride the DR-2 column commit, before the freeze):

1. generating_model + model_licence on every generated voice row — LICENCE LAW: the generating

model's licence is named in every artifact. The Higgs V2-Apache → TTS-3-non-commercial flip is the

proof this must be per-row, not per-doc.

2. The per-line join key, which T99_Translation_Audio Voice Step 5 flags as a genuine GAP ("the

join key itself is not yet specified anywhere"). Proposed interim so the lane is not blocked:

V_<voice_id>__<dialogue_line_id>, falling back to V_<voice_id>__<sha1(line_text)[:12]> until

Quest Writing §6 exists. Content-addressed means a text edit regenerates exactly one file.

QA gate that judges outputs — three teeth, one new:

never a stored claim).

EXCLUDED from every TTS lane (D-GRAND-SAGE-VOICE, 2026-07-23c: a non-lexical sensation asset; no TTS

vendor may synth it). The dispatcher must read the whitelist before it reads the row.

shippable allowlist (Apache-2.0, MIT, ElevenLabs-subscription). This is what stops a benchmark-day

experiment leaking into a shipped path.

---

4. THE AUTONOMY CONTRACT

Dispatch. Director selects T0_Voice_Registry rows where row_class=tts_casting AND

generation_status ∈ {pending, regeneration_required} AND dialogue_dataset_ref resolves; joins the

row's voice_register_ref → its cultural_register row; composes the prompt strictly from

voice_substrate + the register's speech_shape/translation_rule (never freehand, §2.5); hashes;

routes by lane; writes back through the T0.13 manifest.

🛑 BLOCKER — the dispatch key does not exist on disk. `registries/T0_Voice_Registry [DRAFT v0.1]/

T0_Voice_Registry.csv has a 20-column header carrying row_class` TWICE; the 13 register rows are

19 fields with cultural_register landing under the *first* row_class; and **all 10 tts_casting rows

are 18 fields with no row_class value at all** — the recent_changes note claiming "the TTS casting

rows … now carry row_class=tts_casting" is false against the file. Any DictReader-driven dispatch

silently collapses the duplicate and reads every row as classless. Fix the header and populate the 10

rows before the first autonomous voice run. (Also stale: BUILD_PLAN_END_TO_END.md:190/:256 still lists

GRAND_SAGE_REVEAL_VOICE as needing a retire-vs-rescope ruling — it was RULED 2026-07-23c. Director fix,

no Josh needed.)

Promote / demote gate for models. A model enters a lane only after (1) a licence RE-READ that day,

(2) a blind A/B against that lane's reference bar read by the audio critic, (3) zero findings from the

accent-care lens across the 13-register sample set. It is demoted immediately on any licence change,

any care finding, or a QA regression — and demotion re-flags every row it generated to

regeneration_required, which is why generating_model must be per-row.

Unattended vs attended.

construction, zero VRAM, zero consent surface.

Generation runs alone; rows do NOT flip to complete until the care review passes.

The §4.4.1 five-stage authorization gate is a human/director seam and stays one.

CULTURAL CARE — the no-accent-mockery class, stated as a rule not a worry. Accents are never picked

from an ethnicity menu. Only three vectors are permitted: (a) **neutral synthetic delivery + culturally

registered word choice — the register row does the cultural work, the timbre does not; (b) Qwen3

VoiceDesign from a natural-language description whose every clause cites a register field** — and note

the structural property: VoiceDesign generates from *description*, not a reference clip, so like Kokoro

it is incapable of cloning a real person, which is exactly why it can replace Kokoro here without

reopening the consent question; (c) consented community voice talent, cloned only through §4.4.1.

The care review runs as an adversarial critic listening for four defects, before the row flips:

1. Caricature markers — the obvious one.

2. Flattening — two distinct registers rendering as one voice. Measurable: speaker-embedding cosine

distance between register outputs must exceed the distance between two Kokoro presets.

3. Villain-accent correlation — no antagonist role systematically carrying the non-Western register.

This is CVD §17.1's *collective* protection expressed in audio.

4. Care-tier inflation — the timid failure. Under partial-VO, if the 15 voiced hero nodes skew

Western while the culturally-registered chapters stay subtitle-only, the VO *map* carries a bias the

strings do not. Hero-lane VO coverage is checked against the register map, not just the node list.

Silencing a culture is a defect, never the safe option (care-rulings-delegated-bolder-calibration).

---

5. JOSH-MINIMUM (genuinely his hands or account)

1. ElevenLabs account + payment. Commercial licence starts at Starter $6/mo; the 15-node hero budget

likely wants Creator ($11) or Pro ($99). I cannot enter payment details — he creates and pays, then

the API key goes in the env the same way every other credential does.

2. Any voice-talent consent record. If a real person is ever cloned, the §4.4.1 authorization is

Josh's signature, held by Josh — the loop generates, it never authorizes.

3. Ship-time watermark call (small): keep Resemble's Perth watermark in shipped Chatterbox VO as a

provenance asset, or strip it (MIT permits). Recommendation: keep — it costs nothing and it is a

defence if AI-VO provenance is ever questioned.

4. Nothing else. Model selection, register casting, care rulings, and the Kokoro→Qwen3 lane change

are all execution-within-locked-vision and are applied with the path shown, per the cardinal rule.

---

6. BENCHMARK-DAY QUESTIONS (measured before final picks)

1. Kokoro throughput at roster scale — the translation doc's own UNCONFIRMED flag. Lines/min on CPU

vs GPU, and at what concurrency it still leaves UE alone.

2. Does VoiceDesign actually produce 13 distinct, non-caricatured voices from the 13 register

descriptions? Run the flattening metric (§4 defect 2) plus the critic listen. **This is the question

that decides whether the lane change is real.**

3. Chatterbox Multilingual on Swahili / Hindi / Malay vs English+register-word-choice — does actual

target-language delivery beat registered English, or does it read as costume?

4. Blind A/B on the same hero line: ElevenLabs v2.5/v3 vs Qwen3-1.7B vs Chatterbox Turbo. How much

of the hero budget can move local, and is the answer different for the Prologue/Ch-1 ceiling lines

(the King, the mother, the best friend)?

5. Peak VRAM measured with the UE editor resident — replaces the NVIDIA planning figures (2.3/3.5 GB)

with our own numbers, per the runbook's measure-on-day-one rule.

6. Perth watermark survival through our UE compression settings, and whether it colours the audio.

7. Licence re-read on the day for Kokoro, Qwen3-TTS, Chatterbox. Higgs is the proof this is not

ceremony.

8. One line, whole chain: generate → 48 kHz WAV → Tools/import_audio_assets.py → DR-2 writeback →

resolver returns non-placeholder → MetaHuman Animator audio-driven facial. Prove the round trip on a

single line before any batch is queued.

Generated by harness/site/structure_site.py — the URL path is the repo path. review root