pipelines/PIPE_VOICE_2026-07-29.md
Scope: partial-VO (RULED, D-VO-POLICY 2026-07-24) — generated hero/canonical lines only; bulk and
ambient stay subtitle; high-care cultural roles subtitle-or-synthetic. BAKE-TIME only: there is no
runtime TTS lane in canon and this dossier does not create one (ENGINE_OPTIMIZATION_DOCTRINE.md
§2.7.2, U-39). Every model below was checked against a live primary source on 2026-07-29.
---
**Verdict up front: the repo's Kokoro-only local ruling is CORRECT on licence and CANNOT serve the
13 cultural-register rows.** Kokoro v1.0 ships 54 voices across 8 languages — American and British
English plus es/fr/hi/it/ja/pt-br/zh — and *zero* non-Western English accents. Our registers are
English-primary with culturally-registered speech (Manggarai, Balinese, Khmer/Cham, Sinhalese/Vedda,
Tamil, Swahili, Ge'ez, San, Baka/Mbuti/Fang, Yoruba/Dahomey, Shona, fairy-realm otherworld). Kokoro
would voice all thirteen in the same two accents — the flattening defect, and the registry itself
forbids it ("distinct from every neighbouring island register", REGISTER_MANGGARAI_FLORES). The new
finding below closes that hole without touching the cloning-consent posture.
CORRECTED 2026-07-30 to the RULED VO policy (GAP-133; docs/spine/DECISIONS_PENDING_JOSH.md
twenty-sixth sitting item 4, Josh verbatim "Row 4 yes c"). The stack table below assigns a model
to a content class, and its "Bulk / ambient / crowd / UI → Kokoro-82M" row contradicted this
dossier's own scope line — which already said bulk and ambient stay subtitle under D-VO-POLICY.
The ruling resolves it as DATA rather than as a vendor split: T0_Dialogue_Line.vo_tier carries
voiced_performance / synthesis_permitted / subtitle_only, and synthesis_permitted is
confined to NAMED NON-STORY classes — barks, ambient crowd, vendor one-liners — never story
dialogue and never a hard-line beat. Read every model row below as the lane that serves a tier,
not as a licence to synthesize a class: a row is voiced because its vo_tier says so.
| Job | Pick | Licence | VRAM | Source |
|---|---|---|---|---|
| Hero / canonical recurring (15 nodes) | ElevenLabs (cloud), behind the §4.4.1 clone gate | Subscription; commercial licence from Starter $6/mo | n/a | elevenlabs.io/pricing — re-verified: "Everything in free, plus – Commercial License" begins at Starter |
| Cultural-register named NPC — NEW PRIMARY | Qwen3-TTS-12Hz-1.7B-VoiceDesign | Apache 2.0 | ~5-6 GB | HF card · GitHub — released 2026-01-22 |
| Register variant / dialect timbres | Qwen3-TTS-12Hz-1.7B-CustomVoice (9 premium timbres, gender/age/language/dialect) | Apache 2.0 | ~5-6 GB | same GitHub |
Bulk / ambient / crowd / UI — only where vo_tier = synthesis_permitted (named non-story) | Kokoro-82M | Apache 2.0 | ~0-2 GB (CPU-runnable) | HF card — v1.0 2025-01-27, 82M, 54 voices / 8 langs |
| Zero-GPU scratch + temp track | Piper | MIT | 0 (CPU) | licence corroborated across sources; MIT |
| Expressive / cloning-capable (GATED, never default) | Chatterbox-Turbo 350M (EN) · Chatterbox Multilingual (23 langs incl. Swahili, Hindi, Malay) | MIT | 2.3 GB / 3.5 GB (NVIDIA-published) | HF · Turbo · NVIDIA ACE plugin docs |
| Bench-only challengers | Higgs Audio V2 3B (Apache 2.0, 2025-07) · IndexTTS-2 (Apache 2.0) · CosyVoice 3 (Apache 2.0) · StyleTTS2 (MIT) | all permissive | 3-8 GB | see §6 |
LICENCE KILLS — verified, do not install into any shippable path:
no commercial licence is purchasable from anyone. Hard kill; the MPL-2.0 *code* licence is the
trap that makes people think it is clear. (discussion)
deployment/product embedding requires separate commercial licensing. **This is the sharpest trap in
the field**: its predecessor Higgs Audio V2 3B *is* Apache 2.0, so "Higgs is Apache" is now false.
(HF card)
(CC-BY-NC-SA) · DramaBox (LTX-2 Community) — all non-commercial output. Excluded.
VibeVoice in commercial or real-world applications… intended for research and development purposes
only,"* and the TTS code was pulled from the repo once after misuse. Licence permits, vendor posture
and supply-chain volatility exclude. EXCLUDED-BY-POSTURE, revisit only if the wording changes.
(github.com/microsoft/VibeVoice)
Watermark note (honest, not a blocker): every Chatterbox output carries Resemble's Perth neural
watermark by default. Imperceptible, survives MP3 — it makes shipped VO detectably Chatterbox-generated.
MIT permits removal; keeping it is arguably a provenance asset. Ship-time call, flagged in §5.
THE USER'S SEARCH-AI LEADS — adjudicated:
(github.com/ace-step/ACE-Step-1.5). **Music lane, not
this one** — hand to PIPE_MUSIC. The repo's benchmark-gated status stands.
models, not audio. ARDY = autoregressive diffusion interactive human motion, SIGGRAPH 2026, inference
code Apache 2.0 / weights NVIDIA Open Model License (commercial permitted)
(nv-tlabs/ardy); Kimodo = kinematic motion diffusion, 700h
commercially-friendly mocap (nv-tlabs/kimodo). Route to the
animation lane — they are genuinely valuable there.
or ComfyUI node by any of these names exists in any source found. Two plausible mis-hearings, neither
adoptable as named: *SDXL Base / SDXL Turbo* (image lane) or the generic Bass/Treble utility nodes in
niknah/audio-general-ComfyUI. Do not adopt an unverifiable name.
---
Environment lane (per runbook D-1): this lane rides option (b), ComfyUI-on-Windows — no WSL2
needed, all four picks are pure-PyTorch Windows-native. The integration vehicle is
diodiogod/TTS-Audio-Suite v5.6.0 (MIT, ComfyUI custom nodes, 19 engines incl. Chatterbox
classic+multilingual, Qwen3-TTS, IndexTTS-2, Higgs 2/3, CosyVoice3; explicitly "Windows: no additional
system dependencies") — one node graph, one venv, the same ComfyUI the image lane already installs at
Stage 6b. Kokoro + Piper get their own tiny CPU-only venv (.venv-tts-cpu) so bulk/ambient batches
never touch the GPU and can run concurrently with UE.
Disk placement (per the RULED layout): Gen4 4TB — D:\ComfyUI\models\TTS\ (all weights),
D:\tts\out\ (working WAVs). Gen5 stays UE hot path + DDC, no model weights. 8TB SATA —
E:\archive\vo_masters\ (24-bit masters) and E:\archive\vo_reference\ (any consented reference clips,
which never leave that drive).
Install ORDER (audio lane rides its own scheduled window — Stage 6c defers it off day one, correctly):
1. .venv-tts-cpu → pip install kokoro soundfile piper-tts — smoke one line, CPU, zero GPU. (10 min)
2. ComfyUI Manager → install TTS-Audio-Suite → run its installer script (Python 3.12+). (20 min)
3. Qwen3-TTS weights + tokenizer → smoke VoiceDesign on one register description. (20 min)
4. Chatterbox Turbo + Multilingual → smoke, gate-locked: config flag clone_enabled=false by
default so no unattended job can reach cloning. (15 min)
5. Bench challengers only on benchmark day, never at install.
Thursday-night download list (~23 GB total):
| Item | GB |
|---|---|
| Kokoro-82M (weights ~327 MB + voices) | 0.4 |
| Piper voices ×5 (en_US, en_GB) | 0.3 |
| Qwen3-TTS-12Hz-1.7B-VoiceDesign (BF16) | ~4.5 |
| Qwen3-TTS-12Hz-1.7B-CustomVoice | ~4.5 |
| Qwen3-TTS-12Hz-0.6B-Base (fast baseline) | ~2.5 |
| Qwen3-TTS-Tokenizer-12Hz | ~0.2 |
| Chatterbox-Turbo (350M) + Multilingual (500M) | ~1.8 |
| Higgs Audio V2 3B (bench only) | ~6.5 |
| IndexTTS-2 (bench only) | ~2.5 |
CONCURRENCY LAW — this lane's budget and class:
| Sub-lane | VRAM | RAM | Class |
|---|---|---|---|
| Kokoro + Piper (bulk/ambient/UI/scratch) | 0 GB (CPU) | ~4 GB | ALWAYS-ON — safe beside UE and every other lane |
| Qwen3-TTS 1.7B (cultural registers) | ~5-6 GB | ~8 GB | WINDOWED — yields to UE cook / 3D-gen |
| Chatterbox Turbo / Multilingual | 2.3 / 3.5 GB (NVIDIA-published) | ~6 GB | WINDOWED, and attended (clone-capable) |
| Full VO foundry batch (queue drain) | ≤8 GB | ~16 GB | OVERNIGHT |
The whole lane fits under 8 GB VRAM. It is the *cheapest* GPU tenant on the machine and should be
scheduled to fill gaps around the 3D/image lanes, never to contend with them.
---
Formats. Masters: 48 kHz / 24-bit mono WAV → E:\archive\vo_masters\. UE import: 48 kHz /
16-bit mono PCM WAV (matches the 48 kHz SFX standard already in the audio stack). ElevenLabs returns
44.1 kHz — resample to 48 kHz at ingest, never at import.
DR-2 addressing (already contracted). T0_Voice_Registry.voice_id → ue_sound_asset_path, final
drop /Game/Audio/Voice/<cat>/V_<voice_id>, swap = path takeover, tooth = cue-fired telemetry +
QT-9 §17.12 gender-lock. Five-state machine: pending → in_progress (prompt hash set) → complete
(path + timestamp) with rejected / regeneration_required branches. VO regeneration_trigger set:
line-text change, voice_substrate change, bound-register change, or model swap.
Two schema asks this lane needs (both ride the DR-2 column commit, before the freeze):
1. generating_model + model_licence on every generated voice row — LICENCE LAW: the generating
model's licence is named in every artifact. The Higgs V2-Apache → TTS-3-non-commercial flip is the
proof this must be per-row, not per-doc.
2. The per-line join key, which T99_Translation_Audio Voice Step 5 flags as a genuine GAP ("the
join key itself is not yet specified anywhere"). Proposed interim so the lane is not blocked:
V_<voice_id>__<dialogue_line_id>, falling back to V_<voice_id>__<sha1(line_text)[:12]> until
Quest Writing §6 exists. Content-addressed means a text edit regenerates exactly one file.
QA gate that judges outputs — three teeth, one new:
never a stored claim).
grand_sage_silence — GRAND_SAGE_REVEAL_VOICE isEXCLUDED from every TTS lane (D-GRAND-SAGE-VOICE, 2026-07-23c: a non-lexical sensation asset; no TTS
vendor may synth it). The dispatcher must read the whitelist before it reads the row.
model_licence is not on theshippable allowlist (Apache-2.0, MIT, ElevenLabs-subscription). This is what stops a benchmark-day
experiment leaking into a shipped path.
---
Dispatch. Director selects T0_Voice_Registry rows where row_class=tts_casting AND
generation_status ∈ {pending, regeneration_required} AND dialogue_dataset_ref resolves; joins the
row's voice_register_ref → its cultural_register row; composes the prompt strictly from
voice_substrate + the register's speech_shape/translation_rule (never freehand, §2.5); hashes;
routes by lane; writes back through the T0.13 manifest.
🛑 BLOCKER — the dispatch key does not exist on disk. `registries/T0_Voice_Registry [DRAFT v0.1]/
T0_Voice_Registry.csv has a 20-column header carrying row_class` TWICE; the 13 register rows are
19 fields with cultural_register landing under the *first* row_class; and **all 10 tts_casting rows
are 18 fields with no row_class value at all** — the recent_changes note claiming "the TTS casting
rows … now carry row_class=tts_casting" is false against the file. Any DictReader-driven dispatch
silently collapses the duplicate and reads every row as classless. Fix the header and populate the 10
rows before the first autonomous voice run. (Also stale: BUILD_PLAN_END_TO_END.md:190/:256 still lists
GRAND_SAGE_REVEAL_VOICE as needing a retire-vs-rescope ruling — it was RULED 2026-07-23c. Director fix,
no Josh needed.)
Promote / demote gate for models. A model enters a lane only after (1) a licence RE-READ that day,
(2) a blind A/B against that lane's reference bar read by the audio critic, (3) zero findings from the
accent-care lens across the 13-register sample set. It is demoted immediately on any licence change,
any care finding, or a QA regression — and demotion re-flags every row it generated to
regeneration_required, which is why generating_model must be per-row.
Unattended vs attended.
construction, zero VRAM, zero consent surface.
Generation runs alone; rows do NOT flip to complete until the care review passes.
The §4.4.1 five-stage authorization gate is a human/director seam and stays one.
CULTURAL CARE — the no-accent-mockery class, stated as a rule not a worry. Accents are never picked
from an ethnicity menu. Only three vectors are permitted: (a) **neutral synthetic delivery + culturally
registered word choice — the register row does the cultural work, the timbre does not; (b) Qwen3
VoiceDesign from a natural-language description whose every clause cites a register field** — and note
the structural property: VoiceDesign generates from *description*, not a reference clip, so like Kokoro
it is incapable of cloning a real person, which is exactly why it can replace Kokoro here without
reopening the consent question; (c) consented community voice talent, cloned only through §4.4.1.
The care review runs as an adversarial critic listening for four defects, before the row flips:
1. Caricature markers — the obvious one.
2. Flattening — two distinct registers rendering as one voice. Measurable: speaker-embedding cosine
distance between register outputs must exceed the distance between two Kokoro presets.
3. Villain-accent correlation — no antagonist role systematically carrying the non-Western register.
This is CVD §17.1's *collective* protection expressed in audio.
4. Care-tier inflation — the timid failure. Under partial-VO, if the 15 voiced hero nodes skew
Western while the culturally-registered chapters stay subtitle-only, the VO *map* carries a bias the
strings do not. Hero-lane VO coverage is checked against the register map, not just the node list.
Silencing a culture is a defect, never the safe option (care-rulings-delegated-bolder-calibration).
---
1. ElevenLabs account + payment. Commercial licence starts at Starter $6/mo; the 15-node hero budget
likely wants Creator ($11) or Pro ($99). I cannot enter payment details — he creates and pays, then
the API key goes in the env the same way every other credential does.
2. Any voice-talent consent record. If a real person is ever cloned, the §4.4.1 authorization is
Josh's signature, held by Josh — the loop generates, it never authorizes.
3. Ship-time watermark call (small): keep Resemble's Perth watermark in shipped Chatterbox VO as a
provenance asset, or strip it (MIT permits). Recommendation: keep — it costs nothing and it is a
defence if AI-VO provenance is ever questioned.
4. Nothing else. Model selection, register casting, care rulings, and the Kokoro→Qwen3 lane change
are all execution-within-locked-vision and are applied with the path shown, per the cardinal rule.
---
1. Kokoro throughput at roster scale — the translation doc's own UNCONFIRMED flag. Lines/min on CPU
vs GPU, and at what concurrency it still leaves UE alone.
2. Does VoiceDesign actually produce 13 distinct, non-caricatured voices from the 13 register
descriptions? Run the flattening metric (§4 defect 2) plus the critic listen. **This is the question
that decides whether the lane change is real.**
3. Chatterbox Multilingual on Swahili / Hindi / Malay vs English+register-word-choice — does actual
target-language delivery beat registered English, or does it read as costume?
4. Blind A/B on the same hero line: ElevenLabs v2.5/v3 vs Qwen3-1.7B vs Chatterbox Turbo. How much
of the hero budget can move local, and is the answer different for the Prologue/Ch-1 ceiling lines
(the King, the mother, the best friend)?
5. Peak VRAM measured with the UE editor resident — replaces the NVIDIA planning figures (2.3/3.5 GB)
with our own numbers, per the runbook's measure-on-day-one rule.
6. Perth watermark survival through our UE compression settings, and whether it colours the audio.
7. Licence re-read on the day for Kokoro, Qwen3-TTS, Chatterbox. Higgs is the proof this is not
ceremony.
8. One line, whole chain: generate → 48 kHz WAV → Tools/import_audio_assets.py → DR-2 writeback →
resolver returns non-placeholder → MetaHuman Animator audio-driven facial. Prove the round trip on a
single line before any batch is queued.