music/PIPE_AUDIO_MUSIC_2026-07-29.md
Verified live 2026-07-29. VERIFIED = primary source fetched this pass (URL given). CARRIED = from
docs/pipeline_review/tech_research/AUDIO_STACK.md (2026-07-15), not re-fetched. UNVERIFIED-EXCLUDED =
could not be confirmed; never adopted. Ruled state honored: AIVA CLOSED-AS-SKIP, ACE-Step/YuE
benchmark-gated, HeartMuLa candidate, ~~composer-in-loop for canonical themes~~ **[SUPERSEDED
2026-08-05 — see the supersession block below]**, and the sourcing law
no sampling of consecrated performance / reference-COMPOSITION permitted (DESIGN_GAP_REGISTER:1317,
1789; adjudication 18).
"Humans aren't composing the music… what the fuck. The entire music pipeline is yours."
And the correction that re-grounds the lane on this dossier: *"We already developed the music
pipeline… did you forget all the pipelines we created?"*
COMPOSER-IN-LOOP IS RETIRED. The factory owns the score end to end — melodies, hero themes,
leitmotifs, beds, stingers, orchestration, notation. Exactly one row of this dossier changes
(§1's Canonical themes (12) = composer-in-loop, annotated in place below). Everything else here
stands unchanged and binding: the 30-year bar and the hummability test, the per-culture /
per-period scoring law with its wordless-vocal default, the notation-and-concert deliverable, the
sourcing law, the Sonniss no-training/no-derivation constraint, the MusicGen/AudioCraft exclusion,
the VRAM concurrency law, and the DR-2 integration contract.
What replaces the human in the loop is not "nothing" — it is the measurement loop. The bar was
previously held by a composer's ear; it is now held by the factory's own candidate ladder,
instrumented scoring against the doctrine's §2.3 checklist, and fresh-context critics on captured
playback. A generated theme that cannot show its score trail has not passed anything. The honest
tier for everything this lane emits is GENERATED SCORE — ITERATION, never "final", and Josh
hears results on the chapter site's audio players and vetoes by exception.
The one carve-out that survives the retirement, because it is a care line and not a craft line:
§4's ATTENDED list is re-scoped rather than deleted. Cues whose cultural_depiction_required or
hard_line_anchors cell is non-empty, and every realm cue, still take an attended §17.1 read — the
attending party is now a fresh-context critic rather than a human composer, and the read still
cannot be delegated to a prompt. The 132 creature vocalisations remain a sound-design task, not a
generative one.
Josh's quality mandate, given with the AAAAA reaffirmation and refined in his follow-up. This is
the creative bar every music deliverable in this dossier serves:
broad — SNES-era, Zelda, Undertale, Sonic, Super Mario 64, Skyrim, Halo, Pokemon, Animal
Crossing, Donkey Kong — scores that live outside their games. They endure on MELODY and
leitmotif, not production sheen.
cannot be hummed after one hearing, it is not the theme yet. ~~Composer-in-loop is the MECHANISM
for this bar~~ **(§0.0: the mechanism is now the candidate ladder + the measurement rig + critic
listen-checks; the AIVA closed-as-skip ruling stands)**; themes iterate until they pass, judged
with evidence like everything else.
and their time period, using instruments, chants, vocals where appropriate but mostly no vocals
other than hums chants and harmonies, ambient noises." Period-appropriate instrumentation per
region; vocals default WORDLESS (hums / chants / harmonies); ambient texture is part of the
score. This composes with the ruled sourcing law above (reference-COMPOSITION to the culture's
musical register, never sampling of consecrated performance).
VOLUME inside the per-culture register; the identity themes people hum in 2056 are the
composer-in-loop lane's deliverable.~~ **SUPERSEDED (§0.0): the split is no longer human-vs-model,
it is MODEL-vs-MODEL — XL-SFT + the 4B planner own the identity themes on the overnight exclusive
slot, 2B-turbo owns bed and variation VOLUME in shared windows.** The bar never routes down-tier:
a bed may be turbo-generated, an identity theme may not.
playlist at concerts, operas, and Broadway — the Harry Potter / Star Wars / video-game-legends
lineage (Distant Worlds, Symphony of the Goddesses, Symphonic Evolutions). IMPLICATION: the
hero-theme lane's deliverable is a COMPOSED WORK, not just rendered audio — real notation
(MusicXML/MIDI + orchestration) that a human orchestra can sit down and play. Rendered stems
serve the game build; the score serves the concert hall. Both come from the same composition.
to create this, understand the patterns, and how everything comes together." QUEUED ARTIFACT:
THE MUSIC COMPOSITION DOCTRINE — a deep-research foundation doc authored at the composition
program's start, BEFORE any hero-theme authoring: leitmotif architecture (statement /
transformation / combination across the 79 nodes and 22 threads), harmonic + modal language
per cultural register, orchestration and voice-leading practice, form, and the specific
techniques the exemplar lineage uses (how Zelda's overworld theme, Halo's chant, SM64's
bounce actually WORK). The doctrine is the composition lane's runbook; no theme authors
without it.
Consumers: the region-page Section 10 briefs (re-keyed 28ai), T0_Theme_Registry authoring, the
benchmark-day music evaluation rubric (hummability + register fit join the judged axes).
| Lead | Verdict | Source |
|---|---|---|
| "XL SFT" | CONFIRMED — ACE-Step/acestep-v15-xl-sft, licence field mit, 4B DiT, 50 steps, ~20GB on disk | huggingface.co/ACE-Step/acestep-v15-xl-sft |
| "XL Turbo" | CONFIRMED — acestep-v15-xl-turbo, 8 inference steps vs SFT's 50 | huggingface.co/ACE-Step |
| "Excel Bass" | CONFIRMED as XL BASE — acestep-v15-xl-base, pre-trained only. The name is a mis-hear of *XL Base* | huggingface.co/ACE-Step |
| ...as "ComfyUI audio nodes" | KILLED. These are ACE-Step 1.5 checkpoints, not nodes. The real ComfyUI audio nodes are EmptyAceStepLatentAudio, TextEncodeAceStepAudio, LatentOperationTonemapReinhard | docs.comfy.org/tutorials/audio/ace-step/ace-step-v1 |
| "ACE-Step 1.5 / UI" | CONFIRMED — official Gradio UI (uv run acestep) and a REST API server (uv run acestep-api, port 8001) ship in-repo | github.com/ace-step/ACE-Step-1.5 README |
| "ARDY", "Kimodo" | KILLED for this lane — both are NVIDIA motion-generation models (ARDY autoregressive text-to-motion, SIGGRAPH 2026; Kimodo kinematic motion diffusion, nv-tlabs/kimodo). Reroute to the animation lane; zero audio relevance | research.nvidia.com/labs/sil/projects/ardy/ · github.com/nv-tlabs/kimodo |
Repo naming collision to record: ace-step/ACE-Step (the v1 repo, Apache 2.0, 4-min cap) and
ace-step/ACE-Step-1.5 (MIT, 600s cap, XL series) are *different repos with different licences*.
AUDIO_STACK cites 1.5 correctly; anyone searching "ACE-Step" lands on the Apache one. Pin the URL.
| Job | PRIMARY | FALLBACK | Licence |
|---|---|---|---|
| Bulk regional/chapter music (79 nodes) | ACE-Step 1.5 XL-SFT + acestep-5Hz-lm-4B planner | ACE-Step 1.5 2B-turbo (speed lane) | MIT (LICENSE fetched: "MIT License", Copyright ACEStep 2026) |
| Adaptive stems | ~~ACE-Step 1.5 native multi-track / track separation~~ SUPERSEDED-BY-MEASUREMENT 2026-08-05 (§6.0 Q1): successive lego passes on acestep-v15-base. extract/lego/complete are TASK_TYPES_BASE — XL-SFT is excluded by the code, so the stem path and the quality path are DIFFERENT CHECKPOINTS — and extract returned four tracks at cross-stem spectral cosine 0.9998, i.e. the same content under four names. Layers are BUILT additively, never carved out of a mixdown | Demucs htdemucs_6s (analysis + genuine separation) | MIT / MIT |
| Canonical themes (12) | ~~composer-in-loop (ruled) — ACE-Step as reference-composition sketch only~~ SUPERSEDED 2026-08-05 (§0.0): ACE-Step 1.5 XL-SFT + acestep-5Hz-lm-4B planner, run through the candidate ladder + the measurement rig; the bar is held by measurement and iteration, not by a human ear | 2B-turbo candidate ladder; licensed orchestral library render | MIT |
| Bulk SFX / ambience / foley | MOSS-SoundEffect v2.0 (1.3B, DiT+Flow-Matching, 48kHz, 30s) | ElevenLabs SFX endpoint | Apache 2.0 (HF licence field apache-2.0) |
| SFX backbone library | Sonniss GDC bundles (2026 = 7.47GB / 347 WAV; ~200GB historical archive) | BOOM Library buyout | ROYALTY-FREE-OWNED, no attribution |
| Runtime variation (footsteps/UI) | UE 5.8 MetaSounds procedural DSP | — | engine-native, unrestricted |
| Creature vocalisation (132) | Krotos Dehumaniser 2 + Reformer Pro, sound-designer-in-loop | — | CARRIED (repo §2.7); not a generative task |
**Headline correction to AUDIO_STACK — ITSELF SUPERSEDED-BY-MEASUREMENT 2026-08-05; read §6.0 Q1
before routing any stem work.** The paragraph below was a README read, and the benchmark answered it
with numbers: AUDIO_STACK's §1.6 was closer to right than this correction was. ACE-Step 1.5 does not
give the adaptive-layering mechanism a separator would; it gives an ADDITIVE one, on a different
checkpoint, and the design consequence is written out in §6.0. The original text is kept below as the
reasoning that was superseded, not as live routing.
~~Its §1.6 "stem problem" — *"None of YuE/ACE-Step/HeartMuLa/Stable Audio natively output separated
stems"*, which made AIVA the only adaptive-music path — is now false. ACE-Step 1.5's own README
feature table lists "Track Separation — Divide audio into individual stems" and **"Multi-Track
Generation — Add layers like studio features"**, with corroborating reports of per-instrument
(kick/bass/snare/hats/perc/synth/pad) multitrack export. Since AIVA is CLOSED-AS-SKIP, ACE-Step 1.5
is the replacement adaptive-layering mechanism, not a downgrade. **Benchmark-day must confirm
whether multitrack emits true compositional stems or a bundled separator pass** — the whole Quartz
design leans on this.~~
What the benchmark actually found (§6.0 Q1, one paragraph so this table is never read alone):
extract does not isolate — four extracts off one mixdown measured mean cross-stem spectral cosine
0.9998, each ~0.987 identical to the mix, each carrying MORE energy than it (1.44–1.61×), and they
do not sum back (residual −1.29 dB). lego DOES work and is the mechanism the adaptive design is now
built on: the source survives in the output (correlation 0.47, spectral cosine 0.976) and the output
carries substantial content the source never had (residual energy ratio 0.76). So Quartz layers are
generated one per pass on a base checkpoint, at one generation per layer, and hero cues that need
stems render on base rather than xl-sft. Demucs htdemucs_6s remains the real separator and is
what the probe cross-checked against.
Demotions and exclusions (licence-read-before-first-GPU-hour):
2 sessions and states full songs (4+ sessions) need ≥80GB; last update June 4 2025 (~13 months
stale). On a 32GB card it buys ~2-3 minutes. Not a co-equal to ACE-Step.
HeartMuLa/heartlib active (3.8k stars); newest release HeartMuLa-oss-3B-happy-new-year (2026-02-13). The 7B is still an unticked TODO. AUDIO_STACK
Unknown #4 is now settled: no larger checkpoint exists.
a shipped path.
Community License, whose threshold I re-fetched verbatim: *"free for everyone, unless… you or your
organization generate over USD $1M… of annual revenue"* (stability.ai/license). Above it an Enterprise
Licence is required. Open Small = 0.5B / 11s / 44.1kHz; 3 Small SFX = 0.6B, and its own card's code
examples run sample_rate at 16,000 Hz — below game standard. Two independent reasons MOSS wins.
UE 5.8 (current) reality check: import accepts .wav/.ogg/.flac/.aif/.opus/.mp3, **all converted
internally to 16-bit WAV** — so generate at 48kHz/24-bit, deliver 16-bit, and never treat source bit
depth as a shipped-quality lever. New in 5.8: the Windows audio backend switched XAudio2 → WASAPI
(AudioMixerWasapi) — re-run any device/latency assumption from the old machine. MetaSounds + Quartz
both current in 5.8 docs.
D-1 finding — the audio lane needs NO WSL2. ACE-Step 1.5 advertises Mac/AMD/Intel/CUDA support and
ships a Windows-runnable Gradio+REST server; MOSS-TTS installs from a cu128 PyTorch wheel — the exact
sm_120 Blackwell stack the runbook already budgets. **Both run in runbook lane (b), Windows-native, in
their own venvs.** Only MOSS's optional vLLM-Omni/SGLang-Omni backends are Linux-leaning — use the
PyTorch backend on Windows and treat vLLM as a WSL2 option *if* throughput ever demands it. This removes
the audio lane from the D-1 fork entirely.
Disk placement (per the RULED layout). Weights + venvs + HF cache → Gen4 D: (HF_HOME,
HUGGINGFACE_HUB_CACHE, TORCH_HOME already re-pointed at runbook 1.5). Sonniss/BOOM sample libraries →
8TB SATA slot 4 (runbook 1.4 already rules this). Nothing audio touches the Gen5 hot path except the
imported .uasset under C:\dev\Humanity\Humanity\Content\Audio\.
Install order (each step's exit condition is an artifact on disk, not a claim):
1. D:\ai\acestep\ — clone github.com/ace-step/ACE-Step-1.5, own venv, uv run acestep → Gradio up.
2. Pin torch explicitly (torch==2.9.0+cu128) — the runbook's Puget-pack lesson about unpinned torch
applies here identically.
3. uv run acestep-api → port 8001. Add a Stage-6F firewall row (the runbook table lists UE, MCP
9315, MCP 8000, Ollama 11434, ComfyUI 8188 — 8001 is missing and will prompt at 04:00).
4. D:\ai\moss-sfx\ — pip install --extra-index-url https://download.pytorch.org/whl/cu128 -e ".[torch-runtime]"
from github.com/OpenMOSS/MOSS-TTS. First call compiles (torch.compile + Triton CUDA graphs) — budget
several minutes and do not read it as a hang.
5. Demucs (pip install demucs) — small, CPU-viable, fallback only.
6. ComfyUI ACE-Step 1.5 nodes are native (no custom node install): drop
ace_step_1.5_turbo_aio.safetensors into ComfyUI/models/checkpoints/. This is the *turbo* lane only —
XL-SFT quality runs through the native REST API, not ComfyUI. Route accordingly.
Thursday-night download list (measured unless marked):
| Item | Size | Where |
|---|---|---|
acestep-v15-xl-sft (4× 4.99GB shards) | ~20.0 GB | D: HF cache |
acestep-v15-xl-turbo | ~20 GB ESTIMATE (same class) | D: |
acestep-5Hz-lm-4B planner | ~8 GB ESTIMATE | D: |
MOSS-SoundEffect-v2.0 | 11.2 GB | D: |
ComfyUI ace_step_1.5_turbo_aio.safetensors | ~7 GB ESTIMATE | D: ComfyUI/models/checkpoints |
| Sonniss GDC 2026 bundle | 7.47 GB | 8TB SATA |
| Sonniss historical archive (archive.org) | ~200 GB, optional/overnight | 8TB SATA |
Demucs htdemucs_6s | <1 GB | D: TORCH_HOME |
~47 GB for the working set, ~67 GB with the turbo mirror, before the optional 200 GB archive.
Addressing — already ruled, already columned.
T0_SFX_Registry.sfx_id → ue_sound_asset_path, final asset drops at /Game/Audio/SFX/SFX_<sfx_id>,swap = path takeover (idempotent import). Teeth: QT-10 cue-fired telemetry non-zero, QT-AU
muted/wired fixture.
T0_Theme_Registry.theme_id → ue_sound_asset_path, drops at /Game/Audio/Cues/<cat>/A_<theme_id>.stereo, Bink Audio compression for ambience/music, PCM for short high-frequency cues.
Two verified state facts that change the lane's shape:
1. AUDIO_STACK Unknown #6 is CLOSED. It flagged that Theme_Registry carried no tempo/bar/stem-role
fields for Quartz. It now does: the live header carries `tempo_bpm, bar_length, per_stem_role,
ue_sound_asset_path`. 12 rows exist, **all empty on those four fields, and there are zero regional
rows.** The gap is population, not schema.
2. A real DR-2 five-state conformance gap on SFX. T0_Theme_Registry carries `generation_status,
generation_prompt_hash, last_generated_timestamp — but not regeneration_trigger`.
T0_SFX_Registry (30 rows, 20 cols) carries none of the four. Until the assetgen rail lands on
both, generated audio cannot express pending→in_progress→complete→rejected→regeneration_required,
and the prompt-hash regeneration trigger DR-2 depends on has nowhere to live. **This belongs in the
next batched registry-extension pass + one fidelity re-baseline in the same commit.**
Downstream wiring already landed (28-series): the §10.6 arrays exist —
region_footstep_sfx_id_ref_array, region_environmental_ambient_sfx_id_ref_array,
region_ritual_context_sfx_id_ref_array, plus chapter/boss/ability/vril_cast/quest_dialogue/weapon
arrays (docs/registry_extensions.json). Known data defect to repair: LW-7 — Flores'
region_environmental_ambient_sfx_id_ref_array holds Environment-Grammar ids (EG_0001;EG_0002),
not SFX ids.
Stem/loop contract into UE. Each adaptive cue delivers: (a) N separate loopable WAVs, one per
per_stem_role; (b) tempo_bpm; (c) bar_length; (d) clean loop points with no baked-in fades —
ACE-Step's generation must be prompted for, and QA'd against, seamless loop boundaries, because Quartz
triggers layers at Quantization Boundaries and a baked fade destroys the seam. Import via
unreal.SoundFactory + AssetToolsHelpers…import_asset_tasks (the proven bridge —
Content/Audio/SFX/smoke_stone_door already landed this way). MetaSound graph assembly is scriptable via
MetaSoundBuilderSubsystem, but the Builder API has no variables support — expect hand-authoring in
the editor for anything past ~3-4 layers, which is consistent with sound-designer-in-loop for hero cues.
Dispatch. The director never names a model in a work order. It selects registry rows —
T0_Theme_Registry rows with generation_status IN (pending, regeneration_required), or T0_SFX_Registry
rows reachable from a region's §10.6 arrays — and hands them to Tools/compose_audio_prompts.py (the
missing Hop-3 script, ranked BUILD-NOW), which reads the region's Section-10 cue table + Section 17
substrate and emits per-cue prompt objects with a generation_prompt_hash. ~~The lane then POSTs to
ACE-Step's REST server on :8001 (music)~~ **CORRECTED 2026-08-05 (deviation 2 of 2, §6.0): the music
lane calls ACE-Step IN-PROCESS — harness/music_gen/generate.py imports acestep.handler/
acestep.inference directly, so one handler load serves a whole ladder and offload_to_cpu is
controllable per run. No acestep-api server stands (verified: nothing listening on :8001). The
:8001 REST path in §2 step 3 remains an available option, not the dispatch this lane uses** or calls
MOSS locally (SFX), writes the WAV, imports, and writes back through the T0.13 manifest. **Head-of-chain blocker: the Section-10 machine cue table does not exist
yet** — Flores' Section 10 is 100% prose, so check_build_readiness.py lane_audio reports BLOCKED-ON and
the lane has *no addressable unit of work*. That template edit + Flores backfill is the single thing
gating audio autonomy, and it is neither 5090-gated nor vendor-gated.
Scheduling class + VRAM budget (the concurrency law).
| Config | Peak VRAM | Class |
|---|---|---|
| ACE-Step XL-SFT + 4B LM (best quality) | ~24-28 GB ESTIMATE (README tier: ≥24GB) | OVERNIGHT, EXCLUSIVE — cannot co-reside with the UE editor or the image lane on a 32GB card |
| ACE-Step 2B-turbo + 0.6B LM | 6-8 GB (README tier) | WINDOWED — safe alongside UE |
| MOSS-SoundEffect v2.0 (1.3B) | <8 GB ESTIMATE — MEASURE day one | WINDOWED, the always-available bulk-SFX worker |
Demucs htdemucs_6s | modest / CPU-viable | ALWAYS-ON |
The rule this lane commits to: **the XL music model owns the GPU alone, on the overnight slot, after the
soak** — every bulk-SFX and iteration pass runs on the small models in shared windows. A music batch must
never be the reason a QA soak or a UE cook stalls.
Quality gate — promote/demote. Benchmark day scores a fixed set (2-3 regional themes + 1 hero-adjacent
cue + 6 SFX classes: wind bed, ocean, jungle, footfall-per-surface, weapon impact, ritual context) on:
instrumentation fidelity, cultural-substrate accuracy against the Care Doctrine, loop-seam cleanliness,
stem separability, peak VRAM, wall-clock, and whether the licence clears. A model is promoted to
PRIMARY only on that evidence; a model that fails the loop-seam or stem test is demoted to reference-only
regardless of how good the mixdown sounds. Output QA runs through the standing audio teeth (QT-10 event-
fire, QT-AU fixture) plus a listen-check by a critic on captured playback — never a described claim
(the visual-verification directive's audio analogue).
Unattended vs attended. UNATTENDED: bulk regional ambience, per-surface footfall, weapon/item foley,
combat cue variants, and all regeneration triggered by prompt-hash change. **ATTENDED / human-in-loop
always: ~~the 12 canonical themes (composer-in-loop is ruled)~~ SUPERSEDED (§0.0) — the 12
canonical themes are now UNATTENDED-GENERATED and ATTENDED-JUDGED: the ladder runs unattended, and a
fresh-context critic scores the picked candidate against the doctrine's §2.3 checklist on captured
playback before it is served anywhere**, the 132 creature vocalisations (Krotos,
sound-designer-in-loop), any cue whose cultural_depiction_required or hard_line_anchors cell is
non-empty, and every realm cue — the §17.1 read cannot be delegated to a prompt.
The sourcing law, operationalised. *No sampling of consecrated performance; reference-COMPOSITION
permitted.* Concretely: the prompt composer may pass **instrumentation, mode/scale, ensemble shape, and
era descriptors drawn from instrumentation_substrate / cultural_substrate; it may never** pass an
audio file of a consecrated performance as a conditioning input, and no ACE-Step LoRA is ever trained on
one. Two licence walls reinforce this from the other side: **Sonniss's bundle licence explicitly prohibits
using the audio to train AI/ML**, and EastWest's EULA does the same (CARRIED) — so the library backbone is
direct-use-only and can never feed a fine-tune. Record the generating model + licence in every artifact:
ACE-Step-1.5-XL-SFT (MIT) or MOSS-SoundEffect-v2.0 (Apache-2.0), written into the row alongside
generation_prompt_hash.
Genuinely small — AIVA closing as SKIP removed the only account/negotiation item in this lane.
1. Krotos Dehumaniser 2 + Reformer Pro — ~$798 one-time, his card (creature vocalisation; the only
purchase this lane needs). Optional, deferrable to the creature pass.
2. BOOM Library buyout — $99-199, optional, only if benchmark day shows Sonniss+MOSS leaves gaps.
3. Sonniss GDC bundle download — free, but the site is email/form-gated; one sitting, or delegate.
4. Nothing else. No API key to mint (the ElevenLabs key already exists and works —
Tools/generate_sfx.py produces WAV/MP3 today), no licence to negotiate, no consent gate (that is the
voice lane's, already ruled).
**What this section covers and what it does not, stated first so the title can never be read as more
than it is. ANSWERED by measurement: §6 questions 1, 2, 3, 4** — the music-generation questions.
NOT RUN: §6 questions 5, 6, 7, 8 — see the annotations on that list for each one's reason. Against
§4's declared scoring set ("2-3 regional themes + 1 hero-adjacent cue + 6 SFX classes") what actually
ran was one hero-adjacent ladder and 44 ambience beds: zero regional themes (all 12 region
cells are correctly HELD pending the §17.1 read) and zero SFX classes (MOSS is not installed).
The Care-Doctrine question is therefore still open, and nothing here closes it.
Answers below are measurements, each with the artifact that produced it. Tooling:
harness/music_gen/{generate,feature_rig,rank,stem_probe,motif_compare,emit_ladder,emit_picks,to_midi}.py.
Evidence: build/audio/benchmark/stem_probe.json, build/audio/generated/journey_world/ranking.json,
.../family_report.json, and the per-artifact licence + generation records under build/audio/records/
— **65 artifacts, 65 record pairs, zero orphans in either direction, every sha256 re-verified against
the bytes on disk** (the four Demucs separator outputs are included as an explicit DERIVED-ANALYSIS
class carrying licence_class: previz_only, because the pairing law has no analysis exemption and a
model output without a record is the pattern that ships an unrecorded asset the first time a derived
pass produces something that DOES ship).
INSTALLED, MEASURED, AND TWO DEVIATIONS FROM THE INSTALL PLAN RECORDED HONESTLY. ACE-Step 1.5 is
live Windows-native at D:/audio/acestep in its own venv: XL-SFT (19 GB), acestep-5Hz-lm-4B
(8 GB), the main bundle with the 2B turbo DiT + VAE + Qwen3 text encoder (~6 GB), and
acestep-v15-base (4.5 GB) for the stem tasks — 36 GB of checkpoints on the ruled D: placement,
HF_HOME already re-pointed. **The deviation: the plan said pin torch==2.9.0+cu128; the repo's
own pyproject.toml pins torch==2.7.1+cu128 for Windows and resolves it from the
download.pytorch.org/whl/cu128 index BY URL.** The repo's pin was followed rather than the
dossier's, because 2.7.1+cu128 is a cu128 Blackwell wheel and it is the version the project tests
against; verified live at torch 2.7.1+cu128 / cuda 12.8 / sm_120 / RTX 5090. The plan's number
was an estimate and is corrected here rather than quietly satisfied.
Deviation 2 — DISPATCH IS IN-PROCESS, NOT THE :8001 REST SERVER. Install-plan step 3 stands up
uv run acestep-api on port 8001 and §4 described the lane as POSTing to it. generate.py instead
imports acestep.handler and acestep.inference directly and holds one handler across a whole
ladder. This is the better shape for this workload — no HTTP hop per candidate, offload_to_cpu
settable per run, and the handler-load cost paid once instead of per cue — but it is a deviation and
it was undeclared until this line. §4's dispatch sentence is corrected in place. The :8001 firewall
row stays landed in the runbook: the REST server remains a supported path for any consumer that wants
one, and it will not prompt at 04:00 if something stands it up. Nothing else in §2 was departed from.
**Q1 — TRUE COMPOSITIONAL STEMS, OR A BUNDLED SEPARATOR? NEITHER. THE QUESTION HAS THE WRONG
SHAPE, AND THE ANSWER CHANGES THE QUARTZ DESIGN RATHER THAN CONFIRMING IT.** Three findings, in
the order they bite:
1. The quality primary cannot do it at all. extract, lego and complete are
TASK_TYPES_BASE in acestep/constants.py — base checkpoints only. XL-SFT and every turbo
variant are excluded by the code, and the model zoo table marks XL-SFT ❌ on all three. The
adaptive-stem path and the best-quality path are different checkpoints, which the stack
table above did not say.
2. extract did not isolate anything, measured. Four extracts (strings / percussion / guitar /
keyboard) off one mixdown on acestep-v15-base: mean cross-stem spectral cosine 0.9998 —
the four "stems" are the same content returned under four different track names. Each is
spectrally ~0.987 identical to the mixdown, sample-aligned to it (lag −5), and carries more
energy than it (1.44–1.61×), and they do not sum back to it (residual −1.29 dB). So it is not a
mask separator, not a decomposition, and not a stem set. The requested track name did not select
content. *Scope of the claim: one checkpoint, one prompt configuration, duration passed on the
extract call; the retest lever is a base-model run with duration unset and a per-track caption.*
3. lego DOES work, and it is the mechanism the adaptive design should be built on. Its
instruction is *"Generate the {TRACK} track based on the audio context"* — it ADDS rather than
carves. Measured: the source survives in the output (correlation 0.47 at lag −5, spectral cosine
0.976) and the output carries substantial content the source never had (residual energy
ratio 0.76). That is additive compositional layering.
CONSEQUENCE FOR QUARTZ, stated as a design change and not a footnote: adaptive layers are
BUILT by successive lego passes on a base checkpoint, not carved out of a finished mixdown.
That is better than the fallback this dossier feared (pop-tuned 4/6-stem separation of orchestral
material) and worse than the hope (free stems from the quality model). It costs one generation per
layer and it means hero cues render on xl-base/base rather than xl-sft when stems are needed.
**Q2 — LOOP SEAM AND BPM RECOVERABILITY: BPM IS HONOURED, WITH MODEL-DEPENDENT PRECISION, SO
tempo_bpm IS MEASURED FROM THE RENDER AND NEVER ASSUMED FROM THE PROMPT.** Turbo hit 95.7 bpm
against a 96 target (0.3% error). XL-SFT with the 4B planner came back 117.5–129.2 against 112
and 120 targets (2–15% error) — directionally obedient, not exact. So bar_length IS fillable,
which keeps Quartz layering alive, but only off a measured tempo. **Baked fades: zero, across 8
candidates and 44 beds**, once fade_in_duration/fade_out_duration were set to 0 explicitly (the
library defaults are non-zero); the rig's fade detector is positive-controlled against a synthetic
fade, so that zero is a measurement and not an absence of looking. Seam ratios ranged 0.58–2.18 —
some candidates loop cleanly and some do not, so **loop-seam cleanliness is a per-artifact gate,
not a property of the generator.**
**Q3 — XL-SFT VRAM AT HANDLER LOAD, MEASURED. THE QUESTION SAID *PEAK* AND THIS RUN DID NOT MEASURE
ONE — the label is corrected here rather than left standing.** With offload_to_cpu=True, XL-SFT +
the 4B LM added ~9.5 GB over baseline (nvidia-smi 22535 → 32016 MB at handler load) on a card
already holding another lane's UE editor, and generated 32-second candidates in 16.6–40.2 s each.
It fit — with under 1 GB of headroom. The exclusive-overnight class stays the default and this
measurement is why: a windowed XL run is possible but leaves no room for anything to grow, so it is a
deliberate exception, never the schedule. Both conclusions rest on the load-time figure and both
stand.
What did NOT get measured, and the field that wrongly claimed it had. The records carried
nvidia_smi_used_mb_peak_observed, but that value was a single nvidia-smi call made at
record-write time — *after* generation returned. On the eight XL-SFT candidates it read
20765–23376 MB: below the 32016 MB after-load figure and, on three of them, **below the 22535 MB
baseline**. A field named "peak" that can report less than the floor is a label, not a measurement.
The same records do carry a true in-process torch peak — torch_max_allocated_mb 17912–18097 MB —
which is this process's own allocation and not the device total. Fixed on the measured close: the
field is renamed nvidia_smi_used_mb_at_record_write across all 66 landed record and summary files
(a relabel; every value byte-identical, and the generation_prompt_hash is untouched because
build_payload excludes the timing block), and generate.py now samples nvidia-smi on a thread for
the duration of each generate_music call, writing nvidia_smi_used_mb_peak_sampled +
nvidia_smi_peak_samples. The sampler is armed by fixture F17, not merely present. The 61 already-
landed records declare peak_sampling_note: NOT RUN rather than inheriting a number they never
measured — **so the true generation-time device peak for the XL-SFT ladder remains UNMEASURED until
the next XL run**, and the scheduling default above must be cited as a load-time figure, never as a
peak.
**Q4 — INSTRUMENTAL-ORCHESTRAL QUALITY vs THE VOCAL-POP CRITIQUE: the vocal-pop worry is NOT what
bites. The melody is.** With instrumental=True the output is genuinely orchestral and wordless —
no vocal leak in any of 52 artifacts. But measured melodic salience across the eight XL-SFT
candidates was 0.41–0.53, against 1.0 for a synthetic melodic control and 0.0 for both a noise
bed and a held triad. The material is texture-forward: 12–23 note events over 32 seconds. **And the
brief's single licensed anomaly was not delivered** — the winner scored D4 = 1 with *zero* leaps of
a sixth or larger, against a brief that specifies triad-outlining with exactly one unusual leap.
LEITMOTIF_ARCHITECTURE §3.6 already stated that the generator has no melodic conditioning that
could honour a quoted head; this is that statement measured, and it generalises: **text
conditioning does not reliably realise a specified melodic shape.** The 30-year bar is a melody
bar, so this is the lane's real open problem, and the three levers are named rather than assumed:
(a) scale the ladder and let selection pressure do the work, (b) author the head cell and use
lego/repaint to arrange around it — now known to be viable, see Q9, (c) find or build a
melody-conditioning path. The honest tier on everything shipped today is GENERATED SCORE —
ITERATION, and the tier is doing real work.
**[ANSWERED 2026-08-05 — §8 Q10 ran (a) and (b) head to head. (b)'s AUTHORING half wins and is
promoted; (b)'s ARRANGEMENT half does not survive its own decomposition, and (a) at double the
ladder with the contour named as explicitly as words allow still delivered the licensed leap in
2 of 16 and the head figure in 0 of 16. The open problem is no longer a generator-melody problem —
it is an orchestral-REALISATION problem. See §8.]**
Q9 — A QUESTION THIS DOSSIER DID NOT ASK, ANSWERED BECAUSE THE THEME FAMILY NEEDED IT (numbered
9 and not 5: §6 already has a Q5 about MOSS, and two different questions answering to
"PIPE_AUDIO_MUSIC Q5" is exactly the ambiguity that makes a later citation unresolvable) **: DOES
CONDITIONING HOLD A MOTIF? Yes, measurably. A repaint of the statement's back half held the
declared region** (waveform correlation 0.9998 over the untouched 16 s; residual 1.2%, not
byte-identical because the whole file is VAE re-decoded) and the new half ~~**shares real melodic
material with the statement — interval 3-gram Jaccard 0.333, 4-gram 0.238~~ [HALF-WITHDRAWN
2026-08-05 — §8 Q11. The two figures reproduce exactly at HEAD, but they were computed over the
WHOLE FILE, including the 16 s both files share byte-for-byte, so the held region contributed its
own n-grams to both sides and inflated the overlap. With that region excluded the same pair reads
3-gram 0.1250 / 4-gram 0.0000 — this dossier's own WEAK AND UNSAFE band. motif_compare.py now
decides on the generated region; the landed verdict flipped on re-run and the artifact was demoted
to a sibling render.]** That is the finding
that makes lever (b) above credible: the generator cannot be told a tune in words, but it can be
given one in audio and will keep it — **and §8 Q10 measures the limit of that. It KEEPS what it is
handed, exactly (correlation 1.0000 on all four repaints of the melody A/B), and it does not
CONTINUE it: the generated region carries motif 3-gram 0.000 and none of the brief's head figure.**
Status line, 2026-08-05 — read this before treating any item below as open or closed. Items 1-4
are ANSWERED by measurement → §6.0. Items 5-8 are STILL OPEN, each for a stated reason, and
item 6 in particular is a Care-Doctrine gate that nothing in §6.0 touched. §6.0's own extra question
is numbered Q9 so it never collides with item 5 here.
1. ANSWERED → §6.0 Q1 (the question had the wrong shape; the answer was NEITHER).
~~Does ACE-Step 1.5 multitrack emit true compositional stems, or is it a bundled separator?~~ The
entire Quartz adaptive design rests on this. If separator-only, orchestral section splitting is a
known-weak pop-tuned 4/6-stem problem and hero cues fall back to library-rendered stems.
*Measured: extract isolates nothing (cross-stem cosine 0.9998) and is base-checkpoint-only;
lego layers additively and is what Quartz is now built on. §1's stack row is annotated to match.*
2. **ANSWERED → §6.0 Q2 (BPM is honoured; precision is model-dependent, so tempo is measured from the
render).** ~~Loop-seam quality at bar boundaries~~ — generate 8-bar loops at a stated BPM and
measure whether tempo_bpm/bar_length are *recoverable* from the output at all, or must be
imposed by prompting. If ACE-Step will not honour a target BPM, bar_length is unfillable and
Quartz layering is dead. *Measured: turbo 0.3% error, XL-SFT 2-15%; zero baked fades across 52
artifacts; loop-seam cleanliness is a per-artifact gate, not a generator property.*
3. NOW FULLY ANSWERED → §6.0 Q3 for the load-time figure, §8 for the peak. ~~The LOAD-TIME figure
is measured (22535 → 32016 MB); the
GENERATION-TIME PEAK this item literally asks for is STILL UNMEASURED~~ — the field that claimed
it held a post-hoc sample, the sampler that can answer it landed on the measured close, and the
number arrived on the next XL run, which was the melody A/B: **29381-31993 MB sampled peak across
16 XL-SFT candidates, 53-99 samples per call, on a card already holding 12.6 GB of another lane's
work. The exclusive-overnight default is reinforced — the ceiling was reached with under 1 GB of
headroom.** ~~24-28GB is my estimate from the README tier table.~~ *The
scheduling conclusion (exclusive-overnight stays default) rests on the load-time figure and holds.*
4. ANSWERED → §6.0 Q4, and it reframed the lane's real problem. ~~Instrumental-orchestral quality
vs the vocal-pop critique~~ — AUDIO_STACK Unknown #8. *Measured: no vocal leak in any of 52
artifacts, so the vocal-pop worry is closed; melodic salience 0.41-0.53 against 1.0 for a melodic
control, and the brief's one licensed leap was not delivered. The melody is the open problem, and
the honest tier stays GENERATED SCORE — ITERATION.*
5. STILL OPEN — NOT RUN. MOSS-SoundEffect is not installed (no D:/ai/moss-sfx, nothing
MOSS-shaped anywhere on D:), so zero SFX classes were scored. **MOSS-SoundEffect on the specific
hard classes** — per-surface footfall (needs short, dry, variation-friendly one-shots, not
ambience) and ritual-context cues under the sourcing law.
6. STILL OPEN — NOT RUN, AND IT IS THE CARE GATE. All 12 region cells are HELD under the §17.1
carve-out pending the elevated-care read, so no regional theme was generated and nothing tested
this. Cultural-substrate accuracy — do region prompts produce plausibly-registered
instrumentation, or generic "world music" wash? A wash is a Care-Doctrine failure (care as
authenticity), not a taste note. Nothing in §6.0 may be cited as evidence on this question.
7. STILL OPEN — NOT RUN. Stable Audio was never installed or invoked; it remains CAP-RISK
fallback-only and untested. Stable Audio 3 Small SFX true output sample rate — its card's
example reads 16kHz. Confirm or kill before it is trusted even as fallback; a 16kHz asset cannot
ship.
8. STILL OPEN — NOT RUN, and correctly so: its own precondition has not fired. Windows-native
throughput was never the constraint on this batch (32-second candidates in 16.6-40.2 s).
Windows-native throughput vs WSL2+vLLM — only worth measuring if the Windows PyTorch backend
proves throughput-bound on the 79-node batch.
github.com/ace-step/ACE-Step-1.5 (+ /blob/main/README.md, raw…/LICENSE, raw…/README.md) ·
github.com/ace-step/ACE-Step (+ /blob/main/LICENSE) · huggingface.co/ACE-Step ·
huggingface.co/ACE-Step/acestep-v15-xl-sft (+ /tree/main) · arxiv.org/abs/2602.00744 ·
docs.comfy.org/tutorials/audio/ace-step/ace-step-v1 · …/ace-step-v1-5 ·
github.com/multimodal-art-projection/YuE · github.com/HeartMuLa/heartlib · github.com/adefossez/demucs ·
huggingface.co/OpenMOSS-Team/MOSS-SoundEffect-v2.0 (+ /tree/main) · github.com/OpenMOSS/MOSS-TTS ·
huggingface.co/stabilityai/stable-audio-open-small · huggingface.co/stabilityai/stable-audio-3-small-sfx ·
stability.ai/license · elevenlabs.io/docs/overview/capabilities/sound-effects · gdc.sonniss.com ·
sonniss.com/gdc-bundle-license · dev.epicgames.com UE 5.8 release notes / MetaSounds / Quartz ·
research.nvidia.com/labs/sil/projects/ardy · github.com/nv-tlabs/kimodo
---
APPEND-ONLY addendum under the pipe-dossiers-bind law, from the 2026-08-05 comparator-pipeline research
(Josh: "what about their pipelines?"). **Read the scope limit first, because it is the honest finding of
this section: no researcher in that pass was tasked with the AUDIO lane.** Three researchers covered
quests/world, creatures/combat/VFX, and production systems at scale. Audio was not a lane in the fan-out.
So this section is NOT a research pass — it is a routing pass. Its content is the audio-bearing facts
the other three researchers surfaced inside sibling dossiers, where they sat orphaned against a lane that
could not consume them. Each was **re-verified by direct fetch by the synthesizer before being routed
here**; nothing is carried on a sibling's summary alone. The gap itself is boarded at the bottom.
Capcom's RE ENGINE shows captured motion in engine in real time during a shoot, and — the part that
belongs to this lane — sound effects and background music can be played into the same scene, so the
team sees what the player will experience far earlier in the process. Named: animator **Naohiro
Taniguchi (real-time visualization), monster-animation lead Kenji Yamaguchi**. The demonstration
studio carries 36 wall-mounted infrared cameras; a Kyobashi facility has 150. Before Monster Hunter
World, all monster animation was keyframed.
Source — https://www.digitaltrends.com/gaming/monster-hunter-wilds-capcom-studio-tour/
VERIFIED by direct fetch 2026-08-05 (quoted: developers can "see that process work in real time and
continue to polish and refine the animations"; "Capcom can even add sound effects and background music
into the scene").
NO-ADOPT: the capture infrastructure — we have no stage and will not have one; nothing here changes
the ARDY/Kimodo primary in PIPE_ANIMATION §1.
ADOPTED AND ROUTED: audio is not a post-pass. The proving map that PIPE_ART §7.5 and
PIPE_ANIMATION §7.3 both independently converged on must carry cue audio wired at authoring time,
not added after the motion and the VFX are signed off. One proving map, three lanes, one screenshot-and-
listen surface for the visual-verification standing directive. This costs nothing to adopt now and costs
a re-do if adopted after volume authoring.
From the God of War (2018) combat round-table, the concrete audio facts:
ends with a high frequency slash."
the hit pose and holds both Kratos and the target in that first frame for a short duration."
REDUCED shake, "what is reduced by these camera choices is made up for in audio and animation."
Speaker: Christian Wohlwend (Naughty Dog), in the multi-studio round-table.
Source — https://blog.playstation.com/2022/10/04/game-developers-explain-what-makes-god-of-war-2018s-combat-tick/
VERIFIED by direct fetch 2026-08-05. Attribution note: the researcher return framed these as Santa
Monica facts; they are a round-table participant's, and the participant is Naughty Dog's. The claims are
about God of War's combat; the speaker is not a Santa Monica employee. Cite the person, not the studio.
ADOPTED AND ROUTED: the low-end-then-high-frequency envelope is a **cue-design law for the impact
beat**, routed to the SFX half of this lane alongside the four-beat cue grammar. It is also the direct
answer to a question our own stack raises: when the presentation ladder escalates a cast across tiers,
what escalates in the AUDIO? The comparator answer is that weight lives in the low end and legibility
lives in the high end — which makes tier escalation a low-end-mass delta with the high-frequency read
held constant, the exact structural twin of the luminance-versus-hue finding on the VFX lane
(PIPE_ART §7.3). Two lanes, one shape: escalate the intensity axis, hold the identity axis. That
symmetry is this addendum's own synthesis and is flagged as such.
Nintendo built **SLink, a tool letting designers assign sounds by clicking an actor in the game with
the mouse**, as part of the tooling that let a two-designer UI team ship Breath of the Wild's UI without
filing requests with programmers.
Source — CEDEC 2017 (Fujibayashi / Yonezu), Matt Walker translation —
https://gist.github.com/idbrii/e39fe96279aa1670319bfa521d907399
VERIFIED by direct fetch 2026-08-05 (quoted: "SLink, allowing designers to click an actor in game
with their mouse").
ADOPTED AND ROUTED: our equivalent of "click the actor and assign the sound" is **a declared audio
binding on the entity row, authored where the entity is authored**, rather than a separate audio-wiring
step that must re-derive which entity needed which cue. This is the same premise as every
harness/apply_*.py in the repo, pointed at audio, and it is corroboration for the registry-as-authoring-
surface bet rather than a new mechanism.
GDC 2024, **"Tunes of the Kingdom: Evolving Physics and Sounds for 'The Legend of Zelda: Tears of the
Kingdom'" (Dohta / Takayama / Osada**) — https://gdcvault.com/play/1034667/Tunes-of-the-Kingdom-Evolving
The 2026-08-05 pass extracted only the PHYSICS half of this talk (it landed in PIPE_CONTROL_PLANE
§8.2, N7-N8: every moving object made physics-driven; mass and volume auto-derived from shape plus an
assigned material, then de-tuned for feel). **The SOUND half — which the title names and a named audio
speaker presented — was never extracted, and it is the single most on-target first-party source available
to this lane on how audio couples to a systemic physics world.** It is the top re-fetch.
The physics half already implies the audio question, and the question is ours too: when mass and material
are DERIVED per object rather than authored, **the impact sound has to be derived from the same two
values, or every collision in the game sounds like the same wooden knock.** Our element matrix and
material assignments are exactly that kind of derived pair. Whether Nintendo solved it by material-indexed
sample sets, by synthesis, or by something else is precisely what the unread half of that talk would say.
Blizzard's stated Overwatch data-pipeline goal that local changes are seen instantly in engine applies
to audio iteration identically (PIPE_CONTROL_PLANE §7.6, tier SECONDARY with its caveat declared there).
Our generated-audio lanes currently land files in folders and wait for an import step — the same return-
path gap PIPE_CONTROL_PLANE §7.5 item 5 routes to ASSET_DROPIN_CONTRACT.md. **No separate adoption;
the audio lane is a beneficiary of that one.**
No dedicated audio-pipeline comparator research has been run. Everything above is routed from lanes
that were looking at something else and happened to surface an audio fact. The questions a real pass would
answer, none of which this section can: how shipped AAA titles structure an adaptive music system's
authoring surface (our native-multitrack benchmark question, §6, has no comparator anchor); how SFX
libraries are organized and versioned at scale; how VO is routed, batched and QA'd against script churn;
how mix and loudness are validated automatically in CI. Boarded, not done. The comparator-lens program
should carry an audio row the next time it fans out — and the 30-year melodic bar (§0) deserves a
comparator pass of its own, since none of the four teachers studied on 2026-08-05 was chosen for music.
---
What this section answers and what it does not, stated first. ANSWERED by measurement: the open
problem §6.0 Q4 left the lane with — *text conditioning does not reliably realise a specified melodic
shape*. Two of Q4's three named levers were run head to head. NOT ANSWERED and not touched: §6 items
5, 6, 7, 8 remain exactly as §6.0 left them; the Care-Doctrine question (item 6) is still open and
nothing in this section may be cited on it — JOURNEY_WORLD is the home register, not a region
cell, and no region cell was generated here either. Lever (c), a melody-conditioning path, was not
built.
The experiment. ARM 1: the JOURNEY_WORLD motif composed in notes by the factory to the
brief's constraints, realised to audio, then arranged by nine ACE-Step passes conditioning on it
(four repaints holding 24/16/16/8 s, two covers at strength 0.75/0.45, three lego layers on the base
checkpoint). ARM 2, the control: sixteen pure text2music candidates on XL-SFT + the 4B planner —
double the original ladder — with the contour language sharpened as far as words go, the leap named
by size and direction, the head named note by note in scale degrees. Both arms judged by the same
rig against the same measure_spec, and both compared against the **same authored interval
sequence**, so arm 1 got no privileged reference.
Tooling: harness/music_gen/{author_head_cell,ab_judge}.py (new), emit_ladder.py --arm1/--arm2,
feature_rig.py (17/17 controls green first). Evidence:
build/audio/generated/journey_world_ab/ab_report.json, the three job specs beside it, and 50
record files under build/audio/records/journey_world_ab/ — 25 artifacts, zero orphans.
**Q10 — DOES AUTHORING THE HEAD CELL BEAT PROMPTING FOR IT? YES ON THE MELODY DIAGNOSTICS, WITH THE
REGISTER HOLD MET — AND THE AGGREGATE IS THE LEAST INFORMATIVE LINE IN THE RESULT.**
| diagnostic | ARM 1 (n=9) | ARM 2 (n=16) |
|---|---|---|
| melodic salience, median (best) | 0.528 (1.000) | 0.440 (0.538) |
| leap of a sixth or larger delivered | 4/9 | 2/16 |
| the brief's licensed leap, within 1 semitone | 4/9 | 2/16 |
| the head's interval figure present | 4/9 | 0/16 |
| register match cosine, median | 0.739 | 0.704 |
| motif 3-gram Jaccard, median | 0.000 | 0.000 |
| note count, median | 17 | 15.5 |
The register hold is what makes the win admissible rather than merely bigger: arm 2's own median
register match is the floor, arm 1 clears it, and the judge is built so that a melody win below that
floor returns ARM 2 anyway (fixture J14).
THE DECOMPOSITION, WHICH MOVES THE ANSWER. Arm 1 mixes three operations and they do not behave
alike:
register match anywhere in the run — leap 4/4, head figure 4/4, held-region correlation 1.0000
on every one.
figure 0/5. Statistically indistinguishable from arm 2, and slightly worse on register.
Conditioning on authored audio through lego or cover does nothing for the melody.
AND THE REPAINT WIN IS INHERITANCE, NOT COMPOSITION. The head figure sits inside the region the
repaint was told to hold, so a whole-file measurement partly re-reports the input. Scoring the
generated region alone — the model's own contribution, held region excluded — gives, on all
four repaints: motif 3-gram 0.000, no head figure, no mirror of it, no leap, 2-9 note events
across 8-24 seconds, salience 0.000-0.433. The generator holds what it is told to hold, exactly,
and then writes texture. It does not continue the tune. ab_judge.py's fixtures J15-J17 are a
splice control built to make this impossible to miss: authored audio followed by unrelated texture
must still fire the whole-file head test while the generated-region test comes back empty.
**Q11 — A LANDED READING WITHDRAWN ON RE-MEASUREMENT, RECORDED HERE BECAUSE THE CORRECTION MATTERS
MORE THAN THE TIDINESS.** §6.0 Q9 concluded that a repaint's new half "shares real melodic material
with the statement", on interval 3-gram Jaccard 0.333 / 4-gram 0.238. That comparison ran over the
whole file, including the 16 seconds both files shared byte-for-byte, so the held region
contributed its own n-grams to both sides. Re-derived at HEAD on the same landed pair:
| scope | 3-gram | 4-gram |
|---|---|---|
| whole file (what Q9 used) | 0.3333 | 0.2381 |
| held region excluded | 0.1250 | 0.0000 |
The number was right and reproduces exactly. The reading was not: 0.1250/0.0000 is
motif_compare.py's own *WEAK AND UNSAFE* band, not *MOTIF MATERIAL SURVIVES*. Q9's finding is
therefore corrected to: **repaint conditioning HOLDS a motif — measured twice now, at 0.9998 and at
1.0000 — and does not CONTINUE one.** motif_compare.py now computes the overlap on the generated
region and decides the verdict there whenever a held region is declared (7/7 new fixtures, armed in
both directions: a splice must be refused the word variation while its whole-file overlap still
reads 0.857, and a genuine transposed variation must be recognised at 1.000 on the deciding block).
The landed family_report.json verdict flipped on re-run and PICKS.json demoted that artifact
from "variation" to "sibling render" through a branch written for exactly this case that had never
fired.
WHAT IS PROMOTED, AND IT IS NARROWER THAN THE LEVER THAT WAS TESTED. Promoting
"authored-head-cell + lego/repaint" whole would overstate the measurement by the width of the
decomposition above. What the evidence supports:
of the brief's licensed leap and its head figure on authored material, against 12.5% and 0% from
the strongest text-only prompt this lane can write, over sixteen candidates. The tune is authored;
it is not drawn from a prompt. harness/music_gen/author_head_cell.py is the tool, its self-test
checks each constraint against the notes rather than against a comment, and it caught three real
compositional defects on the way in (a response whose registral reset was wider than the licensed
leap; a second deviant arriving from the loop join rather than from the tune; a negative control
that contained what it was controlling for).
cover overwrites; lego does not carry. The arrangement problem is open.
AUTHORED_HEAD_CELL — an authored score in a sketchrealisation.** The site serves it as *"authored score — sketch realisation"* and the tier note says
the tune is the deliverable and the orchestration is not finished. The prompted ladder winner is
demoted to superseded_statement rather than deleted, because deleting it would erase the
comparison that made the promotion legible.
THE NEXT LEVER, NAMED HONESTLY, AND IT IS NOT THE ONE THIS RUNG ASSUMED. Going in, the open
problem looked like a *generator melody* problem. Measured, it splits in two and only one half is
still open:
1. Melody: solved by authoring. No further generator work is required to get the brief's contour
into an artifact. The twelve canonical themes can be composed the same way today.
2. Orchestral realisation: the real blocker, and it is a RENDERER problem, not a generator one.
The authored artifact sounds like additive synthesis because that is what it is. ACE-Step cannot
be handed an authored score and asked to play it — repaint re-decodes the held region unchanged
(residual 1.2%; it is a VAE round-trip, not a re-orchestration) and cover discards the tune. So
the next rung is rendering authored MIDI to orchestral audio — a sampled-orchestra or
synthesis path — with ACE-Step kept to the beds and ambience where §6.0 already measured it
strong. This also re-frames §0's concert-notation deliverable as *closer*, not further: an
authored score is already the thing an orchestra could read, which the prompted lane never was.
3. Still available and untested: Q4's lever (c), a genuine melody-conditioning path. It is now a
*lower* priority than (2), because authoring already produces the melody conditioning was wanted
for.
Two smaller measured facts, recorded rather than dropped. (a) §6 item 3's generation-time VRAM
peak, declared UNMEASURED on the measured close, is now measured: the armed sampler reported
29381-31993 MB peak across the arm-2 XL-SFT ladder, 53-99 samples per call, on a card already
holding 12.6 GB of another lane's work. The exclusive-overnight default is reinforced, not softened —
the ceiling was reached with under 1 GB of headroom. (b) **Loop seam is a live problem for the
authored method**: repaint artifacts are 0/4 seam-clean and the authored cell itself fails the rig's
seam heuristic, because a sentence whose head and tail differ by design wraps badly. Per §6.0 Q2
that is a per-artifact gate, so it does not block the promotion — it is a named item for the
realisation rung, which is where a loop point should be composed anyway.
What this block is and what it explicitly is not. Section 8 closed by naming the next lever:
*rendering authored MIDI to orchestral audio — a sampled-orchestra or synthesis path — with ACE-Step
kept to the beds.* Fourteen authored head cells now exist in notes (commit be9d9520), so that rung
has real input and the blocker is no longer hypothetical. This block is the SURVEY and the LICENCE
READ for it. **Nothing is adopted, nothing is downloaded, and not one sample is rendered through any
of it.** The lane's own rule is the reason: no renders through unread licences. What follows is the
read, so that the rung which runs next runs on a licence somebody has actually looked at.
The reads are FACE-VALUE, per the standing licence law. A licence is read for what it says, not
for what a cautious reading could imagine it saying; AI tooling on properly licensed assets is normal
use; and only an EXPLICITLY TRIGGERED prohibition is flagged. One of the five options below is
refused on that test, and it is refused because its prohibition really does fire.
A. CC0 SAMPLE SETS THROUGH AN OPEN SAMPLER — VSCO 2 Community Edition and VCSL, played by sfizz.
Versilian Studios' VSCO 2 Community Edition and the broader Versilian Community Sample Library are
both released under CC0 — a public-domain dedication: no royalties, no attribution requirement,
no restriction on commercial use, and the publisher states outright that anyone may build on them.
VSCO 2 CE is a chamber-scale set of roughly three thousand samples at 24-bit/44.1 kHz. The player is
sfizz, BSD-2-Clause, explicitly designed to be embedded and driven as a library, with
permissively licensed dependencies throughout. **Triggered prohibitions: none, and there is nothing
in this chain to trigger** — CC0 carries no conditions, and BSD-2 carries only a notice requirement
that applies to redistributing the software, which rendering audio does not do.
B. A COMMERCIAL FREE-TIER LIBRARY — Spitfire Audio's BBC Symphony Orchestra Discover.
Read at face value, the EULA grants what this lane needs: the sounds may be used inside newly created
recordings, commercially, with no ongoing payment, and games and sync are not excluded. Two clauses
matter and they point in opposite directions.
libraries for the purpose of training AI systems or large language models without written consent.
Rendering an authored score through a sampler is not training — it is the creation of new music,
which is the licence's own stated permitted purpose. Under the licence law that is a face-value
read, and the conservative reading that would refuse a sampler because a factory is holding the
mouse is precisely the error the law names.
redistribution of the samples, or of derivatives usable *as samples in a sampler*, is forbidden. So
the rendered stems ship and the sample content never does; no Spitfire byte enters this repo; and —
this is the part that binds the other arm of the pipeline — **no Spitfire sample may ever be a
conditioning, fine-tune or LoRA input to ACE-Step or any other generator.** That is the same rule
the Sonniss bundle already imposes, arriving from a second direction, and it travels on the row the
way the Sonniss rule does rather than sitting in a note nobody was instructed to satisfy.
C. SONATINA SYMPHONIC ORCHESTRA — REFUSED, on a prohibition this product genuinely triggers.
SSO is distributed under Creative Commons Sampling Plus 1.0, a licence Creative Commons
retired in 2011 and no longer recommends. It permits transformative sampling and collage, and it
permits noncommercial sharing of verbatim copies — but it excludes **all advertising and promotional
uses**, except promotion of the derivative work itself. A game soundtrack is used to promote the
game: trailers, store pages, festival reels. That is promotional use of a work which is not the
derivative being promoted, and it is exactly the explicit trigger the licence law says to flag.
Not adopted. Recorded here rather than dropped, because SSO is the library a search for free
orchestral samples returns first, and a later lane would otherwise re-discover it and re-decide it.
D. SYNTHESIS — Surge XT, or extending the additive engine that already exists.
Surge XT is GPL-3.0: the copyleft binds the software, not the audio a user produces with it, so
there is no output-licence question at all. Extending author_head_cell.py's own additive carrier
costs no licence and no download. Triggered prohibitions: none. What it does not do is sound like
an orchestra, which is the entire point of the rung — so this is the honest floor, not the answer.
E. THE SOUNDFONT CLASS — FluidR3_GM through FluidSynth.
FluidR3_GM is MIT (Frank Wen) and is usable for personal or commercial composition; the author's
additional guidance concerns bundling the soundfont itself with commercial software, which shipping
rendered audio does not do. FluidSynth is LGPL-2.1. Triggered prohibitions: none. The
limitation is quality rather than licence — General MIDI sits well below the AAA bar — but it is
deterministic, scriptable and free, which makes it valuable as a reference render the battery can
run against on every commit even if no player ever hears it.
1. RUNG ONE, adoptable with nothing to flag: sfizz (BSD-2) driving VSCO 2 CE and VCSL (CC0).
The whole chain is public-domain or permissive, so every licence record it produces reads
unconditional rather than conditional; it is scriptable and headless, which the battery needs; and
it lifts the sketch from additive synthesis to real recorded instruments in one step. This is the
rung to build.
2. **RUNG TWO, on evidence rather than on schedule: BBC SO Discover for the cues that need the larger
sound**, with the licence re-read and hash-pinned at download time under the existing
DL_0008/DL_0009 precedent, and the no-conditioning rule carried onto every record.
3. REFUSED: SSO. KEPT: the additive sketch, as the honest tier below both.
What the rung must MEASURE, so it cannot be declared done by ear. The composition battery is the
acceptance test and it does not care what produced the audio: an orchestral realisation of an
authored cell must still pass C11 (it reads as a melody, not a texture), C12 (the authored deviant is
recoverable from the render), C13 (the notes sit on the grid the declared tempo implies) and C14 (no
baked fade), against the same negative controls. A realisation that loses the tune inside a beautiful
patch fails C12 — which is exactly the failure a listening-only review would call a success. And the
loop seam should be composed at this rung rather than repaired at it: section 8 named it as a
live per-artifact problem, and a realisation pass is the first point at which a loop point can be
written into the arrangement instead of hunted for in the waveform.
The strongest objection to all of the above: none of these is a professional orchestral library,
so rung one buys realism and does not buy the AAA bar, and a later decision to license a full library
would make some of this work disposable. The answer: the deliverable of the authored method is
the SCORE, and every path above consumes the same MIDI. Changing the renderer changes nothing
upstream of it, which is the whole reason the authored method was promoted over the arranged one.
Sources read for this block (all read 2026-08-05, face value, no download): Versilian Studios
VSCO 2 Community Edition and VCSL product/licence pages (CC0); github.com/sfztools/sfizz
(BSD-2-Clause); Spitfire Audio End User License Agreement (commercial-use grant; section 10
AI-training restriction; redistribution prohibition); Creative Commons Sampling Plus 1.0 legal code
and its retirement notice; github.com/peastman/sso (SSO licence statement); FluidR3_GM README
(MIT, Frank Wen) and FluidSynth (LGPL-2.1); Surge XT (GPL-3.0).
What this block is. §8.1 was a licence read that adopted nothing. This is the rung it recommended,
built and measured: sfizz 1.2.3 (BSD-2) driving VSCO 2 Community Edition and VCSL (both CC0),
rendering the fifteen authored scores that commit be9d9520 landed in notes. Nothing upstream moved.
Every note of every sentence is played exactly as composed; the only material this rung adds is a
written loop turnaround, which is arrangement.
The stack, pinned rather than described. 442 files, 729 MB, every one carrying its own sha256 in
build/audio/stack/sfz_stack_manifest.json; both sample sets pinned by COMMIT SHA with a probe that
reports whether the branch has moved since; three licence bodies stored verbatim — docs/licence_records/sfizz/LICENSE
(BSD-2, pinned at the release tag), docs/licence_records/vsco2_ce/LICENSE and
docs/licence_records/vcsl/LICENSE (both CC0), alongside docs/licence_records/vcsl/README_sfz_branch.md.
Two things the fetch FOUND rather than assumed: **VCSL's sfz branch carries
no LICENSE file at all**, so the row pins the CC0 legal code on master *and* the licence statement in
the README of the branch actually downloaded — which is why that README is stored as a licence record
in its own right rather than paraphrased — and an
upstream packaging defect — the bowed psaltery's SFZ references its samples beside itself while
they are nested inside the Italian harpsichord's folder, and sfizz loads such an instrument without
complaint and renders SILENCE. The repair is declared, fires only on a 404, and is positive-controlled
by a smoke render that returns real energy. Recorded as DL_0016-DL_0018.
**AND THE RECORD OF THAT REPAIR WAS ERASING ITSELF, which a cold re-read caught by reading the landed
manifest instead of the prose about it.** The defect is real — the raw path 404s at the pinned commit
and the declared replacement returns 200 — but the manifest at HEAD recorded
upstream_path_repairs_fired: 0 for both sets and not one of the eleven affected sample rows carried
a repair key. The cause is that a repair is discovered on a 404 at DOWNLOAD time, so on any re-fetch
that finds the files already on disk the download branch never runs, the in-memory repair map comes
back empty, and the counter and rows silently reset to zero. A record whose whole purpose is to make
an upstream defect VISIBLE was quietly deleting its own evidence on every subsequent run. Two fixes:
the prior manifest's repair rows are now carried forward onto files that are skipped, and the fetch
REFUSES outright if sample rows match a declared repair's needle while none of them records the
repair — so the next time this goes to zero it fails loudly, and if upstream ever repackages, the
same refusal is the prompt to retire the declaration rather than leave a stale one in place.
The subset is a CARE decision before it is a download economy. VCSL ships kalimba, mbira, balafon,
darbuka and dan tranh. None of the fifteen artifacts is a region cell; the family heads' allow-lists
say the register they may enter is *each node's own register card*, and those cards belong to region
cells that are explicitly not composed, because the care read for that material has not been run. None
of those instruments is downloaded, and harness/music_gen/sfz_palette.py says why on the row.
Of the 23 instruments that ARE downloaded, the fifteen cards use 17; the six unused are the
glockenspiel, hand chimes, the sustained horn, the vibrato oboe, the bowed psaltery and the tubular
bells. K10 prints that list on every run and the run report carries it, which is how the rung's own
commit message came to say "4 of 23" while its self-test printed six on the same tree — a number
typed once into prose beside a number the tooling derives. The derived one is the record.
THE CARDS ARE HONESTLY UNEVEN — 7 OF 15 REFUSE. The exemplar is the East African Rift Valley card,
whose first sentence is that no instrument is attested in that corpus and whose last is that none may
be invented here. Six family allow-lists defer to cards of exactly that kind, so those heads are
realised on the register-neutral orchestral base — a string frame and a string carrier, claiming
no tradition — and the refusal travels onto the artifact, the record and the site tier note. A second
refusal kind covers the one card that DOES name instruments and may not have them:
FAM_THE_LIVING_VOICE names the Flores gong-waning and the Bali gamelan, both under their own
never-sampled care line, and reaching for an orchestral tam-tam to stand in for a gong-waning is the
"generic world-music wash" the brief pack calls a care failure rather than a taste note. **Seven
cards DEFER part of their register** — a per-age-stage re-voicing, a per-region migration, per-region
re-voicings, the local apparatus registers, a fairy-realm register the realm programme holds until
slice-end, and the ceremonial percussion two rows admit only at an ACTIVATION event — and each says
which half is missing rather than approximating it. Refusal and deferral are INDEPENDENT, and the
correction that made them so is FAMILIAR_BOND: its row asks for two things and the card answered
one. The twenty-two per-species carriers are its REFUSAL — they were previously miscounted into this
deferral list — and the row's second clause, "ceremonial percussion at bond-quest activation", was
neither realised nor deferred. Half a row was silently dropped, which is exactly what the deferral
class exists to prevent, and the identical clause on HOUSE_PARENT_HELD was already carried
correctly. It is now a declared deferral, and the counts live in the self-test's own printout rather
than in a hand-typed sentence: the module docstring said "four" while K09 printed six on the same
run, so the prose no longer carries a number that nothing checks.
THE LOOP SEAM IS COMPOSED AT THIS RUNG, WHICH IS WHAT §8.1 ASKED FOR. Each theme has a written
turnaround under a rule it names for itself: a dominant for the triadic tunes; an open fifth on the
fifth degree for the family whose identity is the absent third, because a dominant triad would destroy
it; a quartal chord whose root moves a fourth for the family built of stacked fourths; the pedal held
under the preparation; the two-chord cycle's own next member held across the join; and, for the
pre-gate cue whose card forbids a swell, **silence, held deliberately, because leaning anywhere is what
would make a listener notice it**. Four score-level teeth check them — no second deviant at the wrap,
no new pitch class, the preparation rule, the harmonic signature — and a fifth was minted when the
measurement demanded it: S6, a seam may not CONTRADICT its own theme's deviant. Its two fixtures
are both real landed defects: a turnaround entering ON the downbeat in the one theme whose deviant is
that it never does, and a turnaround entering on beats 1 and 3 in the theme whose cell always enters a
beat early. The second is the one that matters — it made the deviant **undetectable on the rendered
audio** while every other check stayed green.
THE SEAM'S SECOND HALF LIVES IN THE WAVEFORM AND IS MEASURED. A turnaround makes the loop point
musical; it cannot make it continuous, because the file still stops at the bar line and the
reverberant tail a live orchestra would still be sounding is thrown away. So the overhang is not
discarded: sfizz is allowed to ring past the written end and those samples are ADDED ONTO THE HEAD of
the artifact. It is a sum, not a fade, and C14 still binds and still passes — now with both of its
arms armed, which was not true when this was first written (see correction 5 below). Measured on
JOURNEY_WORLD, the wrap discontinuity against the median interior bar line fell from 28.99 to 7.35
when the release was carried over; on the rig's independent mel-distance ruler the same wrap takes
HOME_AND_LOSS from 5.804 to 0.067 and PATTERN_UNNAMED_PREGATE from 2.725 to 0.888, which is a
second instrument agreeing that the wrap is doing real work. The turnaround's own spectral contribution is small and on several
themes slightly negative, and that is reported rather than hidden: spectral flux is a TIMBRAL ruler and
the turnaround is a HARMONIC device, so S5 gates on the wrap and the turnaround is checked where it
lives — over the notes, in S1-S4 and S6.
THE ACCEPTANCE TEST IS THE BATTERY ON THE MIXDOWN, AND IT DID ITS JOB BY FAILING. §8.1 wrote the
failure mode down before the rung existed: *"a realisation that loses the tune inside a beautiful patch
fails C12 — which is exactly the failure a listening-only review would call a success."* That is
precisely what the first mix did. JOURNEY_WORLD read as a melody (salience 0.9738) and still failed
C12: the tracker recovered 60 notes from a 44-note artifact, the extra sixteen being octave jumps
between the carrier and the string frame, with wide intervals of 12, 19 and 24 semitones drowning the
one authored leap the detector exists to find. Four settings were measured across three themes and the
frame moved from -9 dB to -18 dB, at which the spurious wide intervals stop and the authored leap
is found. A listening review would have called the first mix the better one.
AND THE NOTE-COUNT HALF OF THAT CLAIM WAS OVERSTATED, WHICH A COLD RE-READ CAUGHT. The sweep table
said -18/-16 was the setting "at which the tracker recovers the written note count exactly", and at
HEAD it is exact on 6 of the 15 artifacts. The sweep's salience column reproduces exactly (1.000 /
0.942); its note-count column did not survive the five carrier replacements and the added turnaround
and release wrap, so the same three themes now read 45, 23 and 35 against 44, 22 and 34 written. The
over-recovery signature the mix fix was named against is therefore still present on three published
artifacts — HOME_AND_LOSS 58 against 40 written, HOUSE_PARENT_HELD 43 against 15, and
PATTERN_UNNAMED_PREGATE 25 against 12. It fails nothing, and the reason is a real scoping decision
rather than an accident: those three carry METRIC deviants, and C12 for the metric classes reads the
soloed carrier stem rather than the mixdown, so their whole-mix count is a statement about how
much the ACCOMPANIMENT hands the tracker and not about whether the tune survived — on the mixdown
itself what carries them is C11's predominance number. The correction is that this is now REPORTED
rather than inferable: written_notes and over_recovery_ratio ride on every artifact's diagnostics
and are summarised in the run report, and the sweep table is labelled as the historical reading it
always was.
**THE RULERS THAT HAD TO BE RE-SCOPED — six of them, and the last two were found by a cold re-read
AFTER the rung landed.** The first four corrections are about SAMPLED INSTRUMENTS rather than about
the music: the additive sketch had a 45 ms attack and one steady pitch per note, and a sampled
orchestra has neither. The fifth and sixth are about the MEASURING APPARATUS itself — a positive
control that could not fail, and a second ruler whose disagreement was going unreported.
1. ATTACK LATENCY, MEASURED. Notes written at beats 0.5 / 2.5 / 4.5 / 6.5 came back from the pitch
tracker at 0.789 / 2.949 / 4.992 / 6.920 — late by 0.29-0.45 beats, which is 145-225 ms. C13 now
removes the measured constant offset and reports it, with the 25%-wrong-tempo control computed under
the same correction, because a constant offset cannot fix a tempo error while a tempo error
accumulates across the artifact. JOURNEY_WORLD: median onset distance 0.0217 beats at the
declared tempo against 0.5016 at a tempo 25% wrong.
2. METRIC DEVIANTS READ THE CARRIER'S ATTACKS. offbeat_only measured on pitch-track note starts
reported 75% of the head's onsets sitting on a beat, in a family where every note is written off it.
And a metric property measured on the WHOLE MIX cannot separate the melody's entries from the
accompaniment's bar lines — which is exactly what made HOUSE_PARENT_HELD's early entry unfindable.
The negative control is measured identically, and C11 still proves on the whole mix that the carrier
is the predominant line.
3. ADJACENT SAME-PITCH SEGMENTS ARE COLLAPSED before any interval-shape detector runs. A
predominant-pitch tracker segments on pitch CHANGE, and a sampled string section held for two beats
at 68 bpm makes it close and reopen a segment at the same pitch: FAM_THE_OTHER_SELF's head came
back as [3, 4, 0] and FAM_THE_DEEP_WATER's as [-5, 0], so a self-inversion test and a
descending-triad test were both reading a segmentation artefact as the composition's shape.
Collapsed, those two recover 22/22 and 16/16 notes exactly and both deviants are found. The
collapse is applied identically to the negative control, and never to reciting_repetition, whose
property IS real repeated tones.
4. THE RECITATION TEST BECAME A SEPARATION RATHER THAN AN ABSOLUTE. The sketch's threshold of 2.0
was itself derived from a measured pair (3.5 authored / 1.5 de-repeated control) and it does not
transfer: a sampled cello's own bow attacks put the control at 2.2 while the artifact sat at 4.0.
The rule is now the separation the check was always trying to be — 1.5x over its own control, which
the sketch's own pair (2.33x) also clears — and it cannot be satisfied by a loud renderer, because
the control is rendered the same way.
5. AND THE FADE-IN PROBE BECAME A BOUND RATHER THAN BIT-EQUALITY. C14 asks whether an amplitude
ENVELOPE was applied to the file, and it answered that by comparing the opening samples of the
artifact with the opening samples of its own +1-statement probe. Inside one renderer that is
exact. Across two SEPARATE sfizz invocations rendering different-length MIDI that share a prefix,
it is not: three consecutive themes in a full run came back unequal and every one of them passed
the same probe when run alone, differing in the last bits rather than in the music. The stronger
claim is measured where it IS true — --determinism renders the same plan twice and reports max
abs difference 0.0 on three themes.
**AND THAT PROBE'S OWN POSITIVE CONTROL WAS NOT ARMED, WHICH IS THE MOST SERIOUS THING THIS BLOCK HAS
HAD TO CORRECT.** C14 is two arms — the artifact must show no envelope, and a deliberately faded
control must show one, so that a clean read cannot be a check that is simply unable to speak. The
fade-OUT arm was armed. The fade-IN arm was not: it built its control by ramping the NORMALISED MONO
mix and comparing that against the UN-NORMALISED STEREO head, so what it measured was the normalise
scale (10.92x on JOURNEY_WORLD, 4.06x on FAM_THE_WRITTEN) and the mono fold. Removing the ramp
entirely left it separating just as loudly — 0.4269 and 0.2559 against a bound of 1e-3 — which is the
definition of a control that cannot fail. Two corrections, both measured across all fifteen:
only difference between the artifact arm and the control arm is the envelope. With the ramp
removed the control now reads 0.0 exactly on every one of the fifteen: the arm can fail.
here. Un-normalised head peaks across this roster span 0.00145 (HOME_AND_LOSS) to 0.2176
(FAM_THE_OTHER_SELF) — a factor of 150 — so one constant was a different test on every theme. At
the corrected construction with the old absolute 1e-3, a REAL half-second fade on HOME_AND_LOSS
measured 0.8x of the bound: the control could not clear the threshold it was controlling for.
Relative, every artifact scores 0.0 and every faded control scores 0.5315 to 0.7835, with
the bound at 0.05 between them.
6. **THE RIG'S OWN SEAM HEURISTIC DISAGREES WITH C14 ON TWO ARTIFACTS, AND IT IS PUBLISHED RATHER
THAN QUIETLY DROPPED.** Every artifact also carries rig_loop_seam_heuristic, the older ruler,
printed beside C14 and gating nothing. It reports seam_clean: false on 10 of 15 and
baked_fade_in: true on two PUBLISHED tracks, HOME_AND_LOSS and PATTERN_UNNAMED_PREGATE —
against those artifacts' own no-fade claim. The cause was measured rather than assumed, and the
first hypothesis was WRONG: re-rendering both with the release wrap switched off leaves the flag
unchanged, so the wrap is not it. The real mechanism is in the heuristic's own definition — it
calls a fade-in when the opening 1.5 s is quiet against the middle (head_rms/mid < 0.55) AND
rising (trend > 0.45). Those two artifacts measure 0.1428/0.737 and 0.4352/0.7726. That is a
description of a piece that begins softly and grows, which is composition: HOME_AND_LOSS is
the childhood register and PATTERN_UNNAMED_PREGATE is a pre-gate cue whose card forbids a swell
precisely because leaning anywhere would make a listener notice it. The heuristic cannot separate
a written soft entry from an applied envelope. C14 can, because it compares the head against the
same music rendered one statement longer — where an applied envelope differs and these differ by
0.0 — so C14 remains the authority and the heuristic remains reported. The seam_clean figure is
the same class: it is a MEL-SPECTRUM distance, a timbral ruler reading a harmonic device, which is
the split S5 already makes. Both counts are now in the run report so neither can go quiet.
**AND FIVE CARRIERS WERE REPLACED ON EVIDENCE, WHICH IS WHAT "re-orchestrate, never re-score" MEANS IN
PRACTICE.** The horn recovered all 38 of PROTAGONIST_THEME's notes and then a 39th an octave above the
last — as a horn note decays, its second harmonic outlives the fundamental — and borrowed_degree
refuses any wide interval anywhere, so the carrier moved to the bassoon, whose tenor register is what
the brief's "fragile solo statements" arguably wanted first. The oboe's upper partials turned
COMPANION_HEALERS_CHILD's 15-note artifact into 37 recovered notes with spurious pitches a nineteenth
above the line; the flute recovers 18 and finds the borrowed degree. JOURNEY_WORLD's carrier was
decided by RANGE before anything was rendered: three of the four carriers the old generation table
listed cannot play the score at all, because the tune spans A3-F#5, the head's licensed leap ARRIVES on
the F#5, and the sampled horn stops at F5. The same horn defect then took it off FAM_THE_TURNED as
well, where it put 12s and -14s into a 19-note family's recovered line and left the semitone
neighbour-and-return undetectable; a sustained trombone recovers 19 of 19 exactly and is the more
literal reading of "the court, the company, the station" anyway. And the held parent needed an
ARTICULATION rather than an instrument — a legato trombone cannot state where a metrically-early cell
begins, and its staccato twin (same instrument, same register, same card) recovers 7 of 7 entries, all
a beat before a downbeat.
WHAT IS PROMOTED. The tier is AUTHORED_SCORE_ORCHESTRAL_RUNG_ONE, served as *"authored score —
orchestral rung one"*, and it is not rounded up: the note on every page says outright that this is a
public-domain community sample library rather than a professional scoring session, and the four
doctrine diagnostics that need a human listener are still unrun. The picks belt gained R6 — one
theme is served at ONE rung, the highest whose battery is green, and the superseded sketch is refused
BY NAME rather than deleted, because deleting it would erase the comparison that makes the rung
legible. The site's Sound lede was split at the same time: it carried a hard-coded sentence saying
*"what you hear is a SKETCH realisation: additive synthesis, not an orchestra"*, which the orchestral
rung made false on every page it reaches, so provenance and REALISATION are now separate claims with
separately armed controls.
**AND THE SITE DID NOT ACTUALLY SERVE ANY OF THAT UNTIL NOW, WHICH IS THE FINDING THAT MATTERED
MOST.** The commit that landed the rung said the site served it. The BUILD did; the deployed site did
not. The box that ran the build had no encoder, so every track dropped out of the belt, the pages
served no music tiles at all, and nothing was deployed — leaving the live pages serving 13 and 11
sketch-realisation tiles per chapter and still carrying the exact sentence the rung retired. The
commit body said so honestly and the subject line did not, and under the standing phone-first rule a
tier that is not served is not landed. Two things closed it: the encoder question is now a
dependency (imageio-ffmpeg) rather than an assumption, and — the real fix — **the site's own Sound
lede control refused to call a sweep of zero pages clean.** It had swept zero and reported clean on
its first and only run, because a positive control that proves the PREDICATE can fire says nothing
about whether the sweep had anything to read. run_gates.py had already settled that class ("no
gates ran -- refusing vacuous PASS") and the site control was extended without it. A build that
serves no audio can no longer describe itself as clean, which is what let a silent site look like a
landed one.
THE PAIRING LAW WAS TRUE OF HALF THE ARTIFACTS. "Every artifact carries a paired licence +
generation record" held for the audio and not for the MIDI: the writer emitted a licence record for
both and a generation record for the audio only, so every *_ORCHESTRAL_RUNG_ONE_MIDI was
half-recorded — and the tier BELOW this one does emit *_AUTHORED_HEAD_CELL_MIDI.generation.json, so
this rung was also the odd one out. Nothing read the records back, which is why it survived. The MIDI
record is now written, and it is not a copy of the audio's: the MIDI has no renderer and no samples,
and its battery was measured on the AUDIO, so the record REFERENCES that measurement rather than
restating it as though this file had been measured. The claim is now enforced by O09, which pairs
every landed record id off against its opposite number and has a positive control; it went red on
first arming and named all fifteen missing records. The first version of O09 filtered on the TIER
string, which appears in no record filename — a predicate that could not match, reporting a clean
pairing over an empty set. That is the vacuity the check exists to catch, caught in the check itself.
THE STRONGEST OBJECTION, and it is §8.1's own. None of this is a professional orchestral library,
so rung one buys realism and does not buy the AAA bar. The answer is unchanged and is now demonstrated
rather than argued: the deliverable is the SCORE, every path consumes the same MIDI, and this rung
changed nothing upstream of the renderer. What it did change is four measuring instruments — and those
corrections travel to rung two whatever ends up playing it.