PIPE_AUDIO_MUSIC_2026-07-29.md

music/PIPE_AUDIO_MUSIC_2026-07-29.md

PIPE — MUSIC + SFX (5090 lane dossier)

Verified live 2026-07-29. VERIFIED = primary source fetched this pass (URL given). CARRIED = from

docs/pipeline_review/tech_research/AUDIO_STACK.md (2026-07-15), not re-fetched. UNVERIFIED-EXCLUDED =

could not be confirmed; never adopted. Ruled state honored: AIVA CLOSED-AS-SKIP, ACE-Step/YuE

benchmark-gated, HeartMuLa candidate, ~~composer-in-loop for canonical themes~~ **[SUPERSEDED

2026-08-05 — see the supersession block below]**, and the sourcing law

no sampling of consecrated performance / reference-COMPOSITION permitted (DESIGN_GAP_REGISTER:1317,

1789; adjudication 18).

0.0 SUPERSESSION — JOSH RULING 2026-08-05 (VERBATIM, SUPREME)

"Humans aren't composing the music… what the fuck. The entire music pipeline is yours."

And the correction that re-grounds the lane on this dossier: *"We already developed the music

pipeline… did you forget all the pipelines we created?"*

COMPOSER-IN-LOOP IS RETIRED. The factory owns the score end to end — melodies, hero themes,

leitmotifs, beds, stingers, orchestration, notation. Exactly one row of this dossier changes

(§1's Canonical themes (12) = composer-in-loop, annotated in place below). Everything else here

stands unchanged and binding: the 30-year bar and the hummability test, the per-culture /

per-period scoring law with its wordless-vocal default, the notation-and-concert deliverable, the

sourcing law, the Sonniss no-training/no-derivation constraint, the MusicGen/AudioCraft exclusion,

the VRAM concurrency law, and the DR-2 integration contract.

What replaces the human in the loop is not "nothing" — it is the measurement loop. The bar was

previously held by a composer's ear; it is now held by the factory's own candidate ladder,

instrumented scoring against the doctrine's §2.3 checklist, and fresh-context critics on captured

playback. A generated theme that cannot show its score trail has not passed anything. The honest

tier for everything this lane emits is GENERATED SCORE — ITERATION, never "final", and Josh

hears results on the chapter site's audio players and vetoes by exception.

The one carve-out that survives the retirement, because it is a care line and not a craft line:

§4's ATTENDED list is re-scoped rather than deleted. Cues whose cultural_depiction_required or

hard_line_anchors cell is non-empty, and every realm cue, still take an attended §17.1 read — the

attending party is now a fresh-context critic rather than a human composer, and the read still

cannot be delegated to a prompt. The 132 creature vocalisations remain a sound-design task, not a

generative one.

0. JOSH SCORE DIRECTION (2026-07-29 — BINDS this lane; ledger 28al)

Josh's quality mandate, given with the AAAAA reaffirmation and refined in his follow-up. This is

the creative bar every music deliverable in this dossier serves:

broad — SNES-era, Zelda, Undertale, Sonic, Super Mario 64, Skyrim, Halo, Pokemon, Animal

Crossing, Donkey Kong — scores that live outside their games. They endure on MELODY and

leitmotif, not production sheen.

cannot be hummed after one hearing, it is not the theme yet. ~~Composer-in-loop is the MECHANISM

for this bar~~ **(§0.0: the mechanism is now the candidate ladder + the measurement rig + critic

listen-checks; the AIVA closed-as-skip ruling stands)**; themes iterate until they pass, judged

with evidence like everything else.

and their time period, using instruments, chants, vocals where appropriate but mostly no vocals

other than hums chants and harmonies, ambient noises." Period-appropriate instrumentation per

region; vocals default WORDLESS (hums / chants / harmonies); ambient texture is part of the

score. This composes with the ruled sourcing law above (reference-COMPOSITION to the culture's

musical register, never sampling of consecrated performance).

VOLUME inside the per-culture register; the identity themes people hum in 2056 are the

composer-in-loop lane's deliverable.~~ **SUPERSEDED (§0.0): the split is no longer human-vs-model,

it is MODEL-vs-MODEL — XL-SFT + the 4B planner own the identity themes on the overnight exclusive

slot, 2B-turbo owns bed and variation VOLUME in shared windows.** The bar never routes down-tier:

a bed may be turbo-generated, an identity theme may not.

playlist at concerts, operas, and Broadway — the Harry Potter / Star Wars / video-game-legends

lineage (Distant Worlds, Symphony of the Goddesses, Symphonic Evolutions). IMPLICATION: the

hero-theme lane's deliverable is a COMPOSED WORK, not just rendered audio — real notation

(MusicXML/MIDI + orchestration) that a human orchestra can sit down and play. Rendered stems

serve the game build; the score serves the concert hall. Both come from the same composition.

to create this, understand the patterns, and how everything comes together." QUEUED ARTIFACT:

THE MUSIC COMPOSITION DOCTRINE — a deep-research foundation doc authored at the composition

program's start, BEFORE any hero-theme authoring: leitmotif architecture (statement /

transformation / combination across the 79 nodes and 22 threads), harmonic + modal language

per cultural register, orchestration and voice-leading practice, form, and the specific

techniques the exemplar lineage uses (how Zelda's overworld theme, Halo's chant, SM64's

bounce actually WORK). The doctrine is the composition lane's runbook; no theme authors

without it.

Consumers: the region-page Section 10 briefs (re-keyed 28ai), T0_Theme_Registry authoring, the

benchmark-day music evaluation rubric (hummability + register fit join the judged axes).

THE USER'S LEADS — RULED

LeadVerdictSource
"XL SFT"CONFIRMEDACE-Step/acestep-v15-xl-sft, licence field mit, 4B DiT, 50 steps, ~20GB on diskhuggingface.co/ACE-Step/acestep-v15-xl-sft
"XL Turbo"CONFIRMEDacestep-v15-xl-turbo, 8 inference steps vs SFT's 50huggingface.co/ACE-Step
"Excel Bass"CONFIRMED as XL BASEacestep-v15-xl-base, pre-trained only. The name is a mis-hear of *XL Base*huggingface.co/ACE-Step
...as "ComfyUI audio nodes"KILLED. These are ACE-Step 1.5 checkpoints, not nodes. The real ComfyUI audio nodes are EmptyAceStepLatentAudio, TextEncodeAceStepAudio, LatentOperationTonemapReinharddocs.comfy.org/tutorials/audio/ace-step/ace-step-v1
"ACE-Step 1.5 / UI"CONFIRMED — official Gradio UI (uv run acestep) and a REST API server (uv run acestep-api, port 8001) ship in-repogithub.com/ace-step/ACE-Step-1.5 README
"ARDY", "Kimodo"KILLED for this lane — both are NVIDIA motion-generation models (ARDY autoregressive text-to-motion, SIGGRAPH 2026; Kimodo kinematic motion diffusion, nv-tlabs/kimodo). Reroute to the animation lane; zero audio relevanceresearch.nvidia.com/labs/sil/projects/ardy/ · github.com/nv-tlabs/kimodo

Repo naming collision to record: ace-step/ACE-Step (the v1 repo, Apache 2.0, 4-min cap) and

ace-step/ACE-Step-1.5 (MIT, 600s cap, XL series) are *different repos with different licences*.

AUDIO_STACK cites 1.5 correctly; anyone searching "ACE-Step" lands on the Apache one. Pin the URL.

1. THE VERIFIED STACK

JobPRIMARYFALLBACKLicence
Bulk regional/chapter music (79 nodes)ACE-Step 1.5 XL-SFT + acestep-5Hz-lm-4B plannerACE-Step 1.5 2B-turbo (speed lane)MIT (LICENSE fetched: "MIT License", Copyright ACEStep 2026)
Adaptive stems~~ACE-Step 1.5 native multi-track / track separation~~ SUPERSEDED-BY-MEASUREMENT 2026-08-05 (§6.0 Q1): successive lego passes on acestep-v15-base. extract/lego/complete are TASK_TYPES_BASE — XL-SFT is excluded by the code, so the stem path and the quality path are DIFFERENT CHECKPOINTS — and extract returned four tracks at cross-stem spectral cosine 0.9998, i.e. the same content under four names. Layers are BUILT additively, never carved out of a mixdownDemucs htdemucs_6s (analysis + genuine separation)MIT / MIT
Canonical themes (12)~~composer-in-loop (ruled) — ACE-Step as reference-composition sketch only~~ SUPERSEDED 2026-08-05 (§0.0): ACE-Step 1.5 XL-SFT + acestep-5Hz-lm-4B planner, run through the candidate ladder + the measurement rig; the bar is held by measurement and iteration, not by a human ear2B-turbo candidate ladder; licensed orchestral library renderMIT
Bulk SFX / ambience / foleyMOSS-SoundEffect v2.0 (1.3B, DiT+Flow-Matching, 48kHz, 30s)ElevenLabs SFX endpointApache 2.0 (HF licence field apache-2.0)
SFX backbone librarySonniss GDC bundles (2026 = 7.47GB / 347 WAV; ~200GB historical archive)BOOM Library buyoutROYALTY-FREE-OWNED, no attribution
Runtime variation (footsteps/UI)UE 5.8 MetaSounds procedural DSPengine-native, unrestricted
Creature vocalisation (132)Krotos Dehumaniser 2 + Reformer Pro, sound-designer-in-loopCARRIED (repo §2.7); not a generative task

**Headline correction to AUDIO_STACK — ITSELF SUPERSEDED-BY-MEASUREMENT 2026-08-05; read §6.0 Q1

before routing any stem work.** The paragraph below was a README read, and the benchmark answered it

with numbers: AUDIO_STACK's §1.6 was closer to right than this correction was. ACE-Step 1.5 does not

give the adaptive-layering mechanism a separator would; it gives an ADDITIVE one, on a different

checkpoint, and the design consequence is written out in §6.0. The original text is kept below as the

reasoning that was superseded, not as live routing.

~~Its §1.6 "stem problem" — *"None of YuE/ACE-Step/HeartMuLa/Stable Audio natively output separated
stems"*, which made AIVA the only adaptive-music path — is now false. ACE-Step 1.5's own README
feature table lists "Track Separation — Divide audio into individual stems" and **"Multi-Track
Generation — Add layers like studio features"**, with corroborating reports of per-instrument
(kick/bass/snare/hats/perc/synth/pad) multitrack export. Since AIVA is CLOSED-AS-SKIP, ACE-Step 1.5
is the replacement adaptive-layering mechanism, not a downgrade. **Benchmark-day must confirm
whether multitrack emits true compositional stems or a bundled separator pass** — the whole Quartz
design leans on this.~~

What the benchmark actually found (§6.0 Q1, one paragraph so this table is never read alone):

extract does not isolate — four extracts off one mixdown measured mean cross-stem spectral cosine

0.9998, each ~0.987 identical to the mix, each carrying MORE energy than it (1.44–1.61×), and they

do not sum back (residual −1.29 dB). lego DOES work and is the mechanism the adaptive design is now

built on: the source survives in the output (correlation 0.47, spectral cosine 0.976) and the output

carries substantial content the source never had (residual energy ratio 0.76). So Quartz layers are

generated one per pass on a base checkpoint, at one generation per layer, and hero cues that need

stems render on base rather than xl-sft. Demucs htdemucs_6s remains the real separator and is

what the probe cross-checked against.

Demotions and exclusions (licence-read-before-first-GPU-hour):

2 sessions and states full songs (4+ sessions) need ≥80GB; last update June 4 2025 (~13 months

stale). On a 32GB card it buys ~2-3 minutes. Not a co-equal to ACE-Step.

HeartMuLa-oss-3B-happy-new-year (2026-02-13). The 7B is still an unticked TODO. AUDIO_STACK

Unknown #4 is now settled: no larger checkpoint exists.

a shipped path.

Community License, whose threshold I re-fetched verbatim: *"free for everyone, unless… you or your

organization generate over USD $1M… of annual revenue"* (stability.ai/license). Above it an Enterprise

Licence is required. Open Small = 0.5B / 11s / 44.1kHz; 3 Small SFX = 0.6B, and its own card's code

examples run sample_rate at 16,000 Hz — below game standard. Two independent reasons MOSS wins.

UE 5.8 (current) reality check: import accepts .wav/.ogg/.flac/.aif/.opus/.mp3, **all converted

internally to 16-bit WAV** — so generate at 48kHz/24-bit, deliver 16-bit, and never treat source bit

depth as a shipped-quality lever. New in 5.8: the Windows audio backend switched XAudio2 → WASAPI

(AudioMixerWasapi) — re-run any device/latency assumption from the old machine. MetaSounds + Quartz

both current in 5.8 docs.

2. INSTALL PLAN

D-1 finding — the audio lane needs NO WSL2. ACE-Step 1.5 advertises Mac/AMD/Intel/CUDA support and

ships a Windows-runnable Gradio+REST server; MOSS-TTS installs from a cu128 PyTorch wheel — the exact

sm_120 Blackwell stack the runbook already budgets. **Both run in runbook lane (b), Windows-native, in

their own venvs.** Only MOSS's optional vLLM-Omni/SGLang-Omni backends are Linux-leaning — use the

PyTorch backend on Windows and treat vLLM as a WSL2 option *if* throughput ever demands it. This removes

the audio lane from the D-1 fork entirely.

Disk placement (per the RULED layout). Weights + venvs + HF cache → Gen4 D: (HF_HOME,

HUGGINGFACE_HUB_CACHE, TORCH_HOME already re-pointed at runbook 1.5). Sonniss/BOOM sample libraries →

8TB SATA slot 4 (runbook 1.4 already rules this). Nothing audio touches the Gen5 hot path except the

imported .uasset under C:\dev\Humanity\Humanity\Content\Audio\.

Install order (each step's exit condition is an artifact on disk, not a claim):

1. D:\ai\acestep\ — clone github.com/ace-step/ACE-Step-1.5, own venv, uv run acestep → Gradio up.

2. Pin torch explicitly (torch==2.9.0+cu128) — the runbook's Puget-pack lesson about unpinned torch

applies here identically.

3. uv run acestep-apiport 8001. Add a Stage-6F firewall row (the runbook table lists UE, MCP

9315, MCP 8000, Ollama 11434, ComfyUI 8188 — 8001 is missing and will prompt at 04:00).

4. D:\ai\moss-sfx\pip install --extra-index-url https://download.pytorch.org/whl/cu128 -e ".[torch-runtime]"

from github.com/OpenMOSS/MOSS-TTS. First call compiles (torch.compile + Triton CUDA graphs) — budget

several minutes and do not read it as a hang.

5. Demucs (pip install demucs) — small, CPU-viable, fallback only.

6. ComfyUI ACE-Step 1.5 nodes are native (no custom node install): drop

ace_step_1.5_turbo_aio.safetensors into ComfyUI/models/checkpoints/. This is the *turbo* lane only —

XL-SFT quality runs through the native REST API, not ComfyUI. Route accordingly.

Thursday-night download list (measured unless marked):

ItemSizeWhere
acestep-v15-xl-sft (4× 4.99GB shards)~20.0 GBD: HF cache
acestep-v15-xl-turbo~20 GB ESTIMATE (same class)D:
acestep-5Hz-lm-4B planner~8 GB ESTIMATED:
MOSS-SoundEffect-v2.011.2 GBD:
ComfyUI ace_step_1.5_turbo_aio.safetensors~7 GB ESTIMATED: ComfyUI/models/checkpoints
Sonniss GDC 2026 bundle7.47 GB8TB SATA
Sonniss historical archive (archive.org)~200 GB, optional/overnight8TB SATA
Demucs htdemucs_6s<1 GBD: TORCH_HOME

~47 GB for the working set, ~67 GB with the turbo mirror, before the optional 200 GB archive.

3. THE INTEGRATION CONTRACT (DR-2)

Addressing — already ruled, already columned.

swap = path takeover (idempotent import). Teeth: QT-10 cue-fired telemetry non-zero, QT-AU

muted/wired fixture.

stereo, Bink Audio compression for ambience/music, PCM for short high-frequency cues.

Two verified state facts that change the lane's shape:

1. AUDIO_STACK Unknown #6 is CLOSED. It flagged that Theme_Registry carried no tempo/bar/stem-role

fields for Quartz. It now does: the live header carries `tempo_bpm, bar_length, per_stem_role,

ue_sound_asset_path`. 12 rows exist, **all empty on those four fields, and there are zero regional

rows.** The gap is population, not schema.

2. A real DR-2 five-state conformance gap on SFX. T0_Theme_Registry carries `generation_status,

generation_prompt_hash, last_generated_timestamp — but not regeneration_trigger`.

T0_SFX_Registry (30 rows, 20 cols) carries none of the four. Until the assetgen rail lands on

both, generated audio cannot express pending→in_progress→complete→rejected→regeneration_required,

and the prompt-hash regeneration trigger DR-2 depends on has nowhere to live. **This belongs in the

next batched registry-extension pass + one fidelity re-baseline in the same commit.**

Downstream wiring already landed (28-series): the §10.6 arrays exist —

region_footstep_sfx_id_ref_array, region_environmental_ambient_sfx_id_ref_array,

region_ritual_context_sfx_id_ref_array, plus chapter/boss/ability/vril_cast/quest_dialogue/weapon

arrays (docs/registry_extensions.json). Known data defect to repair: LW-7 — Flores'

region_environmental_ambient_sfx_id_ref_array holds Environment-Grammar ids (EG_0001;EG_0002),

not SFX ids.

Stem/loop contract into UE. Each adaptive cue delivers: (a) N separate loopable WAVs, one per

per_stem_role; (b) tempo_bpm; (c) bar_length; (d) clean loop points with no baked-in fades

ACE-Step's generation must be prompted for, and QA'd against, seamless loop boundaries, because Quartz

triggers layers at Quantization Boundaries and a baked fade destroys the seam. Import via

unreal.SoundFactory + AssetToolsHelpers…import_asset_tasks (the proven bridge —

Content/Audio/SFX/smoke_stone_door already landed this way). MetaSound graph assembly is scriptable via

MetaSoundBuilderSubsystem, but the Builder API has no variables support — expect hand-authoring in

the editor for anything past ~3-4 layers, which is consistent with sound-designer-in-loop for hero cues.

4. THE AUTONOMY CONTRACT

Dispatch. The director never names a model in a work order. It selects registry rows —

T0_Theme_Registry rows with generation_status IN (pending, regeneration_required), or T0_SFX_Registry

rows reachable from a region's §10.6 arrays — and hands them to Tools/compose_audio_prompts.py (the

missing Hop-3 script, ranked BUILD-NOW), which reads the region's Section-10 cue table + Section 17

substrate and emits per-cue prompt objects with a generation_prompt_hash. ~~The lane then POSTs to

ACE-Step's REST server on :8001 (music)~~ **CORRECTED 2026-08-05 (deviation 2 of 2, §6.0): the music

lane calls ACE-Step IN-PROCESS — harness/music_gen/generate.py imports acestep.handler/

acestep.inference directly, so one handler load serves a whole ladder and offload_to_cpu is

controllable per run. No acestep-api server stands (verified: nothing listening on :8001). The

:8001 REST path in §2 step 3 remains an available option, not the dispatch this lane uses** or calls

MOSS locally (SFX), writes the WAV, imports, and writes back through the T0.13 manifest. **Head-of-chain blocker: the Section-10 machine cue table does not exist

yet** — Flores' Section 10 is 100% prose, so check_build_readiness.py lane_audio reports BLOCKED-ON and

the lane has *no addressable unit of work*. That template edit + Flores backfill is the single thing

gating audio autonomy, and it is neither 5090-gated nor vendor-gated.

Scheduling class + VRAM budget (the concurrency law).

ConfigPeak VRAMClass
ACE-Step XL-SFT + 4B LM (best quality)~24-28 GB ESTIMATE (README tier: ≥24GB)OVERNIGHT, EXCLUSIVE — cannot co-reside with the UE editor or the image lane on a 32GB card
ACE-Step 2B-turbo + 0.6B LM6-8 GB (README tier)WINDOWED — safe alongside UE
MOSS-SoundEffect v2.0 (1.3B)<8 GB ESTIMATE — MEASURE day oneWINDOWED, the always-available bulk-SFX worker
Demucs htdemucs_6smodest / CPU-viableALWAYS-ON

The rule this lane commits to: **the XL music model owns the GPU alone, on the overnight slot, after the

soak** — every bulk-SFX and iteration pass runs on the small models in shared windows. A music batch must

never be the reason a QA soak or a UE cook stalls.

Quality gate — promote/demote. Benchmark day scores a fixed set (2-3 regional themes + 1 hero-adjacent

cue + 6 SFX classes: wind bed, ocean, jungle, footfall-per-surface, weapon impact, ritual context) on:

instrumentation fidelity, cultural-substrate accuracy against the Care Doctrine, loop-seam cleanliness,

stem separability, peak VRAM, wall-clock, and whether the licence clears. A model is promoted to

PRIMARY only on that evidence; a model that fails the loop-seam or stem test is demoted to reference-only

regardless of how good the mixdown sounds. Output QA runs through the standing audio teeth (QT-10 event-

fire, QT-AU fixture) plus a listen-check by a critic on captured playback — never a described claim

(the visual-verification directive's audio analogue).

Unattended vs attended. UNATTENDED: bulk regional ambience, per-surface footfall, weapon/item foley,

combat cue variants, and all regeneration triggered by prompt-hash change. **ATTENDED / human-in-loop

always: ~~the 12 canonical themes (composer-in-loop is ruled)~~ SUPERSEDED (§0.0) — the 12

canonical themes are now UNATTENDED-GENERATED and ATTENDED-JUDGED: the ladder runs unattended, and a

fresh-context critic scores the picked candidate against the doctrine's §2.3 checklist on captured

playback before it is served anywhere**, the 132 creature vocalisations (Krotos,

sound-designer-in-loop), any cue whose cultural_depiction_required or hard_line_anchors cell is

non-empty, and every realm cue — the §17.1 read cannot be delegated to a prompt.

The sourcing law, operationalised. *No sampling of consecrated performance; reference-COMPOSITION

permitted.* Concretely: the prompt composer may pass **instrumentation, mode/scale, ensemble shape, and

era descriptors drawn from instrumentation_substrate / cultural_substrate; it may never** pass an

audio file of a consecrated performance as a conditioning input, and no ACE-Step LoRA is ever trained on

one. Two licence walls reinforce this from the other side: **Sonniss's bundle licence explicitly prohibits

using the audio to train AI/ML**, and EastWest's EULA does the same (CARRIED) — so the library backbone is

direct-use-only and can never feed a fine-tune. Record the generating model + licence in every artifact:

ACE-Step-1.5-XL-SFT (MIT) or MOSS-SoundEffect-v2.0 (Apache-2.0), written into the row alongside

generation_prompt_hash.

5. JOSH-MINIMUM

Genuinely small — AIVA closing as SKIP removed the only account/negotiation item in this lane.

1. Krotos Dehumaniser 2 + Reformer Pro — ~$798 one-time, his card (creature vocalisation; the only

purchase this lane needs). Optional, deferrable to the creature pass.

2. BOOM Library buyout — $99-199, optional, only if benchmark day shows Sonniss+MOSS leaves gaps.

3. Sonniss GDC bundle download — free, but the site is email/form-gated; one sitting, or delegate.

4. Nothing else. No API key to mint (the ElevenLabs key already exists and works —

Tools/generate_sfx.py produces WAV/MP3 today), no licence to negotiate, no consent gate (that is the

voice lane's, already ruled).

6.0 BENCHMARK DAY, PART ONE — 2026-08-05: THE MUSIC QUESTIONS ANSWERED, THE SFX AND CARE QUESTIONS NOT RUN

**What this section covers and what it does not, stated first so the title can never be read as more

than it is. ANSWERED by measurement: §6 questions 1, 2, 3, 4** — the music-generation questions.

NOT RUN: §6 questions 5, 6, 7, 8 — see the annotations on that list for each one's reason. Against

§4's declared scoring set ("2-3 regional themes + 1 hero-adjacent cue + 6 SFX classes") what actually

ran was one hero-adjacent ladder and 44 ambience beds: zero regional themes (all 12 region

cells are correctly HELD pending the §17.1 read) and zero SFX classes (MOSS is not installed).

The Care-Doctrine question is therefore still open, and nothing here closes it.

Answers below are measurements, each with the artifact that produced it. Tooling:

harness/music_gen/{generate,feature_rig,rank,stem_probe,motif_compare,emit_ladder,emit_picks,to_midi}.py.

Evidence: build/audio/benchmark/stem_probe.json, build/audio/generated/journey_world/ranking.json,

.../family_report.json, and the per-artifact licence + generation records under build/audio/records/

— **65 artifacts, 65 record pairs, zero orphans in either direction, every sha256 re-verified against

the bytes on disk** (the four Demucs separator outputs are included as an explicit DERIVED-ANALYSIS

class carrying licence_class: previz_only, because the pairing law has no analysis exemption and a

model output without a record is the pattern that ships an unrecorded asset the first time a derived

pass produces something that DOES ship).

INSTALLED, MEASURED, AND TWO DEVIATIONS FROM THE INSTALL PLAN RECORDED HONESTLY. ACE-Step 1.5 is

live Windows-native at D:/audio/acestep in its own venv: XL-SFT (19 GB), acestep-5Hz-lm-4B

(8 GB), the main bundle with the 2B turbo DiT + VAE + Qwen3 text encoder (~6 GB), and

acestep-v15-base (4.5 GB) for the stem tasks — 36 GB of checkpoints on the ruled D: placement,

HF_HOME already re-pointed. **The deviation: the plan said pin torch==2.9.0+cu128; the repo's

own pyproject.toml pins torch==2.7.1+cu128 for Windows and resolves it from the

download.pytorch.org/whl/cu128 index BY URL.** The repo's pin was followed rather than the

dossier's, because 2.7.1+cu128 is a cu128 Blackwell wheel and it is the version the project tests

against; verified live at torch 2.7.1+cu128 / cuda 12.8 / sm_120 / RTX 5090. The plan's number

was an estimate and is corrected here rather than quietly satisfied.

Deviation 2 — DISPATCH IS IN-PROCESS, NOT THE :8001 REST SERVER. Install-plan step 3 stands up

uv run acestep-api on port 8001 and §4 described the lane as POSTing to it. generate.py instead

imports acestep.handler and acestep.inference directly and holds one handler across a whole

ladder. This is the better shape for this workload — no HTTP hop per candidate, offload_to_cpu

settable per run, and the handler-load cost paid once instead of per cue — but it is a deviation and

it was undeclared until this line. §4's dispatch sentence is corrected in place. The :8001 firewall

row stays landed in the runbook: the REST server remains a supported path for any consumer that wants

one, and it will not prompt at 04:00 if something stands it up. Nothing else in §2 was departed from.

**Q1 — TRUE COMPOSITIONAL STEMS, OR A BUNDLED SEPARATOR? NEITHER. THE QUESTION HAS THE WRONG

SHAPE, AND THE ANSWER CHANGES THE QUARTZ DESIGN RATHER THAN CONFIRMING IT.** Three findings, in

the order they bite:

1. The quality primary cannot do it at all. extract, lego and complete are

TASK_TYPES_BASE in acestep/constants.pybase checkpoints only. XL-SFT and every turbo

variant are excluded by the code, and the model zoo table marks XL-SFT ❌ on all three. The

adaptive-stem path and the best-quality path are different checkpoints, which the stack

table above did not say.

2. extract did not isolate anything, measured. Four extracts (strings / percussion / guitar /

keyboard) off one mixdown on acestep-v15-base: mean cross-stem spectral cosine 0.9998

the four "stems" are the same content returned under four different track names. Each is

spectrally ~0.987 identical to the mixdown, sample-aligned to it (lag −5), and carries more

energy than it (1.44–1.61×), and they do not sum back to it (residual −1.29 dB). So it is not a

mask separator, not a decomposition, and not a stem set. The requested track name did not select

content. *Scope of the claim: one checkpoint, one prompt configuration, duration passed on the

extract call; the retest lever is a base-model run with duration unset and a per-track caption.*

3. lego DOES work, and it is the mechanism the adaptive design should be built on. Its

instruction is *"Generate the {TRACK} track based on the audio context"* — it ADDS rather than

carves. Measured: the source survives in the output (correlation 0.47 at lag −5, spectral cosine

0.976) and the output carries substantial content the source never had (residual energy

ratio 0.76). That is additive compositional layering.

CONSEQUENCE FOR QUARTZ, stated as a design change and not a footnote: adaptive layers are

BUILT by successive lego passes on a base checkpoint, not carved out of a finished mixdown.

That is better than the fallback this dossier feared (pop-tuned 4/6-stem separation of orchestral

material) and worse than the hope (free stems from the quality model). It costs one generation per

layer and it means hero cues render on xl-base/base rather than xl-sft when stems are needed.

**Q2 — LOOP SEAM AND BPM RECOVERABILITY: BPM IS HONOURED, WITH MODEL-DEPENDENT PRECISION, SO

tempo_bpm IS MEASURED FROM THE RENDER AND NEVER ASSUMED FROM THE PROMPT.** Turbo hit 95.7 bpm

against a 96 target (0.3% error). XL-SFT with the 4B planner came back 117.5–129.2 against 112

and 120 targets (2–15% error) — directionally obedient, not exact. So bar_length IS fillable,

which keeps Quartz layering alive, but only off a measured tempo. **Baked fades: zero, across 8

candidates and 44 beds**, once fade_in_duration/fade_out_duration were set to 0 explicitly (the

library defaults are non-zero); the rig's fade detector is positive-controlled against a synthetic

fade, so that zero is a measurement and not an absence of looking. Seam ratios ranged 0.58–2.18 —

some candidates loop cleanly and some do not, so **loop-seam cleanliness is a per-artifact gate,

not a property of the generator.**

**Q3 — XL-SFT VRAM AT HANDLER LOAD, MEASURED. THE QUESTION SAID *PEAK* AND THIS RUN DID NOT MEASURE

ONE — the label is corrected here rather than left standing.** With offload_to_cpu=True, XL-SFT +

the 4B LM added ~9.5 GB over baseline (nvidia-smi 22535 → 32016 MB at handler load) on a card

already holding another lane's UE editor, and generated 32-second candidates in 16.6–40.2 s each.

It fit — with under 1 GB of headroom. The exclusive-overnight class stays the default and this

measurement is why: a windowed XL run is possible but leaves no room for anything to grow, so it is a

deliberate exception, never the schedule. Both conclusions rest on the load-time figure and both

stand.

What did NOT get measured, and the field that wrongly claimed it had. The records carried

nvidia_smi_used_mb_peak_observed, but that value was a single nvidia-smi call made at

record-write time — *after* generation returned. On the eight XL-SFT candidates it read

20765–23376 MB: below the 32016 MB after-load figure and, on three of them, **below the 22535 MB

baseline**. A field named "peak" that can report less than the floor is a label, not a measurement.

The same records do carry a true in-process torch peak — torch_max_allocated_mb 17912–18097 MB —

which is this process's own allocation and not the device total. Fixed on the measured close: the

field is renamed nvidia_smi_used_mb_at_record_write across all 66 landed record and summary files

(a relabel; every value byte-identical, and the generation_prompt_hash is untouched because

build_payload excludes the timing block), and generate.py now samples nvidia-smi on a thread for

the duration of each generate_music call, writing nvidia_smi_used_mb_peak_sampled +

nvidia_smi_peak_samples. The sampler is armed by fixture F17, not merely present. The 61 already-

landed records declare peak_sampling_note: NOT RUN rather than inheriting a number they never

measured — **so the true generation-time device peak for the XL-SFT ladder remains UNMEASURED until

the next XL run**, and the scheduling default above must be cited as a load-time figure, never as a

peak.

**Q4 — INSTRUMENTAL-ORCHESTRAL QUALITY vs THE VOCAL-POP CRITIQUE: the vocal-pop worry is NOT what

bites. The melody is.** With instrumental=True the output is genuinely orchestral and wordless —

no vocal leak in any of 52 artifacts. But measured melodic salience across the eight XL-SFT

candidates was 0.41–0.53, against 1.0 for a synthetic melodic control and 0.0 for both a noise

bed and a held triad. The material is texture-forward: 12–23 note events over 32 seconds. **And the

brief's single licensed anomaly was not delivered** — the winner scored D4 = 1 with *zero* leaps of

a sixth or larger, against a brief that specifies triad-outlining with exactly one unusual leap.

LEITMOTIF_ARCHITECTURE §3.6 already stated that the generator has no melodic conditioning that

could honour a quoted head; this is that statement measured, and it generalises: **text

conditioning does not reliably realise a specified melodic shape.** The 30-year bar is a melody

bar, so this is the lane's real open problem, and the three levers are named rather than assumed:

(a) scale the ladder and let selection pressure do the work, (b) author the head cell and use

lego/repaint to arrange around it — now known to be viable, see Q9, (c) find or build a

melody-conditioning path. The honest tier on everything shipped today is GENERATED SCORE —

ITERATION, and the tier is doing real work.

**[ANSWERED 2026-08-05 — §8 Q10 ran (a) and (b) head to head. (b)'s AUTHORING half wins and is

promoted; (b)'s ARRANGEMENT half does not survive its own decomposition, and (a) at double the

ladder with the contour named as explicitly as words allow still delivered the licensed leap in

2 of 16 and the head figure in 0 of 16. The open problem is no longer a generator-melody problem —

it is an orchestral-REALISATION problem. See §8.]**

Q9 — A QUESTION THIS DOSSIER DID NOT ASK, ANSWERED BECAUSE THE THEME FAMILY NEEDED IT (numbered

9 and not 5: §6 already has a Q5 about MOSS, and two different questions answering to

"PIPE_AUDIO_MUSIC Q5" is exactly the ambiguity that makes a later citation unresolvable) **: DOES

CONDITIONING HOLD A MOTIF? Yes, measurably. A repaint of the statement's back half held the

declared region** (waveform correlation 0.9998 over the untouched 16 s; residual 1.2%, not

byte-identical because the whole file is VAE re-decoded) and the new half ~~**shares real melodic

material with the statement — interval 3-gram Jaccard 0.333, 4-gram 0.238~~ [HALF-WITHDRAWN

2026-08-05 — §8 Q11. The two figures reproduce exactly at HEAD, but they were computed over the

WHOLE FILE, including the 16 s both files share byte-for-byte, so the held region contributed its

own n-grams to both sides and inflated the overlap. With that region excluded the same pair reads

3-gram 0.1250 / 4-gram 0.0000 — this dossier's own WEAK AND UNSAFE band. motif_compare.py now

decides on the generated region; the landed verdict flipped on re-run and the artifact was demoted

to a sibling render.]** That is the finding

that makes lever (b) above credible: the generator cannot be told a tune in words, but it can be

given one in audio and will keep it — **and §8 Q10 measures the limit of that. It KEEPS what it is

handed, exactly (correlation 1.0000 on all four repaints of the melody A/B), and it does not

CONTINUE it: the generated region carries motif 3-gram 0.000 and none of the brief's head figure.**

6. BENCHMARK-DAY QUESTIONS

Status line, 2026-08-05 — read this before treating any item below as open or closed. Items 1-4

are ANSWERED by measurement → §6.0. Items 5-8 are STILL OPEN, each for a stated reason, and

item 6 in particular is a Care-Doctrine gate that nothing in §6.0 touched. §6.0's own extra question

is numbered Q9 so it never collides with item 5 here.

1. ANSWERED → §6.0 Q1 (the question had the wrong shape; the answer was NEITHER).

~~Does ACE-Step 1.5 multitrack emit true compositional stems, or is it a bundled separator?~~ The

entire Quartz adaptive design rests on this. If separator-only, orchestral section splitting is a

known-weak pop-tuned 4/6-stem problem and hero cues fall back to library-rendered stems.

*Measured: extract isolates nothing (cross-stem cosine 0.9998) and is base-checkpoint-only;

lego layers additively and is what Quartz is now built on. §1's stack row is annotated to match.*

2. **ANSWERED → §6.0 Q2 (BPM is honoured; precision is model-dependent, so tempo is measured from the

render).** ~~Loop-seam quality at bar boundaries~~ — generate 8-bar loops at a stated BPM and

measure whether tempo_bpm/bar_length are *recoverable* from the output at all, or must be

imposed by prompting. If ACE-Step will not honour a target BPM, bar_length is unfillable and

Quartz layering is dead. *Measured: turbo 0.3% error, XL-SFT 2-15%; zero baked fades across 52

artifacts; loop-seam cleanliness is a per-artifact gate, not a generator property.*

3. NOW FULLY ANSWERED → §6.0 Q3 for the load-time figure, §8 for the peak. ~~The LOAD-TIME figure

is measured (22535 → 32016 MB); the

GENERATION-TIME PEAK this item literally asks for is STILL UNMEASURED~~ — the field that claimed

it held a post-hoc sample, the sampler that can answer it landed on the measured close, and the

number arrived on the next XL run, which was the melody A/B: **29381-31993 MB sampled peak across

16 XL-SFT candidates, 53-99 samples per call, on a card already holding 12.6 GB of another lane's

work. The exclusive-overnight default is reinforced — the ceiling was reached with under 1 GB of

headroom.** ~~24-28GB is my estimate from the README tier table.~~ *The

scheduling conclusion (exclusive-overnight stays default) rests on the load-time figure and holds.*

4. ANSWERED → §6.0 Q4, and it reframed the lane's real problem. ~~Instrumental-orchestral quality

vs the vocal-pop critique~~ — AUDIO_STACK Unknown #8. *Measured: no vocal leak in any of 52

artifacts, so the vocal-pop worry is closed; melodic salience 0.41-0.53 against 1.0 for a melodic

control, and the brief's one licensed leap was not delivered. The melody is the open problem, and

the honest tier stays GENERATED SCORE — ITERATION.*

5. STILL OPEN — NOT RUN. MOSS-SoundEffect is not installed (no D:/ai/moss-sfx, nothing

MOSS-shaped anywhere on D:), so zero SFX classes were scored. **MOSS-SoundEffect on the specific

hard classes** — per-surface footfall (needs short, dry, variation-friendly one-shots, not

ambience) and ritual-context cues under the sourcing law.

6. STILL OPEN — NOT RUN, AND IT IS THE CARE GATE. All 12 region cells are HELD under the §17.1

carve-out pending the elevated-care read, so no regional theme was generated and nothing tested

this. Cultural-substrate accuracy — do region prompts produce plausibly-registered

instrumentation, or generic "world music" wash? A wash is a Care-Doctrine failure (care as

authenticity), not a taste note. Nothing in §6.0 may be cited as evidence on this question.

7. STILL OPEN — NOT RUN. Stable Audio was never installed or invoked; it remains CAP-RISK

fallback-only and untested. Stable Audio 3 Small SFX true output sample rate — its card's

example reads 16kHz. Confirm or kill before it is trusted even as fallback; a 16kHz asset cannot

ship.

8. STILL OPEN — NOT RUN, and correctly so: its own precondition has not fired. Windows-native

throughput was never the constraint on this batch (32-second candidates in 16.6-40.2 s).

Windows-native throughput vs WSL2+vLLM — only worth measuring if the Windows PyTorch backend

proves throughput-bound on the 79-node batch.

Sources fetched this pass (VERIFIED tier)

github.com/ace-step/ACE-Step-1.5 (+ /blob/main/README.md, raw…/LICENSE, raw…/README.md) ·

github.com/ace-step/ACE-Step (+ /blob/main/LICENSE) · huggingface.co/ACE-Step ·

huggingface.co/ACE-Step/acestep-v15-xl-sft (+ /tree/main) · arxiv.org/abs/2602.00744 ·

docs.comfy.org/tutorials/audio/ace-step/ace-step-v1 · …/ace-step-v1-5 ·

github.com/multimodal-art-projection/YuE · github.com/HeartMuLa/heartlib · github.com/adefossez/demucs ·

huggingface.co/OpenMOSS-Team/MOSS-SoundEffect-v2.0 (+ /tree/main) · github.com/OpenMOSS/MOSS-TTS ·

huggingface.co/stabilityai/stable-audio-open-small · huggingface.co/stabilityai/stable-audio-3-small-sfx ·

stability.ai/license · elevenlabs.io/docs/overview/capabilities/sound-effects · gdc.sonniss.com ·

sonniss.com/gdc-bundle-license · dev.epicgames.com UE 5.8 release notes / MetaSounds / Quartz ·

research.nvidia.com/labs/sil/projects/ardy · github.com/nv-tlabs/kimodo

---

7. COMPARATOR-PIPELINE ADDENDUM — WHAT THE TEACHERS' PIPELINES TEACH THE AUDIO LANE (2026-08-05)

APPEND-ONLY addendum under the pipe-dossiers-bind law, from the 2026-08-05 comparator-pipeline research

(Josh: "what about their pipelines?"). **Read the scope limit first, because it is the honest finding of

this section: no researcher in that pass was tasked with the AUDIO lane.** Three researchers covered

quests/world, creatures/combat/VFX, and production systems at scale. Audio was not a lane in the fan-out.

So this section is NOT a research pass — it is a routing pass. Its content is the audio-bearing facts

the other three researchers surfaced inside sibling dossiers, where they sat orphaned against a lane that

could not consume them. Each was **re-verified by direct fetch by the synthesizer before being routed

here**; nothing is carried on a sibling's summary alone. The gap itself is boarded at the bottom.

7.1 Review motion WITH its audio at authoring time — ADOPT the principle

Capcom's RE ENGINE shows captured motion in engine in real time during a shoot, and — the part that

belongs to this lane — sound effects and background music can be played into the same scene, so the

team sees what the player will experience far earlier in the process. Named: animator **Naohiro

Taniguchi (real-time visualization), monster-animation lead Kenji Yamaguchi**. The demonstration

studio carries 36 wall-mounted infrared cameras; a Kyobashi facility has 150. Before Monster Hunter

World, all monster animation was keyframed.

Source — https://www.digitaltrends.com/gaming/monster-hunter-wilds-capcom-studio-tour/

VERIFIED by direct fetch 2026-08-05 (quoted: developers can "see that process work in real time and

continue to polish and refine the animations"; "Capcom can even add sound effects and background music

into the scene").

NO-ADOPT: the capture infrastructure — we have no stage and will not have one; nothing here changes

the ARDY/Kimodo primary in PIPE_ANIMATION §1.

ADOPTED AND ROUTED: audio is not a post-pass. The proving map that PIPE_ART §7.5 and

PIPE_ANIMATION §7.3 both independently converged on must carry cue audio wired at authoring time,

not added after the motion and the VFX are signed off. One proving map, three lanes, one screenshot-and-

listen surface for the visual-verification standing directive. This costs nothing to adopt now and costs

a re-do if adopted after volume authoring.

7.2 Impact audio has a documented STRUCTURE — ADOPT as the cue-design shape

From the God of War (2018) combat round-table, the concrete audio facts:

ends with a high frequency slash."

the hit pose and holds both Kratos and the target in that first frame for a short duration."

REDUCED shake, "what is reduced by these camera choices is made up for in audio and animation."

Speaker: Christian Wohlwend (Naughty Dog), in the multi-studio round-table.

Source — https://blog.playstation.com/2022/10/04/game-developers-explain-what-makes-god-of-war-2018s-combat-tick/

VERIFIED by direct fetch 2026-08-05. Attribution note: the researcher return framed these as Santa

Monica facts; they are a round-table participant's, and the participant is Naughty Dog's. The claims are

about God of War's combat; the speaker is not a Santa Monica employee. Cite the person, not the studio.

ADOPTED AND ROUTED: the low-end-then-high-frequency envelope is a **cue-design law for the impact

beat**, routed to the SFX half of this lane alongside the four-beat cue grammar. It is also the direct

answer to a question our own stack raises: when the presentation ladder escalates a cast across tiers,

what escalates in the AUDIO? The comparator answer is that weight lives in the low end and legibility

lives in the high end — which makes tier escalation a low-end-mass delta with the high-frequency read

held constant, the exact structural twin of the luminance-versus-hue finding on the VFX lane

(PIPE_ART §7.3). Two lanes, one shape: escalate the intensity axis, hold the identity axis. That

symmetry is this addendum's own synthesis and is flagged as such.

7.3 Sound assignment belongs to the person who knows the scene — ADOPT the pattern

Nintendo built **SLink, a tool letting designers assign sounds by clicking an actor in the game with

the mouse**, as part of the tooling that let a two-designer UI team ship Breath of the Wild's UI without

filing requests with programmers.

Source — CEDEC 2017 (Fujibayashi / Yonezu), Matt Walker translation —

https://gist.github.com/idbrii/e39fe96279aa1670319bfa521d907399

VERIFIED by direct fetch 2026-08-05 (quoted: "SLink, allowing designers to click an actor in game

with their mouse").

ADOPTED AND ROUTED: our equivalent of "click the actor and assign the sound" is **a declared audio

binding on the entity row, authored where the entity is authored**, rather than a separate audio-wiring

step that must re-derive which entity needed which cue. This is the same premise as every

harness/apply_*.py in the repo, pointed at audio, and it is corroboration for the registry-as-authoring-

surface bet rather than a new mechanism.

7.4 The one first-party source this lane has NOT read — the top re-fetch

GDC 2024, **"Tunes of the Kingdom: Evolving Physics and Sounds for 'The Legend of Zelda: Tears of the

Kingdom'" (Dohta / Takayama / Osada**) — https://gdcvault.com/play/1034667/Tunes-of-the-Kingdom-Evolving

The 2026-08-05 pass extracted only the PHYSICS half of this talk (it landed in PIPE_CONTROL_PLANE

§8.2, N7-N8: every moving object made physics-driven; mass and volume auto-derived from shape plus an

assigned material, then de-tuned for feel). **The SOUND half — which the title names and a named audio

speaker presented — was never extracted, and it is the single most on-target first-party source available

to this lane on how audio couples to a systemic physics world.** It is the top re-fetch.

The physics half already implies the audio question, and the question is ours too: when mass and material

are DERIVED per object rather than authored, **the impact sound has to be derived from the same two

values, or every collision in the game sounds like the same wooden knock.** Our element matrix and

material assignments are exactly that kind of derived pair. Whether Nintendo solved it by material-indexed

sample sets, by synthesis, or by something else is precisely what the unread half of that talk would say.

7.5 Corroboration, no change

Blizzard's stated Overwatch data-pipeline goal that local changes are seen instantly in engine applies

to audio iteration identically (PIPE_CONTROL_PLANE §7.6, tier SECONDARY with its caveat declared there).

Our generated-audio lanes currently land files in folders and wait for an import step — the same return-

path gap PIPE_CONTROL_PLANE §7.5 item 5 routes to ASSET_DROPIN_CONTRACT.md. **No separate adoption;

the audio lane is a beneficiary of that one.**

7.6 The gap, boarded honestly

No dedicated audio-pipeline comparator research has been run. Everything above is routed from lanes

that were looking at something else and happened to surface an audio fact. The questions a real pass would

answer, none of which this section can: how shipped AAA titles structure an adaptive music system's

authoring surface (our native-multitrack benchmark question, §6, has no comparator anchor); how SFX

libraries are organized and versioned at scale; how VO is routed, batched and QA'd against script churn;

how mix and loudness are validated automatically in CI. Boarded, not done. The comparator-lens program

should carry an audio row the next time it fans out — and the 30-year melodic bar (§0) deserves a

comparator pass of its own, since none of the four teachers studied on 2026-08-05 was chosen for music.

---

8. BENCHMARK DAY, PART TWO — 2026-08-05: THE MELODY A/B. AUTHORING DELIVERS THE TUNE; THE ARRANGER DOES NOT CONTINUE IT

What this section answers and what it does not, stated first. ANSWERED by measurement: the open

problem §6.0 Q4 left the lane with — *text conditioning does not reliably realise a specified melodic

shape*. Two of Q4's three named levers were run head to head. NOT ANSWERED and not touched: §6 items

5, 6, 7, 8 remain exactly as §6.0 left them; the Care-Doctrine question (item 6) is still open and

nothing in this section may be cited on it — JOURNEY_WORLD is the home register, not a region

cell, and no region cell was generated here either. Lever (c), a melody-conditioning path, was not

built.

The experiment. ARM 1: the JOURNEY_WORLD motif composed in notes by the factory to the

brief's constraints, realised to audio, then arranged by nine ACE-Step passes conditioning on it

(four repaints holding 24/16/16/8 s, two covers at strength 0.75/0.45, three lego layers on the base

checkpoint). ARM 2, the control: sixteen pure text2music candidates on XL-SFT + the 4B planner —

double the original ladder — with the contour language sharpened as far as words go, the leap named

by size and direction, the head named note by note in scale degrees. Both arms judged by the same

rig against the same measure_spec, and both compared against the **same authored interval

sequence**, so arm 1 got no privileged reference.

Tooling: harness/music_gen/{author_head_cell,ab_judge}.py (new), emit_ladder.py --arm1/--arm2,

feature_rig.py (17/17 controls green first). Evidence:

build/audio/generated/journey_world_ab/ab_report.json, the three job specs beside it, and 50

record files under build/audio/records/journey_world_ab/ — 25 artifacts, zero orphans.

**Q10 — DOES AUTHORING THE HEAD CELL BEAT PROMPTING FOR IT? YES ON THE MELODY DIAGNOSTICS, WITH THE

REGISTER HOLD MET — AND THE AGGREGATE IS THE LEAST INFORMATIVE LINE IN THE RESULT.**

diagnosticARM 1 (n=9)ARM 2 (n=16)
melodic salience, median (best)0.528 (1.000)0.440 (0.538)
leap of a sixth or larger delivered4/92/16
the brief's licensed leap, within 1 semitone4/92/16
the head's interval figure present4/90/16
register match cosine, median0.7390.704
motif 3-gram Jaccard, median0.0000.000
note count, median1715.5

The register hold is what makes the win admissible rather than merely bigger: arm 2's own median

register match is the floor, arm 1 clears it, and the judge is built so that a melody win below that

floor returns ARM 2 anyway (fixture J14).

THE DECOMPOSITION, WHICH MOVES THE ANSWER. Arm 1 mixes three operations and they do not behave

alike:

register match anywhere in the run — leap 4/4, head figure 4/4, held-region correlation 1.0000

on every one.

figure 0/5. Statistically indistinguishable from arm 2, and slightly worse on register.

Conditioning on authored audio through lego or cover does nothing for the melody.

AND THE REPAINT WIN IS INHERITANCE, NOT COMPOSITION. The head figure sits inside the region the

repaint was told to hold, so a whole-file measurement partly re-reports the input. Scoring the

generated region alone — the model's own contribution, held region excluded — gives, on all

four repaints: motif 3-gram 0.000, no head figure, no mirror of it, no leap, 2-9 note events

across 8-24 seconds, salience 0.000-0.433. The generator holds what it is told to hold, exactly,

and then writes texture. It does not continue the tune. ab_judge.py's fixtures J15-J17 are a

splice control built to make this impossible to miss: authored audio followed by unrelated texture

must still fire the whole-file head test while the generated-region test comes back empty.

**Q11 — A LANDED READING WITHDRAWN ON RE-MEASUREMENT, RECORDED HERE BECAUSE THE CORRECTION MATTERS

MORE THAN THE TIDINESS.** §6.0 Q9 concluded that a repaint's new half "shares real melodic material

with the statement", on interval 3-gram Jaccard 0.333 / 4-gram 0.238. That comparison ran over the

whole file, including the 16 seconds both files shared byte-for-byte, so the held region

contributed its own n-grams to both sides. Re-derived at HEAD on the same landed pair:

scope3-gram4-gram
whole file (what Q9 used)0.33330.2381
held region excluded0.12500.0000

The number was right and reproduces exactly. The reading was not: 0.1250/0.0000 is

motif_compare.py's own *WEAK AND UNSAFE* band, not *MOTIF MATERIAL SURVIVES*. Q9's finding is

therefore corrected to: **repaint conditioning HOLDS a motif — measured twice now, at 0.9998 and at

1.0000 — and does not CONTINUE one.** motif_compare.py now computes the overlap on the generated

region and decides the verdict there whenever a held region is declared (7/7 new fixtures, armed in

both directions: a splice must be refused the word variation while its whole-file overlap still

reads 0.857, and a genuine transposed variation must be recognised at 1.000 on the deciding block).

The landed family_report.json verdict flipped on re-run and PICKS.json demoted that artifact

from "variation" to "sibling render" through a branch written for exactly this case that had never

fired.

WHAT IS PROMOTED, AND IT IS NARROWER THAN THE LEVER THAT WAS TESTED. Promoting

"authored-head-cell + lego/repaint" whole would overstate the measurement by the width of the

decomposition above. What the evidence supports:

of the brief's licensed leap and its head figure on authored material, against 12.5% and 0% from

the strongest text-only prompt this lane can write, over sixteen candidates. The tune is authored;

it is not drawn from a prompt. harness/music_gen/author_head_cell.py is the tool, its self-test

checks each constraint against the notes rather than against a comment, and it caught three real

compositional defects on the way in (a response whose registral reset was wider than the licensed

leap; a second deviant arriving from the loop join rather than from the tune; a negative control

that contained what it was controlling for).

cover overwrites; lego does not carry. The arrangement problem is open.

realisation.** The site serves it as *"authored score — sketch realisation"* and the tier note says

the tune is the deliverable and the orchestration is not finished. The prompted ladder winner is

demoted to superseded_statement rather than deleted, because deleting it would erase the

comparison that made the promotion legible.

THE NEXT LEVER, NAMED HONESTLY, AND IT IS NOT THE ONE THIS RUNG ASSUMED. Going in, the open

problem looked like a *generator melody* problem. Measured, it splits in two and only one half is

still open:

1. Melody: solved by authoring. No further generator work is required to get the brief's contour

into an artifact. The twelve canonical themes can be composed the same way today.

2. Orchestral realisation: the real blocker, and it is a RENDERER problem, not a generator one.

The authored artifact sounds like additive synthesis because that is what it is. ACE-Step cannot

be handed an authored score and asked to play it — repaint re-decodes the held region unchanged

(residual 1.2%; it is a VAE round-trip, not a re-orchestration) and cover discards the tune. So

the next rung is rendering authored MIDI to orchestral audio — a sampled-orchestra or

synthesis path — with ACE-Step kept to the beds and ambience where §6.0 already measured it

strong. This also re-frames §0's concert-notation deliverable as *closer*, not further: an

authored score is already the thing an orchestra could read, which the prompted lane never was.

3. Still available and untested: Q4's lever (c), a genuine melody-conditioning path. It is now a

*lower* priority than (2), because authoring already produces the melody conditioning was wanted

for.

Two smaller measured facts, recorded rather than dropped. (a) §6 item 3's generation-time VRAM

peak, declared UNMEASURED on the measured close, is now measured: the armed sampler reported

29381-31993 MB peak across the arm-2 XL-SFT ladder, 53-99 samples per call, on a card already

holding 12.6 GB of another lane's work. The exclusive-overnight default is reinforced, not softened —

the ceiling was reached with under 1 GB of headroom. (b) **Loop seam is a live problem for the

authored method**: repaint artifacts are 0/4 seam-clean and the authored cell itself fails the rig's

seam heuristic, because a sentence whose head and tail differ by design wraps badly. Per §6.0 Q2

that is a per-artifact gate, so it does not block the promotion — it is a named item for the

realisation rung, which is where a loop point should be composed anyway.

8.1 THE REALISATION RUNG, SURVEYED AND READ — 2026-08-05. A LICENCE READ, NOT AN ADOPTION

What this block is and what it explicitly is not. Section 8 closed by naming the next lever:

*rendering authored MIDI to orchestral audio — a sampled-orchestra or synthesis path — with ACE-Step

kept to the beds.* Fourteen authored head cells now exist in notes (commit be9d9520), so that rung

has real input and the blocker is no longer hypothetical. This block is the SURVEY and the LICENCE

READ for it. **Nothing is adopted, nothing is downloaded, and not one sample is rendered through any

of it.** The lane's own rule is the reason: no renders through unread licences. What follows is the

read, so that the rung which runs next runs on a licence somebody has actually looked at.

The reads are FACE-VALUE, per the standing licence law. A licence is read for what it says, not

for what a cautious reading could imagine it saying; AI tooling on properly licensed assets is normal

use; and only an EXPLICITLY TRIGGERED prohibition is flagged. One of the five options below is

refused on that test, and it is refused because its prohibition really does fire.

The five options, and what each licence actually says

A. CC0 SAMPLE SETS THROUGH AN OPEN SAMPLER — VSCO 2 Community Edition and VCSL, played by sfizz.

Versilian Studios' VSCO 2 Community Edition and the broader Versilian Community Sample Library are

both released under CC0 — a public-domain dedication: no royalties, no attribution requirement,

no restriction on commercial use, and the publisher states outright that anyone may build on them.

VSCO 2 CE is a chamber-scale set of roughly three thousand samples at 24-bit/44.1 kHz. The player is

sfizz, BSD-2-Clause, explicitly designed to be embedded and driven as a library, with

permissively licensed dependencies throughout. **Triggered prohibitions: none, and there is nothing

in this chain to trigger** — CC0 carries no conditions, and BSD-2 carries only a notice requirement

that applies to redistributing the software, which rendering audio does not do.

B. A COMMERCIAL FREE-TIER LIBRARY — Spitfire Audio's BBC Symphony Orchestra Discover.

Read at face value, the EULA grants what this lane needs: the sounds may be used inside newly created

recordings, commercially, with no ongoing payment, and games and sync are not excluded. Two clauses

matter and they point in opposite directions.

libraries for the purpose of training AI systems or large language models without written consent.

Rendering an authored score through a sampler is not training — it is the creation of new music,

which is the licence's own stated permitted purpose. Under the licence law that is a face-value

read, and the conservative reading that would refuse a sampler because a factory is holding the

mouse is precisely the error the law names.

redistribution of the samples, or of derivatives usable *as samples in a sampler*, is forbidden. So

the rendered stems ship and the sample content never does; no Spitfire byte enters this repo; and —

this is the part that binds the other arm of the pipeline — **no Spitfire sample may ever be a

conditioning, fine-tune or LoRA input to ACE-Step or any other generator.** That is the same rule

the Sonniss bundle already imposes, arriving from a second direction, and it travels on the row the

way the Sonniss rule does rather than sitting in a note nobody was instructed to satisfy.

C. SONATINA SYMPHONIC ORCHESTRA — REFUSED, on a prohibition this product genuinely triggers.

SSO is distributed under Creative Commons Sampling Plus 1.0, a licence Creative Commons

retired in 2011 and no longer recommends. It permits transformative sampling and collage, and it

permits noncommercial sharing of verbatim copies — but it excludes **all advertising and promotional

uses**, except promotion of the derivative work itself. A game soundtrack is used to promote the

game: trailers, store pages, festival reels. That is promotional use of a work which is not the

derivative being promoted, and it is exactly the explicit trigger the licence law says to flag.

Not adopted. Recorded here rather than dropped, because SSO is the library a search for free

orchestral samples returns first, and a later lane would otherwise re-discover it and re-decide it.

D. SYNTHESIS — Surge XT, or extending the additive engine that already exists.

Surge XT is GPL-3.0: the copyleft binds the software, not the audio a user produces with it, so

there is no output-licence question at all. Extending author_head_cell.py's own additive carrier

costs no licence and no download. Triggered prohibitions: none. What it does not do is sound like

an orchestra, which is the entire point of the rung — so this is the honest floor, not the answer.

E. THE SOUNDFONT CLASS — FluidR3_GM through FluidSynth.

FluidR3_GM is MIT (Frank Wen) and is usable for personal or commercial composition; the author's

additional guidance concerns bundling the soundfont itself with commercial software, which shipping

rendered audio does not do. FluidSynth is LGPL-2.1. Triggered prohibitions: none. The

limitation is quality rather than licence — General MIDI sits well below the AAA bar — but it is

deterministic, scriptable and free, which makes it valuable as a reference render the battery can

run against on every commit even if no player ever hears it.

THE RECOMMENDATION, and it is two rungs rather than one

1. RUNG ONE, adoptable with nothing to flag: sfizz (BSD-2) driving VSCO 2 CE and VCSL (CC0).

The whole chain is public-domain or permissive, so every licence record it produces reads

unconditional rather than conditional; it is scriptable and headless, which the battery needs; and

it lifts the sketch from additive synthesis to real recorded instruments in one step. This is the

rung to build.

2. **RUNG TWO, on evidence rather than on schedule: BBC SO Discover for the cues that need the larger

sound**, with the licence re-read and hash-pinned at download time under the existing

DL_0008/DL_0009 precedent, and the no-conditioning rule carried onto every record.

3. REFUSED: SSO. KEPT: the additive sketch, as the honest tier below both.

What the rung must MEASURE, so it cannot be declared done by ear. The composition battery is the

acceptance test and it does not care what produced the audio: an orchestral realisation of an

authored cell must still pass C11 (it reads as a melody, not a texture), C12 (the authored deviant is

recoverable from the render), C13 (the notes sit on the grid the declared tempo implies) and C14 (no

baked fade), against the same negative controls. A realisation that loses the tune inside a beautiful

patch fails C12 — which is exactly the failure a listening-only review would call a success. And the

loop seam should be composed at this rung rather than repaired at it: section 8 named it as a

live per-artifact problem, and a realisation pass is the first point at which a loop point can be

written into the arrangement instead of hunted for in the waveform.

The strongest objection to all of the above: none of these is a professional orchestral library,

so rung one buys realism and does not buy the AAA bar, and a later decision to license a full library

would make some of this work disposable. The answer: the deliverable of the authored method is

the SCORE, and every path above consumes the same MIDI. Changing the renderer changes nothing

upstream of it, which is the whole reason the authored method was promoted over the arranged one.

Sources read for this block (all read 2026-08-05, face value, no download): Versilian Studios

VSCO 2 Community Edition and VCSL product/licence pages (CC0); github.com/sfztools/sfizz

(BSD-2-Clause); Spitfire Audio End User License Agreement (commercial-use grant; section 10

AI-training restriction; redistribution prohibition); Creative Commons Sampling Plus 1.0 legal code

and its retirement notice; github.com/peastman/sso (SSO licence statement); FluidR3_GM README

(MIT, Frank Wen) and FluidSynth (LGPL-2.1); Surge XT (GPL-3.0).

8.2 THE REALISATION RUNG, BUILT — 2026-08-05. AN AUTHORED SCORE PLAYED BY AN ORCHESTRA, AND THE FOUR RULERS THE SAMPLER BROKE

What this block is. §8.1 was a licence read that adopted nothing. This is the rung it recommended,

built and measured: sfizz 1.2.3 (BSD-2) driving VSCO 2 Community Edition and VCSL (both CC0),

rendering the fifteen authored scores that commit be9d9520 landed in notes. Nothing upstream moved.

Every note of every sentence is played exactly as composed; the only material this rung adds is a

written loop turnaround, which is arrangement.

The stack, pinned rather than described. 442 files, 729 MB, every one carrying its own sha256 in

build/audio/stack/sfz_stack_manifest.json; both sample sets pinned by COMMIT SHA with a probe that

reports whether the branch has moved since; three licence bodies stored verbatim — docs/licence_records/sfizz/LICENSE

(BSD-2, pinned at the release tag), docs/licence_records/vsco2_ce/LICENSE and

docs/licence_records/vcsl/LICENSE (both CC0), alongside docs/licence_records/vcsl/README_sfz_branch.md.

Two things the fetch FOUND rather than assumed: **VCSL's sfz branch carries

no LICENSE file at all**, so the row pins the CC0 legal code on master *and* the licence statement in

the README of the branch actually downloaded — which is why that README is stored as a licence record

in its own right rather than paraphrased — and an

upstream packaging defect — the bowed psaltery's SFZ references its samples beside itself while

they are nested inside the Italian harpsichord's folder, and sfizz loads such an instrument without

complaint and renders SILENCE. The repair is declared, fires only on a 404, and is positive-controlled

by a smoke render that returns real energy. Recorded as DL_0016-DL_0018.

**AND THE RECORD OF THAT REPAIR WAS ERASING ITSELF, which a cold re-read caught by reading the landed

manifest instead of the prose about it.** The defect is real — the raw path 404s at the pinned commit

and the declared replacement returns 200 — but the manifest at HEAD recorded

upstream_path_repairs_fired: 0 for both sets and not one of the eleven affected sample rows carried

a repair key. The cause is that a repair is discovered on a 404 at DOWNLOAD time, so on any re-fetch

that finds the files already on disk the download branch never runs, the in-memory repair map comes

back empty, and the counter and rows silently reset to zero. A record whose whole purpose is to make

an upstream defect VISIBLE was quietly deleting its own evidence on every subsequent run. Two fixes:

the prior manifest's repair rows are now carried forward onto files that are skipped, and the fetch

REFUSES outright if sample rows match a declared repair's needle while none of them records the

repair — so the next time this goes to zero it fails loudly, and if upstream ever repackages, the

same refusal is the prompt to retire the declaration rather than leave a stale one in place.

The subset is a CARE decision before it is a download economy. VCSL ships kalimba, mbira, balafon,

darbuka and dan tranh. None of the fifteen artifacts is a region cell; the family heads' allow-lists

say the register they may enter is *each node's own register card*, and those cards belong to region

cells that are explicitly not composed, because the care read for that material has not been run. None

of those instruments is downloaded, and harness/music_gen/sfz_palette.py says why on the row.

Of the 23 instruments that ARE downloaded, the fifteen cards use 17; the six unused are the

glockenspiel, hand chimes, the sustained horn, the vibrato oboe, the bowed psaltery and the tubular

bells. K10 prints that list on every run and the run report carries it, which is how the rung's own

commit message came to say "4 of 23" while its self-test printed six on the same tree — a number

typed once into prose beside a number the tooling derives. The derived one is the record.

THE CARDS ARE HONESTLY UNEVEN — 7 OF 15 REFUSE. The exemplar is the East African Rift Valley card,

whose first sentence is that no instrument is attested in that corpus and whose last is that none may

be invented here. Six family allow-lists defer to cards of exactly that kind, so those heads are

realised on the register-neutral orchestral base — a string frame and a string carrier, claiming

no tradition — and the refusal travels onto the artifact, the record and the site tier note. A second

refusal kind covers the one card that DOES name instruments and may not have them:

FAM_THE_LIVING_VOICE names the Flores gong-waning and the Bali gamelan, both under their own

never-sampled care line, and reaching for an orchestral tam-tam to stand in for a gong-waning is the

"generic world-music wash" the brief pack calls a care failure rather than a taste note. **Seven

cards DEFER part of their register** — a per-age-stage re-voicing, a per-region migration, per-region

re-voicings, the local apparatus registers, a fairy-realm register the realm programme holds until

slice-end, and the ceremonial percussion two rows admit only at an ACTIVATION event — and each says

which half is missing rather than approximating it. Refusal and deferral are INDEPENDENT, and the

correction that made them so is FAMILIAR_BOND: its row asks for two things and the card answered

one. The twenty-two per-species carriers are its REFUSAL — they were previously miscounted into this

deferral list — and the row's second clause, "ceremonial percussion at bond-quest activation", was

neither realised nor deferred. Half a row was silently dropped, which is exactly what the deferral

class exists to prevent, and the identical clause on HOUSE_PARENT_HELD was already carried

correctly. It is now a declared deferral, and the counts live in the self-test's own printout rather

than in a hand-typed sentence: the module docstring said "four" while K09 printed six on the same

run, so the prose no longer carries a number that nothing checks.

THE LOOP SEAM IS COMPOSED AT THIS RUNG, WHICH IS WHAT §8.1 ASKED FOR. Each theme has a written

turnaround under a rule it names for itself: a dominant for the triadic tunes; an open fifth on the

fifth degree for the family whose identity is the absent third, because a dominant triad would destroy

it; a quartal chord whose root moves a fourth for the family built of stacked fourths; the pedal held

under the preparation; the two-chord cycle's own next member held across the join; and, for the

pre-gate cue whose card forbids a swell, **silence, held deliberately, because leaning anywhere is what

would make a listener notice it**. Four score-level teeth check them — no second deviant at the wrap,

no new pitch class, the preparation rule, the harmonic signature — and a fifth was minted when the

measurement demanded it: S6, a seam may not CONTRADICT its own theme's deviant. Its two fixtures

are both real landed defects: a turnaround entering ON the downbeat in the one theme whose deviant is

that it never does, and a turnaround entering on beats 1 and 3 in the theme whose cell always enters a

beat early. The second is the one that matters — it made the deviant **undetectable on the rendered

audio** while every other check stayed green.

THE SEAM'S SECOND HALF LIVES IN THE WAVEFORM AND IS MEASURED. A turnaround makes the loop point

musical; it cannot make it continuous, because the file still stops at the bar line and the

reverberant tail a live orchestra would still be sounding is thrown away. So the overhang is not

discarded: sfizz is allowed to ring past the written end and those samples are ADDED ONTO THE HEAD of

the artifact. It is a sum, not a fade, and C14 still binds and still passes — now with both of its

arms armed, which was not true when this was first written (see correction 5 below). Measured on

JOURNEY_WORLD, the wrap discontinuity against the median interior bar line fell from 28.99 to 7.35

when the release was carried over; on the rig's independent mel-distance ruler the same wrap takes

HOME_AND_LOSS from 5.804 to 0.067 and PATTERN_UNNAMED_PREGATE from 2.725 to 0.888, which is a

second instrument agreeing that the wrap is doing real work. The turnaround's own spectral contribution is small and on several

themes slightly negative, and that is reported rather than hidden: spectral flux is a TIMBRAL ruler and

the turnaround is a HARMONIC device, so S5 gates on the wrap and the turnaround is checked where it

lives — over the notes, in S1-S4 and S6.

THE ACCEPTANCE TEST IS THE BATTERY ON THE MIXDOWN, AND IT DID ITS JOB BY FAILING. §8.1 wrote the

failure mode down before the rung existed: *"a realisation that loses the tune inside a beautiful patch

fails C12 — which is exactly the failure a listening-only review would call a success."* That is

precisely what the first mix did. JOURNEY_WORLD read as a melody (salience 0.9738) and still failed

C12: the tracker recovered 60 notes from a 44-note artifact, the extra sixteen being octave jumps

between the carrier and the string frame, with wide intervals of 12, 19 and 24 semitones drowning the

one authored leap the detector exists to find. Four settings were measured across three themes and the

frame moved from -9 dB to -18 dB, at which the spurious wide intervals stop and the authored leap

is found. A listening review would have called the first mix the better one.

AND THE NOTE-COUNT HALF OF THAT CLAIM WAS OVERSTATED, WHICH A COLD RE-READ CAUGHT. The sweep table

said -18/-16 was the setting "at which the tracker recovers the written note count exactly", and at

HEAD it is exact on 6 of the 15 artifacts. The sweep's salience column reproduces exactly (1.000 /

0.942); its note-count column did not survive the five carrier replacements and the added turnaround

and release wrap, so the same three themes now read 45, 23 and 35 against 44, 22 and 34 written. The

over-recovery signature the mix fix was named against is therefore still present on three published

artifacts — HOME_AND_LOSS 58 against 40 written, HOUSE_PARENT_HELD 43 against 15, and

PATTERN_UNNAMED_PREGATE 25 against 12. It fails nothing, and the reason is a real scoping decision

rather than an accident: those three carry METRIC deviants, and C12 for the metric classes reads the

soloed carrier stem rather than the mixdown, so their whole-mix count is a statement about how

much the ACCOMPANIMENT hands the tracker and not about whether the tune survived — on the mixdown

itself what carries them is C11's predominance number. The correction is that this is now REPORTED

rather than inferable: written_notes and over_recovery_ratio ride on every artifact's diagnostics

and are summarised in the run report, and the sweep table is labelled as the historical reading it

always was.

**THE RULERS THAT HAD TO BE RE-SCOPED — six of them, and the last two were found by a cold re-read

AFTER the rung landed.** The first four corrections are about SAMPLED INSTRUMENTS rather than about

the music: the additive sketch had a 45 ms attack and one steady pitch per note, and a sampled

orchestra has neither. The fifth and sixth are about the MEASURING APPARATUS itself — a positive

control that could not fail, and a second ruler whose disagreement was going unreported.

1. ATTACK LATENCY, MEASURED. Notes written at beats 0.5 / 2.5 / 4.5 / 6.5 came back from the pitch

tracker at 0.789 / 2.949 / 4.992 / 6.920 — late by 0.29-0.45 beats, which is 145-225 ms. C13 now

removes the measured constant offset and reports it, with the 25%-wrong-tempo control computed under

the same correction, because a constant offset cannot fix a tempo error while a tempo error

accumulates across the artifact. JOURNEY_WORLD: median onset distance 0.0217 beats at the

declared tempo against 0.5016 at a tempo 25% wrong.

2. METRIC DEVIANTS READ THE CARRIER'S ATTACKS. offbeat_only measured on pitch-track note starts

reported 75% of the head's onsets sitting on a beat, in a family where every note is written off it.

And a metric property measured on the WHOLE MIX cannot separate the melody's entries from the

accompaniment's bar lines — which is exactly what made HOUSE_PARENT_HELD's early entry unfindable.

The negative control is measured identically, and C11 still proves on the whole mix that the carrier

is the predominant line.

3. ADJACENT SAME-PITCH SEGMENTS ARE COLLAPSED before any interval-shape detector runs. A

predominant-pitch tracker segments on pitch CHANGE, and a sampled string section held for two beats

at 68 bpm makes it close and reopen a segment at the same pitch: FAM_THE_OTHER_SELF's head came

back as [3, 4, 0] and FAM_THE_DEEP_WATER's as [-5, 0], so a self-inversion test and a

descending-triad test were both reading a segmentation artefact as the composition's shape.

Collapsed, those two recover 22/22 and 16/16 notes exactly and both deviants are found. The

collapse is applied identically to the negative control, and never to reciting_repetition, whose

property IS real repeated tones.

4. THE RECITATION TEST BECAME A SEPARATION RATHER THAN AN ABSOLUTE. The sketch's threshold of 2.0

was itself derived from a measured pair (3.5 authored / 1.5 de-repeated control) and it does not

transfer: a sampled cello's own bow attacks put the control at 2.2 while the artifact sat at 4.0.

The rule is now the separation the check was always trying to be — 1.5x over its own control, which

the sketch's own pair (2.33x) also clears — and it cannot be satisfied by a loud renderer, because

the control is rendered the same way.

5. AND THE FADE-IN PROBE BECAME A BOUND RATHER THAN BIT-EQUALITY. C14 asks whether an amplitude

ENVELOPE was applied to the file, and it answered that by comparing the opening samples of the

artifact with the opening samples of its own +1-statement probe. Inside one renderer that is

exact. Across two SEPARATE sfizz invocations rendering different-length MIDI that share a prefix,

it is not: three consecutive themes in a full run came back unequal and every one of them passed

the same probe when run alone, differing in the last bits rather than in the music. The stronger

claim is measured where it IS true — --determinism renders the same plan twice and reports max

abs difference 0.0 on three themes.

**AND THAT PROBE'S OWN POSITIVE CONTROL WAS NOT ARMED, WHICH IS THE MOST SERIOUS THING THIS BLOCK HAS

HAD TO CORRECT.** C14 is two arms — the artifact must show no envelope, and a deliberately faded

control must show one, so that a clean read cannot be a check that is simply unable to speak. The

fade-OUT arm was armed. The fade-IN arm was not: it built its control by ramping the NORMALISED MONO

mix and comparing that against the UN-NORMALISED STEREO head, so what it measured was the normalise

scale (10.92x on JOURNEY_WORLD, 4.06x on FAM_THE_WRITTEN) and the mono fold. Removing the ramp

entirely left it separating just as loudly — 0.4269 and 0.2559 against a bound of 1e-3 — which is the

definition of a control that cannot fail. Two corrections, both measured across all fifteen:

only difference between the artifact arm and the control arm is the envelope. With the ramp

removed the control now reads 0.0 exactly on every one of the fifteen: the arm can fail.

here. Un-normalised head peaks across this roster span 0.00145 (HOME_AND_LOSS) to 0.2176

(FAM_THE_OTHER_SELF) — a factor of 150 — so one constant was a different test on every theme. At

the corrected construction with the old absolute 1e-3, a REAL half-second fade on HOME_AND_LOSS

measured 0.8x of the bound: the control could not clear the threshold it was controlling for.

Relative, every artifact scores 0.0 and every faded control scores 0.5315 to 0.7835, with

the bound at 0.05 between them.

6. **THE RIG'S OWN SEAM HEURISTIC DISAGREES WITH C14 ON TWO ARTIFACTS, AND IT IS PUBLISHED RATHER

THAN QUIETLY DROPPED.** Every artifact also carries rig_loop_seam_heuristic, the older ruler,

printed beside C14 and gating nothing. It reports seam_clean: false on 10 of 15 and

baked_fade_in: true on two PUBLISHED tracks, HOME_AND_LOSS and PATTERN_UNNAMED_PREGATE

against those artifacts' own no-fade claim. The cause was measured rather than assumed, and the

first hypothesis was WRONG: re-rendering both with the release wrap switched off leaves the flag

unchanged, so the wrap is not it. The real mechanism is in the heuristic's own definition — it

calls a fade-in when the opening 1.5 s is quiet against the middle (head_rms/mid < 0.55) AND

rising (trend > 0.45). Those two artifacts measure 0.1428/0.737 and 0.4352/0.7726. That is a

description of a piece that begins softly and grows, which is composition: HOME_AND_LOSS is

the childhood register and PATTERN_UNNAMED_PREGATE is a pre-gate cue whose card forbids a swell

precisely because leaning anywhere would make a listener notice it. The heuristic cannot separate

a written soft entry from an applied envelope. C14 can, because it compares the head against the

same music rendered one statement longer — where an applied envelope differs and these differ by

0.0 — so C14 remains the authority and the heuristic remains reported. The seam_clean figure is

the same class: it is a MEL-SPECTRUM distance, a timbral ruler reading a harmonic device, which is

the split S5 already makes. Both counts are now in the run report so neither can go quiet.

**AND FIVE CARRIERS WERE REPLACED ON EVIDENCE, WHICH IS WHAT "re-orchestrate, never re-score" MEANS IN

PRACTICE.** The horn recovered all 38 of PROTAGONIST_THEME's notes and then a 39th an octave above the

last — as a horn note decays, its second harmonic outlives the fundamental — and borrowed_degree

refuses any wide interval anywhere, so the carrier moved to the bassoon, whose tenor register is what

the brief's "fragile solo statements" arguably wanted first. The oboe's upper partials turned

COMPANION_HEALERS_CHILD's 15-note artifact into 37 recovered notes with spurious pitches a nineteenth

above the line; the flute recovers 18 and finds the borrowed degree. JOURNEY_WORLD's carrier was

decided by RANGE before anything was rendered: three of the four carriers the old generation table

listed cannot play the score at all, because the tune spans A3-F#5, the head's licensed leap ARRIVES on

the F#5, and the sampled horn stops at F5. The same horn defect then took it off FAM_THE_TURNED as

well, where it put 12s and -14s into a 19-note family's recovered line and left the semitone

neighbour-and-return undetectable; a sustained trombone recovers 19 of 19 exactly and is the more

literal reading of "the court, the company, the station" anyway. And the held parent needed an

ARTICULATION rather than an instrument — a legato trombone cannot state where a metrically-early cell

begins, and its staccato twin (same instrument, same register, same card) recovers 7 of 7 entries, all

a beat before a downbeat.

WHAT IS PROMOTED. The tier is AUTHORED_SCORE_ORCHESTRAL_RUNG_ONE, served as *"authored score —

orchestral rung one"*, and it is not rounded up: the note on every page says outright that this is a

public-domain community sample library rather than a professional scoring session, and the four

doctrine diagnostics that need a human listener are still unrun. The picks belt gained R6 — one

theme is served at ONE rung, the highest whose battery is green, and the superseded sketch is refused

BY NAME rather than deleted, because deleting it would erase the comparison that makes the rung

legible. The site's Sound lede was split at the same time: it carried a hard-coded sentence saying

*"what you hear is a SKETCH realisation: additive synthesis, not an orchestra"*, which the orchestral

rung made false on every page it reaches, so provenance and REALISATION are now separate claims with

separately armed controls.

**AND THE SITE DID NOT ACTUALLY SERVE ANY OF THAT UNTIL NOW, WHICH IS THE FINDING THAT MATTERED

MOST.** The commit that landed the rung said the site served it. The BUILD did; the deployed site did

not. The box that ran the build had no encoder, so every track dropped out of the belt, the pages

served no music tiles at all, and nothing was deployed — leaving the live pages serving 13 and 11

sketch-realisation tiles per chapter and still carrying the exact sentence the rung retired. The

commit body said so honestly and the subject line did not, and under the standing phone-first rule a

tier that is not served is not landed. Two things closed it: the encoder question is now a

dependency (imageio-ffmpeg) rather than an assumption, and — the real fix — **the site's own Sound

lede control refused to call a sweep of zero pages clean.** It had swept zero and reported clean on

its first and only run, because a positive control that proves the PREDICATE can fire says nothing

about whether the sweep had anything to read. run_gates.py had already settled that class ("no

gates ran -- refusing vacuous PASS") and the site control was extended without it. A build that

serves no audio can no longer describe itself as clean, which is what let a silent site look like a

landed one.

THE PAIRING LAW WAS TRUE OF HALF THE ARTIFACTS. "Every artifact carries a paired licence +

generation record" held for the audio and not for the MIDI: the writer emitted a licence record for

both and a generation record for the audio only, so every *_ORCHESTRAL_RUNG_ONE_MIDI was

half-recorded — and the tier BELOW this one does emit *_AUTHORED_HEAD_CELL_MIDI.generation.json, so

this rung was also the odd one out. Nothing read the records back, which is why it survived. The MIDI

record is now written, and it is not a copy of the audio's: the MIDI has no renderer and no samples,

and its battery was measured on the AUDIO, so the record REFERENCES that measurement rather than

restating it as though this file had been measured. The claim is now enforced by O09, which pairs

every landed record id off against its opposite number and has a positive control; it went red on

first arming and named all fifteen missing records. The first version of O09 filtered on the TIER

string, which appears in no record filename — a predicate that could not match, reporting a clean

pairing over an empty set. That is the vacuity the check exists to catch, caught in the check itself.

THE STRONGEST OBJECTION, and it is §8.1's own. None of this is a professional orchestral library,

so rung one buys realism and does not buy the AAA bar. The answer is unchanged and is now demonstrated

rather than argued: the deliverable is the SCORE, every path consumes the same MIDI, and this rung

changed nothing upstream of the renderer. What it did change is four measuring instruments — and those

corrections travel to rung two whatever ends up playing it.

Generated by harness/site/structure_site.py — the URL path is the repo path. review root