PIPE_MUSIC_MODELS_2026-08-06.md

music/PIPE_MUSIC_MODELS_2026-08-06.md

PIPE_MUSIC_MODELS — the music-generation model survey, benched

CANON SUBORDINATION — this document is PROPOSAL-TIER: it serves canon and never outranks it.
Authority order: docs/DOC_MAP.md §0. If this document disagrees with canon, CANON WINS and this
document is the defect. Under the pipe-dossiers-bind law it is the **in-place supersession
candidate** for docs/pipeline_review/tech_research/PIPE_AUDIO_MUSIC_2026-07-29.md §1's generator
rows; nothing here is applied to that dossier until the director ratifies it.

Research date 2026-08-06. Every model, licence and figure below is either VERIFIED (fetched

live this pass, URL inline), MEASURED (run on this box tonight, artifact on disk), ON-DISK

(read from the vendor's own files already installed at D:/audio/acestep/repo), or

LANE-REPORTED (returned by one of three parallel research lanes and marked as such). Licences

are read at FACE VALUE per the standing licence law: only an EXPLICITLY triggered prohibition is

flagged, and it is flagged with its clause quoted.

Dispatched by Josh, 2026-08-06, verbatim, after listening to the three loop-proof beds:

*"that music is still terrible."*

---

DERIVATION

DERIVED FROM:
  - docs/spine/DECISIONS_PENDING_JOSH.md § "## RULED 2026-08-06 (remote sitting, in-chat) -
    THE MUSIC GENERATOR VERDICT" — the authority for this lane: "the generator itself is the
    candidate for replacement, the measurement rig and cards carry over unchanged whatever
    generates. The melody-first ruling stands."
  - docs/spine/CH_02.md § Asset anchors · the music_mood bullet — "survival-horror-opener mood,
    low-and-tense rising to the crater-lake climax; instrumentation drawn from the Flores
    gong-waning tradition at reference register (Thread 21 ...); combat-cue register for the
    stalk [POPULATE-> T0_Theme_Registry]" — the canon the rejected samples were realizing, and
    the canon any replacement must realize instead
  - build/audio/proof/TABLES.md — the 21 measured card dimensions per sample per revision; the
    source of the PINNED set this bench is required to test
  - docs/proposals/music/MEASURED_LOOP_PROOF.md §6 ("8-11 of the 21 scalar dimensions moved ...
    10-13 did not move at all", "the stuck ones carry roughly half the residual distance") and
    §4.2 (the wrong-checkpoint defect) — the claim under test
  - docs/proposals/music/MASTERPIECE_PROGRAM.md §11.1-§11.10 — the card vocabulary a candidate
    must be steerable against (lane analysis, event grammar, tempo function, onset/flow law,
    complexity law)
  - build/audio/exemplars/CORPUS.json :: legal_posture — "The bytes NEVER become model input: no
    training, no fine-tuning, no conditioning, no derivation." Binds every candidate: conditioned
    on TEXT / SPEC / our own MIDI only.
  - docs/pipeline_review/tech_research/PIPE_AUDIO_MUSIC_2026-07-29.md §0 (the melody-first and
    30-year bars), §1 (the stack rows this supersedes), §6.0 Q4 and §8 (the melody finding), §8.1
    and §8.2 (the realisation rung already built) — lane law
  - D:/audio/acestep/repo/README.md L130-137, L255-272 (ON-DISK, the vendor's own model zoo and
    VRAM tier table) and
    D:/audio/acestep/repo/acestep/core/generation/handler/generate_music.py L286-292 (ON-DISK,
    the turbo CFG override) — the incumbent's documented operating point
  - build/audio/proof/spec_r2.json, run_r1.json, run_r2.json, control_surface.json — what the
    rejected samples were actually generated with

NOT DERIVED (authored judgment, and why it had no canon home):
  - The whole model roster and every licence read. Canon names the music the game needs; it has
    never named a generator, and `harness/route.py music` routes to composition inputs
    (CH_02 Asset anchors, the region page's Section 10 brief, T0_Theme_Registry,
    T99_Translation_Audio, MUSIC_COMPOSITION_DOCTRINE), not to tooling selection. There is no
    work kind for a model survey; the canon node read instead is CH_02 L234 (the music_mood
    bullet quoted above), plus the ruling that dispatched this lane.
  - The three-arm bench design (§6). The rig and the cards are canon-adjacent tooling; how to
    point them at a candidate is an engineering choice and is grounded in the probe mechanism
    that already existed at masterpiece_proof.py `probe` / `probe-report`.

---

0. THE HEADLINE — findings 1 and 2 say the incumbent was never properly tested. Finding 0 is what happened when it was

**Finding 0, because it is the one that decides the lane and it was produced by an experiment that

failed.** Two config defects (findings 1 and 2) made the case that ACE-Step had never been run at its

own quality tier, so the tier was run tonight: acestep-v15-xl-sft + acestep-5Hz-lm-4B, CFG live,

against the identical 8-dimension two-pole probe. **It scored 0 of 8 responding dimensions against

the pinned turbo config's 3 of 8**, with three dimensions returning literally identical values at

both poles, and supplying the 4B LM moved the numbers barely at all off the LM-less run the lane had

already condemned. **The generator's ceiling is confirmed on measurement rather than on inference,

and Josh's verdict is not a config accident.** §6.2 is the table; §7.3 is what it decides.

---

1. **The rejected music was generated on the configuration ACE-Step's own table prescribes for a

≤6GB card — on a 32GB card — and two of the spec's knobs were being discarded in flight.** The

vendor's VRAM table's bottom row reads *"≤6GB · 2B turbo · None (DiT only)"*, and that is exactly

the checkpoint-and-planner pair the lane ran (the row's INT8 + full-offload half was not used, so

the match is on the model choice, not the whole row). build/audio/proof/run_r2.json

records dit: acestep-v15-turbo, lm: None — the 2B turbo checkpoint with no LM planner — while

spec_r2.json asked every job for guidance_scale: 7.5 and inference_steps: 60. ACE-Step's own

model zoo (ON-DISK, README.md L262) marks acestep-v15-turbo CFG ❌, Step 8; its handler

(L286-292) logs the override on every single call, and it fired 16/16 times in tonight's run:

*"Turbo model detected: overriding guidance_scale 7.5 -> 1.0 (turbo does not use CFG)."* The

vendor's VRAM table (L137) routes a ≥24GB card to **XL sft + acestep-5Hz-lm-4B — "Best

quality, all models fit without offload."** This box has 32GB. So classifier-free guidance — the

mechanism by which a caption steers a diffusion model — was structurally unavailable on the

checkpoint that produced everything Josh has heard, and the step count was 7.5x the distilled

schedule.

2. The landed control-surface probe was inadmissible and says so on its own face.

build/audio/proof/control_surface.json carries positive_controls_responded: false,

admissible: false, and generator: "ACE-Step 1.5 XL-SFT, no LM" — a config MEASURED_LOOP_PROOF

§4.2 had already measured as not separating at all (dark pole 1980 Hz vs bright pole 1976 Hz).

6 of 8 dimensions were measured; the run stopped. **ACE-Step's control surface has therefore never

been measured on any config this lane would ship.** "Half the dimensions don't respond" was a real

observation about a caption-revision loop; it was not a measurement of the generator.

3. No open-weights model closes the fidelity gap, and the leaderboard is unambiguous about it.

VERIFIED: the Artificial Analysis instrumental leaderboard's top ten are Suno V5.5 (1188), Mureka

V9 (1183), Mureka V8 (1165), Suno V5 (1158), Lyria 3 Pro (1117), Suno V4.5 (1084), Music 2.6

(1067), Eleven Music v2 (1062), MiniMax Music 2.5+ (1052), FUZZ-1.1 Pro (1036) — **every one of

them closed or API-only**. And in the one open-model blind arena the survey found (766 votes, the

Khala paper's own table), the **single open model that beats ACE-Step — Khala at Elo 1510.9

against 1470.9 — is CC-BY-NC-4.0 and cannot be used** (§4); the other 2026 quality claim, LeVo 2,

rests on a blog comparative rather than an arena and is academic-licence-only. **Replacing the

generator with another open generator cannot buy AAA fidelity in August 2026. What it can buy is

CONTROL.**

4. A second conditioning channel with a 2048-token budget has been carrying one word. ON-DISK,

conditioning_text.py tokenises the DiT text prompt at max_length=256 (the truncation defect

§4.1 of the loop proof already closed) and the lyric stream separately at max_length=2048

(L142). harness/music_gen/generate.py L452 puts "[Instrumental]" in it. In ACE-Step's training

format that stream is where structure lives. Whether structure tags in it steer form without

summoning vocals is an experiment, not a claim — it is boarded as the cheapest unrun lever in §9,

not asserted here.

5. The strongest replacement candidate is gated behind one click that only Josh can make.

Stable Audio 3.0 Medium (Stability AI, 2026-05-20) is the only surveyed waveform model that is

simultaneously licence-clean at this revenue tier, output-owning, ~6 GB, Windows-native, and

structurally steerable — multi-region mask inpainting is a timeline event grammar, which is

exactly the class of control the pinned dimensions need. `huggingface.co/api/models/stabilityai/

stable-audio-3-medium returns gated: auto`, and an authenticated fetch with the box's own token

returns 403: the licence has not been accepted on Josh's account. Accepting a licence

agreement on his behalf is not this lane's to do. It is §8's one-line Josh-minimum.

---

1. THE RULING THIS ANSWERS

docs/spine/DECISIONS_PENDING_JOSH.md L3884-3893, verbatim:

"that music is still terrible." Recorded as the ruled quality verdict on ACE-Step output at
draft-bed tier. ... CONSEQUENCE: the $900 digital bulk STAYS HELD; the music-model survey launches
... the generator itself is the candidate for replacement, the measurement rig and cards carry over
unchanged whatever generates. The melody-first ruling stands.

Two boundaries this lane honours. The rig and the cards are not on trial — every number below is

produced by acquire_exemplar.py --no-corpus and masterpiece_proof.py probe-report, unchanged.

Candidates are judged as BED / UNDERSCORE generators, per the melody-first ruling; the leitmotif

layer is the authored lane's (author_head_cell.py), which §8 of the lane dossier already promoted

on measurement and which nothing in this survey disturbs.

---

2. WHAT THE INCUMBENT ACTUALLY RAN — the operating point, from the vendor's own files

Vendor's own table (ON-DISK D:/audio/acestep/repo/README.md)SaysThe lane ran
Model zoo L262 — acestep-v15-turboCFG · Step 8 · Quality "Very High" · 2Bthis, at guidance_scale 7.5 (discarded) and inference_steps 60
Model zoo L271 — acestep-v15-xl-sftCFG · Step 50 · Quality "Very High" · 4Bnever, on any probe or shipped bed
VRAM tier L137 — ≥24GB"XL sft (or xl-base for extract/lego/complete) · acestep-5Hz-lm-4B · backend vllm · Best quality, all models fit without offload"2B turbo, no LM, backend pt
Handler L286-292if self.is_turbo_model() and guidance_scale != 1.0: guidance_scale = 1.0fired 16/16 in tonight's probe run (MEASURED, D:/audio/bench/acestep_turbo/gen.log)

The guard-run refusal that pins acestep-v15-turbo is not wrong — it correctly refuses a run that

did not use the config the evidence was gathered on. The defect is one level up: **the config the

evidence was gathered on was chosen to escape a broken XL-SFT run** (`--dit acestep-v15-xl-sft

--no-lm`, whose caption conditioning §6.0 of the lane dossier states plainly "does not arrive without

its LM"), and the repair was to abandon the checkpoint rather than to supply the LM. That

substitution has been carrying every claim about ACE-Step's ceiling since.

Nothing in this section says the music is good. It says the ceiling was measured on the floor.

---

3. THE SURVEY

Three parallel research lanes, each required to cite a page fetched this pass. Steerability is scored

against our card vocabulary (MASTERPIECE_PROGRAM §11): tempo FUNCTION, section/form, lane and

instrumentation control, and the event grammar (drops · raises · breaks · tuttis · silence budget).

A model with only a text prompt scores LOW no matter how good it sounds.

3.1 The waveform generators — song and bed class

ModelOrg · dateLicenceSteerability against our cardsQuality evidence5090 fitVerdict
ACE-Step 1.5 (xl-base/xl-sft/xl-turbo + 5Hz-lm-0.6B/1.7B/4B)ACE Studio · last commit 2026-06-26, releases stop at v0.1.8MITduration · BPM · key/scale · time signature · reference audio · repaint (held-region editing) · cover · retake/flow-edit · lego additive layering · LoRA. No tempo curve, no section map, no dynamics envelopeBT-Elo 1470.9, 2nd open-source in a 766-vote blind arena (Khala paper); does not appear in the Artificial Analysis instrumental top 10≥24GB tier = XL-SFT + 4B LM, no offloadincumbent; never run at its own quality tier — §6
Stable Audio 3.0 Medium (1.4B)Stability AI · 2026-05-20Stability AI Community License — VERIFIED verbatim: *"free for everyone, unless you're using the Core Models for a commercial purpose and you or your organization generate over USD $1M of annual revenue"*; *"You own outputs generated from the Core Models or Derivative Works and therefore can use those outputs at your discretion"*duration conditioning · binary-mask inpainting: single-region, multi-region, causal continuation · audio-to-audio · official LoRA fine-tuning. Mask regions over a timeline ARE an event grammar. No BPM field (BPM goes in the prompt text)FAD 0.107 / CLAP 0.390 (tech report, LANE-REPORTED); 44.1kHz stereo to 6m20s5.07-6.52 GB; MIT code, PyTorch 2.7.1/CUDA 12.6, explicit Windows PowerShell bootstrap, ComfyUI day-0SHORTLIST 1 — blocked on a licence gate only Josh can accept (§8)
Magenta RealTime 2 (2.4B / 230M)Google DeepMind · 2026-06-04Apache 2.0 code + CC-BY-4.0 weights; VERIFIED: *"Google claims no rights in outputs you generate using Magenta RealTime 2."* The cleanest licence in the surveytext · audio-example · MIDI (128-dim multihot pitch state per frame) · ~200 ms control latency. But 2 s chunks over a 10 s context: it is a live-performance instrument, not a form generator — no section map over an 88 s cue*"Model evaluation metrics ... will be shared in our forthcoming technical report"* — no published benchmark; trained on ~71k h of mostly-instrumental stock musicUNGATED. JAX/MLX only; NVIDIA is offline-only, no real-time. Windows-native path UNVERIFIED (WSL2 Ubuntu-24.04 is present on this box)SHORTLIST 3 — licence-safe research track, not the bed factory
Khala 1.0CCoM + Tsinghua · 2026-05-01CC BY-NC 4.0text · lyrics · duration onlybest open-source: Elo 1510.9, 4th overall in its own arena≥24GBREFUSED §4
LeVo 2 / SongGeneration v2-large (4B)Tencent AI Lab · 2026-03-01Tencent custom, academic-onlystructured lyrics + style promptLANE-REPORTED "most natural-sounding clips" in a 2026-05 comparative4BREFUSED §4
DiffRhythm v1.2 / DiffRhythm2ASLP@NPUApache 2.0style prompt · ref audio · instrumental mode · LRC · duration. No tempo/section/lane controlnone published8GB min, Windows supportedpass — no control gain over the incumbent
HeartMuLa-oss-3BHeartMuLaApache 2.0lyrics + tags + length onlyElo 1421.8, below ACE-Steppass; the 7B is still unreleased
YuEM-A-P · last update 2025-06-04Apache 2.0section tags, genre/instrument tags, dual-track ICL2025-era24GB = 2 sessions; 80GB for a full songpass — 14 months stale, worse fit than the 2026-07-29 read
Muse (0.6B)Fudan NLP · 2026-01-11MITadvertises segment-level style control; mechanism not enumeratednonesmallwatch — right control axis, 119 stars
MusicGen / AudioCraft / MusicGen-Stem / Coco-Mulla / JASCOMetaweights CC-BY-NC-4.0JASCO is best-in-class grammar (text + chords + drums + melody, temporally aligned)REFUSED §4 — and JASCO is the painful one
NVIDIA FugattoNVIDIANO WEIGHTS as of 2026-08

HYPE-ONLY or API-ONLY, named so nobody re-discovers them: DiffRhythm+ (paper, no weights);

HeartMuLa-7B (claimed, internal); ByteDance Seed-Music (no public weights); Mureka v8/v9, MiniMax

Music, Suno, Lyria 3 Pro / Lyria 3.5 (Gemini API only, no weights, no model id published);

Stable Audio 3.0 Large (enterprise API); SegTune (ACL 2026, segment-prompt timeline control — exactly

the grammar we want, no weights); MuQ (a representation model, not a generator). **There is no

ACE-Step 2.0 or v1.6** — the lineage's growth since May is a community control layer (chord-progression

editors, LoRA trainers, stem tooling), not a new checkpoint.

3.2 The symbolic route — a different class, scored on the same axes

The two-stage route (a model or our own code emits MIDI; a renderer plays it) scores differently by

construction, because every dimension the cards measure is a symbolic quantity: tempo function,

section count, simultaneous lanes, dynamic range, tutti arrivals and silence budget are all *set* in

a score and only *inferred* from a waveform.

ModelLicenceControl surfaceVerdict
Anticipatory Music Transformer (Stanford CRFM, 780M)Apache 2.0infilling · accompaniment conditioned on a supplied melody · continuation. No tempo/instrumentation/form conditioningthe cleanest fit — it fills bars inside a scaffold our code owns
MuseCoco (Microsoft muzic)MITthe only fetched model whose declared attributes match our suite: instrument, bar count, time signature, key, tempo, pitch range, rhythm intensityno 2025-26 activity
MIDI-GPT (Metacreation, AAAI'25)repo MIT but weights CC-BY-NC-4.0strongest control list found: note density, polyphony, duration, key, pitch range, silence, pitch-class set + bar/track infilling. Tempo and instrument NOT controllableREFUSED §4 — and it is the best control surface in the class
MIDI-LLM (ISMIR 2026, arXiv v2 2026-08-04)HF licence tag llama3.2 (terms UNVERIFIED)text → multitrack MIDI: genre, mood, instrumentation, key, time signature, chords, tempo *adjectives*16GB+, fits; licence read outstanding
NotaGen / NotaGen-XMITperiod-composer-instrumentation only — no tempo, no bar count, no dynamics, no formmaterial, not control
Aria (EleutherAI) · MMT · SkyTNT midi-model · Composer's Assistant 2Apache 2.0 / MIT / MITpiano-only · multi-instrument continuation · instrument+tempo+key seeding (Windows app) · note density + polyphony + instrument-per-track + bar count + tempo (needs REAPER)useful, none decisive
SymphonyGen (arXiv 2026-04-28)code not released32-bar hierarchy, multi-voice skeleton — tempo hard-fixed at 120 BPMno weights

The lane's own finding, stated as it was returned: *"Not one model fetched exposes a tempo curve,

a section map, or a dynamics envelope."* The neural symbolic models buy material, not control. So

the symbolic route's control does not come from a model at all — it comes from author_head_cell.py

and orchestrate.py, which already own tempo, form, lane count, arrivals and silence by construction,

and which §8 of the lane dossier already promoted on measured evidence.

3.3 The renderers — where fidelity actually lives on the symbolic route

OptionLicenceVerdict
sfizz (BSD-2) + VSCO 2 CE + VCSL (CC0)unconditionalincumbent, keep — rung one, already built and measured (PIPE_AUDIO_MUSIC §8.2)
BBC SO DiscoverSpitfire EULA: grant is *"only within your own newly-created sound recording(s)"*; no games/advertising prohibition found; no 2026 EULA change foundrung two on evidence — but it is a proprietary plugin, not SFZ, so sfizz cannot drive it: it costs a plugin host in the chain
Virtual Playing Orchestra v3.3maintainer says *"I enforce no restrictions ... even for commercial purposes"* — but the library is built from Sonatina Symphonic Orchestra, listed on the same page under CC Sampling Plus 1.0REFUSED §4 — this is the licence PIPE_AUDIO_MUSIC §8.1 already refused, arriving under a new name
MIDI-DDSP (Magenta)Apache 2.0repo archived 2024-02-01, TensorFlow 2.7 / Python 3.8 — a dead stack on Blackwell
Symphony Rendering (ICASSP 2026)not specifiedonset F1 0.409-0.477 — it does not faithfully play the notes it is given, which destroys arrival placement and silence budget, the exact dimensions in question
DDSP-VSTfree pluginmonophonic timbre transfer, not an orchestral renderer

---

4. THE REFUSALS — each with the clause that actually fires

The licence law: face value, no conservatism, flag only an EXPLICITLY triggered prohibition.

Five fire, and three of them hurt.

highest-scoring open model in the survey (Elo 1510.9) and it cannot be used. *No workaround: the

NC term is on the weights.*

for academic, research and education purposes, and refrain from using it for any commercial or

production purposes under any circumstances."* Its HF card reads license: unknown, which is not a

grant. **Flagged as LANE-REPORTED rather than VERIFIED — the repo and card returned 401/404 to

direct fetch this pass, and this refusal should be re-verified before it is cited as settled.**

from the 2026-07-29 read. JASCO is the painful one: text + chords + drums + melody, temporally

aligned, is the closest thing in the survey to our event grammar, and it is unusable.

symbolic control surface found, refused.

Plus 1.0**, whose advertising/promotional-use exclusion PIPE_AUDIO_MUSIC §8.1 already ruled fires on a game

soundtrack. A maintainer's blanket permission cannot relicence upstream samples. Recorded here

because a search for "free orchestral SFZ" returns it first, and a later lane would otherwise

re-decide it.

Not a refusal but a second licence chain nobody had flagged: Stable Audio 3 ships a

LICENSE_GEMMA.md, because text conditioning runs through t5gemma-b-b-ul2. VERIFIED from the model

card: *"The Gemma Terms of Use apply; users agree to those terms as well, including the use

restrictions in Section 3.2."* Adopting SA3 means adopting the Gemma prohibited-use policy as

well as the Stability Community License. That read is outstanding and belongs before the first GPU

hour, not after.

---

5. THE SHORTLIST

1. Stable Audio 3.0 Medium — the only candidate that is licence-clean at this revenue tier,

output-owning, Windows-native today, ~6 GB, and structurally steerable via multi-region mask

inpainting, which is the closest thing available to authoring an event grammar over a timeline.

Blocked tonight on a licence gate (§8).

2. ~~ACE-Step 1.5 XL-SFT + acestep-5Hz-lm-4B — the incumbent at the operating point its vendor

recommends for this exact card, which has never been run. Zero install, zero licence, zero

download.~~ **BENCHED AND FAILED TONIGHT — 0 of 8 dimensions respond against the pinned turbo

config's 3 of 8 (§6.2).** It is struck from the shortlist rather than deleted, because the run

that removed it is the evidence that promotes item 1.

3. Magenta RealTime 2 — the best licence in the survey and genuine MIDI conditioning, kept as the

adaptive/interactive research track rather than the bed factory. VERIFIED from Google's own page:

the model generates in 40 ms frames under a local sliding-window attention, its steering is

*"Text, Audio, MIDI"* with *"note and drums on/off control"*, and the hardware framing is

Apple-Silicon-first — *"While the original Magenta RealTime required a high-power GPU or TPU,

Magenta RealTime 2 brings live generation to the hardware musicians actually use"*, with real-time

tiers quoted per MacBook model and no NVIDIA real-time claim anywhere on the page. A

frame-wise live instrument is not a cue-form generator: nothing in it addresses section count or

time-to-full-texture across an 88-second cue.

Why no third candidate was installed tonight, stated rather than left as an absence. Stable Audio

3 is gated (§8). Magenta RT2 is ungated and WSL2 Ubuntu-24.04 is present on this box, so a JAX/CUDA

build is possible — it was declined on evidence, not on effort: its own documentation gives it no

form control, no published benchmark, and no NVIDIA real-time path, so a night spent on the build

would most likely have produced a candidate that scores WORSE on precisely the pinned dimensions this

bench exists to test. The night went instead to the arm with the largest measurable consequence and

zero install cost — §2's finding, run as an experiment.

---

6. THE BENCH — three arms, one rig, the same 8-dimension two-pole probe

The method is the one that already existed, so nothing here is a bespoke measurement built to

flatter a conclusion. masterpiece_proof.py probe emits 16 jobs: eight dimensions, two maximally

opposed captions each, the same seed at both poles, neutral on everything else. Each render is

measured by acquire_exemplar.py --no-corpus — the identical rig that measured the exemplars — and

probe-report rules per dimension: a dimension RESPONDS when the two poles separate by at least a

quarter of its own scale in the correct direction. spectral_centroid_hz and global_bpm

are the declared positive controls; if they do not separate, the whole report is inadmissible,

because a probe that cannot detect brightness tells us nothing by staying silent about the rest.

ArmWhat it isInstall costStatus
A — turbo, no LMthe PINNED INCUMBENT: acestep-v15-turbo, lm: None, 60 steps, guidance_scale 7.5 requested and overridden to 1.0. The exact config behind the three samples Josh rejectednoneMEASURED, §6.1
B — XL-SFT + 4B LMthe vendor's own ≥24GB recommendation: acestep-v15-xl-sft + acestep-5Hz-lm-4B, 60 steps, CFG live at 7.5. Same spec file, same seeds, same captionsnone — both checkpoints already on disk (19 GB + 8 GB)MEASURED, §6.2
C — symbolic + sfizzthe authored route already promoted in §8/§8.2 of the lane dossier, measured on the same rig as a REACH testnoneMEASURED, §6.4

6.1 ARM A — the incumbent's control surface, measured for the first time on a config it ships

build/audio/bench/acestep_turbo/control_surface_turbo.json (16/16 measured; audio at

D:/audio/bench/acestep_turbo/audio).

Dimensionlow polehigh polenormalised spreaddirectionresponds
spectral_centroid_hz *(positive control)*1004.51334.80.2475correctno — by 0.0025
global_bpm *(positive control)*107.67107.670.0no
silence_fraction0.08360.08930.057correctno
mean_lanes5.7535.7830.0052correctno
dynamic_range_db18.6131.330.406correctYES
time_to_full_texture_s2.9021.2770.065wrongno
note_rate_per_s1.6893.0440.4451correctYES
tuttis_per_min0.02.66670.8889correctYES

Read it two ways, because both are true. Formally the report is INADMISSIBLE

(positive_controls_responded: false) and refuses to certify anything. Substantively it is not

silent, because a broken probe does not produce three strong responders: what it says is that on the

pinned config, the caption cannot move tempo at all (identical to two decimal places across

"about 50 bpm" and "about 190 bpm"), brightness misses the bar by a hair, arrangement density and

withholding do not respond, and three dimensions do.

And one of those three corrects a landed claim. MEASURED_LOOP_PROOF §6 lists tuttis_per_min

among the dimensions "pinned at full error in every run — the generator never gathers the full

ensemble." Asked directly, it gathers it: 0.0 → 2.667 tuttis/min, the widest spread in the arm.

The dimension was never unreachable; the revision loop's captions never asked for it in those words.

That is a defect in the revision rules, not in the generator, and it was invisible while the probe

sat inadmissible.

One precision the probe cannot give and must not be read as giving. For global_bpm the probe

deliberately passes bpm: None so the caption is the only route in. Tempo IS reachable on this

generator through the meta parameter — §6.0 Q2 measured turbo at 0.3% error against a stated

target. So the correct statement is narrow: *the caption cannot set tempo; the meta can.*

6.2 ARM B — the vendor-recommended tier, and it is WORSE. The experiment failed, and that is the answer

build/audio/bench/acestep_xlsft_lm/control_surface_xlsft_lm.jsonacestep-v15-xl-sft +

acestep-5Hz-lm-4B, thinking=True, CFG live at 7.5, 60 steps, 16/16 jobs, 0 errors, peak

sampled VRAM 22667–32074 MB on a 32607 MB card (533 MB of headroom — the exclusive-overnight

class is reinforced again, not softened).

Dimensionlow polehigh polenormalised spreadrespondsARM A (turbo) for comparison
spectral_centroid_hz *(pos. control)*1467.61377.00.0617no (wrong direction)0.2475, correct direction
global_bpm *(pos. control)*123.05123.050.0no0.0
silence_fraction0.00.00.0no0.057
mean_lanes6.6276.6060.0032no0.0052
dynamic_range_db14.5814.100.0329no0.406 YES
time_to_full_texture_s1.0680.1860.0353no0.065
note_rate_per_s0.4220.2440.1187no0.4451 YES
tuttis_per_min0.00.00.0no0.8889 YES

0 of 8 respond. Three dimensions returned literally identical values at both poles. The right

reading of "every direction is wrong" is not that the model inverts instructions — it is that at

spreads of 0.003 to 0.12 of scale **the two renders are effectively the same audio, and direction is

noise on a near-constant output.** The captions were "a single unaccompanied solo instrument, alone,

nothing else at any point" against "full orchestra and choir at once, every section playing together

throughout, a wall of sound"; mean_lanes moved by 0.021.

And it refutes the hypothesis the pinned config was chosen on. PIPE_AUDIO_MUSIC §6.0 states

that XL-SFT's "caption conditioning does not arrive without its LM." The 4B LM was supplied, with

thinking on, and the numbers barely moved off the old LM-less run:

XL-SFT no LM (the landed inadmissible probe)XL-SFT + 4B LM (tonight)
global_bpm123.05 / 123.05123.05 / 123.05
silence_fraction0.0 / 0.00.0 / 0.0
dynamic_range_db14.10 / 14.1114.58 / 14.10
mean_lanes7.05 / 6.9776.627 / 6.606
note_rate_per_s0.222 / 0.2220.422 / 0.244

Supplying the LM changed essentially nothing. **The 8-step distilled 2B turbo — the checkpoint whose

CFG is structurally disabled — is measurably MORE caption-responsive than the 4B quality checkpoint

with its planner and CFG live.** That is the opposite of what §2's argument predicted, and the

prediction was made in §7.3 before the result landed, so it cannot be fitted to it.

The one alternative explanation, named rather than argued away. This measures ACE-Step *through

generate.py's in-process invocation* — pt backend where the vendor's ≥24GB row recommends

vllm, 60 steps against a documented 50, and a GenerationParams path this repo wrote. Either

XL-SFT is caption-deaf, or our integration of it is, and this probe cannot separate those. The

discriminating test is cheap and is the named next step: run the same two poles through ACE-Step's

own Gradio UI or its :8001 REST server, where none of our code is in the path. Until that runs,

the claim is bounded to *"XL-SFT does not steer through the path this factory uses"* — which is

still decisive for routing, because that path is the factory.

6.3 THE EAR SAMPLES — the three rejected cues, re-rendered with only the checkpoint changed

build/audio/bench/ear_xlsft_lm/spec_r2_xlsft.json is build/audio/proof/spec_r2.json job for job

— same captions, same seeds (7101 / 7201 / 7301), same bpm, keyscale, duration, steps and guidance.

The only changed field is required_generator, so the checkpoint is the single moving part

against the three tracks Josh rejected. guard-run still has teeth on this arm: the bench spec

declares the XL-SFT config as its requirement rather than bypassing the check.

Audio, for Josh's ear, at D:/audio/bench/ear_xlsft_lm/audio/

MPX_CH02_EXPLORE_r2.wav (88 s, 263 s to generate) · MPX_CH02_BATTLE_r2.wav (81 s, 158 s) ·

MPX_TITLE_MAIN_r2.wav (142 s, 321 s). The turbo originals sit beside them at

build/audio/proof/audio/ for the A/B. Not deployed: serving audible content is the director's

call, and emit_proof_picks.py is the seam when he wants it.

The measured result: further on all three, by the same rig and the same target cards.

Sampleturbo r2 (landed)XL-SFT + 4B LM (tonight)change
MPX_CH02_EXPLORE0.37940.4512+0.0718 further
MPX_CH02_BATTLE0.32860.4579+0.1293 further
MPX_TITLE_MAIN0.35890.3960+0.0371 further
mean0.35560.4350+0.0794 further

And the per-dimension read is more interesting than the composite, so it is not hidden behind it.

XL-SFT is not uniformly worse — it wins decisively on exactly the dimension the loop proof named as

permanently pinned, and loses catastrophically on ones turbo had right (MPX_CH02_BATTLE, error per

dimension):

DimensiontargetturboXL-SFT+LMwho wins
time_to_full_texture_s12.37673.909 (err 1.0)5.201 (err 0.287)XL-SFT, by a mile
mean_lanes6.9633.808 (err 0.453)6.401 (err 0.081)XL-SFT
flow_peak_position0.4560.9309 (err 0.475)0.6316 (err 0.176)XL-SFT
silence_fraction0.01060.0398 (err 0.292)0.0 (err 0.106)XL-SFT
note_rate_per_s2.4992.765 (err 0.106)0.235 (err 0.906)turbo
time_to_hook_s0.1160.116 (err 0.0)58.015 (err 1.0)turbo
onset_typefade_infade_in (err 0.0)cold_statement (err 1.0)turbo
section_count88 (err 0.0)11 (err 0.375)turbo
tuttis_per_min5.9390.0 (err 1.0)0.0 (err 1.0)neither

A DEFECT THIS RUN FOUND BY HITTING IT, and it was hitting the proof lane's own evidence.

masterpiece_proof.py diff accepts --generated-card from anywhere and wrote its output to a

hard-wired build/audio/proof/diffs/<sample>_r<rev>.diff.json. So the first bench diff silently

overwrote three landed, git-tracked diff files — the exact files TABLES.md is derived from — and

nothing complained. They were restored from git, the bench diffs now live in

build/audio/bench/ear_xlsft_lm/diffs/, and the tool gained an --out flag with the reason written

on the argument. A read-from-anywhere flag with a write-to-one-place default is a shared-tree trap;

it had simply never been used from outside the proof lane before. 25/25 controls still pass and the

default path is unchanged.

The shape of XL-SFT's material is consistent across all three samples and across the probe: **a

dense, slow, drifting wash that fills its texture early and then barely moves** — 0.235 to 0.38 note

events per second against targets of 0.9 to 2.5, a hook that arrives 58 seconds in, and a cold

statement where the card wants a fade. That is what a checkpoint the caption cannot steer produces:

its own prior, at length. **The composite distance and the ear are likely to agree here, and the

per-dimension wins do not rescue it** — a cue whose first identifiable material lands a minute in is

not a cue, whatever its lane count measures.

6.4 ARM C — the symbolic route, measured as a REACH test (not a distance test)

Five landed *_ORCHESTRAL_RUNG_ONE artifacts — authored MIDI played by sfizz through VSCO 2 CE and

VCSL — put through the identical rig. **This is not a distance comparison and must not be read as

one:** these are their own themes against their own briefs, not renditions of the BATTLE card. The

question is narrower and it is the one that matters: *do the dimensions ACE-Step pins take

non-degenerate values on this route?*

DimensionFAM_THE_TURNEDHOME_AND_LOSSJOURNEY_WORLDPROTAGONIST_THEMETHE_RECURRENCEACE-Step turbo probe rangeExemplar targets
tuttis_per_min10.2210.590.09.8811.110.0 – 2.671.36 · 5.94 · 12.29
dynamic_range_db4.7024.9210.976.6822.6018.61 – 31.336.71 · 10.59 · 35.84
time_to_full_texture_s0.0460.7899.0790.01.8341.28 – 2.903.81 · 12.38 · 22.11
section_count465737 · 8 · 8
global_bpm129.2123.05117.4595.7103.36107.67 (pinned)99.4 · 103.4 · 143.6
onset_typecold_statementostinato_firstcold_statementcold_statementfade_inostinato_first / fade_in onlyfade_in
silence_fraction0.00.00.00.00.00.084 – 0.0890.0106 · 0.0307 · 0.0618
complexity_steps_per_min0.035.291.760.00.05.20 · 9.32 · 19.02

What it says, in both directions. The dimensions ACE-Step cannot reach — full-ensemble arrivals

at exemplar magnitude, a real dynamic range spread, a controlled withholding time, a chosen onset

type, an actual tempo that differs per cue — **all take live, well-separated values on the symbolic

route, because they are written rather than hoped for.** onset_type is the cleanest single

illustration: ACE-Step returned only two of the eight onset classes across every artifact this lane

has produced, and it is scored at a full 1.0 miss on every revision; the authored route produces

three different classes across five cues, including the fade_in the cards actually target.

And the honest negative, which is a build item and not a ceiling. silence_fraction is **0.0 on

all five and complexity_steps_per_min is 0.0 on three of five**. The silence budget is a

first-class card parameter (MASTERPIECE_PROGRAM §11.2 — *"restraint is a masterpiece mechanism and

the current factory has no vocabulary for it"*) and the authoring code does not write rests or

step complexity at section boundaries. Those two dimensions are not unreachable on this route; they

are unauthored. That is a named defect in orchestrate.py / author_head_cell.py, and it is

the kind of defect the route can actually fix — which is the whole difference between the two

architectures.

---

7. THE RECOMMENDATION

**Split the lane by what each half is actually for. The SCORE stops using a waveform generator at

all. The BED lane keeps one, and its replacement is named — but the incumbent's own trial had to run

first, and running it is what made the replacement case evidence instead of assertion.**

7.1 — The SCORE lane's generator is not a waveform model, and the survey settles it. Every

dimension the cards measure is a symbolic quantity: a tempo function, a section map, a lane count, a

dynamics envelope, an arrival schedule, a silence budget. In a score they are set by construction and

measurable exactly; in a waveform they are neither settable nor reliably measurable. The survey

found **no model of any class — waveform or symbolic — that exposes a tempo curve, a section map, or

a dynamics envelope.** So anything with FORM (the twelve canonical themes, chapter cues, boss

phases, everything the 12-20-tracks-per-chapter target means) is authored and rendered, exactly as

§8 and §8.2 of the lane dossier already promoted on measurement. §6.4 is that route measured against

the pinned dimensions, and it reaches them.

7.2 — Fidelity on that lane is bought with a LIBRARY, not with a model. This is the part that

answers Josh's actual complaint rather than deflecting it. The reason rung one still sounds like a

mockup is that VSCO 2 CE and VCSL are public-domain community sample sets, not a scoring session.

The ladder above them is already read and licence-cleared (§8.1's rung two, BBC SO Discover) and its

cost is a plugin host in the chain, not a research programme. **No generator improving changes this,

and no library improving depends on a generator.**

Two different things are called "rung two" and this document conflates neither. §8.1's rung two

is a RENDERER upgrade — a larger sample library under the same authored score. The *recompose* rung

two is a COMPOSITION upgrade, and it already ran: docs/proposals/music/RECOMPOSE_RUNG_TWO_FINDINGS.md

records a cold verification that returned NOT-YET on eight findings, the pick demoted, the

apparatus kept, three findings marked BINDING on any future rung. **This survey disturbs none of

that.** It is evidence about generators; that board is the score lane's own open work, and its

depth-law freeze behind the world pipeline is unaffected by anything here.

**7.3 — The BED/AMBIENCE lane keeps a waveform generator, and the incumbent got exactly one honest

trial at the operating point its vendor recommends. It failed.** The decision rule was written into

this section before the result landed, so it cannot be fitted to it: *if arm B's positive controls

respond and it reaches materially more of the eight dimensions than arm A, the pinned config moves to

XL-SFT + the 4B LM; if it does not, ACE-Step's control surface is confirmed as the ceiling on

evidence rather than on inference.* Arm B scored 0 of 8 against arm A's 3 of 8 (§6.2). So:

needs no change — but its control_surface_evidence should now point at

build/audio/bench/acestep_turbo/control_surface_turbo.json, which is the first probe of it that

was ever run, and the landed build/audio/proof/control_surface.json should be marked as what it

is: a probe of a config this lane refuses, superseded.

config (0.0 → 2.667) and was written off as unreachable. Whatever else is true of this generator,

the loop has been leaving a responding dimension on the table because no revision clause ever asked

for a full-ensemble arrival in the words the probe used.

inpainting — the only timeline event grammar in the survey available under a licence that ships —

and it is blocked on §8's one click.

through ACE-Step's own Gradio UI or :8001 REST server, with none of our code in the path. If

XL-SFT steers there and not here, the defect is ours and it is worth more than a new model.

7.4 — Nothing ships to Josh's ear until it passes his ear, and the $900 stays held. The ruling

already held it; nothing in this survey argues for releasing it. §8 of the lane dossier's own caution

stands and is now stronger: buying more exemplar cards sharpens the TARGETS, and the targets were

never the problem.

THE STRONGEST OBJECTION

"You were told the music sounds terrible and you have answered with a control argument." That is

a fair reading and it is the one to beat. Josh's verdict is about how the audio SOUNDS. Every number

in §6 is structural — the rig measures form, density, events and spectrum, and it is blind to timbre,

legato, room, ensemble realism and whether a phrase is beautiful. A cue can score perfectly on all 21

dimensions and still be a wash. The leaderboard finding cuts the same way: the models that sound good

are closed, and none of the open ones will close that gap by being steered better. **If the failure

is timbral rather than structural, most of §6 is beside the point and the metric suite is one step

from being Goodharted.**

The answer concedes the half that is true and changes the plan rather than defending it.

(a) The fidelity path is named, and it is §7.2 — a real orchestral library played from an authored

score is the only route in this entire survey that can sound like an orchestra, and it depends on no

model improving. (b) Control is not offered as a substitute for fidelity; it is the precondition for

volume. Twelve to twenty cues per chapter that each have to FIT a scene cannot come from a generator

that cannot be told when the ensemble arrives — however good any individual sample sounds. (c) The

gap the objection exposes is real and is **boarded, not waved off: a blind timbre/realism critic that

is not computed from the score.** Every current diagnostic is derived from what was authored, so

nothing in the loop can currently fail an artifact for sounding fake. Until that exists, §6's numbers

are NECESSARY and never SUFFICIENT, and this document does not claim otherwise.

---

8. JOSH-MINIMUM — one click, and it is the only thing this lane is blocked on

Accept the licence on https://huggingface.co/stabilityai/stable-audio-3-medium. The repo is

gated: auto; an authenticated fetch with the box's own token returns 403. Accepting a licence

agreement and sharing contact details on his account is not this lane's to do under the standing

action rules, so the survey's first-ranked candidate could not be downloaded, installed or benched

tonight. It is one click plus "agree to share your contact information".

Nothing else. No purchase, no API key, no negotiation. The BBC SO Discover rung is a free download

whose EULA is already read (§3.3), and it is not blocking anything today.

---

9. WHAT THIS DOCUMENT DOES NOT ANSWER

Audio 3 is gated; Magenta RT2 was declined on evidence (§5). The survey's rankings for those two

are LANE-REPORTED and VERIFIED reads, never measurements.

carrying one word. Whether structure tags in it steer form without summoning vocals — and whether

they can coexist with the [instrumental] flag that server_utils.py L143 derives from that exact

field — is a 4-generation experiment nobody has run. It is the cheapest unexplored lever in the

lane.

dossier remains exactly as it was: all 12 region cells are HELD pending the §17.1 read, and

nothing in this document may be cited on cultural-substrate accuracy.

deploys of audible content are the director's.

401/404 to direct fetch this pass and the GitHub API rate-limited. Re-verify before citing it as

settled.

for eight; the other thirteen (section_count, drops_per_min, breaks_per_min,

boundary_step_share, complexity_steps_per_min, flow_peak_position, dynamics_peak_position,

melodic_salience, time_to_hook_s, max_lanes, total_length_s, onset_type, raises_per_min)

have no pole pair written, so "3 of 8" and "0 of 8" are counts over the probed subset and must

never be restated as "3 of 21". Writing the missing thirteen pole pairs is a cheap, mechanical

extension and is the honest way to make the headline number the one the ruling actually asked for.

§6.3's per-dimension table covers all 21 on real cue specs, which is the nearest thing this run has

to the full picture.

and names the discriminating test (the same poles through ACE-Step's own UI or REST server). Until

that runs, no sentence anywhere should say "XL-SFT cannot be steered" without the qualifier

"through the path this factory uses".

---

10. THE ARTIFACTS THIS RUN LANDED

PathWhat it is
build/audio/bench/acestep_turbo/probe_spec.json · run_probe.json · control_surface_turbo.json · cards/ARM A — the incumbent's first admissibility-tested control-surface probe
build/audio/bench/acestep_xlsft_lm/probe_spec.json · run_probe.json · control_surface_xlsft_lm.json · cards/ARM B — the same probe at the vendor-recommended tier
build/audio/bench/ear_xlsft_lm/spec_r2_xlsft.json · run_r2_xlsft.json · cards/ · diffs/the ear A/B — spec_r2.json job-for-job, only required_generator changed
build/audio/bench/symbolic_sfizz/cards/ARM C — the authored route's reach on the same rig
D:/audio/bench/{acestep_turbo,acestep_xlsft_lm,ear_xlsft_lm,symbolic_sfizz}/audio/the audio, off-repo per the media contract
harness/music_gen/masterpiece_proof.pydiff --outthe shared-tree fix §6.3 names, with the reason on the argument; 25/25 controls green, default path unchanged
docs/line_citation_baseline.jsonre-emitted SCOPED (--own-surface on this file alone) in the same commit, per the C5 same-commit re-emit law. A plain emit would have declared a concurrent lane's uncommitted rows: the run reported *"3 owned citations merged, 5922 foreign entries carried through as committed"*, and the diff is 5 insertions / 2 deletions.

11. SOURCES

VERIFIED — fetched live 2026-08-06 by the synthesizer: stability.ai/license ·

huggingface.co/stabilityai/stable-audio-3-medium (+ its /api/models/... gating state) ·

huggingface.co/google/magenta-realtime-2 · github.com/magenta/magenta-realtime ·

magenta.withgoogle.com/magenta-realtime-2 · magenta.github.io/magenta-realtime ·

artificialanalysis.ai/music/leaderboard/instrumental.

ON-DISK — read from the installed vendor files: D:/audio/acestep/repo/README.md (model zoo

L255-272, VRAM tiers L130-145) · acestep/core/generation/handler/generate_music.py L180-292 ·

acestep/core/generation/handler/conditioning_text.py L125-145 · acestep/api/server_utils.py L143.

LANE-REPORTED — three parallel research lanes, each citing pages fetched this pass: the ACE-Step

lineage sweep (github.com/ace-step/ACE-Step-1.5 releases/commits, huggingface.co/ACE-Step,

awesome-ace-step, Khala paper arXiv 2605.01790 with its 766-vote arena table, github.com/Khala-Music-AI/Khala,

github.com/HeartMuLa/heartlib, github.com/multimodal-art-projection/YuE, github.com/ASLP-lab/DiffRhythm,

github.com/yuhui1038/Muse, it-jim.com comparative 2026-05-21) · the diffusion/big-lab sweep

(github.com/Stability-AI/stable-audio-3, huggingface.co/stabilityai/stable-audio-3-small-sfx,

stability.ai community-license-agreement, huggingface.co/facebook/jasco-chords-drums-melody-1B,

github.com/fundwotsai2001/MuseControlLite, github.com/juhayna-zh/AudioControlNet,

FUGATTO publication PDF, marktechpost/techcrunch SA3 coverage) · the symbolic+renderer sweep

(github.com/ElectricAlexis/NotaGen, github.com/jthickstun/anticipation +

huggingface.co/stanford-crfm/music-large-100k, github.com/Metacreation-Lab/MIDI-GPT +

huggingface.co/Metacreation/MIDI-GPT, github.com/slSeanWU/MIDI-LLM, github.com/AMAAI-Lab/Text2midi,

github.com/m-malandro/composers-assistant-REAPER, microsoft.github.io/muzic/musecoco,

github.com/EleutherAI/aria, github.com/SkyTNT/midi-model, arXiv 2604.25498 SymphonyGen,

symphony-rendering.github.io, github.com/magenta/midi-ddsp, sfzinstruments.github.io/orchestra,

virtualplaying.com/virtual-playing-orchestra, spitfireaudio.com EULA).

UNVERIFIED, named as such: tencent-ailab/SongGeneration licence text (401/404 + API rate limit) ·

MIDI-LLM's llama3.2 licence terms · Magenta RT2's Windows/WSL2 CUDA build.

---

12. THE GATE OPENED, AND THE CANDIDATE WAS BENCHED — 2026-08-06, same sitting

IN-PLACE SUPERSESSION with provenance. §8 named one blocker: a licence gate only Josh could

accept. He accepted it (form submitted on the account matching the box token; authenticated download

now succeeds). §5's shortlist item 1 and §7.3's "becomes the replacement candidate" are no longer

predictions — they are measured. Nothing above this line is edited; the earlier text is the

record of what was believed before the arm ran, and this section is what the arm returned. Where

they disagree, this section wins and the disagreement is named.

Honest register, first. These are BEDS. The SCORE lane's verdict in §7.1 is unchanged and

this section does not touch it. No claim is made anywhere below that any of this music is good;

§12.4 is the closest thing to a quality instrument in the repo and its own first finding is that it

admits the tracks Josh rejected.

12.1 What was installed, and the two corrections the first-hand licence read forces

WeightsD:/models/stable-audio-3-mediummodel.safetensors 8.79 GB, sha256 48d9c65e290e7bcd5194e0633bfc2424a59ee9683f5c2d58762d997b7d8ce0b5 (matches the hub's own content hash), plus t5gemma-b-b-ul2 1.18 GB, sha256 9b05ea5a4f211d023832f706fb2c0e83e4fc721b6da35ab69ceb0b55eb7800d3
Licence records, pinneddocs/licence_records/stable_audio_3_medium/{LICENSE,LICENSE_GEMMA,NOTICE} — sha256 d6f6b1a4…, e77acc0d…, 66f856d7…. Read FIRST-HAND from the pinned files, not from the web summary §3.1 carried. Extensionless on purpose (matching sfizz/, vcsl/, vsco2_ce/): landed as .md they became doc-graph ISLANDS and the ledger gate went red on L3 island growth, correctly — a licence BODY is evidence, not a document with a consumer. Do not rename them back
StackD:/audio/bench/sa3/venv — Python 3.12.10, torch 2.7.1+cu128, CUDA 12.8, sm_120 on the RTX 5090, stable-audio-3 0.1.0 from the MIT repo. Windows-native, no WSL2
Adapterharness/music_gen/sa3_generate.py, 9/9 controls

1. A genuinely triggered OBLIGATION was missed. §3.1 quoted the revenue threshold and the

output-ownership sentence and stopped there. Section III of the pinned text also says: *"If You

are using or distributing the Stability AI Materials for a Commercial Purpose, You must register

with Stability AI at (https://stability.ai/community-license)."* Making music for a game that

ships is a Commercial Purpose, so registration is required and it is not satisfied. It is an

account action; §12.6 boards it. Every generation record this lane writes carries the requirement

so it travels on the artifact instead of living in a note.

2. **The attribution clause, read at face value, does NOT fire on shipping audio — and the reason is

worth writing down because it is the clause people assume fires.** IV(a) obliges attribution when

You distribute "the Stability AI Materials or a Derivative Work … or a product or service that

uses any portion of them". Section V defines Derivative Work and then excludes model output

explicitly: *"but do not include the output of any Model."* A shipped game carries rendered audio

and no portion of the weights. So IV(a) is not triggered by audio-only shipping, IV(c)(iii) gives

us the outputs, and Gemma 3.3 says the same on its chain (*"Google claims no rights in Outputs"*).

**What IS forbidden and must travel on every row: no Stability weight or output may ever become a

conditioning, fine-tune or LoRA input to a foundational model** — IV(b) — and the corpus posture

bars audio conditioning from the other direction.

One thing the config file said that the survey did not know. model_config.json declares the

prompt conditioner as {"type": "t5gemma", "max_length": 256}. **Stable Audio 3 has the same

256-token caption ceiling ACE-Step does.** The budget discipline MEASURED_LOOP_PROOF §4.1 built for

one generator is not generator-specific, and it carries over unchanged.

12.2 THE CONTROL SURFACE — 4 of 8 respond, and 7 of 8 MOVE

Identical probe. Same eight dimensions, same two opposed captions, same seed 8800 at both poles,

same 45 s, measured by the same acquire_exemplar.py --no-corpus, ruled by the same probe-report.

The one fairness adjustment is declared on every artifact: SA3 has no bpm meta channel, so a declared

bpm is appended to the caption as the model card's own convention (bpm_route: caption_text) — and

for the global_bpm dimension the spec declares bpm: null, so **both generators were asked that

question through the caption alone.**

Dimensionlow polehigh polespreadrespondsACE-Step turboACE-Step XL-SFT+LM
spectral_centroid_hz *(pos. control)*1497.82618.60.428YES0.2475 no0.0617 no
global_bpm *(pos. control)*129.295.70.2593no — inverted0.00.0
silence_fraction0.02790.32151.000YES0.057 no0.0 no
mean_lanes5.0867.2500.2985YES0.0052 no0.0032 no
dynamic_range_db13.2127.060.5118YES0.406 YES0.0329 no
time_to_full_texture_s19.2261.2770.718no — inverted0.065 no0.0353 no
note_rate_per_s2.8441.3110.539no — inverted0.4451 YES0.1187 no
tuttis_per_min8.08.00.0no — pinned0.8889 YES0.0 no
RESPOND4 / 83 / 80 / 8
MOVE AT ALL (spread ≥ 0.25)7 / 84 / 80 / 8

The formal verdict is still INADMISSIBLE, and for a smaller reason than before. One of the two

positive controls now fires — brightness separates by 0.428, the first time any arm has moved it —

but global_bpm moves 0.2593 in the *wrong* direction, so probe-report refuses to certify, and

that refusal is reported rather than argued around.

The failure mode is different in kind, and that is the finding. ACE-Step mostly did not move:

on turbo four dimensions sat under 0.07, and on XL-SFT all eight sat under 0.12. Stable Audio 3

moves hard and sometimes backwards — three dimensions (global_bpm, time_to_full_texture_s,

note_rate_per_s) travel 0.26 to 0.72 of their own scale in the opposite direction to the

instruction. A model that hears an instruction and inverts it is a prompt-phrasing and

negative-prompt problem. A model that does not move is a ceiling. **Those are not the same problem,

and only one of them is ours to fix.**

And tuttis_per_min says something neither number alone does. It is pinned — but pinned *at

8.0*, where ACE-Step XL-SFT sat pinned at 0.0 and turbo ranged 0.0–2.67 against exemplar targets of

1.36 / 5.94 / 12.29. Stable Audio 3 gathers the full ensemble by default and cannot be talked out of

it; ACE-Step's quality tier never gathers it at all. For a BED the SA3 default is arguably the wrong

one — but it is a DEFAULT, not a ceiling, and the difference matters.

12.3 THROUGHPUT — the number that changes what a ladder costs

Arm45 s probe renderthe three cues (88 s + 81 s + 142 s)VRAM
ACE-Step turbo~10 s~3 min8.4 GB peak
ACE-Step XL-SFT + 4B LM14–95 s741 s22.7–32.1 GB (533 MB headroom)
Stable Audio 3 Medium1.4–6.5 s9.4 s~9.5 GB over baseline

The whole 16-job probe generated in about 30 seconds of GPU time, on a card already holding 20 GB

of the art lane's gen-4 face refine — no queueing was needed and nothing was killed. XL-SFT's same

probe took roughly twenty minutes and owned the card. **A candidate ladder that costs seconds rather

than minutes is a different kind of tool**: selection pressure becomes affordable, which is the one

lever PIPE_AUDIO_MUSIC §6.0 Q4 named that this lane has never been able to pull at scale.

12.4 THE BLIND TIMBRE CRITIC — built, and its first finding is against this survey's own argument

harness/music_gen/timbre_critic.py, 9/9 controls. §9 boarded it in one line: *"a blind

timbre/realism critic that is not computed from the score."* It reads only a measured card's

timbre and register blocks — spectral centroid, 85% rolloff, spectral flatness, and the ten-band

octave profile — and scores them by robust z against the 197 measured exemplar cards, real

released game soundtracks. It never opens a MIDI file, and --self-test T5 proves that structurally

by parsing its own reader and asserting that no structural block of the card is even named. The bands

are not chosen: they are the leave-one-out 90th and 98th percentiles of the corpus scoring itself

(1.658 / 2.669), and white noise and a single-band sine stack both read OUTSIDE.

Armnmedian timbre distanceverdictsmedian spectral flatness
exemplars (the population)1970.663 (p50)90.4% INSIDE by construction0.01695
ACE-Step XL-SFT + 4B LM (probe)160.23416 INSIDE0.00853
Stable Audio 3 Medium (probe)160.46316 INSIDE0.00730
ACE-Step turbo (probe)161.14316 INSIDE0.00055
ACE-Step turbo — the three cues Josh rejected60.9826 INSIDE0.00055
authored + sfizz, orchestral rung one51.9325 MARGINAL0.00008

THE FIRST THING TO SAY ABOUT THIS INSTRUMENT IS WHAT IT GETS WRONG. It admits, at 0.982, the

exact six artifacts whose ruling launched this survey. That is not a defect to be tuned away — it is

the calibration of what a PASS is worth, printed in the same table as the PASS. **A pass here means

"the spectrum is not alien", and nothing more.** Josh's ear overruled a set of tracks this critic

calls INSIDE, so the critic can never be cited to argue with him.

THE SECOND THING CUTS AGAINST §7.1 AND §7.2, AND IT IS THE MOST USEFUL RESULT OF THE SITTING. The

authored + sfizz route — the one §6.4 showed reaching every structural dimension the generators pin —

is the arm furthest from real released game music, the only one not INSIDE, and the gap is

overwhelmingly one feature: its median spectral flatness is **0.00008 against the corpus's 0.01695,

roughly two hundred times more tonal**, a robust z of 4.298. That is what a small public-domain

sample library playing sparse authored lines with no percussion, no room and no dense mix measures

like. §8.2 of the lane dossier said those words as a caveat; this is the number.

So the two rulers disagree, and the disagreement is the honest shape of the lane:

bought with a LIBRARY" is now measured as not yet delivered by rung one's library. That

upgrades §8.1's rung two from an optional improvement to the named lever, and gives it an

acceptance test it did not have: *a rung-two render must move the flatness term toward the

population, and this critic will say by how much.*

Two caveats on the instrument, published rather than buried. Its measured blind spot: a

band-limited mid-only spectrum reads INSIDE at 0.75, so it screens for an alien SPECTRUM and NOT for

a missing bottom or top end — control T4b asserts that value rather than deleting the control,

because a limit you cannot fix today should be pinned where it cannot drift. And across every arm the

deciding term was spectral_flatness_mean, so in practice this is mostly a tonality detector and

should be described as one.

12.5 THE EAR SAMPLES — the same three cues, a fourth generator

build/audio/bench/sa3/spec_r2_sa3.json is build/audio/proof/spec_r2.json job for job: same

captions, same seeds (7101 / 7201 / 7301), same durations. Audio for Josh at

D:/audio/bench/sa3/audio_ear/MPX_CH02_EXPLORE_r2.wav · MPX_CH02_BATTLE_r2.wav ·

MPX_TITLE_MAIN_r2.wav, beside the turbo originals at build/audio/proof/audio/ and the XL-SFT arm

at D:/audio/bench/ear_xlsft_lm/audio/. Four generators, three cues, one spec. Not deployed

serving audible content is the director's call, and emit_proof_picks.py is the seam when he wants it.

12.6 THE DIFFS — Stable Audio 3 is closer on all three cues, and it wins where the misses were LARGE

Same rig, same target cards, same distance function. build/audio/bench/sa3/diffs/.

SampleACE-Step turbo (the rejected set)ACE-Step XL-SFT + 4B LMStable Audio 3 Mediumchange vs turbo
MPX_CH02_EXPLORE0.37940.45120.2832−0.0962 (25% closer)
MPX_CH02_BATTLE0.32860.45790.2881−0.0405 (12% closer)
MPX_TITLE_MAIN0.35890.39600.2617−0.0972 (27% closer)
mean0.35560.43500.2777−0.0779 (22% closer)

**That is the best distance any arm has produced, and the per-dimension read says exactly why —

which is more interesting than the total.** On MPX_CH02_BATTLE, counting per dimension, turbo

actually wins 6, Stable Audio 3 wins 5 and 6 tie. The composite still moves decisively because **the

wins are not the same size**:

DimensiontargetturboSA3
time_to_full_texture_s12.37673.909 (err 1.0)18.646 (err 0.251)SA3, by a mile
mean_lanes6.9633.808 (err 0.453)6.536 (err 0.061)SA3
tuttis_per_min5.9390.0 (err 1.0)0.741 (err 0.875)SA3
spectral_centroid_hz1059.7824.9 (err 0.196)1068.8 (err 0.008)SA3
flow_peak_position0.4560.9309 (err 0.475)0.2305 (err 0.226)SA3
onset_typefade_infade_in (err 0.0)cold_statement (err 1.0)turbo
section_count88 (err 0.0)9 (err 0.125)turbo
note_rate_per_s2.4992.765 (err 0.106)1.444 (err 0.422)turbo
silence_fraction0.01060.0398 (err 0.292)0.0553 (err 0.447)turbo
curve distance0.27850.2073SA3

**Stable Audio 3 wins the dimensions that were failing at full error and loses the ones that were

already nearly right.** time_to_full_texture_s — pinned at a full 1.0 miss on every ACE-Step

revision this lane has ever run, and named in MEASURED_LOOP_PROOF §6 as one of the dimensions

carrying half the residual distance — lands at 0.251. The six curves, which no caption revision has

ever moved much, come down 0.2785 → 0.2073.

Its two real regressions are both nameable and one is cheap. onset_type goes from a correct

fade_in to cold_statement at a full 1.0 miss on all three cues — SA3 starts hard. The engine

forbids a baked fade (the stem/loop contract), so this is an ARRANGEMENT ask, not a mastering one,

and it is exactly the kind of thing an 8-step, 2-second generator can be laddered against. And it is

thinner: note_rate_per_s 1.444 against a 2.499 target.

12.7 THE TIMBRE CRITIC ON THE CUES — and it FIRES on a real artifact, not just on a control

Armnmedianverdicts
SA3 (probe)160.46816 INSIDE
SA3 (the three cues)30.7812 INSIDE, 1 OUTSIDE
ACE-Step XL-SFT+LM (probe)160.23116 INSIDE
ACE-Step XL-SFT+LM (3 cues)30.3223 INSIDE
ACE-Step turbo (probe)161.13516 INSIDE
ACE-Step turbo (rejected cues)60.9746 INSIDE
authored + sfizz (rung one)51.9225 MARGINAL

MPX_CH02_EXPLORE_r2 reads OUTSIDE at 2.79 — flatness z 4.64, centroid z 2.96, rolloff z 2.63.

**This is the first time the critic has failed a real generated artifact rather than a synthetic

control, and the reading is genuinely ambiguous in a way that must be stated rather than resolved by

preference.** The EXPLORE spec asks for the sparsest thing this lane generates — two or three lanes

at once, upper register left empty, a 6% silence budget — so a very tonal, quiet, narrow-band result

is what the *instruction* asks for. And MEASURED_LOOP_PROOF §8 already recorded that the exemplar

corpus is thinnest exactly here ("almost nothing for melancholic/tragic"). **So the OUTSIDE verdict

is either the artifact being spectrally odd or the population being too thin to judge sparse cues,

and this run cannot separate them.** The resolution is corpus, not tuning: it is one more argument

for the class anchors the purchase run was built to buy, and the critic must not be cited against a

sparse cue until that population exists.

12.8 WHAT THIS CHANGES, AND WHAT IT DOES NOT

CHANGES — the bed lane's generator, on evidence. §7.3 said Stable Audio 3 "becomes the

replacement candidate" if arm B failed. Arm B failed and the candidate then beat the incumbent on

every axis this lane can measure: **4/8 responding dimensions against 3/8 and 0/8; 7/8 dimensions

moving at all against 4/8 and 0/8; the best distance on all three cues (mean 0.2777 against 0.3556);

16/16 INSIDE on the blind timbre screen; and roughly eighty times the throughput.** The recommendation

is therefore PROMOTE Stable Audio 3 Medium to the bed/ambience lane's generator, with ACE-Step

retained rather than deleted for the two things it still does better — a correct fade_in onset and

a denser note rate — and because deleting it would erase the comparison that makes this legible.

DOES NOT CHANGE — the score lane, the honest tier, or the $900. §7.1 stands untouched: the SCORE

is authored and rendered, because every card dimension is a symbolic quantity and no generator in

this survey exposes a tempo curve, a section map or a dynamics envelope. §12.4 strengthens that

split rather than weakening it. Nothing here is a quality claim: the honest tier on every artifact

this section produced is BENCH, the generation records say so in those words, and **Josh's ear is

still the only gate that has ever overruled anything.**

BOARDED, with the reason each is not done:

1. Register with Stability AI — Community Licence Section III, triggered by commercial use, not

satisfied. An account action; §12.9.

2. The three inverted dimensions. global_bpm, time_to_full_texture_s and note_rate_per_s

move hard in the wrong direction. The lever is prompt phrasing and SA3's negative_prompt

argument, which this run did not touch at all — a two-pole probe with negatives is the next cheap

experiment and it is now affordable at 2 s a render.

3. onset_type is a full 1.0 miss on all three cues. SA3 starts cold; the cards want fade_in

and the engine forbids a baked one. This is an arrangement instruction to ladder against.

4. The multi-region mask was never used. The whole reason SA3 was shortlisted in §3.1 is

inpaint_mask_start_seconds / ..._end_seconds accepting LISTS — a timeline event grammar. This

run measured plain text-to-audio only, so **the candidate's headline control feature is still

unmeasured** and the promotion above rests on its text conditioning alone.

5. The probe still covers 8 of 21 dimensions. Unchanged from §9. "4 of 8" must never be restated

as "4 of 21".

6. No region cue, so the Care-Doctrine question is still untouched on this generator too.

12.9 JOSH-MINIMUM, UPDATED

§8's licence-gate item is DONE — he accepted it and the arm ran. One item replaces it, and it is

smaller:

Register the commercial use at https://stability.ai/community-license. The Community Licence's

Section III requires it of any Commercial Purpose user; the grant is still royalty-free and still

free below USD $1M annual revenue. It is a form, not a purchase. Until it is done, everything this

lane generates on Stable Audio 3 carries triggered_obligation: NOT SATISFIED on its own licence

record, which is the honest state and not a blocker for benching.

Generated by harness/site/structure_site.py — the URL path is the repo path. review root