music/PIPE_MUSIC_MODELS_2026-08-06.md
CANON SUBORDINATION — this document is PROPOSAL-TIER: it serves canon and never outranks it.
Authority order: docs/DOC_MAP.md §0. If this document disagrees with canon, CANON WINS and this
document is the defect. Under the pipe-dossiers-bind law it is the **in-place supersession
candidate** for docs/pipeline_review/tech_research/PIPE_AUDIO_MUSIC_2026-07-29.md §1's generator
rows; nothing here is applied to that dossier until the director ratifies it.
Research date 2026-08-06. Every model, licence and figure below is either VERIFIED (fetched
live this pass, URL inline), MEASURED (run on this box tonight, artifact on disk), ON-DISK
(read from the vendor's own files already installed at D:/audio/acestep/repo), or
LANE-REPORTED (returned by one of three parallel research lanes and marked as such). Licences
are read at FACE VALUE per the standing licence law: only an EXPLICITLY triggered prohibition is
flagged, and it is flagged with its clause quoted.
Dispatched by Josh, 2026-08-06, verbatim, after listening to the three loop-proof beds:
*"that music is still terrible."*
---
DERIVED FROM:
- docs/spine/DECISIONS_PENDING_JOSH.md § "## RULED 2026-08-06 (remote sitting, in-chat) -
THE MUSIC GENERATOR VERDICT" — the authority for this lane: "the generator itself is the
candidate for replacement, the measurement rig and cards carry over unchanged whatever
generates. The melody-first ruling stands."
- docs/spine/CH_02.md § Asset anchors · the music_mood bullet — "survival-horror-opener mood,
low-and-tense rising to the crater-lake climax; instrumentation drawn from the Flores
gong-waning tradition at reference register (Thread 21 ...); combat-cue register for the
stalk [POPULATE-> T0_Theme_Registry]" — the canon the rejected samples were realizing, and
the canon any replacement must realize instead
- build/audio/proof/TABLES.md — the 21 measured card dimensions per sample per revision; the
source of the PINNED set this bench is required to test
- docs/proposals/music/MEASURED_LOOP_PROOF.md §6 ("8-11 of the 21 scalar dimensions moved ...
10-13 did not move at all", "the stuck ones carry roughly half the residual distance") and
§4.2 (the wrong-checkpoint defect) — the claim under test
- docs/proposals/music/MASTERPIECE_PROGRAM.md §11.1-§11.10 — the card vocabulary a candidate
must be steerable against (lane analysis, event grammar, tempo function, onset/flow law,
complexity law)
- build/audio/exemplars/CORPUS.json :: legal_posture — "The bytes NEVER become model input: no
training, no fine-tuning, no conditioning, no derivation." Binds every candidate: conditioned
on TEXT / SPEC / our own MIDI only.
- docs/pipeline_review/tech_research/PIPE_AUDIO_MUSIC_2026-07-29.md §0 (the melody-first and
30-year bars), §1 (the stack rows this supersedes), §6.0 Q4 and §8 (the melody finding), §8.1
and §8.2 (the realisation rung already built) — lane law
- D:/audio/acestep/repo/README.md L130-137, L255-272 (ON-DISK, the vendor's own model zoo and
VRAM tier table) and
D:/audio/acestep/repo/acestep/core/generation/handler/generate_music.py L286-292 (ON-DISK,
the turbo CFG override) — the incumbent's documented operating point
- build/audio/proof/spec_r2.json, run_r1.json, run_r2.json, control_surface.json — what the
rejected samples were actually generated with
NOT DERIVED (authored judgment, and why it had no canon home):
- The whole model roster and every licence read. Canon names the music the game needs; it has
never named a generator, and `harness/route.py music` routes to composition inputs
(CH_02 Asset anchors, the region page's Section 10 brief, T0_Theme_Registry,
T99_Translation_Audio, MUSIC_COMPOSITION_DOCTRINE), not to tooling selection. There is no
work kind for a model survey; the canon node read instead is CH_02 L234 (the music_mood
bullet quoted above), plus the ruling that dispatched this lane.
- The three-arm bench design (§6). The rig and the cards are canon-adjacent tooling; how to
point them at a candidate is an engineering choice and is grounded in the probe mechanism
that already existed at masterpiece_proof.py `probe` / `probe-report`.
---
**Finding 0, because it is the one that decides the lane and it was produced by an experiment that
failed.** Two config defects (findings 1 and 2) made the case that ACE-Step had never been run at its
own quality tier, so the tier was run tonight: acestep-v15-xl-sft + acestep-5Hz-lm-4B, CFG live,
against the identical 8-dimension two-pole probe. **It scored 0 of 8 responding dimensions against
the pinned turbo config's 3 of 8**, with three dimensions returning literally identical values at
both poles, and supplying the 4B LM moved the numbers barely at all off the LM-less run the lane had
already condemned. **The generator's ceiling is confirmed on measurement rather than on inference,
and Josh's verdict is not a config accident.** §6.2 is the table; §7.3 is what it decides.
---
1. **The rejected music was generated on the configuration ACE-Step's own table prescribes for a
≤6GB card — on a 32GB card — and two of the spec's knobs were being discarded in flight.** The
vendor's VRAM table's bottom row reads *"≤6GB · 2B turbo · None (DiT only)"*, and that is exactly
the checkpoint-and-planner pair the lane ran (the row's INT8 + full-offload half was not used, so
the match is on the model choice, not the whole row). build/audio/proof/run_r2.json
records dit: acestep-v15-turbo, lm: None — the 2B turbo checkpoint with no LM planner — while
spec_r2.json asked every job for guidance_scale: 7.5 and inference_steps: 60. ACE-Step's own
model zoo (ON-DISK, README.md L262) marks acestep-v15-turbo CFG ❌, Step 8; its handler
(L286-292) logs the override on every single call, and it fired 16/16 times in tonight's run:
*"Turbo model detected: overriding guidance_scale 7.5 -> 1.0 (turbo does not use CFG)."* The
vendor's VRAM table (L137) routes a ≥24GB card to **XL sft + acestep-5Hz-lm-4B — "Best
quality, all models fit without offload."** This box has 32GB. So classifier-free guidance — the
mechanism by which a caption steers a diffusion model — was structurally unavailable on the
checkpoint that produced everything Josh has heard, and the step count was 7.5x the distilled
schedule.
2. The landed control-surface probe was inadmissible and says so on its own face.
build/audio/proof/control_surface.json carries positive_controls_responded: false,
admissible: false, and generator: "ACE-Step 1.5 XL-SFT, no LM" — a config MEASURED_LOOP_PROOF
§4.2 had already measured as not separating at all (dark pole 1980 Hz vs bright pole 1976 Hz).
6 of 8 dimensions were measured; the run stopped. **ACE-Step's control surface has therefore never
been measured on any config this lane would ship.** "Half the dimensions don't respond" was a real
observation about a caption-revision loop; it was not a measurement of the generator.
3. No open-weights model closes the fidelity gap, and the leaderboard is unambiguous about it.
VERIFIED: the Artificial Analysis instrumental leaderboard's top ten are Suno V5.5 (1188), Mureka
V9 (1183), Mureka V8 (1165), Suno V5 (1158), Lyria 3 Pro (1117), Suno V4.5 (1084), Music 2.6
(1067), Eleven Music v2 (1062), MiniMax Music 2.5+ (1052), FUZZ-1.1 Pro (1036) — **every one of
them closed or API-only**. And in the one open-model blind arena the survey found (766 votes, the
Khala paper's own table), the **single open model that beats ACE-Step — Khala at Elo 1510.9
against 1470.9 — is CC-BY-NC-4.0 and cannot be used** (§4); the other 2026 quality claim, LeVo 2,
rests on a blog comparative rather than an arena and is academic-licence-only. **Replacing the
generator with another open generator cannot buy AAA fidelity in August 2026. What it can buy is
CONTROL.**
4. A second conditioning channel with a 2048-token budget has been carrying one word. ON-DISK,
conditioning_text.py tokenises the DiT text prompt at max_length=256 (the truncation defect
§4.1 of the loop proof already closed) and the lyric stream separately at max_length=2048
(L142). harness/music_gen/generate.py L452 puts "[Instrumental]" in it. In ACE-Step's training
format that stream is where structure lives. Whether structure tags in it steer form without
summoning vocals is an experiment, not a claim — it is boarded as the cheapest unrun lever in §9,
not asserted here.
5. The strongest replacement candidate is gated behind one click that only Josh can make.
Stable Audio 3.0 Medium (Stability AI, 2026-05-20) is the only surveyed waveform model that is
simultaneously licence-clean at this revenue tier, output-owning, ~6 GB, Windows-native, and
structurally steerable — multi-region mask inpainting is a timeline event grammar, which is
exactly the class of control the pinned dimensions need. `huggingface.co/api/models/stabilityai/
stable-audio-3-medium returns gated: auto`, and an authenticated fetch with the box's own token
returns 403: the licence has not been accepted on Josh's account. Accepting a licence
agreement on his behalf is not this lane's to do. It is §8's one-line Josh-minimum.
---
docs/spine/DECISIONS_PENDING_JOSH.md L3884-3893, verbatim:
"that music is still terrible." Recorded as the ruled quality verdict on ACE-Step output at
draft-bed tier. ... CONSEQUENCE: the $900 digital bulk STAYS HELD; the music-model survey launches
... the generator itself is the candidate for replacement, the measurement rig and cards carry over
unchanged whatever generates. The melody-first ruling stands.
Two boundaries this lane honours. The rig and the cards are not on trial — every number below is
produced by acquire_exemplar.py --no-corpus and masterpiece_proof.py probe-report, unchanged.
Candidates are judged as BED / UNDERSCORE generators, per the melody-first ruling; the leitmotif
layer is the authored lane's (author_head_cell.py), which §8 of the lane dossier already promoted
on measurement and which nothing in this survey disturbs.
---
Vendor's own table (ON-DISK D:/audio/acestep/repo/README.md) | Says | The lane ran |
|---|---|---|
Model zoo L262 — acestep-v15-turbo | CFG ❌ · Step 8 · Quality "Very High" · 2B | this, at guidance_scale 7.5 (discarded) and inference_steps 60 |
Model zoo L271 — acestep-v15-xl-sft | CFG ✅ · Step 50 · Quality "Very High" · 4B | never, on any probe or shipped bed |
| VRAM tier L137 — ≥24GB | "XL sft (or xl-base for extract/lego/complete) · acestep-5Hz-lm-4B · backend vllm · Best quality, all models fit without offload" | 2B turbo, no LM, backend pt |
| Handler L286-292 | if self.is_turbo_model() and guidance_scale != 1.0: guidance_scale = 1.0 | fired 16/16 in tonight's probe run (MEASURED, D:/audio/bench/acestep_turbo/gen.log) |
The guard-run refusal that pins acestep-v15-turbo is not wrong — it correctly refuses a run that
did not use the config the evidence was gathered on. The defect is one level up: **the config the
evidence was gathered on was chosen to escape a broken XL-SFT run** (`--dit acestep-v15-xl-sft
--no-lm`, whose caption conditioning §6.0 of the lane dossier states plainly "does not arrive without
its LM"), and the repair was to abandon the checkpoint rather than to supply the LM. That
substitution has been carrying every claim about ACE-Step's ceiling since.
Nothing in this section says the music is good. It says the ceiling was measured on the floor.
---
Three parallel research lanes, each required to cite a page fetched this pass. Steerability is scored
against our card vocabulary (MASTERPIECE_PROGRAM §11): tempo FUNCTION, section/form, lane and
instrumentation control, and the event grammar (drops · raises · breaks · tuttis · silence budget).
A model with only a text prompt scores LOW no matter how good it sounds.
| Model | Org · date | Licence | Steerability against our cards | Quality evidence | 5090 fit | Verdict |
|---|---|---|---|---|---|---|
ACE-Step 1.5 (xl-base/xl-sft/xl-turbo + 5Hz-lm-0.6B/1.7B/4B) | ACE Studio · last commit 2026-06-26, releases stop at v0.1.8 | MIT | duration · BPM · key/scale · time signature · reference audio · repaint (held-region editing) · cover · retake/flow-edit · lego additive layering · LoRA. No tempo curve, no section map, no dynamics envelope | BT-Elo 1470.9, 2nd open-source in a 766-vote blind arena (Khala paper); does not appear in the Artificial Analysis instrumental top 10 | ≥24GB tier = XL-SFT + 4B LM, no offload | incumbent; never run at its own quality tier — §6 |
| Stable Audio 3.0 Medium (1.4B) | Stability AI · 2026-05-20 | Stability AI Community License — VERIFIED verbatim: *"free for everyone, unless you're using the Core Models for a commercial purpose and you or your organization generate over USD $1M of annual revenue"*; *"You own outputs generated from the Core Models or Derivative Works and therefore can use those outputs at your discretion"* | duration conditioning · binary-mask inpainting: single-region, multi-region, causal continuation · audio-to-audio · official LoRA fine-tuning. Mask regions over a timeline ARE an event grammar. No BPM field (BPM goes in the prompt text) | FAD 0.107 / CLAP 0.390 (tech report, LANE-REPORTED); 44.1kHz stereo to 6m20s | 5.07-6.52 GB; MIT code, PyTorch 2.7.1/CUDA 12.6, explicit Windows PowerShell bootstrap, ComfyUI day-0 | SHORTLIST 1 — blocked on a licence gate only Josh can accept (§8) |
| Magenta RealTime 2 (2.4B / 230M) | Google DeepMind · 2026-06-04 | Apache 2.0 code + CC-BY-4.0 weights; VERIFIED: *"Google claims no rights in outputs you generate using Magenta RealTime 2."* The cleanest licence in the survey | text · audio-example · MIDI (128-dim multihot pitch state per frame) · ~200 ms control latency. But 2 s chunks over a 10 s context: it is a live-performance instrument, not a form generator — no section map over an 88 s cue | *"Model evaluation metrics ... will be shared in our forthcoming technical report"* — no published benchmark; trained on ~71k h of mostly-instrumental stock music | UNGATED. JAX/MLX only; NVIDIA is offline-only, no real-time. Windows-native path UNVERIFIED (WSL2 Ubuntu-24.04 is present on this box) | SHORTLIST 3 — licence-safe research track, not the bed factory |
| Khala 1.0 | CCoM + Tsinghua · 2026-05-01 | CC BY-NC 4.0 | text · lyrics · duration only | best open-source: Elo 1510.9, 4th overall in its own arena | ≥24GB | REFUSED §4 |
| LeVo 2 / SongGeneration v2-large (4B) | Tencent AI Lab · 2026-03-01 | Tencent custom, academic-only | structured lyrics + style prompt | LANE-REPORTED "most natural-sounding clips" in a 2026-05 comparative | 4B | REFUSED §4 |
| DiffRhythm v1.2 / DiffRhythm2 | ASLP@NPU | Apache 2.0 | style prompt · ref audio · instrumental mode · LRC · duration. No tempo/section/lane control | none published | 8GB min, Windows supported | pass — no control gain over the incumbent |
| HeartMuLa-oss-3B | HeartMuLa | Apache 2.0 | lyrics + tags + length only | Elo 1421.8, below ACE-Step | — | pass; the 7B is still unreleased |
| YuE | M-A-P · last update 2025-06-04 | Apache 2.0 | section tags, genre/instrument tags, dual-track ICL | 2025-era | 24GB = 2 sessions; 80GB for a full song | pass — 14 months stale, worse fit than the 2026-07-29 read |
| Muse (0.6B) | Fudan NLP · 2026-01-11 | MIT | advertises segment-level style control; mechanism not enumerated | none | small | watch — right control axis, 119 stars |
| MusicGen / AudioCraft / MusicGen-Stem / Coco-Mulla / JASCO | Meta | weights CC-BY-NC-4.0 | JASCO is best-in-class grammar (text + chords + drums + melody, temporally aligned) | — | — | REFUSED §4 — and JASCO is the painful one |
| NVIDIA Fugatto | NVIDIA | — | — | — | — | NO WEIGHTS as of 2026-08 |
HYPE-ONLY or API-ONLY, named so nobody re-discovers them: DiffRhythm+ (paper, no weights);
HeartMuLa-7B (claimed, internal); ByteDance Seed-Music (no public weights); Mureka v8/v9, MiniMax
Music, Suno, Lyria 3 Pro / Lyria 3.5 (Gemini API only, no weights, no model id published);
Stable Audio 3.0 Large (enterprise API); SegTune (ACL 2026, segment-prompt timeline control — exactly
the grammar we want, no weights); MuQ (a representation model, not a generator). **There is no
ACE-Step 2.0 or v1.6** — the lineage's growth since May is a community control layer (chord-progression
editors, LoRA trainers, stem tooling), not a new checkpoint.
The two-stage route (a model or our own code emits MIDI; a renderer plays it) scores differently by
construction, because every dimension the cards measure is a symbolic quantity: tempo function,
section count, simultaneous lanes, dynamic range, tutti arrivals and silence budget are all *set* in
a score and only *inferred* from a waveform.
| Model | Licence | Control surface | Verdict |
|---|---|---|---|
| Anticipatory Music Transformer (Stanford CRFM, 780M) | Apache 2.0 | infilling · accompaniment conditioned on a supplied melody · continuation. No tempo/instrumentation/form conditioning | the cleanest fit — it fills bars inside a scaffold our code owns |
| MuseCoco (Microsoft muzic) | MIT | the only fetched model whose declared attributes match our suite: instrument, bar count, time signature, key, tempo, pitch range, rhythm intensity | no 2025-26 activity |
| MIDI-GPT (Metacreation, AAAI'25) | repo MIT but weights CC-BY-NC-4.0 | strongest control list found: note density, polyphony, duration, key, pitch range, silence, pitch-class set + bar/track infilling. Tempo and instrument NOT controllable | REFUSED §4 — and it is the best control surface in the class |
| MIDI-LLM (ISMIR 2026, arXiv v2 2026-08-04) | HF licence tag llama3.2 (terms UNVERIFIED) | text → multitrack MIDI: genre, mood, instrumentation, key, time signature, chords, tempo *adjectives* | 16GB+, fits; licence read outstanding |
| NotaGen / NotaGen-X | MIT | period-composer-instrumentation only — no tempo, no bar count, no dynamics, no form | material, not control |
| Aria (EleutherAI) · MMT · SkyTNT midi-model · Composer's Assistant 2 | Apache 2.0 / MIT / MIT | piano-only · multi-instrument continuation · instrument+tempo+key seeding (Windows app) · note density + polyphony + instrument-per-track + bar count + tempo (needs REAPER) | useful, none decisive |
| SymphonyGen (arXiv 2026-04-28) | code not released | 32-bar hierarchy, multi-voice skeleton — tempo hard-fixed at 120 BPM | no weights |
The lane's own finding, stated as it was returned: *"Not one model fetched exposes a tempo curve,
a section map, or a dynamics envelope."* The neural symbolic models buy material, not control. So
the symbolic route's control does not come from a model at all — it comes from author_head_cell.py
and orchestrate.py, which already own tempo, form, lane count, arrivals and silence by construction,
and which §8 of the lane dossier already promoted on measured evidence.
| Option | Licence | Verdict |
|---|---|---|
| sfizz (BSD-2) + VSCO 2 CE + VCSL (CC0) | unconditional | incumbent, keep — rung one, already built and measured (PIPE_AUDIO_MUSIC §8.2) |
| BBC SO Discover | Spitfire EULA: grant is *"only within your own newly-created sound recording(s)"*; no games/advertising prohibition found; no 2026 EULA change found | rung two on evidence — but it is a proprietary plugin, not SFZ, so sfizz cannot drive it: it costs a plugin host in the chain |
| Virtual Playing Orchestra v3.3 | maintainer says *"I enforce no restrictions ... even for commercial purposes"* — but the library is built from Sonatina Symphonic Orchestra, listed on the same page under CC Sampling Plus 1.0 | REFUSED §4 — this is the licence PIPE_AUDIO_MUSIC §8.1 already refused, arriving under a new name |
| MIDI-DDSP (Magenta) | Apache 2.0 | repo archived 2024-02-01, TensorFlow 2.7 / Python 3.8 — a dead stack on Blackwell |
| Symphony Rendering (ICASSP 2026) | not specified | onset F1 0.409-0.477 — it does not faithfully play the notes it is given, which destroys arrival placement and silence budget, the exact dimensions in question |
| DDSP-VST | free plugin | monophonic timbre transfer, not an orchestral renderer |
---
The licence law: face value, no conservatism, flag only an EXPLICITLY triggered prohibition.
Five fire, and three of them hurt.
highest-scoring open model in the survey (Elo 1510.9) and it cannot be used. *No workaround: the
NC term is on the weights.*
for academic, research and education purposes, and refrain from using it for any commercial or
production purposes under any circumstances."* Its HF card reads license: unknown, which is not a
grant. **Flagged as LANE-REPORTED rather than VERIFIED — the repo and card returned 401/404 to
direct fetch this pass, and this refusal should be re-verified before it is cited as settled.**
from the 2026-07-29 read. JASCO is the painful one: text + chords + drums + melody, temporally
aligned, is the closest thing in the survey to our event grammar, and it is unusable.
symbolic control surface found, refused.
Plus 1.0**, whose advertising/promotional-use exclusion PIPE_AUDIO_MUSIC §8.1 already ruled fires on a game
soundtrack. A maintainer's blanket permission cannot relicence upstream samples. Recorded here
because a search for "free orchestral SFZ" returns it first, and a later lane would otherwise
re-decide it.
Not a refusal but a second licence chain nobody had flagged: Stable Audio 3 ships a
LICENSE_GEMMA.md, because text conditioning runs through t5gemma-b-b-ul2. VERIFIED from the model
card: *"The Gemma Terms of Use apply; users agree to those terms as well, including the use
restrictions in Section 3.2."* Adopting SA3 means adopting the Gemma prohibited-use policy as
well as the Stability Community License. That read is outstanding and belongs before the first GPU
hour, not after.
---
1. Stable Audio 3.0 Medium — the only candidate that is licence-clean at this revenue tier,
output-owning, Windows-native today, ~6 GB, and structurally steerable via multi-region mask
inpainting, which is the closest thing available to authoring an event grammar over a timeline.
Blocked tonight on a licence gate (§8).
2. ~~ACE-Step 1.5 XL-SFT + acestep-5Hz-lm-4B — the incumbent at the operating point its vendor
recommends for this exact card, which has never been run. Zero install, zero licence, zero
download.~~ **BENCHED AND FAILED TONIGHT — 0 of 8 dimensions respond against the pinned turbo
config's 3 of 8 (§6.2).** It is struck from the shortlist rather than deleted, because the run
that removed it is the evidence that promotes item 1.
3. Magenta RealTime 2 — the best licence in the survey and genuine MIDI conditioning, kept as the
adaptive/interactive research track rather than the bed factory. VERIFIED from Google's own page:
the model generates in 40 ms frames under a local sliding-window attention, its steering is
*"Text, Audio, MIDI"* with *"note and drums on/off control"*, and the hardware framing is
Apple-Silicon-first — *"While the original Magenta RealTime required a high-power GPU or TPU,
Magenta RealTime 2 brings live generation to the hardware musicians actually use"*, with real-time
tiers quoted per MacBook model and no NVIDIA real-time claim anywhere on the page. A
frame-wise live instrument is not a cue-form generator: nothing in it addresses section count or
time-to-full-texture across an 88-second cue.
Why no third candidate was installed tonight, stated rather than left as an absence. Stable Audio
3 is gated (§8). Magenta RT2 is ungated and WSL2 Ubuntu-24.04 is present on this box, so a JAX/CUDA
build is possible — it was declined on evidence, not on effort: its own documentation gives it no
form control, no published benchmark, and no NVIDIA real-time path, so a night spent on the build
would most likely have produced a candidate that scores WORSE on precisely the pinned dimensions this
bench exists to test. The night went instead to the arm with the largest measurable consequence and
zero install cost — §2's finding, run as an experiment.
---
The method is the one that already existed, so nothing here is a bespoke measurement built to
flatter a conclusion. masterpiece_proof.py probe emits 16 jobs: eight dimensions, two maximally
opposed captions each, the same seed at both poles, neutral on everything else. Each render is
measured by acquire_exemplar.py --no-corpus — the identical rig that measured the exemplars — and
probe-report rules per dimension: a dimension RESPONDS when the two poles separate by at least a
quarter of its own scale in the correct direction. spectral_centroid_hz and global_bpm
are the declared positive controls; if they do not separate, the whole report is inadmissible,
because a probe that cannot detect brightness tells us nothing by staying silent about the rest.
| Arm | What it is | Install cost | Status |
|---|---|---|---|
| A — turbo, no LM | the PINNED INCUMBENT: acestep-v15-turbo, lm: None, 60 steps, guidance_scale 7.5 requested and overridden to 1.0. The exact config behind the three samples Josh rejected | none | MEASURED, §6.1 |
| B — XL-SFT + 4B LM | the vendor's own ≥24GB recommendation: acestep-v15-xl-sft + acestep-5Hz-lm-4B, 60 steps, CFG live at 7.5. Same spec file, same seeds, same captions | none — both checkpoints already on disk (19 GB + 8 GB) | MEASURED, §6.2 |
| C — symbolic + sfizz | the authored route already promoted in §8/§8.2 of the lane dossier, measured on the same rig as a REACH test | none | MEASURED, §6.4 |
build/audio/bench/acestep_turbo/control_surface_turbo.json (16/16 measured; audio at
D:/audio/bench/acestep_turbo/audio).
| Dimension | low pole | high pole | normalised spread | direction | responds |
|---|---|---|---|---|---|
spectral_centroid_hz *(positive control)* | 1004.5 | 1334.8 | 0.2475 | correct | no — by 0.0025 |
global_bpm *(positive control)* | 107.67 | 107.67 | 0.0 | — | no |
silence_fraction | 0.0836 | 0.0893 | 0.057 | correct | no |
mean_lanes | 5.753 | 5.783 | 0.0052 | correct | no |
dynamic_range_db | 18.61 | 31.33 | 0.406 | correct | YES |
time_to_full_texture_s | 2.902 | 1.277 | 0.065 | wrong | no |
note_rate_per_s | 1.689 | 3.044 | 0.4451 | correct | YES |
tuttis_per_min | 0.0 | 2.6667 | 0.8889 | correct | YES |
Read it two ways, because both are true. Formally the report is INADMISSIBLE
(positive_controls_responded: false) and refuses to certify anything. Substantively it is not
silent, because a broken probe does not produce three strong responders: what it says is that on the
pinned config, the caption cannot move tempo at all (identical to two decimal places across
"about 50 bpm" and "about 190 bpm"), brightness misses the bar by a hair, arrangement density and
withholding do not respond, and three dimensions do.
And one of those three corrects a landed claim. MEASURED_LOOP_PROOF §6 lists tuttis_per_min
among the dimensions "pinned at full error in every run — the generator never gathers the full
ensemble." Asked directly, it gathers it: 0.0 → 2.667 tuttis/min, the widest spread in the arm.
The dimension was never unreachable; the revision loop's captions never asked for it in those words.
That is a defect in the revision rules, not in the generator, and it was invisible while the probe
sat inadmissible.
One precision the probe cannot give and must not be read as giving. For global_bpm the probe
deliberately passes bpm: None so the caption is the only route in. Tempo IS reachable on this
generator through the meta parameter — §6.0 Q2 measured turbo at 0.3% error against a stated
target. So the correct statement is narrow: *the caption cannot set tempo; the meta can.*
build/audio/bench/acestep_xlsft_lm/control_surface_xlsft_lm.json — acestep-v15-xl-sft +
acestep-5Hz-lm-4B, thinking=True, CFG live at 7.5, 60 steps, 16/16 jobs, 0 errors, peak
sampled VRAM 22667–32074 MB on a 32607 MB card (533 MB of headroom — the exclusive-overnight
class is reinforced again, not softened).
| Dimension | low pole | high pole | normalised spread | responds | ARM A (turbo) for comparison |
|---|---|---|---|---|---|
spectral_centroid_hz *(pos. control)* | 1467.6 | 1377.0 | 0.0617 | no (wrong direction) | 0.2475, correct direction |
global_bpm *(pos. control)* | 123.05 | 123.05 | 0.0 | no | 0.0 |
silence_fraction | 0.0 | 0.0 | 0.0 | no | 0.057 |
mean_lanes | 6.627 | 6.606 | 0.0032 | no | 0.0052 |
dynamic_range_db | 14.58 | 14.10 | 0.0329 | no | 0.406 YES |
time_to_full_texture_s | 1.068 | 0.186 | 0.0353 | no | 0.065 |
note_rate_per_s | 0.422 | 0.244 | 0.1187 | no | 0.4451 YES |
tuttis_per_min | 0.0 | 0.0 | 0.0 | no | 0.8889 YES |
0 of 8 respond. Three dimensions returned literally identical values at both poles. The right
reading of "every direction is wrong" is not that the model inverts instructions — it is that at
spreads of 0.003 to 0.12 of scale **the two renders are effectively the same audio, and direction is
noise on a near-constant output.** The captions were "a single unaccompanied solo instrument, alone,
nothing else at any point" against "full orchestra and choir at once, every section playing together
throughout, a wall of sound"; mean_lanes moved by 0.021.
And it refutes the hypothesis the pinned config was chosen on. PIPE_AUDIO_MUSIC §6.0 states
that XL-SFT's "caption conditioning does not arrive without its LM." The 4B LM was supplied, with
thinking on, and the numbers barely moved off the old LM-less run:
| XL-SFT no LM (the landed inadmissible probe) | XL-SFT + 4B LM (tonight) | |
|---|---|---|
global_bpm | 123.05 / 123.05 | 123.05 / 123.05 |
silence_fraction | 0.0 / 0.0 | 0.0 / 0.0 |
dynamic_range_db | 14.10 / 14.11 | 14.58 / 14.10 |
mean_lanes | 7.05 / 6.977 | 6.627 / 6.606 |
note_rate_per_s | 0.222 / 0.222 | 0.422 / 0.244 |
Supplying the LM changed essentially nothing. **The 8-step distilled 2B turbo — the checkpoint whose
CFG is structurally disabled — is measurably MORE caption-responsive than the 4B quality checkpoint
with its planner and CFG live.** That is the opposite of what §2's argument predicted, and the
prediction was made in §7.3 before the result landed, so it cannot be fitted to it.
The one alternative explanation, named rather than argued away. This measures ACE-Step *through
generate.py's in-process invocation* — pt backend where the vendor's ≥24GB row recommends
vllm, 60 steps against a documented 50, and a GenerationParams path this repo wrote. Either
XL-SFT is caption-deaf, or our integration of it is, and this probe cannot separate those. The
discriminating test is cheap and is the named next step: run the same two poles through ACE-Step's
own Gradio UI or its :8001 REST server, where none of our code is in the path. Until that runs,
the claim is bounded to *"XL-SFT does not steer through the path this factory uses"* — which is
still decisive for routing, because that path is the factory.
build/audio/bench/ear_xlsft_lm/spec_r2_xlsft.json is build/audio/proof/spec_r2.json job for job
— same captions, same seeds (7101 / 7201 / 7301), same bpm, keyscale, duration, steps and guidance.
The only changed field is required_generator, so the checkpoint is the single moving part
against the three tracks Josh rejected. guard-run still has teeth on this arm: the bench spec
declares the XL-SFT config as its requirement rather than bypassing the check.
Audio, for Josh's ear, at D:/audio/bench/ear_xlsft_lm/audio/ —
MPX_CH02_EXPLORE_r2.wav (88 s, 263 s to generate) · MPX_CH02_BATTLE_r2.wav (81 s, 158 s) ·
MPX_TITLE_MAIN_r2.wav (142 s, 321 s). The turbo originals sit beside them at
build/audio/proof/audio/ for the A/B. Not deployed: serving audible content is the director's
call, and emit_proof_picks.py is the seam when he wants it.
The measured result: further on all three, by the same rig and the same target cards.
| Sample | turbo r2 (landed) | XL-SFT + 4B LM (tonight) | change |
|---|---|---|---|
MPX_CH02_EXPLORE | 0.3794 | 0.4512 | +0.0718 further |
MPX_CH02_BATTLE | 0.3286 | 0.4579 | +0.1293 further |
MPX_TITLE_MAIN | 0.3589 | 0.3960 | +0.0371 further |
| mean | 0.3556 | 0.4350 | +0.0794 further |
And the per-dimension read is more interesting than the composite, so it is not hidden behind it.
XL-SFT is not uniformly worse — it wins decisively on exactly the dimension the loop proof named as
permanently pinned, and loses catastrophically on ones turbo had right (MPX_CH02_BATTLE, error per
dimension):
| Dimension | target | turbo | XL-SFT+LM | who wins |
|---|---|---|---|---|
time_to_full_texture_s | 12.376 | 73.909 (err 1.0) | 5.201 (err 0.287) | XL-SFT, by a mile |
mean_lanes | 6.963 | 3.808 (err 0.453) | 6.401 (err 0.081) | XL-SFT |
flow_peak_position | 0.456 | 0.9309 (err 0.475) | 0.6316 (err 0.176) | XL-SFT |
silence_fraction | 0.0106 | 0.0398 (err 0.292) | 0.0 (err 0.106) | XL-SFT |
note_rate_per_s | 2.499 | 2.765 (err 0.106) | 0.235 (err 0.906) | turbo |
time_to_hook_s | 0.116 | 0.116 (err 0.0) | 58.015 (err 1.0) | turbo |
onset_type | fade_in | fade_in (err 0.0) | cold_statement (err 1.0) | turbo |
section_count | 8 | 8 (err 0.0) | 11 (err 0.375) | turbo |
tuttis_per_min | 5.939 | 0.0 (err 1.0) | 0.0 (err 1.0) | neither |
A DEFECT THIS RUN FOUND BY HITTING IT, and it was hitting the proof lane's own evidence.
masterpiece_proof.py diff accepts --generated-card from anywhere and wrote its output to a
hard-wired build/audio/proof/diffs/<sample>_r<rev>.diff.json. So the first bench diff silently
overwrote three landed, git-tracked diff files — the exact files TABLES.md is derived from — and
nothing complained. They were restored from git, the bench diffs now live in
build/audio/bench/ear_xlsft_lm/diffs/, and the tool gained an --out flag with the reason written
on the argument. A read-from-anywhere flag with a write-to-one-place default is a shared-tree trap;
it had simply never been used from outside the proof lane before. 25/25 controls still pass and the
default path is unchanged.
The shape of XL-SFT's material is consistent across all three samples and across the probe: **a
dense, slow, drifting wash that fills its texture early and then barely moves** — 0.235 to 0.38 note
events per second against targets of 0.9 to 2.5, a hook that arrives 58 seconds in, and a cold
statement where the card wants a fade. That is what a checkpoint the caption cannot steer produces:
its own prior, at length. **The composite distance and the ear are likely to agree here, and the
per-dimension wins do not rescue it** — a cue whose first identifiable material lands a minute in is
not a cue, whatever its lane count measures.
Five landed *_ORCHESTRAL_RUNG_ONE artifacts — authored MIDI played by sfizz through VSCO 2 CE and
VCSL — put through the identical rig. **This is not a distance comparison and must not be read as
one:** these are their own themes against their own briefs, not renditions of the BATTLE card. The
question is narrower and it is the one that matters: *do the dimensions ACE-Step pins take
non-degenerate values on this route?*
| Dimension | FAM_THE_TURNED | HOME_AND_LOSS | JOURNEY_WORLD | PROTAGONIST_THEME | THE_RECURRENCE | ACE-Step turbo probe range | Exemplar targets |
|---|---|---|---|---|---|---|---|
tuttis_per_min | 10.22 | 10.59 | 0.0 | 9.88 | 11.11 | 0.0 – 2.67 | 1.36 · 5.94 · 12.29 |
dynamic_range_db | 4.70 | 24.92 | 10.97 | 6.68 | 22.60 | 18.61 – 31.33 | 6.71 · 10.59 · 35.84 |
time_to_full_texture_s | 0.046 | 0.789 | 9.079 | 0.0 | 1.834 | 1.28 – 2.90 | 3.81 · 12.38 · 22.11 |
section_count | 4 | 6 | 5 | 7 | 3 | — | 7 · 8 · 8 |
global_bpm | 129.2 | 123.05 | 117.45 | 95.7 | 103.36 | 107.67 (pinned) | 99.4 · 103.4 · 143.6 |
onset_type | cold_statement | ostinato_first | cold_statement | cold_statement | fade_in | ostinato_first / fade_in only | fade_in |
silence_fraction | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.084 – 0.089 | 0.0106 · 0.0307 · 0.0618 |
complexity_steps_per_min | 0.0 | 35.29 | 1.76 | 0.0 | 0.0 | — | 5.20 · 9.32 · 19.02 |
What it says, in both directions. The dimensions ACE-Step cannot reach — full-ensemble arrivals
at exemplar magnitude, a real dynamic range spread, a controlled withholding time, a chosen onset
type, an actual tempo that differs per cue — **all take live, well-separated values on the symbolic
route, because they are written rather than hoped for.** onset_type is the cleanest single
illustration: ACE-Step returned only two of the eight onset classes across every artifact this lane
has produced, and it is scored at a full 1.0 miss on every revision; the authored route produces
three different classes across five cues, including the fade_in the cards actually target.
And the honest negative, which is a build item and not a ceiling. silence_fraction is **0.0 on
all five and complexity_steps_per_min is 0.0 on three of five**. The silence budget is a
first-class card parameter (MASTERPIECE_PROGRAM §11.2 — *"restraint is a masterpiece mechanism and
the current factory has no vocabulary for it"*) and the authoring code does not write rests or
step complexity at section boundaries. Those two dimensions are not unreachable on this route; they
are unauthored. That is a named defect in orchestrate.py / author_head_cell.py, and it is
the kind of defect the route can actually fix — which is the whole difference between the two
architectures.
---
**Split the lane by what each half is actually for. The SCORE stops using a waveform generator at
all. The BED lane keeps one, and its replacement is named — but the incumbent's own trial had to run
first, and running it is what made the replacement case evidence instead of assertion.**
7.1 — The SCORE lane's generator is not a waveform model, and the survey settles it. Every
dimension the cards measure is a symbolic quantity: a tempo function, a section map, a lane count, a
dynamics envelope, an arrival schedule, a silence budget. In a score they are set by construction and
measurable exactly; in a waveform they are neither settable nor reliably measurable. The survey
found **no model of any class — waveform or symbolic — that exposes a tempo curve, a section map, or
a dynamics envelope.** So anything with FORM (the twelve canonical themes, chapter cues, boss
phases, everything the 12-20-tracks-per-chapter target means) is authored and rendered, exactly as
§8 and §8.2 of the lane dossier already promoted on measurement. §6.4 is that route measured against
the pinned dimensions, and it reaches them.
7.2 — Fidelity on that lane is bought with a LIBRARY, not with a model. This is the part that
answers Josh's actual complaint rather than deflecting it. The reason rung one still sounds like a
mockup is that VSCO 2 CE and VCSL are public-domain community sample sets, not a scoring session.
The ladder above them is already read and licence-cleared (§8.1's rung two, BBC SO Discover) and its
cost is a plugin host in the chain, not a research programme. **No generator improving changes this,
and no library improving depends on a generator.**
Two different things are called "rung two" and this document conflates neither. §8.1's rung two
is a RENDERER upgrade — a larger sample library under the same authored score. The *recompose* rung
two is a COMPOSITION upgrade, and it already ran: docs/proposals/music/RECOMPOSE_RUNG_TWO_FINDINGS.md
records a cold verification that returned NOT-YET on eight findings, the pick demoted, the
apparatus kept, three findings marked BINDING on any future rung. **This survey disturbs none of
that.** It is evidence about generators; that board is the score lane's own open work, and its
depth-law freeze behind the world pipeline is unaffected by anything here.
**7.3 — The BED/AMBIENCE lane keeps a waveform generator, and the incumbent got exactly one honest
trial at the operating point its vendor recommends. It failed.** The decision rule was written into
this section before the result landed, so it cannot be fitted to it: *if arm B's positive controls
respond and it reaches materially more of the eight dimensions than arm A, the pinned config moves to
XL-SFT + the 4B LM; if it does not, ACE-Step's control surface is confirmed as the ceiling on
evidence rather than on inference.* Arm B scored 0 of 8 against arm A's 3 of 8 (§6.2). So:
acestep-v15-turbo. masterpiece_proof.py GENERATOR is correct and needs no change — but its control_surface_evidence should now point at
build/audio/bench/acestep_turbo/control_surface_turbo.json, which is the first probe of it that
was ever run, and the landed build/audio/proof/control_surface.json should be marked as what it
is: a probe of a config this lane refuses, superseded.
tuttis_per_min responds on the pinnedconfig (0.0 → 2.667) and was written off as unreachable. Whatever else is true of this generator,
the loop has been leaving a responding dimension on the table because no revision clause ever asked
for a full-ensemble arrival in the words the probe used.
inpainting — the only timeline event grammar in the survey available under a licence that ships —
and it is blocked on §8's one click.
through ACE-Step's own Gradio UI or :8001 REST server, with none of our code in the path. If
XL-SFT steers there and not here, the defect is ours and it is worth more than a new model.
7.4 — Nothing ships to Josh's ear until it passes his ear, and the $900 stays held. The ruling
already held it; nothing in this survey argues for releasing it. §8 of the lane dossier's own caution
stands and is now stronger: buying more exemplar cards sharpens the TARGETS, and the targets were
never the problem.
"You were told the music sounds terrible and you have answered with a control argument." That is
a fair reading and it is the one to beat. Josh's verdict is about how the audio SOUNDS. Every number
in §6 is structural — the rig measures form, density, events and spectrum, and it is blind to timbre,
legato, room, ensemble realism and whether a phrase is beautiful. A cue can score perfectly on all 21
dimensions and still be a wash. The leaderboard finding cuts the same way: the models that sound good
are closed, and none of the open ones will close that gap by being steered better. **If the failure
is timbral rather than structural, most of §6 is beside the point and the metric suite is one step
from being Goodharted.**
The answer concedes the half that is true and changes the plan rather than defending it.
(a) The fidelity path is named, and it is §7.2 — a real orchestral library played from an authored
score is the only route in this entire survey that can sound like an orchestra, and it depends on no
model improving. (b) Control is not offered as a substitute for fidelity; it is the precondition for
volume. Twelve to twenty cues per chapter that each have to FIT a scene cannot come from a generator
that cannot be told when the ensemble arrives — however good any individual sample sounds. (c) The
gap the objection exposes is real and is **boarded, not waved off: a blind timbre/realism critic that
is not computed from the score.** Every current diagnostic is derived from what was authored, so
nothing in the loop can currently fail an artifact for sounding fake. Until that exists, §6's numbers
are NECESSARY and never SUFFICIENT, and this document does not claim otherwise.
---
Accept the licence on https://huggingface.co/stabilityai/stable-audio-3-medium. The repo is
gated: auto; an authenticated fetch with the box's own token returns 403. Accepting a licence
agreement and sharing contact details on his account is not this lane's to do under the standing
action rules, so the survey's first-ranked candidate could not be downloaded, installed or benched
tonight. It is one click plus "agree to share your contact information".
Nothing else. No purchase, no API key, no negotiation. The BBC SO Discover rung is a free download
whose EULA is already read (§3.3), and it is not blocking anything today.
---
Audio 3 is gated; Magenta RT2 was declined on evidence (§5). The survey's rankings for those two
are LANE-REPORTED and VERIFIED reads, never measurements.
carrying one word. Whether structure tags in it steer form without summoning vocals — and whether
they can coexist with the [instrumental] flag that server_utils.py L143 derives from that exact
field — is a 4-generation experiment nobody has run. It is the cheapest unexplored lever in the
lane.
dossier remains exactly as it was: all 12 region cells are HELD pending the §17.1 read, and
nothing in this document may be cited on cultural-substrate accuracy.
deploys of audible content are the director's.
401/404 to direct fetch this pass and the GitHub API rate-limited. Re-verify before citing it as
settled.
PROBE_POLES defines opposed captions for eight; the other thirteen (section_count, drops_per_min, breaks_per_min,
boundary_step_share, complexity_steps_per_min, flow_peak_position, dynamics_peak_position,
melodic_salience, time_to_hook_s, max_lanes, total_length_s, onset_type, raises_per_min)
have no pole pair written, so "3 of 8" and "0 of 8" are counts over the probed subset and must
never be restated as "3 of 21". Writing the missing thirteen pole pairs is a cheap, mechanical
extension and is the honest way to make the headline number the one the ruling actually asked for.
§6.3's per-dimension table covers all 21 on real cue specs, which is the nearest thing this run has
to the full picture.
and names the discriminating test (the same poles through ACE-Step's own UI or REST server). Until
that runs, no sentence anywhere should say "XL-SFT cannot be steered" without the qualifier
"through the path this factory uses".
---
| Path | What it is |
|---|---|
build/audio/bench/acestep_turbo/probe_spec.json · run_probe.json · control_surface_turbo.json · cards/ | ARM A — the incumbent's first admissibility-tested control-surface probe |
build/audio/bench/acestep_xlsft_lm/probe_spec.json · run_probe.json · control_surface_xlsft_lm.json · cards/ | ARM B — the same probe at the vendor-recommended tier |
build/audio/bench/ear_xlsft_lm/spec_r2_xlsft.json · run_r2_xlsft.json · cards/ · diffs/ | the ear A/B — spec_r2.json job-for-job, only required_generator changed |
build/audio/bench/symbolic_sfizz/cards/ | ARM C — the authored route's reach on the same rig |
D:/audio/bench/{acestep_turbo,acestep_xlsft_lm,ear_xlsft_lm,symbolic_sfizz}/audio/ | the audio, off-repo per the media contract |
harness/music_gen/masterpiece_proof.py — diff --out | the shared-tree fix §6.3 names, with the reason on the argument; 25/25 controls green, default path unchanged |
docs/line_citation_baseline.json | re-emitted SCOPED (--own-surface on this file alone) in the same commit, per the C5 same-commit re-emit law. A plain emit would have declared a concurrent lane's uncommitted rows: the run reported *"3 owned citations merged, 5922 foreign entries carried through as committed"*, and the diff is 5 insertions / 2 deletions. |
VERIFIED — fetched live 2026-08-06 by the synthesizer: stability.ai/license ·
huggingface.co/stabilityai/stable-audio-3-medium (+ its /api/models/... gating state) ·
huggingface.co/google/magenta-realtime-2 · github.com/magenta/magenta-realtime ·
magenta.withgoogle.com/magenta-realtime-2 · magenta.github.io/magenta-realtime ·
artificialanalysis.ai/music/leaderboard/instrumental.
ON-DISK — read from the installed vendor files: D:/audio/acestep/repo/README.md (model zoo
L255-272, VRAM tiers L130-145) · acestep/core/generation/handler/generate_music.py L180-292 ·
acestep/core/generation/handler/conditioning_text.py L125-145 · acestep/api/server_utils.py L143.
LANE-REPORTED — three parallel research lanes, each citing pages fetched this pass: the ACE-Step
lineage sweep (github.com/ace-step/ACE-Step-1.5 releases/commits, huggingface.co/ACE-Step,
awesome-ace-step, Khala paper arXiv 2605.01790 with its 766-vote arena table, github.com/Khala-Music-AI/Khala,
github.com/HeartMuLa/heartlib, github.com/multimodal-art-projection/YuE, github.com/ASLP-lab/DiffRhythm,
github.com/yuhui1038/Muse, it-jim.com comparative 2026-05-21) · the diffusion/big-lab sweep
(github.com/Stability-AI/stable-audio-3, huggingface.co/stabilityai/stable-audio-3-small-sfx,
stability.ai community-license-agreement, huggingface.co/facebook/jasco-chords-drums-melody-1B,
github.com/fundwotsai2001/MuseControlLite, github.com/juhayna-zh/AudioControlNet,
FUGATTO publication PDF, marktechpost/techcrunch SA3 coverage) · the symbolic+renderer sweep
(github.com/ElectricAlexis/NotaGen, github.com/jthickstun/anticipation +
huggingface.co/stanford-crfm/music-large-100k, github.com/Metacreation-Lab/MIDI-GPT +
huggingface.co/Metacreation/MIDI-GPT, github.com/slSeanWU/MIDI-LLM, github.com/AMAAI-Lab/Text2midi,
github.com/m-malandro/composers-assistant-REAPER, microsoft.github.io/muzic/musecoco,
github.com/EleutherAI/aria, github.com/SkyTNT/midi-model, arXiv 2604.25498 SymphonyGen,
symphony-rendering.github.io, github.com/magenta/midi-ddsp, sfzinstruments.github.io/orchestra,
virtualplaying.com/virtual-playing-orchestra, spitfireaudio.com EULA).
UNVERIFIED, named as such: tencent-ailab/SongGeneration licence text (401/404 + API rate limit) ·
MIDI-LLM's llama3.2 licence terms · Magenta RT2's Windows/WSL2 CUDA build.
---
IN-PLACE SUPERSESSION with provenance. §8 named one blocker: a licence gate only Josh could
accept. He accepted it (form submitted on the account matching the box token; authenticated download
now succeeds). §5's shortlist item 1 and §7.3's "becomes the replacement candidate" are no longer
predictions — they are measured. Nothing above this line is edited; the earlier text is the
record of what was believed before the arm ran, and this section is what the arm returned. Where
they disagree, this section wins and the disagreement is named.
Honest register, first. These are BEDS. The SCORE lane's verdict in §7.1 is unchanged and
this section does not touch it. No claim is made anywhere below that any of this music is good;
§12.4 is the closest thing to a quality instrument in the repo and its own first finding is that it
admits the tracks Josh rejected.
| Weights | D:/models/stable-audio-3-medium — model.safetensors 8.79 GB, sha256 48d9c65e290e7bcd5194e0633bfc2424a59ee9683f5c2d58762d997b7d8ce0b5 (matches the hub's own content hash), plus t5gemma-b-b-ul2 1.18 GB, sha256 9b05ea5a4f211d023832f706fb2c0e83e4fc721b6da35ab69ceb0b55eb7800d3 |
|---|---|
| Licence records, pinned | docs/licence_records/stable_audio_3_medium/{LICENSE,LICENSE_GEMMA,NOTICE} — sha256 d6f6b1a4…, e77acc0d…, 66f856d7…. Read FIRST-HAND from the pinned files, not from the web summary §3.1 carried. Extensionless on purpose (matching sfizz/, vcsl/, vsco2_ce/): landed as .md they became doc-graph ISLANDS and the ledger gate went red on L3 island growth, correctly — a licence BODY is evidence, not a document with a consumer. Do not rename them back |
| Stack | D:/audio/bench/sa3/venv — Python 3.12.10, torch 2.7.1+cu128, CUDA 12.8, sm_120 on the RTX 5090, stable-audio-3 0.1.0 from the MIT repo. Windows-native, no WSL2 |
| Adapter | harness/music_gen/sa3_generate.py, 9/9 controls |
1. A genuinely triggered OBLIGATION was missed. §3.1 quoted the revenue threshold and the
output-ownership sentence and stopped there. Section III of the pinned text also says: *"If You
are using or distributing the Stability AI Materials for a Commercial Purpose, You must register
with Stability AI at (https://stability.ai/community-license)."* Making music for a game that
ships is a Commercial Purpose, so registration is required and it is not satisfied. It is an
account action; §12.6 boards it. Every generation record this lane writes carries the requirement
so it travels on the artifact instead of living in a note.
2. **The attribution clause, read at face value, does NOT fire on shipping audio — and the reason is
worth writing down because it is the clause people assume fires.** IV(a) obliges attribution when
You distribute "the Stability AI Materials or a Derivative Work … or a product or service that
uses any portion of them". Section V defines Derivative Work and then excludes model output
explicitly: *"but do not include the output of any Model."* A shipped game carries rendered audio
and no portion of the weights. So IV(a) is not triggered by audio-only shipping, IV(c)(iii) gives
us the outputs, and Gemma 3.3 says the same on its chain (*"Google claims no rights in Outputs"*).
**What IS forbidden and must travel on every row: no Stability weight or output may ever become a
conditioning, fine-tune or LoRA input to a foundational model** — IV(b) — and the corpus posture
bars audio conditioning from the other direction.
One thing the config file said that the survey did not know. model_config.json declares the
prompt conditioner as {"type": "t5gemma", "max_length": 256}. **Stable Audio 3 has the same
256-token caption ceiling ACE-Step does.** The budget discipline MEASURED_LOOP_PROOF §4.1 built for
one generator is not generator-specific, and it carries over unchanged.
Identical probe. Same eight dimensions, same two opposed captions, same seed 8800 at both poles,
same 45 s, measured by the same acquire_exemplar.py --no-corpus, ruled by the same probe-report.
The one fairness adjustment is declared on every artifact: SA3 has no bpm meta channel, so a declared
bpm is appended to the caption as the model card's own convention (bpm_route: caption_text) — and
for the global_bpm dimension the spec declares bpm: null, so **both generators were asked that
question through the caption alone.**
| Dimension | low pole | high pole | spread | responds | ACE-Step turbo | ACE-Step XL-SFT+LM |
|---|---|---|---|---|---|---|
spectral_centroid_hz *(pos. control)* | 1497.8 | 2618.6 | 0.428 | YES | 0.2475 no | 0.0617 no |
global_bpm *(pos. control)* | 129.2 | 95.7 | 0.2593 | no — inverted | 0.0 | 0.0 |
silence_fraction | 0.0279 | 0.3215 | 1.000 | YES | 0.057 no | 0.0 no |
mean_lanes | 5.086 | 7.250 | 0.2985 | YES | 0.0052 no | 0.0032 no |
dynamic_range_db | 13.21 | 27.06 | 0.5118 | YES | 0.406 YES | 0.0329 no |
time_to_full_texture_s | 19.226 | 1.277 | 0.718 | no — inverted | 0.065 no | 0.0353 no |
note_rate_per_s | 2.844 | 1.311 | 0.539 | no — inverted | 0.4451 YES | 0.1187 no |
tuttis_per_min | 8.0 | 8.0 | 0.0 | no — pinned | 0.8889 YES | 0.0 no |
| RESPOND | 4 / 8 | 3 / 8 | 0 / 8 | |||
| MOVE AT ALL (spread ≥ 0.25) | 7 / 8 | 4 / 8 | 0 / 8 |
The formal verdict is still INADMISSIBLE, and for a smaller reason than before. One of the two
positive controls now fires — brightness separates by 0.428, the first time any arm has moved it —
but global_bpm moves 0.2593 in the *wrong* direction, so probe-report refuses to certify, and
that refusal is reported rather than argued around.
The failure mode is different in kind, and that is the finding. ACE-Step mostly did not move:
on turbo four dimensions sat under 0.07, and on XL-SFT all eight sat under 0.12. Stable Audio 3
moves hard and sometimes backwards — three dimensions (global_bpm, time_to_full_texture_s,
note_rate_per_s) travel 0.26 to 0.72 of their own scale in the opposite direction to the
instruction. A model that hears an instruction and inverts it is a prompt-phrasing and
negative-prompt problem. A model that does not move is a ceiling. **Those are not the same problem,
and only one of them is ours to fix.**
And tuttis_per_min says something neither number alone does. It is pinned — but pinned *at
8.0*, where ACE-Step XL-SFT sat pinned at 0.0 and turbo ranged 0.0–2.67 against exemplar targets of
1.36 / 5.94 / 12.29. Stable Audio 3 gathers the full ensemble by default and cannot be talked out of
it; ACE-Step's quality tier never gathers it at all. For a BED the SA3 default is arguably the wrong
one — but it is a DEFAULT, not a ceiling, and the difference matters.
| Arm | 45 s probe render | the three cues (88 s + 81 s + 142 s) | VRAM |
|---|---|---|---|
| ACE-Step turbo | ~10 s | ~3 min | 8.4 GB peak |
| ACE-Step XL-SFT + 4B LM | 14–95 s | 741 s | 22.7–32.1 GB (533 MB headroom) |
| Stable Audio 3 Medium | 1.4–6.5 s | 9.4 s | ~9.5 GB over baseline |
The whole 16-job probe generated in about 30 seconds of GPU time, on a card already holding 20 GB
of the art lane's gen-4 face refine — no queueing was needed and nothing was killed. XL-SFT's same
probe took roughly twenty minutes and owned the card. **A candidate ladder that costs seconds rather
than minutes is a different kind of tool**: selection pressure becomes affordable, which is the one
lever PIPE_AUDIO_MUSIC §6.0 Q4 named that this lane has never been able to pull at scale.
harness/music_gen/timbre_critic.py, 9/9 controls. §9 boarded it in one line: *"a blind
timbre/realism critic that is not computed from the score."* It reads only a measured card's
timbre and register blocks — spectral centroid, 85% rolloff, spectral flatness, and the ten-band
octave profile — and scores them by robust z against the 197 measured exemplar cards, real
released game soundtracks. It never opens a MIDI file, and --self-test T5 proves that structurally
by parsing its own reader and asserting that no structural block of the card is even named. The bands
are not chosen: they are the leave-one-out 90th and 98th percentiles of the corpus scoring itself
(1.658 / 2.669), and white noise and a single-band sine stack both read OUTSIDE.
| Arm | n | median timbre distance | verdicts | median spectral flatness |
|---|---|---|---|---|
| exemplars (the population) | 197 | 0.663 (p50) | 90.4% INSIDE by construction | 0.01695 |
| ACE-Step XL-SFT + 4B LM (probe) | 16 | 0.234 | 16 INSIDE | 0.00853 |
| Stable Audio 3 Medium (probe) | 16 | 0.463 | 16 INSIDE | 0.00730 |
| ACE-Step turbo (probe) | 16 | 1.143 | 16 INSIDE | 0.00055 |
| ACE-Step turbo — the three cues Josh rejected | 6 | 0.982 | 6 INSIDE | 0.00055 |
| authored + sfizz, orchestral rung one | 5 | 1.932 | 5 MARGINAL | 0.00008 |
THE FIRST THING TO SAY ABOUT THIS INSTRUMENT IS WHAT IT GETS WRONG. It admits, at 0.982, the
exact six artifacts whose ruling launched this survey. That is not a defect to be tuned away — it is
the calibration of what a PASS is worth, printed in the same table as the PASS. **A pass here means
"the spectrum is not alien", and nothing more.** Josh's ear overruled a set of tracks this critic
calls INSIDE, so the critic can never be cited to argue with him.
THE SECOND THING CUTS AGAINST §7.1 AND §7.2, AND IT IS THE MOST USEFUL RESULT OF THE SITTING. The
authored + sfizz route — the one §6.4 showed reaching every structural dimension the generators pin —
is the arm furthest from real released game music, the only one not INSIDE, and the gap is
overwhelmingly one feature: its median spectral flatness is **0.00008 against the corpus's 0.01695,
roughly two hundred times more tonal**, a robust z of 4.298. That is what a small public-domain
sample library playing sparse authored lines with no percussion, no room and no dense mix measures
like. §8.2 of the lane dossier said those words as a caveat; this is the number.
So the two rulers disagree, and the disagreement is the honest shape of the lane:
bought with a LIBRARY" is now measured as not yet delivered by rung one's library. That
upgrades §8.1's rung two from an optional improvement to the named lever, and gives it an
acceptance test it did not have: *a rung-two render must move the flatness term toward the
population, and this critic will say by how much.*
Two caveats on the instrument, published rather than buried. Its measured blind spot: a
band-limited mid-only spectrum reads INSIDE at 0.75, so it screens for an alien SPECTRUM and NOT for
a missing bottom or top end — control T4b asserts that value rather than deleting the control,
because a limit you cannot fix today should be pinned where it cannot drift. And across every arm the
deciding term was spectral_flatness_mean, so in practice this is mostly a tonality detector and
should be described as one.
build/audio/bench/sa3/spec_r2_sa3.json is build/audio/proof/spec_r2.json job for job: same
captions, same seeds (7101 / 7201 / 7301), same durations. Audio for Josh at
D:/audio/bench/sa3/audio_ear/ — MPX_CH02_EXPLORE_r2.wav · MPX_CH02_BATTLE_r2.wav ·
MPX_TITLE_MAIN_r2.wav, beside the turbo originals at build/audio/proof/audio/ and the XL-SFT arm
at D:/audio/bench/ear_xlsft_lm/audio/. Four generators, three cues, one spec. Not deployed —
serving audible content is the director's call, and emit_proof_picks.py is the seam when he wants it.
Same rig, same target cards, same distance function. build/audio/bench/sa3/diffs/.
| Sample | ACE-Step turbo (the rejected set) | ACE-Step XL-SFT + 4B LM | Stable Audio 3 Medium | change vs turbo |
|---|---|---|---|---|
MPX_CH02_EXPLORE | 0.3794 | 0.4512 | 0.2832 | −0.0962 (25% closer) |
MPX_CH02_BATTLE | 0.3286 | 0.4579 | 0.2881 | −0.0405 (12% closer) |
MPX_TITLE_MAIN | 0.3589 | 0.3960 | 0.2617 | −0.0972 (27% closer) |
| mean | 0.3556 | 0.4350 | 0.2777 | −0.0779 (22% closer) |
**That is the best distance any arm has produced, and the per-dimension read says exactly why —
which is more interesting than the total.** On MPX_CH02_BATTLE, counting per dimension, turbo
actually wins 6, Stable Audio 3 wins 5 and 6 tie. The composite still moves decisively because **the
wins are not the same size**:
| Dimension | target | turbo | SA3 | |
|---|---|---|---|---|
time_to_full_texture_s | 12.376 | 73.909 (err 1.0) | 18.646 (err 0.251) | SA3, by a mile |
mean_lanes | 6.963 | 3.808 (err 0.453) | 6.536 (err 0.061) | SA3 |
tuttis_per_min | 5.939 | 0.0 (err 1.0) | 0.741 (err 0.875) | SA3 |
spectral_centroid_hz | 1059.7 | 824.9 (err 0.196) | 1068.8 (err 0.008) | SA3 |
flow_peak_position | 0.456 | 0.9309 (err 0.475) | 0.2305 (err 0.226) | SA3 |
onset_type | fade_in | fade_in (err 0.0) | cold_statement (err 1.0) | turbo |
section_count | 8 | 8 (err 0.0) | 9 (err 0.125) | turbo |
note_rate_per_s | 2.499 | 2.765 (err 0.106) | 1.444 (err 0.422) | turbo |
silence_fraction | 0.0106 | 0.0398 (err 0.292) | 0.0553 (err 0.447) | turbo |
| curve distance | 0.2785 | 0.2073 | SA3 |
**Stable Audio 3 wins the dimensions that were failing at full error and loses the ones that were
already nearly right.** time_to_full_texture_s — pinned at a full 1.0 miss on every ACE-Step
revision this lane has ever run, and named in MEASURED_LOOP_PROOF §6 as one of the dimensions
carrying half the residual distance — lands at 0.251. The six curves, which no caption revision has
ever moved much, come down 0.2785 → 0.2073.
Its two real regressions are both nameable and one is cheap. onset_type goes from a correct
fade_in to cold_statement at a full 1.0 miss on all three cues — SA3 starts hard. The engine
forbids a baked fade (the stem/loop contract), so this is an ARRANGEMENT ask, not a mastering one,
and it is exactly the kind of thing an 8-step, 2-second generator can be laddered against. And it is
thinner: note_rate_per_s 1.444 against a 2.499 target.
| Arm | n | median | verdicts |
|---|---|---|---|
| SA3 (probe) | 16 | 0.468 | 16 INSIDE |
| SA3 (the three cues) | 3 | 0.781 | 2 INSIDE, 1 OUTSIDE |
| ACE-Step XL-SFT+LM (probe) | 16 | 0.231 | 16 INSIDE |
| ACE-Step XL-SFT+LM (3 cues) | 3 | 0.322 | 3 INSIDE |
| ACE-Step turbo (probe) | 16 | 1.135 | 16 INSIDE |
| ACE-Step turbo (rejected cues) | 6 | 0.974 | 6 INSIDE |
| authored + sfizz (rung one) | 5 | 1.922 | 5 MARGINAL |
MPX_CH02_EXPLORE_r2 reads OUTSIDE at 2.79 — flatness z 4.64, centroid z 2.96, rolloff z 2.63.
**This is the first time the critic has failed a real generated artifact rather than a synthetic
control, and the reading is genuinely ambiguous in a way that must be stated rather than resolved by
preference.** The EXPLORE spec asks for the sparsest thing this lane generates — two or three lanes
at once, upper register left empty, a 6% silence budget — so a very tonal, quiet, narrow-band result
is what the *instruction* asks for. And MEASURED_LOOP_PROOF §8 already recorded that the exemplar
corpus is thinnest exactly here ("almost nothing for melancholic/tragic"). **So the OUTSIDE verdict
is either the artifact being spectrally odd or the population being too thin to judge sparse cues,
and this run cannot separate them.** The resolution is corpus, not tuning: it is one more argument
for the class anchors the purchase run was built to buy, and the critic must not be cited against a
sparse cue until that population exists.
CHANGES — the bed lane's generator, on evidence. §7.3 said Stable Audio 3 "becomes the
replacement candidate" if arm B failed. Arm B failed and the candidate then beat the incumbent on
every axis this lane can measure: **4/8 responding dimensions against 3/8 and 0/8; 7/8 dimensions
moving at all against 4/8 and 0/8; the best distance on all three cues (mean 0.2777 against 0.3556);
16/16 INSIDE on the blind timbre screen; and roughly eighty times the throughput.** The recommendation
is therefore PROMOTE Stable Audio 3 Medium to the bed/ambience lane's generator, with ACE-Step
retained rather than deleted for the two things it still does better — a correct fade_in onset and
a denser note rate — and because deleting it would erase the comparison that makes this legible.
DOES NOT CHANGE — the score lane, the honest tier, or the $900. §7.1 stands untouched: the SCORE
is authored and rendered, because every card dimension is a symbolic quantity and no generator in
this survey exposes a tempo curve, a section map or a dynamics envelope. §12.4 strengthens that
split rather than weakening it. Nothing here is a quality claim: the honest tier on every artifact
this section produced is BENCH, the generation records say so in those words, and **Josh's ear is
still the only gate that has ever overruled anything.**
BOARDED, with the reason each is not done:
1. Register with Stability AI — Community Licence Section III, triggered by commercial use, not
satisfied. An account action; §12.9.
2. The three inverted dimensions. global_bpm, time_to_full_texture_s and note_rate_per_s
move hard in the wrong direction. The lever is prompt phrasing and SA3's negative_prompt
argument, which this run did not touch at all — a two-pole probe with negatives is the next cheap
experiment and it is now affordable at 2 s a render.
3. onset_type is a full 1.0 miss on all three cues. SA3 starts cold; the cards want fade_in
and the engine forbids a baked one. This is an arrangement instruction to ladder against.
4. The multi-region mask was never used. The whole reason SA3 was shortlisted in §3.1 is
inpaint_mask_start_seconds / ..._end_seconds accepting LISTS — a timeline event grammar. This
run measured plain text-to-audio only, so **the candidate's headline control feature is still
unmeasured** and the promotion above rests on its text conditioning alone.
5. The probe still covers 8 of 21 dimensions. Unchanged from §9. "4 of 8" must never be restated
as "4 of 21".
6. No region cue, so the Care-Doctrine question is still untouched on this generator too.
§8's licence-gate item is DONE — he accepted it and the arm ran. One item replaces it, and it is
smaller:
Register the commercial use at https://stability.ai/community-license. The Community Licence's
Section III requires it of any Commercial Purpose user; the grant is still royalty-free and still
free below USD $1M annual revenue. It is a form, not a purchase. Until it is done, everything this
lane generates on Stable Audio 3 carries triggered_obligation: NOT SATISFIED on its own licence
record, which is the honest state and not a blocker for benching.