music/STEPBACK_OPEN_STACK.md
CANON SUBORDINATION — this document is PROPOSAL-TIER: it serves canon and never outranks it.
Canon served: Josh's four recorded music grades (docs/review_candidates.json), the
melody-first / thirty-year-bar direction (memory music-direction-melody-first-30-year-bar), the
composed-never-prompted rule, the no-human-composer ruling (2026-08-08 — "composer_in_loop" is the
factory's own pass ladder, never a hire), the CVD §17.1 care line, local-first on the 5090, and
proof-before-spend.
If this document disagrees with canon, CANON WINS and this document is the defect.
TIER: PROPOSAL — ENGINE, not CANON. It sets no canon, names no region content, changes no spine
or registry row, and adopts nothing. It is a component survey with a licence read and three
candidate pipelines. Nothing in it has been installed or run this pass.
Written: 2026-08-09. Lane: music open-stack step-back. Consumer: the round-5
architecture ruling, alongside STEPBACK_POSTMORTEM.md (what our own stack did) and
STEPBACK_COMMERCIAL_ENGINES.md (how the products people pay for are built). This is the third
leg: what is actually available, open, and runnable on the box.
---
DERIVED FROM:
docs/review_candidates.json :: candidates[11..13] — PASS7_FLORES_{ROAD,FALLS,ROUNDS}, fields note, comparison_card.lines[], comparison_card.musical_axes_all_three_rounds[],
grade_verbatim (round 4, graded 2026-08-09) — the grade this doc is answering
docs/pipeline_review/tech_research/PIPE_MUSIC_MODELS_2026-08-06.md §3.1, §3.2, §3.3, §4, §5, §12 — the owning PIPE dossier; it binds this lane (memory pipe-dossiers-bind-generation-lanes).
This doc extends it on one axis and corrects nothing in it.
docs/proposals/music/STEPBACK_POSTMORTEM.md §0 (the survives/retires list, including *"theCC0-sfz chain as the final realisation path"* under RETIRES)
docs/proposals/music/STEPBACK_COMMERCIAL_ENGINES.md §0, §0.1 (the rulings block)harness/music_gen/sfz_palette.py :: PLAYER, SETS — the incumbent renderer stack, read at HEAD: sfizz 1.2.3 (BSD-2-Clause) driving VSCO 2 CE, VCSL, legato_vocal_tutorial, body_percussion
(all CC0-1.0), each pinned by commit SHA
harness/music_gen/sa3_generate.py L1-L45 — the Stable Audio 3 adapter and, critically, itsdeclared legal posture on audio conditioning (§8.3 below turns on the exact wording)
harness/music_gen/fetch_sfz_stack.py L1-L28 — the licence-body-verbatim discipline this doc followsdocs/licence_records/ — the pinned licence bodies already on disk (sfizz, vsco2_ce, vcsl, ace_step, stable_audio_3_medium, legato_vocal, body_percussion)
CLAUDE.md — the content bar, the decision protocol, the do-not-invent reframelicense-reads-not-conservative-solo-dev-tooling — face-value licence reads; flag only anEXPLICITLY triggered prohibition
5090-box-live-remote-stack — Milks5090, RTX 5090 32GB, compute-clean, 24/7NOT DERIVED (authored judgment, and why it had no canon home):
this doc feeds, not one it may make.
none exists the column says so in those words. The judgment of what a benchmark is WORTH for our
use (orchestral cue, 60-120 s, melody-bearing) is mine and is declared as such.
sa3_generate.py's legal posture as scoped to the EXEMPLAR corpus rather than toour own renders. The script's docstring states the posture; the scoping inference is mine, and it
is the director's to veto because a pipeline depends on it.
---
Round 4 was graded, verbatim (docs/review_candidates.json, PASS7_FLORES_ROAD.grade_verbatim,
2026-08-09):
you are using. I hate how slow the melody is and there is no variation or changes or movement
throughout the songs. Everything sucks."
Two of those clauses are about COMPOSITION and two are about REALISATION. A model survey cannot fix
composition — STEPBACK_POSTMORTEM.md already ruled that the melodies are typed pitch tuples
developed by rules, and no amount of new renderer changes that. So this survey is organised around
the half a model CAN answer, and it asks every candidate the same question:
**Given a leitmotif that Josh owns and our engine composed — a specific sequence of specific
pitches at specific times — can this component realise THAT, or does it only produce something of
its own that sounds a bit like the prompt?**
That question is not decoration. It is the composed-never-prompted ruling stated as an engineering
requirement. A model that cannot be given a melody is not a candidate at any quality level, because
adopting it means Josh stops owning the themes. **Every table below carries a MELODY CONDITIONING
column and it is the first column that can disqualify.**
realise a theme it is GIVEN. A generator that invents the theme is prompting.
ladder.
bought thing is the remaining gap.
component whose weights were learned from a living tradition's recordings is a care question, not
only a licence question.
It is not a re-run of PIPE_MUSIC_MODELS_2026-08-06.md. That dossier surveyed the generator field
three days ago and benched three arms on this box; its findings stand and are cited, not
repeated. Its own §9 ("what this document does not answer") left the axis this doc takes:
realisation of a supplied melody, and the renderer ladder above the CC0 sfz chain. Where a fact
appears in both, this doc marks it and moves on.
It also adopts nothing, installs nothing and spends nothing. Every "candidate" in §9 ends at a first
proof step with a cost.
---
Every claim below carries one of two marks, and the mark is the point.
The URL is in §11.
source was NOT fetched this pass. It may be true. It is not evidence yet.
No claim below is marked VERIFIED on the strength of a search-result snippet. That distinction is
the thing fetch_sfz_stack.py's docstring calls out — *"the licence BODY of each component is stored
verbatim ... not summarised — the summary is what the critic rejected."* Nothing here has had its
licence body pinned to docs/licence_records/, which is the step BEFORE any of it renders a byte.
---
The box: RTX 5090, 32 GB GDDR7, compute-clean, 24/7 (memory 5090-box-live-remote-stack).
sm_120 / compute capability 12.0, and it needs CUDA 12.8+. PyTorch 2.7.0 was the first stable release shipping pre-built cu128 wheels with sm_120 kernels. *(LANE-REPORTED —
PyTorch forums + issue 159207.)* Our own SA3 adapter already runs torch 2.7.1+cu128 on this box
(PIPE_MUSIC_MODELS §12.1), so the PyTorch half of the ecosystem is proven here.
this survey — is JAX/MLX only. A jax[cuda12] build on sm_120 is plausible and unverified. That is
a named risk with a cheap resolution (one install night), not a blocker.
while we were not looking:
| Component | Params | Peak VRAM | Fits 32 GB? |
|---|---|---|---|
| Stable Audio 3 Medium | 1.4B | 1.69–6.52 GB (chunked decode 5.14 GB at 380 s) | trivially — VERIFIED |
| ACE-Step 1.5 XL-SFT + 4B LM | 4B + 4B | vendor table routes ≥24 GB to this tier | yes, no offload — VERIFIED |
Magenta RT 2 base | 2.4B | not published | almost certainly; offline-only on NVIDIA — VERIFIED (offline-only) |
| Anticipatory Music Transformer | ~780M (medium) | <4 GB | trivially |
| NotaGen-large | 516M | <4 GB | trivially |
| sfizz + sample libraries | — | 0 (CPU) | — |
**The consequence: we can run a SYMBOLIC model, an AUDIO model and a SAMPLER concurrently on one
card with room left.** A pipeline does not have to choose one of the three families. That is the
single most useful hardware fact in this document, and it is what makes §9's stacked pipelines
buildable rather than aspirational.
---
This family is structurally aligned with composed-never-prompted: it operates on notes, so a note
we authored can be pinned and the model asked to work around it. That is the whole reason to look
here first.
| Model | Org · date | Licence (code / weights) | MELODY CONDITIONING — can it take OUR theme? | VRAM | Quality ceiling, honestly |
|---|---|---|---|---|---|
Anticipatory Music Transformer (stanford-crfm/music-medium-800k) | Stanford CRFM · 2023, repo live | Apache-2.0 code — VERIFIED. Weights licence NOT stated in the repo — VERIFIED that it is unstated; HF card read outstanding | YES, and it is the cleanest instance of the primitive in the whole survey. The repo exposes extract_instruments() to split a melody line out of a score and then generates the REST conditioned on it; the paper's own framing is infilling and accompaniment conditioned on supplied events. Josh's tune goes in as fixed control events; the model writes around it and never over it | <4 GB | Human evaluators rated its accompaniments comparably musical to human-composed ones over 20-second clips *(LANE-REPORTED — paper abstract)*. Twenty seconds is the honest number. It is a bar-filler inside a scaffold our code owns, not a cue writer. Provenance flag: trained on the Lakh MIDI dataset, which is scraped MIDI of largely copyrighted popular songs. For a shipped commercial game that is an outstanding read, not a settled clearance |
| NotaGen / NotaGen-X | Central Conservatory / IJCAI 2025 · sizes 110M / 244M / 516M | MIT — LANE-REPORTED (HF card and repo LICENSE both read MIT in search results) | NO — it is a MATERIAL source, not a realiser. Conditioning is period-composer-instrumentation only. It writes classical ABC from a style prompt; there is no slot for a supplied theme | <4 GB | Its own A/B tests report it beating baselines against human compositions in the classical idiom *(LANE-REPORTED)*. Genuinely the strongest "writes like a trained musician" symbolic model found. And it is the one that most directly collides with composed-never-prompted — see §8.5 |
| MuseCoco | Microsoft muzic | MIT — *(carried from PIPE_MUSIC_MODELS §3.2)* | Attribute conditioning only (instrument, bars, time signature, key, tempo, pitch range, rhythm intensity). No melody slot | small | No 2025-26 activity. Attribute surface is the best match to our card dimensions and the model is stale |
| MIDI-RWKV | 2025 · RWKV-7 | not established this pass | Infilling model — masked-region regeneration with numerical attribute controls. Infilling IS a melody-preserving primitive if the melody track is left unmasked *(LANE-REPORTED — paper only)* | small | Paper-tier. No adoption signal found |
| MIDI-GPT | Metacreation · AAAI'25 | repo MIT, weights CC-BY-NC-4.0 | strongest control list in the class; bar/track infilling preserves a supplied line | small | REFUSED — the NC term is on the weights and a shipped game is commercial. Unchanged from PIPE_MUSIC_MODELS §4 |
| text2midi | AAAI 2025 | not established this pass | text → MIDI. No melody slot | small | Beats MuseCoco on tempo/key adherence and in listening tests *(LANE-REPORTED)*. Wrong shape for us regardless |
| MIDI-LLM | ISMIR 2026 (arXiv v2 2026-08-04) | HF tag llama3.2, terms UNVERIFIED *(carried from PIPE_MUSIC_MODELS §3.2)* | text → multitrack MIDI. No melody slot | 16 GB+ | Newest in the class; licence read outstanding |
Exactly one open, permissively-licensed symbolic model implements the primitive we need, and it
is the Anticipatory Music Transformer: *give it the notes Josh owns, get back everything else.* Every
other member of the family either has no melody slot (NotaGen, MuseCoco, text2midi, MIDI-LLM) or is
licence-refused (MIDI-GPT).
And the second half of that finding, stated plainly: AMT's control surface is thin. It has no
tempo curve, no section map, no dynamics envelope, no instrumentation control worth the name. Which
is precisely the gap PIPE_MUSIC_MODELS §3.2 already closed by pointing at our own code — *"the
symbolic route's control does not come from a model at all — it comes from author_head_cell.py and
orchestrate.py, which already own tempo, form, lane count, arrivals and silence by construction."*
That remains true. AMT is a candidate for ONE job: **writing the inner voices, counter-lines and
accompaniment that round 3's grade said were missing** ("never any harmonies added on"), inside bars
whose tempo, length, instrumentation and dynamic shape our code already fixed.
---
| Model | Org · date | Licence (code / weights) | MELODY CONDITIONING | VRAM | Quality ceiling, honestly |
|---|---|---|---|---|---|
| ACE-Step 1.5 (2B + XL 4B DiT; 0.6B/1.7B/4B LM) | ACE Studio · v1.5 late Jan 2026, XL 2026-04-02 | MIT (code) — VERIFIED. Weights MIT — LANE-REPORTED (HF ACE-Step/Ace-Step1.5 card). Our incumbent; licence body already pinned at docs/licence_records/ace_step/ | Indirectly, and this is the underused finding. No MIDI input anywhere — VERIFIED by reading the repo. But it ships Reference Audio Input ("use reference audio to guide generation style"), Cover Generation ("create covers from existing audio"), Repaint & Edit ("selective local audio editing and regeneration"), Vocal2BGM ("auto-generate accompaniment for vocal tracks") and Track Separation — all VERIFIED verbatim from the repo. Feed it a rough render of OUR score and every one of those becomes a melody-preserving operation | VERIFIED vendor table: ≤6 GB 2B-turbo; 8-16 GB turbo/sft + 1.7B LM; 20-24 GB XL; ≥24 GB XL-sft + 4B LM = best quality | Benched on this box and its ceiling is confirmed — PIPE_MUSIC_MODELS §6.2: the vendor-recommended XL-SFT + 4B LM tier reached 0 of 8 probe dimensions against pinned-turbo's 3 of 8, and landed FURTHER on all three ear cues. As a text-to-song GENERATOR it is measured and it is not enough. It has never been measured as a refiner over our own audio |
| Stable Audio 3 (Small 433M / Medium 1.4B / Large 2.7B) | Stability AI · Medium 2026-05-20 | Stability AI Community License — VERIFIED: free commercial use under USD $1M annual revenue; above that, registration and possibly a paid Enterprise licence. "You own outputs generated from the Core Models" — VERIFIED. Second chain: Gemma Terms, via the t5gemma-b-b-ul2 text encoder. Both bodies pinned at docs/licence_records/stable_audio_3_medium/ | YES via init_audio + init_noise_level — VERIFIED from the repo docs. *"init_noise_level controls how much the init audio influences the output (0.0–1.0, default 1.0). At 1.0 the init audio is fully replaced by noise... Lower values preserve more of the original."* Plus inpainting over multiple non-contiguous regions and causal continuation | VERIFIED: 1.69–6.52 GB peak, 5.14 GB with chunked decode at 380 s. 44.1 kHz stereo, up to 380 s (Medium) | Our best-measured generator. PIPE_MUSIC_MODELS §12.6: 4/8 responding dimensions, 7/8 moving, best distance on all three cues (mean 0.3556 → 0.2777), ~80× the throughput. Also the honest counterweight in §12.7 — the blind timbre critic reads our authored+sfizz route as the arm FURTHEST from real released game music, and ADMITS the six tracks Josh rejected. A PASS from our instruments is worth exactly that much |
Magenta RealTime 2 (base 2.4B / small 230M) | Google DeepMind · 2026-06-04 | Apache-2.0 code + CC-BY-4.0 weights — VERIFIED from the HF card. *"Google claims no rights in outputs you generate"* — VERIFIED. The cleanest licence in the survey, and the only one with no revenue threshold, no registration and no second chain | YES — DIRECT, FRAME-WISE MIDI CONDITIONING. Verbatim from the model card: *"128-dim multihot vector representing the state of each MIDI pitch during this frame (0 = Off, 1 = Sustain, 2 = Onset, 3 = Sustain or onset, model decides)"* — VERIFIED. Style comes separately as 12 MusicCoCa tokens from a text OR audio prompt. This is the primitive stated exactly: our notes drive the pitches, a prompt drives the timbre world | not published; 2.4B at bf16 ≈ 5 GB — fits with enormous headroom | Unknown, and that is the honest word. *"Model evaluation metrics ... will be shared in our forthcoming technical report"* — no published benchmark. Trained on ~71k h of mostly-instrumental stock music, which is a real prior against orchestral character. 48 kHz stereo. Architecture cost: 25 Hz frames under a ~20 s effective receptive field — it is a live instrument, so long-form continuity is OUR problem, and it emits a MIXDOWN, not stems |
| YuE (7B) | M-A-P / HKUST · Jan 2025 | Apache-2.0 including weights — LANE-REPORTED; attribution to "YuE by HKUST/M-A-P" requested | genre/instrument/section tags, dual-track ICL. No melody slot | 24 GB = 2 sessions; 80 GB for a full song *(carried from PIPE_MUSIC_MODELS)* | 14 months stale by this field's clock. Lyrics-to-song shape, wrong for instrumental cues |
| DiffRhythm v1.2 / DiffRhythm 2 | ASLP@NPU | Apache-2.0 — LANE-REPORTED | style prompt, ref audio, instrumental mode, LRC, duration. No note-level slot | 8 GB min | Fast (full song ~10 s) and controllable only at the style level. No control gain over the incumbent |
| HeartMuLa-oss-3B | HeartMuLa · 3B released 2026-01-14, RL variant 01-23 | Apache-2.0 — VERIFIED from the repo | lyrics + tags + length. No melody slot | fits | Elo 1421.8, below ACE-Step *(carried from PIPE_MUSIC_MODELS §3.1)*. The claimed-better 7B is unreleased |
| LeVo 2 / SongGeneration 2 (4B) | Tencent AI Lab · 2026-03-01 | NOASSERTION / Tencent custom, academic-only — LANE-REPORTED, and PIPE_MUSIC_MODELS §4 already flagged this refusal as needing re-verification | structured lyrics + style prompt | 4B | Reported to outperform all open baselines on all six dimensions of its own eval *(LANE-REPORTED)*. Refused on licence; the refusal itself is still unverified |
| Khala 1.0 | CCoM + Tsinghua · 2026-05-01 | CC-BY-NC-4.0 | text, lyrics, duration | ≥24 GB | The highest-scoring open model in the field (Elo 1510.9) and it is unusable. No workaround: the NC term is on the weights |
| MusicGen / AudioCraft / JASCO | Meta | code MIT, weights CC-BY-NC-4.0 — VERIFIED from the model card | JASCO is the best melody+chord+drum conditioning surface in the entire survey and MusicGen-melody's chromagram conditioning is the canonical implementation of the primitive | 3.3B max | REFUSED, and it is the painful one. Every guide on the open web still recommends MusicGen-melody for exactly our use case. The weights licence forecloses it for a commercial game |
| Stable Audio Open 1.0 | Stability AI | Community License; training data 472,618 Freesound + 13,874 FMA, all CC0/CC-BY/CC-Sampling+ — VERIFIED, with Audible Magic screening | text-to-audio only; 47 s max | small | Superseded by Stable Audio 3 for our purposes. Kept in the survey for one reason: it has the cleanest declared training-data provenance of any audio model here, which matters if provenance ever becomes the deciding axis |
Three of these can be handed our music. None of them can be handed our SCORE.
refined. The melody survives because it is physically present in the input.
it happens to carry the best licence.
And the field-level fact worth stating once: **the two best-scoring open models in the world right
now (Khala, and arguably LeVo 2) are both licence-refused, and the best conditioning surface ever
built for this exact problem (JASCO) is licence-refused.** The open audio field's quality frontier
and its usable frontier are not the same frontier. Anyone re-running this survey will re-discover
that; it is written here so they do not have to.
---
STEPBACK_POSTMORTEM.md put *"the CC0-sfz chain as the final realisation path"* under RETIRES. This
section is what replaces it. There are two ladders — sampled and neural — and they fail in opposite
directions.
Read at HEAD from harness/music_gen/sfz_palette.py:
sfizz_render.exe, offline.6dd651d5), VCSL (CC0-1.0, commit dfcf4a49), legato_vocal_tutorial (CC0-1.0), body_percussion (CC0-1.0).
comparison_card.honest_limits_short):*"THE PLAYERS ARE A FREE CC0 COMMUNITY SAMPLE LIBRARY played by an offline sampler, one dynamic
layer a note."*
"One dynamic layer a note" is the load-bearing clause and it is OURS, not the library's. VSCO 2
CE ships multiple velocity layers on many instruments. A renderer that plays one layer per note
produces a dead performance from ANY library, free or bought. **Which means the first rung of the
render ladder is not a library at all — it is emitting expression.** See §8.2.
| Option | Licence — face-value read | What it costs to use | Quality ceiling |
|---|---|---|---|
| VSCO 2 CE + VCSL (incumbent) | CC0-1.0. Unconditional. No attribution, no restriction | zero | Community-recorded, sparse articulations, thin dynamic layering. The measured floor: our own blind timbre critic reads this route as the arm furthest from real released game music (PIPE_MUSIC_MODELS §12.7) |
| Spitfire BBC SO Discover | Spitfire EULA, quoted from the official EULA page: *"provided that you use the purchased Sound File(s) only within your own newly-created sound recording(s) and/or performances in a manner that renders the Sound File(s) substantially dissimilar to the original sound of the Sound File"* — LANE-REPORTED (search snippet of the official EULA; body not pinned). Face-value read (memory license-reads-not-conservative-solo-dev-tooling): a multi-instrument orchestral cue, mixed and reverbed, IS a newly-created recording substantially dissimilar to any one sample. The grant fires. Nothing in the FAQ or EULA found this pass prohibits games. The one genuinely triggered prohibition is the obvious one: we may not ship or redistribute the samples themselves, and a near-raw single held note is the edge case to avoid | Free, but it is a proprietary VST3/AU/AAX plugin — VERIFIED from the Spitfire FAQ, which also states a 64-bit DAW is required and the licence is single-user (two machines). sfizz cannot drive it. This is the whole cost: a plugin host enters the chain | Recorded at Air Studios by the BBC Symphony Orchestra. It is a *reduced* tier — one dynamic layer per articulation on many patches, no round-robins — but the SOURCE is a professional orchestra in a professional hall. That is a categorical step above community samples, not an incremental one |
| Virtual Playing Orchestra | REFUSED, and the refusal is already on the record (PIPE_MUSIC_MODELS §4): the maintainer grants unrestricted commercial use, but the library re-imports Sonatina Symphonic Orchestra under CC Sampling Plus 1.0, whose advertising/promotional-use exclusion PIPE_AUDIO_MUSIC §8.1 already ruled fires on a game soundtrack. A downstream maintainer cannot relicence upstream samples. Recorded here because a search for "free orchestral SFZ" returns it first | — | — |
| Sonatina Symphonic Orchestra | CC Sampling Plus 1.0 — same clause, same refusal | — | — |
The host, which is the actual new component:
| Host | Licence | Fit |
|---|---|---|
| DawDreamer | GPLv3 — LANE-REPORTED. Built on JUCE; also drags in the Steinberg VST3 SDK terms | Python VST2/VST3/AU host with offline batch rendering, MIDI playback, per-parameter automation at audio rate and PPQN, and multi-processor graphs. It is a BUILD-TIME TOOL, not a shipped component — we distribute rendered WAV, not the host. GPLv3 binds redistribution of the software; it does not reach the audio it renders. That is the same reasoning sfz_palette.py already records for sfizz's BSD-2 notice condition (*"the BSD-2 notice condition binds redistribution of the software, and this lane redistributes rendered audio only"*). Face-value: clean for our use. Risk: some plugins misbehave headless — the project documents this |
| REAPER (CLI) | commercial, DRM-free, cheap; 60-day full evaluation — LANE-REPORTED | -batchconvert with a filelist, plus ReaScript in Lua/Python/EEL. The industry-proven batch path. Costs money, so proof-before-spend puts it second |
| Carla / kuriborosu | GPL | headless plugin host, simpler API, less automation control |
This is the class that would render a score with learned human phrasing rather than sample playback.
It is the most exciting family in the survey and the least usable.
| System | Licence | What it renders | Verdict |
|---|---|---|---|
| Magenta RT 2 | Apache-2.0 + CC-BY-4.0 weights — VERIFIED | audio conditioned on frame-wise MIDI pitch state + a style embedding | The ONLY commercially-usable neural MIDI realiser found. See §4. Its constraints are architectural, not legal |
| Multi-Aspect Conditioning for Diffusion-Based Music Synthesis (benadar293, TASLP 2024) | CC-BY-NC-SA-4.0 — VERIFIED from the project page | MIDI → audio for solo instruments, chamber ensembles, orchestras, jazz/rock bands, drums; trained on ~58 h of REAL audio; T5 backbone + SoundStream vocoder | REFUSED. It is the closest thing in existence to "play our orchestral score with real recorded character," and NC forecloses it |
| RenderBox (2025) | *"distributed for non-commercial research purposes only"* — LANE-REPORTED | text + score → expressive multi-instrument performance audio; diffusion transformer with cross-attention joint conditioning | REFUSED. Same shape, same problem |
| nii-yamagishilab/midi-to-audio | Apache-2.0 code + CC-BY-4.0 weights — VERIFIED | Piano only (MAESTRO). Transformer-TTS MIDI→mel + HiFiGAN mel→audio | Licence-clean and the wrong instrument. Useful only if a cue is piano-led |
| MIDI-VALLE (ISMIR 2025) · Pianist Transformer (135M) | not established this pass | expressive piano performance synthesis / rendering | Piano-only; the class's centre of gravity is piano because MAESTRO exists and no orchestral equivalent does |
| MIDI-DDSP (Magenta) | Apache-2.0 | monophonic-instrument synthesis with note-expression control | Repo archived 2024-02-01, TF 2.7 / Python 3.8 — a dead stack on Blackwell *(carried from PIPE_MUSIC_MODELS §3.3)* |
Google music-spectrogram-diffusion | Apache-2.0 (weights on HF) | multi-instrument MIDI → spectrogram → audio | 2022-era quality; superseded by everything above it in this table |
The neural rendering class is real, it is good, and for a commercial game it is one model wide.
Every orchestral-capable neural renderer found is CC-BY-NC or research-only except Magenta RT 2;
every licence-clean one except Magenta RT 2 is piano.
Meanwhile the sampled ladder has a free rung sitting unclimbed: **a professional orchestra recorded
at Air Studios, free, with a licence that on a face-value read grants exactly what we need — and the
only thing between us and it is a plugin host in the render chain.**
---
Products change every eight weeks. The primitives do not, and naming them is what lets a later lane
evaluate something this survey never saw.
1. Symbolic conditioning (notes in, notes out). A supplied note stream is held fixed and the
model completes around it. *Instances:* AMT extract_instruments() + control events; MIDI-RWKV
infilling; MIDI-GPT bar/track infilling (refused). **Fidelity to our melody: EXACT — the notes are
not regenerated, they are held.**
2. Frame-wise symbolic → audio conditioning. Note state is fed to an audio model per frame.
*Instance:* Magenta RT 2's 128-dim multihot pitch vector. **Fidelity: high by construction — the
model is told which pitches sound in every frame.** The 3 = model decides state is the honest
nuance: it can be given latitude where we want latitude.
3. Chromagram / melody-audio conditioning. A melody's pitch-class energy over time steers
generation. *Instance:* MusicGen-melody, JASCO (both refused). **Fidelity: pitch-class only —
octave and rhythm can drift.**
4. Audio-to-audio / init-audio (the img2img analogue). A rough render is partially re-noised and
denoised under a text prompt. *Instances:* Stable Audio 3 init_noise_level (VERIFIED
semantics — lower preserves more); ACE-Step Reference Audio / Cover / Vocal2BGM. **Fidelity:
TUNABLE and that is the point — there is a strength band where structure survives and timbre
changes, and finding it is a sweep, not a guess.**
5. Masked inpainting / repainting over a timeline. Regenerate bars 17-24 and hold everything
else. *Instances:* Stable Audio 3 multi-region binary masks (still UNMEASURED per
PIPE_MUSIC_MODELS §12.9); ACE-Step Repaint & Edit. **This is the primitive that lets a cue be
FIXED rather than re-rolled**, and it is the closest thing in the audio family to an event grammar.
6. Style embedding from a reference. Timbre/genre is set by an embedding rather than by words.
*Instances:* Magenta RT 2's MusicCoCa audio prompt; ACE-Step reference audio; ACE-Step LoRA from a
handful of songs. **CARE + LICENCE GATE: sa3_generate.py already declares the corpus posture —
the exemplar bytes NEVER become model input. Any use of this primitive must source its reference
from audio we own.**
7. Expression automation into a sampler. CC1/CC11/velocity/keyswitch curves shape a sampled
performance. *Instances:* DawDreamer PPQN + audio-rate automation; sfizz opcodes. **The
zero-model, zero-VRAM primitive we are currently not using, and the one that "one dynamic layer a
note" names as absent.**
---
Face-value reads per memory license-reads-not-conservative-solo-dev-tooling. **Nothing here is
cleared until its body is pinned to docs/licence_records/.**
| Component | Verdict for a commercial shipped game | The clause that decides it |
|---|---|---|
| ACE-Step 1.5 | GO | MIT (code VERIFIED; weights LANE-REPORTED). Already pinned |
| Stable Audio 3 Medium | GO, with one live obligation | Community License: free under $1M revenue; outputs owned. Section III commercial-use REGISTRATION is triggered and NOT SATISFIED (PIPE_MUSIC_MODELS §12.9 — Josh's account action). Second chain: Gemma Terms |
| Magenta RealTime 2 | GO — cleanest in the survey | Apache-2.0 code + CC-BY-4.0 weights; *"Google claims no rights in outputs you generate."* No revenue threshold, no registration, no second chain. CC-BY attribution is the only obligation |
| Anticipatory Music Transformer | CONDITIONAL | Apache-2.0 code VERIFIED; weights licence unstated and Lakh MIDI training provenance is an open question for a shipped game. Two reads outstanding |
| NotaGen | GO on licence (MIT, LANE-REPORTED) | but see §8.5 — the constraint that binds it is composed-never-prompted, not the licence |
| HeartMuLa-oss-3B · YuE · DiffRhythm | GO on licence (Apache-2.0) | no melody slot; no reason to adopt |
| Khala 1.0 | NO | CC-BY-NC-4.0 on the weights |
| MusicGen / AudioCraft / JASCO | NO | CC-BY-NC-4.0 on the weights |
| MIDI-GPT | NO | CC-BY-NC-4.0 on the weights (repo MIT — the split is the trap) |
| LeVo 2 / SongGeneration | NO (refusal itself still LANE-REPORTED) | Tencent custom, academic-only |
| benadar293 multi-aspect · RenderBox | NO | CC-BY-NC-SA-4.0 / non-commercial research only |
| VSCO 2 CE · VCSL · legato_vocal · body_percussion | GO | CC0-1.0, unconditional. Pinned |
| BBC SO Discover | GO on a face-value read; body not yet pinned | *"only within your own newly-created sound recording(s) ... substantially dissimilar to the original sound"* — an orchestral cue satisfies it. Do not ship samples; avoid near-raw one-shots |
| Virtual Playing Orchestra · Sonatina SO | NO | CC Sampling Plus 1.0 upstream, advertising/promotional exclusion. Already ruled |
| DawDreamer | GO as a build-time tool | GPLv3 binds redistribution of the SOFTWARE. We redistribute rendered audio. Same reasoning already recorded for sfizz |
| REAPER | GO, costs money | commercial licence, DRM-free. Proof-before-spend defers it |
---
Josh said *"I hate the instruments you are using."* We are using community CC0 samples. **A
professional orchestra recorded at Air Studios is available for free, under a licence that on a
face-value read grants exactly our use.** The only thing standing between the current chain and that
one is that BBC SO Discover is a VST3 plugin and sfizz plays SFZ. **DawDreamer closes that gap in
Python, offline, at zero cost and zero VRAM.** No model, no spend, no new licence risk beyond one
body to pin. This is the highest ratio of grade-clause-answered to effort in the entire document.
comparison_card.honest_limits_short says *"one dynamic layer a note."* That is a property of our
REALISER, not our library. Hand a world-class orchestral library a MIDI file with static velocities,
no CC1 expression curve, no CC11 dynamics, no articulation keyswitches and no legato intervals, and
it will sound like a better sampler playing a dead performance. **Rung one of the render ladder is
emitting expression from orchestrate.py — velocity shaping, CC1/CC11 curves, articulation
selection per phrase.** The library is rung two. Attempting rung two first will produce a null result
and a wrong conclusion, and that is a prediction this document is making on the record so it can be
checked.
harness/music_gen/sa3_generate.py declares, correctly and deliberately:
*"NO AUDIO EVER CONDITIONS A GENERATION. ... The corpus legal posture
(build/audio/exemplars/CORPUS.json :: legal_posture) is absolute — 'The bytes NEVER become model
input: no training, no fine-tuning, no conditioning, no derivation.'"*
**That posture is about the EXEMPLAR CORPUS — commercial recordings we do not own, held for
measurement only. It is not, and was never, a rule against conditioning on audio we rendered
ourselves from a score we composed.** Feeding SA3 or ACE-Step our own sfz render is categorically
outside the thing that posture forbids.
This inference is mine and is flagged as NOT DERIVED. It is load-bearing: candidate pipeline C
depends entirely on it, and if the director reads the posture as blanket, C dies and should die
cleanly rather than be built on a misreading. The safe implementation is explicit — a separate entry
point with a positive control asserting the init audio's provenance is build/audio/ and never
build/audio/exemplars/.
PIPE_MUSIC_MODELS §5 shortlisted Magenta RT 2 third and declined to install it on evidence:
*"its own documentation gives it no form control, no published benchmark, and no NVIDIA real-time
path."* Every word of that is true and one of the three objections does not apply to us.
"No form control" is only a defect for a system that must own form. Our system already owns form.
author_head_cell.py and orchestrate.py fix tempo, section map, lane count, arrivals and silence
by construction — PIPE_MUSIC_MODELS §3.2 says so in those words. What we lack is not form; it is
SOUND. Magenta RT 2 is the only component in this survey that takes our notes directly and returns
audio under a licence with no threshold, no registration and no second chain.
The two objections that DO stand, and they are serious:
possibility that our leitmotif comes back sounding like production library music, which is the
precise aesthetic Josh has rejected four times.
measurement rig operate on lanes. A mixdown-only realiser costs us those instruments — it is not
a drop-in for the sampler, it is a different architecture with a different evidence surface.
The resolution is cheap and it is a proof, not an argument: one install night (JAX on sm_120,
unproven on this box), one leitmotif, one 30-second render, judged by the blind timbre critic
(harness/music_gen/timbre_critic.py) and by motif preservation. That is a night, not a program.
STEPBACK_POSTMORTEM.md already named it: **is a learned model PROPOSING melodic material
composition, or is it prompting under composed-never-prompted?** That fork is Josh's.
This survey sharpens it into a clean three-way, because the candidates now separate cleanly:
they are given. Compatible with composed-never-prompted under any reading.
Compatible under most readings; it is the "develops, does not invent" case, and it is the direct
answer to round 3's *"never any harmonies added on."*
classical-craft symbolic model found, so the cost of the strict reading is real and should be
named rather than absorbed silently.
No pipeline in §9 depends on the third case. All three are buildable under the strictest reading
of the rule. That is deliberate.
Re-verified this pass: no new open checkpoint has landed in the three days since the PIPE dossier.
ACE-Step's lineage still stops at v1.5/XL (there is no v1.6 or 2.0). Khala, LeVo 2 and JASCO are
still licence-refused. The instrumental quality frontier is still closed-weights. **The survey did
not go stale; the axis this doc reads was simply not the axis that dossier read.**
---
Every pipeline below starts at the same place — **a leitmotif Josh owns, a score our engine
composed** — and differs only in how that score becomes sound. None invents melody. Each ends at a
first proof step with a cost, per proof-before-spend.
They are stackable, not exclusive. A is the floor and both others sit on top of it.
---
orchestrate.py ──emit MIDI + CC1/CC11/velocity/keyswitch──► DawDreamer (GPLv3, build-time)
│ hosts VST3
▼
BBC SO Discover ──► per-lane stems
│
existing mix_policy.py + battery
Steinberg VST3 SDK terms. No model licence at all.
This is the only pipeline with zero evidence-surface cost.
sampler playing the same dead performance. **Rung one is orchestrate.py emitting expression;
the library is rung two.** Build them in that order or the result is a null.
hosting of authorised commercial plugins is where DawDreamer's documented plugin-compatibility
caveats bite. If it does not host, REAPER CLI is the fallback and it costs money — which is
exactly the kind of spend proof-before-spend permits, because A's proof would have earned it.
expression curve and velocity shaping, render it BOTH ways — sfizz/VSCO2 and DawDreamer/BBC SO
Discover — and run both through timbre_critic.py and the existing battery. **Two numbers and two
files. If the timbre critic does not move, the instrument hypothesis is wrong and we learned it in
a night.**
---
orchestrate.py ──► note stream ──► 128-dim multihot pitch state per frame ─┐
├─► Magenta RT 2 ──► 48 kHz stereo
style reference (audio WE own, or text) ──► 12 MusicCoCa tokens ───────────┘ │
▼
our seam/continuity layer (loop_seams.py)
sampler closes — because the character comes from a model trained on recorded performance.
base (2.4B) or small (230M). VRAM: ~5 GB. Spend: zero. frame; the 3 = model decides state is deliberate latitude we control per-note.
registration, no second chain, *"Google claims no rights in outputs."*
1. Mixdown, not stems. Our mix policy, lane census and several instruments operate on lanes.
This pipeline costs us that evidence surface, or forces per-lane rendering and a re-mix — which
is untested and may not sum coherently.
2. ~20 s effective receptive field, 25 Hz frames. Long-form continuity over an 88-second cue is
OUR problem to solve by stitching. We already own loop_seams.py, so this is work, not a wall.
3. ~71k h of stock-music training prior, and no published benchmark. The genuine risk that our
leitmotif returns sounding like production library music.
present (PIPE_MUSIC_MODELS §5), so the path exists; nobody has walked it.
through mrt2_base with a text style prompt, and score it against the same cue rendered through
Pipeline A. Judged by timbre_critic.py (blind) plus a motif-preservation check. **The install is
the risk; the render is minutes.**
---
orchestrate.py ──► MIDI ──► sfizz or Pipeline A render (structure is now physically present)
│
▼ init_audio, init_noise_level swept
Stable Audio 3 Medium (MIT-ish/Community) ──or── ACE-Step 1.5 (MIT)
│
▼
motif_compare.py + timbre_critic.py, scored across the sweep
cohesion, performance grit — while the composition, form, arrivals and silence survive because they
are physically present in the init audio.
already installed and licence-pinned on this box.**
*"lower values preserve more of the original."* At a low noise level you get the same dead
performance under a haze; at a high one the melody drifts away from Josh's theme. There is a band
between and it is found by a sweep scored on melody preservation, not by a guess.
registration obligation**; ACE-Step carries none.
posture. If the director reads that posture as blanket, this pipeline does not exist. That read
is the cheapest thing in this document and should happen first.
arrival-placement and silence-budget dimensions our cards measure — measurable with instruments we
already have; (ii) it produces a mixdown, same evidence-surface cost as B, though here we keep
the pre-refinement stems as a fallback mix.
render, sweep init_noise_level across 0.1 / 0.2 / 0.35 / 0.5 under one caption, and plot melody
preservation against timbre-critic distance. **The output is a curve, and the curve either has a
usable band or it does not.** This is the cheapest proof in the document because every component
is already on disk.
---
They stack: **A is the renderer floor. C rides on A's output. B is an independent track that
bypasses the sampler entirely.**
| Answers "hate the instruments" | Answers "sounds like a mockup" | Keeps stems | Melody fidelity | Install cost | Spend | |
|---|---|---|---|---|---|---|
| A | yes, directly | partly | yes | exact | none | zero |
| B | yes | potentially, fully | no | high | JAX/sm_120 unproven | zero |
| C | yes | yes, if a band exists | no (A's stems survive as fallback) | tunable, unproven | none — both installed | zero |
RECOMMENDATION (this lane's, and the director's to overturn): run C's sweep first because it
is hours on components already on disk and it either finds a band or eliminates a whole family;
then A rung one (expression from orchestrate.py) because §8.2 predicts everything else in A
depends on it; then A rung two (BBC SO Discover through DawDreamer); then B as the
independent track once the JAX install has a night to spare.
THE STRONGEST OBJECTION to all three: none of them touches *"the melody sucks"* or *"no variation
or changes or movement."* Those are composition defects, STEPBACK_POSTMORTEM.md located them in
pass3_flores.py:114-144, and a renderer cannot fix a tune. **If round 5 ships better instruments
playing the same melody, it will be graded down again, and this document will have been part of the
reason. The only component here that addresses the composition side at all is AMT** (§3.1) —
harmonising and counter-lining a theme it never alters — and it belongs in the round-5 architecture
alongside whichever renderer wins, not after it.
---
verbatim body under docs/licence_records/. **BBC SO Discover, DawDreamer, AMT's weights and
NotaGen all need that step before a byte renders.**
the provenance question is the more serious of the two.
orchestrate.py (§8.2) is scoped as a prediction, not designed.It is the actual engineering in Pipeline A and it is not costed here.
UNMEASURED (PIPE_MUSIC_MODELS §12.9) and is not exercised by any pipeline above.
has not been run. Neither vendor enumerates its corpus at the granularity that question needs.
game — is out of scope here and boarded.
---
All fetched or searched 2026-08-09.
Audio generation
https://github.com/ace-step/ACE-Step-1.5/blob/main/README.md · paper:
https://arxiv.org/abs/2602.00744 · weights: https://huggingface.co/ACE-Step/Ace-Step1.5
https://github.com/Stability-AI/stable-audio-3/blob/main/docs/workflows/inference.md
https://huggingface.co/google/magenta-realtime-2 · docs:
https://magenta.github.io/magenta-realtime/ · original announcement:
https://magenta.withgoogle.com/magenta-realtime
https://github.com/facebookresearch/audiocraft/blob/main/model_cards/MUSICGEN_MODEL_CARD.md
https://huggingface.co/m-a-p/YuE-s1-7B-anneal-en-icl
https://arxiv.org/html/2606.30642v1
Symbolic / MIDI
https://arxiv.org/abs/2306.08620 · weights: stanford-crfm/music-medium-800k
https://huggingface.co/ElectricAlexis/NotaGen · https://arxiv.org/abs/2502.18008
MIDI-to-audio rendering
https://benadar293.github.io/multi-aspect-conditioning/ ·
https://github.com/benadar293/multi-aspect-conditioning
Renderers, hosts and libraries
https://support.spitfireaudio.com/en/articles/11815834-bbc-symphony-orchestra-discover-faq ·
EULA: https://www.spitfireaudio.com/info/eula/
plugin compatibility: https://dirt.design/DawDreamer/compatibility.html · paper:
https://archives.ismir.net/ismir2021/latebreaking/000001.pdf
Platform
https://docs.salad.com/container-engine/tutorials/machine-learning/pytorch-rtx5090
Repo-internal (canon and lane-canon)
docs/pipeline_review/tech_research/PIPE_MUSIC_MODELS_2026-08-06.md §3.1-§3.3, §4, §5, §12docs/proposals/music/STEPBACK_POSTMORTEM.md · docs/proposals/music/STEPBACK_COMMERCIAL_ENGINES.mdharness/music_gen/sfz_palette.py · harness/music_gen/sa3_generate.py · harness/music_gen/fetch_sfz_stack.py
docs/review_candidates.json :: PASS7_FLORES_{ROAD,FALLS,ROUNDS}