STEPBACK_OPEN_STACK.md

music/STEPBACK_OPEN_STACK.md

STEPBACK — the open stack, read against one question: can it play OUR tune

CANON SUBORDINATION — this document is PROPOSAL-TIER: it serves canon and never outranks it.
Canon served: Josh's four recorded music grades (docs/review_candidates.json), the
melody-first / thirty-year-bar direction (memory music-direction-melody-first-30-year-bar), the
composed-never-prompted rule, the no-human-composer ruling (2026-08-08 — "composer_in_loop" is the
factory's own pass ladder, never a hire), the CVD §17.1 care line, local-first on the 5090, and
proof-before-spend.
If this document disagrees with canon, CANON WINS and this document is the defect.

TIER: PROPOSAL — ENGINE, not CANON. It sets no canon, names no region content, changes no spine

or registry row, and adopts nothing. It is a component survey with a licence read and three

candidate pipelines. Nothing in it has been installed or run this pass.

Written: 2026-08-09. Lane: music open-stack step-back. Consumer: the round-5

architecture ruling, alongside STEPBACK_POSTMORTEM.md (what our own stack did) and

STEPBACK_COMMERCIAL_ENGINES.md (how the products people pay for are built). This is the third

leg: what is actually available, open, and runnable on the box.

---

DERIVATION

DERIVED FROM:

fields note, comparison_card.lines[], comparison_card.musical_axes_all_three_rounds[],

grade_verbatim (round 4, graded 2026-08-09) — the grade this doc is answering

§12 — the owning PIPE dossier; it binds this lane (memory pipe-dossiers-bind-generation-lanes).

This doc extends it on one axis and corrects nothing in it.

CC0-sfz chain as the final realisation path"* under RETIRES)

sfizz 1.2.3 (BSD-2-Clause) driving VSCO 2 CE, VCSL, legato_vocal_tutorial, body_percussion

(all CC0-1.0), each pinned by commit SHA

declared legal posture on audio conditioning (§8.3 below turns on the exact wording)

ace_step, stable_audio_3_medium, legato_vocal, body_percussion)

EXPLICITLY triggered prohibition

NOT DERIVED (authored judgment, and why it had no canon home):

this doc feeds, not one it may make.

none exists the column says so in those words. The judgment of what a benchmark is WORTH for our

use (orchestral cue, 60-120 s, melody-bearing) is mine and is declared as such.

our own renders. The script's docstring states the posture; the scoping inference is mine, and it

is the director's to veto because a pipeline depends on it.

---

0. The one question

Round 4 was graded, verbatim (docs/review_candidates.json, PASS7_FLORES_ROAD.grade_verbatim,

2026-08-09):

you are using. I hate how slow the melody is and there is no variation or changes or movement

throughout the songs. Everything sucks."

Two of those clauses are about COMPOSITION and two are about REALISATION. A model survey cannot fix

composition — STEPBACK_POSTMORTEM.md already ruled that the melodies are typed pitch tuples

developed by rules, and no amount of new renderer changes that. So this survey is organised around

the half a model CAN answer, and it asks every candidate the same question:

**Given a leitmotif that Josh owns and our engine composed — a specific sequence of specific
pitches at specific times — can this component realise THAT, or does it only produce something of
its own that sounds a bit like the prompt?**

That question is not decoration. It is the composed-never-prompted ruling stated as an engineering

requirement. A model that cannot be given a melody is not a candidate at any quality level, because

adopting it means Josh stops owning the themes. **Every table below carries a MELODY CONDITIONING

column and it is the first column that can disqualify.**

0.1 The bindings, restated so no reader has to fetch them

realise a theme it is GIVEN. A generator that invents the theme is prompting.

ladder.

bought thing is the remaining gap.

component whose weights were learned from a living tradition's recordings is a care question, not

only a licence question.

0.2 What this doc is NOT

It is not a re-run of PIPE_MUSIC_MODELS_2026-08-06.md. That dossier surveyed the generator field

three days ago and benched three arms on this box; its findings stand and are cited, not

repeated. Its own §9 ("what this document does not answer") left the axis this doc takes:

realisation of a supplied melody, and the renderer ladder above the CC0 sfz chain. Where a fact

appears in both, this doc marks it and moves on.

It also adopts nothing, installs nothing and spends nothing. Every "candidate" in §9 ends at a first

proof step with a cost.

---

1. Honesty register

Every claim below carries one of two marks, and the mark is the point.

The URL is in §11.

source was NOT fetched this pass. It may be true. It is not evidence yet.

No claim below is marked VERIFIED on the strength of a search-result snippet. That distinction is

the thing fetch_sfz_stack.py's docstring calls out — *"the licence BODY of each component is stored

verbatim ... not summarised — the summary is what the critic rejected."* Nothing here has had its

licence body pinned to docs/licence_records/, which is the step BEFORE any of it renders a byte.

---

2. The hardware envelope — what 32 GB of Blackwell actually buys

The box: RTX 5090, 32 GB GDDR7, compute-clean, 24/7 (memory 5090-box-live-remote-stack).

first stable release shipping pre-built cu128 wheels with sm_120 kernels. *(LANE-REPORTED —

PyTorch forums + issue 159207.)* Our own SA3 adapter already runs torch 2.7.1+cu128 on this box

(PIPE_MUSIC_MODELS §12.1), so the PyTorch half of the ecosystem is proven here.

this survey — is JAX/MLX only. A jax[cuda12] build on sm_120 is plausible and unverified. That is

a named risk with a cheap resolution (one install night), not a blocker.

while we were not looking:

ComponentParamsPeak VRAMFits 32 GB?
Stable Audio 3 Medium1.4B1.69–6.52 GB (chunked decode 5.14 GB at 380 s)trivially — VERIFIED
ACE-Step 1.5 XL-SFT + 4B LM4B + 4Bvendor table routes ≥24 GB to this tieryes, no offload — VERIFIED
Magenta RT 2 base2.4Bnot publishedalmost certainly; offline-only on NVIDIA — VERIFIED (offline-only)
Anticipatory Music Transformer~780M (medium)<4 GBtrivially
NotaGen-large516M<4 GBtrivially
sfizz + sample libraries0 (CPU)

**The consequence: we can run a SYMBOLIC model, an AUDIO model and a SAMPLER concurrently on one

card with room left.** A pipeline does not have to choose one of the three families. That is the

single most useful hardware fact in this document, and it is what makes §9's stacked pipelines

buildable rather than aspirational.

---

3. Symbolic / MIDI generation — the family that can be GIVEN a melody

This family is structurally aligned with composed-never-prompted: it operates on notes, so a note

we authored can be pinned and the model asked to work around it. That is the whole reason to look

here first.

ModelOrg · dateLicence (code / weights)MELODY CONDITIONING — can it take OUR theme?VRAMQuality ceiling, honestly
Anticipatory Music Transformer (stanford-crfm/music-medium-800k)Stanford CRFM · 2023, repo liveApache-2.0 code — VERIFIED. Weights licence NOT stated in the repo — VERIFIED that it is unstated; HF card read outstandingYES, and it is the cleanest instance of the primitive in the whole survey. The repo exposes extract_instruments() to split a melody line out of a score and then generates the REST conditioned on it; the paper's own framing is infilling and accompaniment conditioned on supplied events. Josh's tune goes in as fixed control events; the model writes around it and never over it<4 GBHuman evaluators rated its accompaniments comparably musical to human-composed ones over 20-second clips *(LANE-REPORTED — paper abstract)*. Twenty seconds is the honest number. It is a bar-filler inside a scaffold our code owns, not a cue writer. Provenance flag: trained on the Lakh MIDI dataset, which is scraped MIDI of largely copyrighted popular songs. For a shipped commercial game that is an outstanding read, not a settled clearance
NotaGen / NotaGen-XCentral Conservatory / IJCAI 2025 · sizes 110M / 244M / 516MMIT — LANE-REPORTED (HF card and repo LICENSE both read MIT in search results)NO — it is a MATERIAL source, not a realiser. Conditioning is period-composer-instrumentation only. It writes classical ABC from a style prompt; there is no slot for a supplied theme<4 GBIts own A/B tests report it beating baselines against human compositions in the classical idiom *(LANE-REPORTED)*. Genuinely the strongest "writes like a trained musician" symbolic model found. And it is the one that most directly collides with composed-never-prompted — see §8.5
MuseCocoMicrosoft muzicMIT — *(carried from PIPE_MUSIC_MODELS §3.2)*Attribute conditioning only (instrument, bars, time signature, key, tempo, pitch range, rhythm intensity). No melody slotsmallNo 2025-26 activity. Attribute surface is the best match to our card dimensions and the model is stale
MIDI-RWKV2025 · RWKV-7not established this passInfilling model — masked-region regeneration with numerical attribute controls. Infilling IS a melody-preserving primitive if the melody track is left unmasked *(LANE-REPORTED — paper only)*smallPaper-tier. No adoption signal found
MIDI-GPTMetacreation · AAAI'25repo MIT, weights CC-BY-NC-4.0strongest control list in the class; bar/track infilling preserves a supplied linesmallREFUSED — the NC term is on the weights and a shipped game is commercial. Unchanged from PIPE_MUSIC_MODELS §4
text2midiAAAI 2025not established this passtext → MIDI. No melody slotsmallBeats MuseCoco on tempo/key adherence and in listening tests *(LANE-REPORTED)*. Wrong shape for us regardless
MIDI-LLMISMIR 2026 (arXiv v2 2026-08-04)HF tag llama3.2, terms UNVERIFIED *(carried from PIPE_MUSIC_MODELS §3.2)*text → multitrack MIDI. No melody slot16 GB+Newest in the class; licence read outstanding

3.1 The finding this family produces

Exactly one open, permissively-licensed symbolic model implements the primitive we need, and it

is the Anticipatory Music Transformer: *give it the notes Josh owns, get back everything else.* Every

other member of the family either has no melody slot (NotaGen, MuseCoco, text2midi, MIDI-LLM) or is

licence-refused (MIDI-GPT).

And the second half of that finding, stated plainly: AMT's control surface is thin. It has no

tempo curve, no section map, no dynamics envelope, no instrumentation control worth the name. Which

is precisely the gap PIPE_MUSIC_MODELS §3.2 already closed by pointing at our own code — *"the

symbolic route's control does not come from a model at all — it comes from author_head_cell.py and

orchestrate.py, which already own tempo, form, lane count, arrivals and silence by construction."*

That remains true. AMT is a candidate for ONE job: **writing the inner voices, counter-lines and

accompaniment that round 3's grade said were missing** ("never any harmonies added on"), inside bars

whose tempo, length, instrumentation and dynamic shape our code already fixed.

---

4. Audio generation — the family that sounds recorded and mostly cannot be given a tune

ModelOrg · dateLicence (code / weights)MELODY CONDITIONINGVRAMQuality ceiling, honestly
ACE-Step 1.5 (2B + XL 4B DiT; 0.6B/1.7B/4B LM)ACE Studio · v1.5 late Jan 2026, XL 2026-04-02MIT (code) — VERIFIED. Weights MIT — LANE-REPORTED (HF ACE-Step/Ace-Step1.5 card). Our incumbent; licence body already pinned at docs/licence_records/ace_step/Indirectly, and this is the underused finding. No MIDI input anywhere — VERIFIED by reading the repo. But it ships Reference Audio Input ("use reference audio to guide generation style"), Cover Generation ("create covers from existing audio"), Repaint & Edit ("selective local audio editing and regeneration"), Vocal2BGM ("auto-generate accompaniment for vocal tracks") and Track Separation — all VERIFIED verbatim from the repo. Feed it a rough render of OUR score and every one of those becomes a melody-preserving operationVERIFIED vendor table: ≤6 GB 2B-turbo; 8-16 GB turbo/sft + 1.7B LM; 20-24 GB XL; ≥24 GB XL-sft + 4B LM = best qualityBenched on this box and its ceiling is confirmedPIPE_MUSIC_MODELS §6.2: the vendor-recommended XL-SFT + 4B LM tier reached 0 of 8 probe dimensions against pinned-turbo's 3 of 8, and landed FURTHER on all three ear cues. As a text-to-song GENERATOR it is measured and it is not enough. It has never been measured as a refiner over our own audio
Stable Audio 3 (Small 433M / Medium 1.4B / Large 2.7B)Stability AI · Medium 2026-05-20Stability AI Community License — VERIFIED: free commercial use under USD $1M annual revenue; above that, registration and possibly a paid Enterprise licence. "You own outputs generated from the Core Models" — VERIFIED. Second chain: Gemma Terms, via the t5gemma-b-b-ul2 text encoder. Both bodies pinned at docs/licence_records/stable_audio_3_medium/YES via init_audio + init_noise_level — VERIFIED from the repo docs. *"init_noise_level controls how much the init audio influences the output (0.0–1.0, default 1.0). At 1.0 the init audio is fully replaced by noise... Lower values preserve more of the original."* Plus inpainting over multiple non-contiguous regions and causal continuationVERIFIED: 1.69–6.52 GB peak, 5.14 GB with chunked decode at 380 s. 44.1 kHz stereo, up to 380 s (Medium)Our best-measured generator. PIPE_MUSIC_MODELS §12.6: 4/8 responding dimensions, 7/8 moving, best distance on all three cues (mean 0.3556 → 0.2777), ~80× the throughput. Also the honest counterweight in §12.7 — the blind timbre critic reads our authored+sfizz route as the arm FURTHEST from real released game music, and ADMITS the six tracks Josh rejected. A PASS from our instruments is worth exactly that much
Magenta RealTime 2 (base 2.4B / small 230M)Google DeepMind · 2026-06-04Apache-2.0 code + CC-BY-4.0 weights — VERIFIED from the HF card. *"Google claims no rights in outputs you generate"* — VERIFIED. The cleanest licence in the survey, and the only one with no revenue threshold, no registration and no second chainYES — DIRECT, FRAME-WISE MIDI CONDITIONING. Verbatim from the model card: *"128-dim multihot vector representing the state of each MIDI pitch during this frame (0 = Off, 1 = Sustain, 2 = Onset, 3 = Sustain or onset, model decides)"* — VERIFIED. Style comes separately as 12 MusicCoCa tokens from a text OR audio prompt. This is the primitive stated exactly: our notes drive the pitches, a prompt drives the timbre worldnot published; 2.4B at bf16 ≈ 5 GB — fits with enormous headroomUnknown, and that is the honest word. *"Model evaluation metrics ... will be shared in our forthcoming technical report"* — no published benchmark. Trained on ~71k h of mostly-instrumental stock music, which is a real prior against orchestral character. 48 kHz stereo. Architecture cost: 25 Hz frames under a ~20 s effective receptive field — it is a live instrument, so long-form continuity is OUR problem, and it emits a MIXDOWN, not stems
YuE (7B)M-A-P / HKUST · Jan 2025Apache-2.0 including weights — LANE-REPORTED; attribution to "YuE by HKUST/M-A-P" requestedgenre/instrument/section tags, dual-track ICL. No melody slot24 GB = 2 sessions; 80 GB for a full song *(carried from PIPE_MUSIC_MODELS)*14 months stale by this field's clock. Lyrics-to-song shape, wrong for instrumental cues
DiffRhythm v1.2 / DiffRhythm 2ASLP@NPUApache-2.0 — LANE-REPORTEDstyle prompt, ref audio, instrumental mode, LRC, duration. No note-level slot8 GB minFast (full song ~10 s) and controllable only at the style level. No control gain over the incumbent
HeartMuLa-oss-3BHeartMuLa · 3B released 2026-01-14, RL variant 01-23Apache-2.0 — VERIFIED from the repolyrics + tags + length. No melody slotfitsElo 1421.8, below ACE-Step *(carried from PIPE_MUSIC_MODELS §3.1)*. The claimed-better 7B is unreleased
LeVo 2 / SongGeneration 2 (4B)Tencent AI Lab · 2026-03-01NOASSERTION / Tencent custom, academic-only — LANE-REPORTED, and PIPE_MUSIC_MODELS §4 already flagged this refusal as needing re-verificationstructured lyrics + style prompt4BReported to outperform all open baselines on all six dimensions of its own eval *(LANE-REPORTED)*. Refused on licence; the refusal itself is still unverified
Khala 1.0CCoM + Tsinghua · 2026-05-01CC-BY-NC-4.0text, lyrics, duration≥24 GBThe highest-scoring open model in the field (Elo 1510.9) and it is unusable. No workaround: the NC term is on the weights
MusicGen / AudioCraft / JASCOMetacode MIT, weights CC-BY-NC-4.0 — VERIFIED from the model cardJASCO is the best melody+chord+drum conditioning surface in the entire survey and MusicGen-melody's chromagram conditioning is the canonical implementation of the primitive3.3B maxREFUSED, and it is the painful one. Every guide on the open web still recommends MusicGen-melody for exactly our use case. The weights licence forecloses it for a commercial game
Stable Audio Open 1.0Stability AICommunity License; training data 472,618 Freesound + 13,874 FMA, all CC0/CC-BY/CC-Sampling+ — VERIFIED, with Audible Magic screeningtext-to-audio only; 47 s maxsmallSuperseded by Stable Audio 3 for our purposes. Kept in the survey for one reason: it has the cleanest declared training-data provenance of any audio model here, which matters if provenance ever becomes the deciding axis

4.1 The finding this family produces

Three of these can be handed our music. None of them can be handed our SCORE.

refined. The melody survives because it is physically present in the input.

it happens to carry the best licence.

And the field-level fact worth stating once: **the two best-scoring open models in the world right

now (Khala, and arguably LeVo 2) are both licence-refused, and the best conditioning surface ever

built for this exact problem (JASCO) is licence-refused.** The open audio field's quality frontier

and its usable frontier are not the same frontier. Anyone re-running this survey will re-discover

that; it is written here so they do not have to.

---

5. Rendering above CC0 sfz — where "I hate the instruments" actually lives

STEPBACK_POSTMORTEM.md put *"the CC0-sfz chain as the final realisation path"* under RETIRES. This

section is what replaces it. There are two ladders — sampled and neural — and they fail in opposite

directions.

5.1 The incumbent, named precisely

Read at HEAD from harness/music_gen/sfz_palette.py:

dfcf4a49), legato_vocal_tutorial (CC0-1.0), body_percussion (CC0-1.0).

*"THE PLAYERS ARE A FREE CC0 COMMUNITY SAMPLE LIBRARY played by an offline sampler, one dynamic

layer a note."*

"One dynamic layer a note" is the load-bearing clause and it is OURS, not the library's. VSCO 2

CE ships multiple velocity layers on many instruments. A renderer that plays one layer per note

produces a dead performance from ANY library, free or bought. **Which means the first rung of the

render ladder is not a library at all — it is emitting expression.** See §8.2.

5.2 The sampled ladder — free, pro-grade, and it needs a host

OptionLicence — face-value readWhat it costs to useQuality ceiling
VSCO 2 CE + VCSL (incumbent)CC0-1.0. Unconditional. No attribution, no restrictionzeroCommunity-recorded, sparse articulations, thin dynamic layering. The measured floor: our own blind timbre critic reads this route as the arm furthest from real released game music (PIPE_MUSIC_MODELS §12.7)
Spitfire BBC SO DiscoverSpitfire EULA, quoted from the official EULA page: *"provided that you use the purchased Sound File(s) only within your own newly-created sound recording(s) and/or performances in a manner that renders the Sound File(s) substantially dissimilar to the original sound of the Sound File"* — LANE-REPORTED (search snippet of the official EULA; body not pinned). Face-value read (memory license-reads-not-conservative-solo-dev-tooling): a multi-instrument orchestral cue, mixed and reverbed, IS a newly-created recording substantially dissimilar to any one sample. The grant fires. Nothing in the FAQ or EULA found this pass prohibits games. The one genuinely triggered prohibition is the obvious one: we may not ship or redistribute the samples themselves, and a near-raw single held note is the edge case to avoidFree, but it is a proprietary VST3/AU/AAX plugin — VERIFIED from the Spitfire FAQ, which also states a 64-bit DAW is required and the licence is single-user (two machines). sfizz cannot drive it. This is the whole cost: a plugin host enters the chainRecorded at Air Studios by the BBC Symphony Orchestra. It is a *reduced* tier — one dynamic layer per articulation on many patches, no round-robins — but the SOURCE is a professional orchestra in a professional hall. That is a categorical step above community samples, not an incremental one
Virtual Playing OrchestraREFUSED, and the refusal is already on the record (PIPE_MUSIC_MODELS §4): the maintainer grants unrestricted commercial use, but the library re-imports Sonatina Symphonic Orchestra under CC Sampling Plus 1.0, whose advertising/promotional-use exclusion PIPE_AUDIO_MUSIC §8.1 already ruled fires on a game soundtrack. A downstream maintainer cannot relicence upstream samples. Recorded here because a search for "free orchestral SFZ" returns it first
Sonatina Symphonic OrchestraCC Sampling Plus 1.0 — same clause, same refusal

The host, which is the actual new component:

HostLicenceFit
DawDreamerGPLv3 — LANE-REPORTED. Built on JUCE; also drags in the Steinberg VST3 SDK termsPython VST2/VST3/AU host with offline batch rendering, MIDI playback, per-parameter automation at audio rate and PPQN, and multi-processor graphs. It is a BUILD-TIME TOOL, not a shipped component — we distribute rendered WAV, not the host. GPLv3 binds redistribution of the software; it does not reach the audio it renders. That is the same reasoning sfz_palette.py already records for sfizz's BSD-2 notice condition (*"the BSD-2 notice condition binds redistribution of the software, and this lane redistributes rendered audio only"*). Face-value: clean for our use. Risk: some plugins misbehave headless — the project documents this
REAPER (CLI)commercial, DRM-free, cheap; 60-day full evaluation — LANE-REPORTED-batchconvert with a filelist, plus ReaScript in Lua/Python/EEL. The industry-proven batch path. Costs money, so proof-before-spend puts it second
Carla / kuriborosuGPLheadless plugin host, simpler API, less automation control

5.3 The neural ladder — and it is almost entirely licence-blocked

This is the class that would render a score with learned human phrasing rather than sample playback.

It is the most exciting family in the survey and the least usable.

SystemLicenceWhat it rendersVerdict
Magenta RT 2Apache-2.0 + CC-BY-4.0 weights — VERIFIEDaudio conditioned on frame-wise MIDI pitch state + a style embeddingThe ONLY commercially-usable neural MIDI realiser found. See §4. Its constraints are architectural, not legal
Multi-Aspect Conditioning for Diffusion-Based Music Synthesis (benadar293, TASLP 2024)CC-BY-NC-SA-4.0 — VERIFIED from the project pageMIDI → audio for solo instruments, chamber ensembles, orchestras, jazz/rock bands, drums; trained on ~58 h of REAL audio; T5 backbone + SoundStream vocoderREFUSED. It is the closest thing in existence to "play our orchestral score with real recorded character," and NC forecloses it
RenderBox (2025)*"distributed for non-commercial research purposes only"* — LANE-REPORTEDtext + score → expressive multi-instrument performance audio; diffusion transformer with cross-attention joint conditioningREFUSED. Same shape, same problem
nii-yamagishilab/midi-to-audioApache-2.0 code + CC-BY-4.0 weights — VERIFIEDPiano only (MAESTRO). Transformer-TTS MIDI→mel + HiFiGAN mel→audioLicence-clean and the wrong instrument. Useful only if a cue is piano-led
MIDI-VALLE (ISMIR 2025) · Pianist Transformer (135M)not established this passexpressive piano performance synthesis / renderingPiano-only; the class's centre of gravity is piano because MAESTRO exists and no orchestral equivalent does
MIDI-DDSP (Magenta)Apache-2.0monophonic-instrument synthesis with note-expression controlRepo archived 2024-02-01, TF 2.7 / Python 3.8 — a dead stack on Blackwell *(carried from PIPE_MUSIC_MODELS §3.3)*
Google music-spectrogram-diffusionApache-2.0 (weights on HF)multi-instrument MIDI → spectrogram → audio2022-era quality; superseded by everything above it in this table

5.4 The finding this section produces

The neural rendering class is real, it is good, and for a commercial game it is one model wide.

Every orchestral-capable neural renderer found is CC-BY-NC or research-only except Magenta RT 2;

every licence-clean one except Magenta RT 2 is piano.

Meanwhile the sampled ladder has a free rung sitting unclimbed: **a professional orchestra recorded

at Air Studios, free, with a licence that on a face-value read grants exactly what we need — and the

only thing between us and it is a plugin host in the render chain.**

---

6. Steering and conditioning — the primitives, separated from the products

Products change every eight weeks. The primitives do not, and naming them is what lets a later lane

evaluate something this survey never saw.

1. Symbolic conditioning (notes in, notes out). A supplied note stream is held fixed and the

model completes around it. *Instances:* AMT extract_instruments() + control events; MIDI-RWKV

infilling; MIDI-GPT bar/track infilling (refused). **Fidelity to our melody: EXACT — the notes are

not regenerated, they are held.**

2. Frame-wise symbolic → audio conditioning. Note state is fed to an audio model per frame.

*Instance:* Magenta RT 2's 128-dim multihot pitch vector. **Fidelity: high by construction — the

model is told which pitches sound in every frame.** The 3 = model decides state is the honest

nuance: it can be given latitude where we want latitude.

3. Chromagram / melody-audio conditioning. A melody's pitch-class energy over time steers

generation. *Instance:* MusicGen-melody, JASCO (both refused). **Fidelity: pitch-class only —

octave and rhythm can drift.**

4. Audio-to-audio / init-audio (the img2img analogue). A rough render is partially re-noised and

denoised under a text prompt. *Instances:* Stable Audio 3 init_noise_level (VERIFIED

semantics — lower preserves more); ACE-Step Reference Audio / Cover / Vocal2BGM. **Fidelity:

TUNABLE and that is the point — there is a strength band where structure survives and timbre

changes, and finding it is a sweep, not a guess.**

5. Masked inpainting / repainting over a timeline. Regenerate bars 17-24 and hold everything

else. *Instances:* Stable Audio 3 multi-region binary masks (still UNMEASURED per

PIPE_MUSIC_MODELS §12.9); ACE-Step Repaint & Edit. **This is the primitive that lets a cue be

FIXED rather than re-rolled**, and it is the closest thing in the audio family to an event grammar.

6. Style embedding from a reference. Timbre/genre is set by an embedding rather than by words.

*Instances:* Magenta RT 2's MusicCoCa audio prompt; ACE-Step reference audio; ACE-Step LoRA from a

handful of songs. **CARE + LICENCE GATE: sa3_generate.py already declares the corpus posture —

the exemplar bytes NEVER become model input. Any use of this primitive must source its reference

from audio we own.**

7. Expression automation into a sampler. CC1/CC11/velocity/keyswitch curves shape a sampled

performance. *Instances:* DawDreamer PPQN + audio-rate automation; sfizz opcodes. **The

zero-model, zero-VRAM primitive we are currently not using, and the one that "one dynamic layer a

note" names as absent.**

---

7. The licence board — one table, go / no-go

Face-value reads per memory license-reads-not-conservative-solo-dev-tooling. **Nothing here is

cleared until its body is pinned to docs/licence_records/.**

ComponentVerdict for a commercial shipped gameThe clause that decides it
ACE-Step 1.5GOMIT (code VERIFIED; weights LANE-REPORTED). Already pinned
Stable Audio 3 MediumGO, with one live obligationCommunity License: free under $1M revenue; outputs owned. Section III commercial-use REGISTRATION is triggered and NOT SATISFIED (PIPE_MUSIC_MODELS §12.9 — Josh's account action). Second chain: Gemma Terms
Magenta RealTime 2GO — cleanest in the surveyApache-2.0 code + CC-BY-4.0 weights; *"Google claims no rights in outputs you generate."* No revenue threshold, no registration, no second chain. CC-BY attribution is the only obligation
Anticipatory Music TransformerCONDITIONALApache-2.0 code VERIFIED; weights licence unstated and Lakh MIDI training provenance is an open question for a shipped game. Two reads outstanding
NotaGenGO on licence (MIT, LANE-REPORTED)but see §8.5 — the constraint that binds it is composed-never-prompted, not the licence
HeartMuLa-oss-3B · YuE · DiffRhythmGO on licence (Apache-2.0)no melody slot; no reason to adopt
Khala 1.0NOCC-BY-NC-4.0 on the weights
MusicGen / AudioCraft / JASCONOCC-BY-NC-4.0 on the weights
MIDI-GPTNOCC-BY-NC-4.0 on the weights (repo MIT — the split is the trap)
LeVo 2 / SongGenerationNO (refusal itself still LANE-REPORTED)Tencent custom, academic-only
benadar293 multi-aspect · RenderBoxNOCC-BY-NC-SA-4.0 / non-commercial research only
VSCO 2 CE · VCSL · legato_vocal · body_percussionGOCC0-1.0, unconditional. Pinned
BBC SO DiscoverGO on a face-value read; body not yet pinned*"only within your own newly-created sound recording(s) ... substantially dissimilar to the original sound"* — an orchestral cue satisfies it. Do not ship samples; avoid near-raw one-shots
Virtual Playing Orchestra · Sonatina SONOCC Sampling Plus 1.0 upstream, advertising/promotional exclusion. Already ruled
DawDreamerGO as a build-time toolGPLv3 binds redistribution of the SOFTWARE. We redistribute rendered audio. Same reasoning already recorded for sfizz
REAPERGO, costs moneycommercial licence, DRM-free. Proof-before-spend defers it

---

8. WHAT IT MEANS FOR OUR STACK

8.1 The instrument complaint has a free answer that we have not tried

Josh said *"I hate the instruments you are using."* We are using community CC0 samples. **A

professional orchestra recorded at Air Studios is available for free, under a licence that on a

face-value read grants exactly our use.** The only thing standing between the current chain and that

one is that BBC SO Discover is a VST3 plugin and sfizz plays SFZ. **DawDreamer closes that gap in

Python, offline, at zero cost and zero VRAM.** No model, no spend, no new licence risk beyond one

body to pin. This is the highest ratio of grade-clause-answered to effort in the entire document.

8.2 …and it will NOT work on its own, which must be said before it is tried

comparison_card.honest_limits_short says *"one dynamic layer a note."* That is a property of our

REALISER, not our library. Hand a world-class orchestral library a MIDI file with static velocities,

no CC1 expression curve, no CC11 dynamics, no articulation keyswitches and no legato intervals, and

it will sound like a better sampler playing a dead performance. **Rung one of the render ladder is

emitting expression from orchestrate.py — velocity shaping, CC1/CC11 curves, articulation

selection per phrase.** The library is rung two. Attempting rung two first will produce a null result

and a wrong conclusion, and that is a prediction this document is making on the record so it can be

checked.

8.3 Audio-to-audio is available to us and a lane might wrongly believe it is banned

harness/music_gen/sa3_generate.py declares, correctly and deliberately:

*"NO AUDIO EVER CONDITIONS A GENERATION. ... The corpus legal posture
(build/audio/exemplars/CORPUS.json :: legal_posture) is absolute — 'The bytes NEVER become model
input: no training, no fine-tuning, no conditioning, no derivation.'"*

**That posture is about the EXEMPLAR CORPUS — commercial recordings we do not own, held for

measurement only. It is not, and was never, a rule against conditioning on audio we rendered

ourselves from a score we composed.** Feeding SA3 or ACE-Step our own sfz render is categorically

outside the thing that posture forbids.

This inference is mine and is flagged as NOT DERIVED. It is load-bearing: candidate pipeline C

depends entirely on it, and if the director reads the posture as blanket, C dies and should die

cleanly rather than be built on a misreading. The safe implementation is explicit — a separate entry

point with a positive control asserting the init audio's provenance is build/audio/ and never

build/audio/exemplars/.

8.4 There is exactly one licence-clean neural realiser and we have not looked at it properly

PIPE_MUSIC_MODELS §5 shortlisted Magenta RT 2 third and declined to install it on evidence:

*"its own documentation gives it no form control, no published benchmark, and no NVIDIA real-time

path."* Every word of that is true and one of the three objections does not apply to us.

"No form control" is only a defect for a system that must own form. Our system already owns form.

author_head_cell.py and orchestrate.py fix tempo, section map, lane count, arrivals and silence

by construction — PIPE_MUSIC_MODELS §3.2 says so in those words. What we lack is not form; it is

SOUND. Magenta RT 2 is the only component in this survey that takes our notes directly and returns

audio under a licence with no threshold, no registration and no second chain.

The two objections that DO stand, and they are serious:

possibility that our leitmotif comes back sounding like production library music, which is the

precise aesthetic Josh has rejected four times.

measurement rig operate on lanes. A mixdown-only realiser costs us those instruments — it is not

a drop-in for the sampler, it is a different architecture with a different evidence surface.

The resolution is cheap and it is a proof, not an argument: one install night (JAX on sm_120,

unproven on this box), one leitmotif, one 30-second render, judged by the blind timbre critic

(harness/music_gen/timbre_critic.py) and by motif preservation. That is a night, not a program.

8.5 The one fork this document may not decide, restated because a new candidate sharpens it

STEPBACK_POSTMORTEM.md already named it: **is a learned model PROPOSING melodic material

composition, or is it prompting under composed-never-prompted?** That fork is Josh's.

This survey sharpens it into a clean three-way, because the candidates now separate cleanly:

they are given. Compatible with composed-never-prompted under any reading.

Compatible under most readings; it is the "develops, does not invent" case, and it is the direct

answer to round 3's *"never any harmonies added on."*

classical-craft symbolic model found, so the cost of the strict reading is real and should be

named rather than absorbed silently.

No pipeline in §9 depends on the third case. All three are buildable under the strictest reading

of the rule. That is deliberate.

8.6 What did NOT change since 2026-08-06

Re-verified this pass: no new open checkpoint has landed in the three days since the PIPE dossier.

ACE-Step's lineage still stops at v1.5/XL (there is no v1.6 or 2.0). Khala, LeVo 2 and JASCO are

still licence-refused. The instrumental quality frontier is still closed-weights. **The survey did

not go stale; the axis this doc reads was simply not the axis that dossier read.**

---

9. Three candidate pipelines, under the composed-skeleton constraint

Every pipeline below starts at the same place — **a leitmotif Josh owns, a score our engine

composed** — and differs only in how that score becomes sound. None invents melody. Each ends at a

first proof step with a cost, per proof-before-spend.

They are stackable, not exclusive. A is the floor and both others sit on top of it.

---

PIPELINE A — "Better players" · composed MIDI → pro-grade sampled orchestra

orchestrate.py  ──emit MIDI + CC1/CC11/velocity/keyswitch──►  DawDreamer (GPLv3, build-time)
                                                                    │  hosts VST3
                                                                    ▼
                                                     BBC SO Discover  ──►  per-lane stems
                                                                    │
                                                        existing mix_policy.py + battery

Steinberg VST3 SDK terms. No model licence at all.

This is the only pipeline with zero evidence-surface cost.

sampler playing the same dead performance. **Rung one is orchestrate.py emitting expression;

the library is rung two.** Build them in that order or the result is a null.

hosting of authorised commercial plugins is where DawDreamer's documented plugin-compatibility

caveats bite. If it does not host, REAPER CLI is the fallback and it costs money — which is

exactly the kind of spend proof-before-spend permits, because A's proof would have earned it.

expression curve and velocity shaping, render it BOTH ways — sfizz/VSCO2 and DawDreamer/BBC SO

Discover — and run both through timbre_critic.py and the existing battery. **Two numbers and two

files. If the timbre critic does not move, the instrument hypothesis is wrong and we learned it in

a night.**

---

PIPELINE B — "Neural realiser" · composed MIDI → Magenta RT 2 → audio, our code owns form

orchestrate.py ──► note stream ──► 128-dim multihot pitch state per frame ─┐
                                                                            ├─► Magenta RT 2 ──► 48 kHz stereo
style reference (audio WE own, or text) ──► 12 MusicCoCa tokens ───────────┘         │
                                                                                     ▼
                                                            our seam/continuity layer (loop_seams.py)

sampler closes — because the character comes from a model trained on recorded performance.

frame; the 3 = model decides state is deliberate latitude we control per-note.

registration, no second chain, *"Google claims no rights in outputs."*

1. Mixdown, not stems. Our mix policy, lane census and several instruments operate on lanes.

This pipeline costs us that evidence surface, or forces per-lane rendering and a re-mix — which

is untested and may not sum coherently.

2. ~20 s effective receptive field, 25 Hz frames. Long-form continuity over an 88-second cue is

OUR problem to solve by stitching. We already own loop_seams.py, so this is work, not a wall.

3. ~71k h of stock-music training prior, and no published benchmark. The genuine risk that our

leitmotif returns sounding like production library music.

present (PIPE_MUSIC_MODELS §5), so the path exists; nobody has walked it.

through mrt2_base with a text style prompt, and score it against the same cue rendered through

Pipeline A. Judged by timbre_critic.py (blind) plus a motif-preservation check. **The install is

the risk; the render is minutes.**

---

PIPELINE C — "Refiner ladder" · composed score → rough render → audio-to-audio at a swept strength

orchestrate.py ──► MIDI ──► sfizz or Pipeline A render (structure is now physically present)
                                              │
                                              ▼   init_audio, init_noise_level swept
                        Stable Audio 3 Medium (MIT-ish/Community) ──or── ACE-Step 1.5 (MIT)
                                              │
                                              ▼
                          motif_compare.py + timbre_critic.py, scored across the sweep

cohesion, performance grit — while the composition, form, arrivals and silence survive because they

are physically present in the init audio.

already installed and licence-pinned on this box.**

*"lower values preserve more of the original."* At a low noise level you get the same dead

performance under a haze; at a high one the melody drifts away from Josh's theme. There is a band

between and it is found by a sweep scored on melody preservation, not by a guess.

registration obligation**; ACE-Step carries none.

posture. If the director reads that posture as blanket, this pipeline does not exist. That read

is the cheapest thing in this document and should happen first.

arrival-placement and silence-budget dimensions our cards measure — measurable with instruments we

already have; (ii) it produces a mixdown, same evidence-surface cost as B, though here we keep

the pre-refinement stems as a fallback mix.

render, sweep init_noise_level across 0.1 / 0.2 / 0.35 / 0.5 under one caption, and plot melody

preservation against timbre-critic distance. **The output is a curve, and the curve either has a

usable band or it does not.** This is the cheapest proof in the document because every component

is already on disk.

---

9.1 How they compose, and the recommended order

They stack: **A is the renderer floor. C rides on A's output. B is an independent track that

bypasses the sampler entirely.**

Answers "hate the instruments"Answers "sounds like a mockup"Keeps stemsMelody fidelityInstall costSpend
Ayes, directlypartlyyesexactnonezero
Byespotentially, fullynohighJAX/sm_120 unprovenzero
Cyesyes, if a band existsno (A's stems survive as fallback)tunable, unprovennone — both installedzero

RECOMMENDATION (this lane's, and the director's to overturn): run C's sweep first because it

is hours on components already on disk and it either finds a band or eliminates a whole family;

then A rung one (expression from orchestrate.py) because §8.2 predicts everything else in A

depends on it; then A rung two (BBC SO Discover through DawDreamer); then B as the

independent track once the JAX install has a night to spare.

THE STRONGEST OBJECTION to all three: none of them touches *"the melody sucks"* or *"no variation

or changes or movement."* Those are composition defects, STEPBACK_POSTMORTEM.md located them in

pass3_flores.py:114-144, and a renderer cannot fix a tune. **If round 5 ships better instruments

playing the same melody, it will be graded down again, and this document will have been part of the

reason. The only component here that addresses the composition side at all is AMT** (§3.1) —

harmonising and counter-lining a theme it never alters — and it belongs in the round-5 architecture

alongside whichever renderer wins, not after it.

---

10. What this document does not answer

verbatim body under docs/licence_records/. **BBC SO Discover, DawDreamer, AMT's weights and

NotaGen all need that step before a byte renders.**

the provenance question is the more serious of the two.

It is the actual engineering in Pipeline A and it is not costed here.

UNMEASURED (PIPE_MUSIC_MODELS §12.9) and is not exercised by any pipeline above.

has not been run. Neither vendor enumerates its corpus at the granularity that question needs.

game — is out of scope here and boarded.

---

11. SOURCES

All fetched or searched 2026-08-09.

Audio generation

https://github.com/ace-step/ACE-Step-1.5/blob/main/README.md · paper:

https://arxiv.org/abs/2602.00744 · weights: https://huggingface.co/ACE-Step/Ace-Step1.5

https://github.com/Stability-AI/stable-audio-3/blob/main/docs/workflows/inference.md

https://huggingface.co/google/magenta-realtime-2 · docs:

https://magenta.github.io/magenta-realtime/ · original announcement:

https://magenta.withgoogle.com/magenta-realtime

https://github.com/facebookresearch/audiocraft/blob/main/model_cards/MUSICGEN_MODEL_CARD.md

https://huggingface.co/m-a-p/YuE-s1-7B-anneal-en-icl

https://arxiv.org/html/2606.30642v1

Symbolic / MIDI

https://arxiv.org/abs/2306.08620 · weights: stanford-crfm/music-medium-800k

https://huggingface.co/ElectricAlexis/NotaGen · https://arxiv.org/abs/2502.18008

MIDI-to-audio rendering

https://benadar293.github.io/multi-aspect-conditioning/ ·

https://github.com/benadar293/multi-aspect-conditioning

Renderers, hosts and libraries

https://support.spitfireaudio.com/en/articles/11815834-bbc-symphony-orchestra-discover-faq ·

EULA: https://www.spitfireaudio.com/info/eula/

plugin compatibility: https://dirt.design/DawDreamer/compatibility.html · paper:

https://archives.ismir.net/ismir2021/latebreaking/000001.pdf

Platform

https://docs.salad.com/container-engine/tutorials/machine-learning/pytorch-rtx5090

Repo-internal (canon and lane-canon)

harness/music_gen/fetch_sfz_stack.py

Generated by harness/site/structure_site.py — the URL path is the repo path. review root