music/MUSIC_ARCHITECTURE_DECISION_BRIEF.md
CANON SUBORDINATION — this document is PROPOSAL-TIER: it serves canon and never outranks it.
Canon served: melody-first with hummable thirty-year leitmotifs; composed-never-prompted;
the no-human-composer-hire ruling (2026-08-08 — "composer_in_loop" is the factory's own pass
ladder, never a hire); the CVD §17.1 care line and the §17 floor; local-first on the 5090;
proof-before-spend.
If this document disagrees with canon, CANON WINS and this document is the defect.
TIER: DECISION BRIEF — PROPOSAL. ENGINE, not CANON. It sets no canon, names no region content,
changes no spine or registry row, installs nothing and spends nothing. It presents four steelmanned
architectures with costs, risks and one-wave proofs, recommends one, states the strongest objection to
that recommendation, and carries the one genuine fork the post-mortem ruled was Josh's and not this
lane's (§7).
Written 2026-08-09. REVISION 2 — re-adjudicated in full against
docs/proposals/music/STEPBACK_OPEN_STACK.md, which did not exist when revision 1 was written and
whose absence revision 1 declared as a hole in its own §1.3 and §8. **That hole is now closed, and
closing it changed three of the four options.** Lane: music architecture comparator. Consumer: the
round-5 architecture ruling.
---
Josh, on round 4, verbatim (docs/review_candidates.json, PASS7_FLORES_ROAD.grade_verbatim, read
at HEAD for this brief):
you are using. I hate how slow the melody is and there is no variation or changes or movement
throughout the songs. Everything sucks."
Four rounds, four turndowns, six recorded negative verdicts and zero positive ones. Our own
instruments went from 7-of-57 to 57-of-57 over the same rounds his grade went down. The three
step-back docs establish why, and they agree: **every engine that satisfies listeners has four layers
— generation, rendering, curation, delivery — and we own one of them, at its floor in one case and
absent in two.** Four rounds of work went into deepening the one layer we already owned.
**The recommendation is Option A, run as four stages in proof-cost order, with two of the stages
running concurrently rather than in series.** The change from revision 1 is that the symbolic-infill
proof moves INTO wave one instead of waiting behind the rendering and selection work. Both sibling
docs now predict, on the record, that a round 5 shipping better instruments playing the same melody
gets graded down again. Revision 1's serial ordering walked into that prediction. This one does not.
Three verification results drove the re-adjudication, and each is a direct read of a primary source
this pass, not a search summary:
1. The Spitfire EULA — revision 1's named blocker — is now READ, and the grant fires. It also
carries a Section 10 prohibition on AI training that neither step-back doc has, and that clause
forecloses a path revision 1 left open (§2.1).
2. STEPBACK_OPEN_STACK.md §3.1 is wrong in a way that strengthens the recommendation. It finds
"exactly one open, permissively-licensed symbolic model implements the primitive we need." There
are two, and the one it missed has strictly better training provenance and is the component
revision 1 already recommended (§2.2).
3. A genuinely new component entered the option space — Magenta RealTime 2 — the only neural
realiser in the survey that takes our notes directly and carries a licence with no revenue
threshold, no registration and no second chain (§2.3). It re-scopes Option B.
---
Each was fetched and read this pass. Where a read contradicted a sibling document, that is stated at
the point of use and again in §8.
https://www.spitfireaudio.com/en-us/pages/spitfire-audio-end-user-license-agreement
(**note: the URL both step-back docs cite, https://www.spitfireaudio.com/info/eula/ , served a
PRIVACY POLICY on fetch this pass, not the EULA. The licence body is at the /en-us/pages/ URL
above.**)
https://github.com/m-malandro/composers-assistant-REAPER
https://github.com/Stability-AI/stable-audio-3/blob/main/docs/workflows/inference.md
https://support.spitfireaudio.com/en/articles/11815834-bbc-symphony-orchestra-discover-faq
NotaGen (https://arxiv.org/abs/2502.18008), Meta Audiobox Aesthetics
(https://arxiv.org/abs/2502.05139), CLaMP 3 (https://arxiv.org/abs/2502.10362), SMART
(https://arxiv.org/pdf/2504.16839), the annotation-free MIDI-to-audio refinement paper
(https://arxiv.org/pdf/2410.16785), ACE-Step 1.5 (https://github.com/ace-step/ACE-Step-1.5), AIVA
legal terms (https://www.aiva.ai/legal/1), ReaRender (https://github.com/YatingMusic/ReaRender), the
SWAM String Sections review (https://www.soundonsound.com/reviews/audio-modeling-swam-string-sections).
docs/proposals/music/STEPBACK_POSTMORTEM.md — the self-audit. §2 the anti-correlation and itsthree self-corrections, §3 the typed pitch literals, §5 the four structural limits, §6 the
grade-to-layer apportionment, §8.1 survives, §8.2 retires, §8.3 the six requirements and the fork.
docs/proposals/music/STEPBACK_COMMERCIAL_ENGINES.md — the outside read. §2 the four-layer frame,§3 AIVA, §4 Suno/Udio, §5 loop assembly, §6 middleware delivery, §7 why rule systems fail.
docs/proposals/music/STEPBACK_OPEN_STACK.md — the leg revision 1 was missing. §2 the hardwareenvelope, §3 symbolic models and their melody-conditioning surfaces, §4 audio models, §5 the render
ladder above CC0-sfz, §6 the seven conditioning primitives, §7 the licence board, §8 the stack
consequences, §9 three candidate pipelines.
docs/pipeline_review/tech_research/PIPE_MUSIC_MODELS_2026-08-06.md — the owning PIPE dossier, which binds this lane (memory pipe-dossiers-bind-generation-lanes).
docs/review_candidates.json — the round-4 verbatim grade quoted in §0.harness/music_gen/emit_review_picks.py — the generic ranked path exists (sorts by rank_in_theme, counts kept_in_theme); the three round emitters hardcode both to 1. **The
curation machinery is built and bypassed.**
harness/music_gen/ import census for timbre_critic — imported by slice_exemplars.py and instr_noise_structure.py only. No battery and no realiser imports it.
harness/music_gen/sfz_palette.py, sa3_generate.py, fetch_sfz_stack.py — the incumbentrenderer, the SA3 adapter's declared legal posture, and the licence-body-verbatim discipline.
---
Revision 1 called this "the specific blocker — not on the product page, and it gates the free proof's
publishability." STEPBACK_OPEN_STACK.md §5.2 quoted the key clause but marked it LANE-REPORTED from
a search snippet. It is now read directly. Verbatim from the licence:
and/or performances in a manner that renders the Sound File(s) substantially dissimilar to the
original sound"** of the Sound File in each case.
sampler; mixing, combining, filtering, re-synthesising or editing for use as sounds, multi-samples,
effects or patches in samplers or playback devices; giving Products to other persons; uploading to
file-sharing sites; renting, leasing, sublicensing or transferring copies.
restriction.**
systems or large language models (LLMs) without obtaining express written consent."**
Face-value read, per memory license-reads-not-conservative-solo-dev-tooling — face value, flag
only an EXPLICITLY triggered prohibition:
sound recording substantially dissimilar to any single Sound File. The grant fires. Nothing in the
licence prohibits a video-game soundtrack.** The prohibitions are all about redistributing the
SAMPLES — making a competing library, reformatting for another sampler, sharing the Products. We
redistribute rendered cues. Option C's licence gate is CLEARED on a face-value read.
an AI system on Spitfire-derived audio without written consent. That does not touch rendering, and
it does not touch conditioning a model at inference on our own render — training is the named act.
**But it forecloses one thing revision 1 left open: bootstrapping or fine-tuning any local model,
LoRA included, on a corpus of our own BBC-SO-rendered cues.** If a later wave wants a house
renderer trained on our own output, its training audio may not come from a Spitfire library. That
constraint belongs in the lane contract now, before anyone builds a corpus that cannot be used.
(LABS, BBC SO Discover) sit under identical terms — the safe and face-value read is that they do,
since it is the same agreement page, but it is not stated. And a near-raw single sustained note,
unmixed, is the one output shape where "substantially dissimilar" is arguable. Neither is a blocker;
both are avoidable by construction.
STEPBACK_OPEN_STACK.md §3.1 states: *"Exactly one open, permissively-licensed symbolic model
implements the primitive we need"* — the Anticipatory Music Transformer — and its §3 table surveys
NotaGen, MuseCoco, MIDI-RWKV, MIDI-GPT, text2midi and MIDI-LLM. **Composer's Assistant 2 does not
appear in that table at all.** Verified this pass, CA2 implements the same primitive with strictly
better provenance:
| axis | Anticipatory Music Transformer | Composer's Assistant 2 |
|---|---|---|
| Melody held fixed while the model writes around it | YES — extract_instruments() isolates the line, generate(inputs, controls) writes accompaniment conditioned on supplied melody events, ops.combine() recombines — VERIFIED from the repo | YES — *"Users place empty MIDI items in REAPER to tell the model in which measures to write notes, and track names to tell the model what instrument is on each track."* Selected items are targets; unselected items remain fixed context for the encoder — VERIFIED from the paper |
| Code licence | Apache-2.0 — VERIFIED | MIT — VERIFIED |
| Weights licence | not stated — VERIFIED that it is unstated | referenced in the repo, not detailed in the read — an outstanding read, one line |
| Training provenance | the full Lakh MIDI dataset — VERIFIED. Scraped MIDI of largely copyrighted popular songs. An open clearance question for a shipped commercial game | "copyright-free and permissively-licensed" MIDI files — VERIFIED from the paper; the repo says *"trained only on public domain and permissively-licensed MIDI files."* The cleanest provenance of any generative component in any of the three step-back docs |
| Control surface | thin — no tempo curve, no section map, no dynamics envelope | 34 controls, VERIFIED: horizontal onset density 6 bins (half notes to 4.5+ onsets per quarter), vertical density 5 bins, step propensity 7 bins, leap propensity 7 bins, pitch range strict and loose, rhythmic interest 3 bins, pitch classes per onset 5 bins, DNOC token preventing octave-shifted duplicates |
| Local / offline | local | local — *"The server is just a separate process that runs on your computer. It does NOT send any information over the internet."* Python 3.9+, CPU works, NVIDIA GPU optional |
| Size | ~780M medium | T5-like; large is 512-dim with 16 encoder and 16 decoder layers, small is 384-dim with 10 and 10 |
| Honest limitation | 20-second clips is the honest evaluation horizon | "most proficient at infilling classical, choral, and folk music" — VERIFIED, and it is §6's strongest objection. Medium rhythmic-interest control succeeds 42% against a 35.2% baseline; 4+ pitch classes per onset succeeds 8.92% |
The consequence for this brief: the open-stack leg does not overturn revision 1's stage-3 pick. It
independently discovered the same primitive in a weaker instance and concluded the field was one model
wide. It is two models wide, and CA2 is the better one on provenance, control surface and host
integration. AMT is the fallback, and its Lakh provenance is the price of using it.
Verified from the model card this pass, every line:
claims no rights in outputs you generate using Magenta RealTime 2. You and your users are solely
responsible for outputs and their subsequent uses."* **No revenue threshold, no registration, no
second licence chain. The cleanest licence attached to any generative component in the survey.**
receives a **128-dimensional multihot vector per frame representing MIDI pitch state — 0 = Off,
1 = Sustain, 2 = Onset, 3 = Sustain or onset (model decides)**. Our notes drive which pitches sound
in every frame. The 3 state is deliberate latitude we control per note.
derived from a text description OR an audio example, plus the audio context itself.
sources, mostly instrumental"** — a genuine prior against orchestral character and toward exactly
the production-library aesthetic Josh has rejected four times. **"Model evaluation metrics and
results will be shared in our forthcoming technical report"** — there is no published benchmark.
25 Hz frames, 20-second effective receptive field for both models, 48 kHz stereo.
MIXDOWN, not stems"* as a settled fact. The model card does not state output format. The
mixdown reading is architecturally near-certain for a single-stream audio decoder, but it is an
inference, not a verified fact, and the evidence-surface cost it implies for our lane census and mix
policy should be confirmed by the install night rather than assumed.
STEPBACK_OPEN_STACK.md §5.4 concludes: *"The neural rendering class is real, it is good, and for a
commercial game it is one model wide"* — every orchestral-capable neural renderer refused on licence
except Magenta RT 2. Its §5.3 table lists benadar293, RenderBox, nii-yamagishilab, MIDI-VALLE,
Pianist Transformer, MIDI-DDSP and music-spectrogram-diffusion. **Symphony Rendering (ICASSP 2026) is
not in it, and STEPBACK_COMMERCIAL_ENGINES.md §9.2 builds a whole pattern on it.** Verified this
pass:
reduction to symphonic arrangement), both *"in the style of a target composer."*
model, plus a global composer label embedding.
pretrained audio VAE latent at ~25 Hz, decoding to 48 kHz.
Shostakovich, Dvořák, Schubert, Rachmaninoff, Prokofiev, Tchaikovsky, Mahler — **sourced from
YouTube.**
and the transcription model."* No licence terms are stated anywhere on the page.
The correct narrowing: Magenta RT 2 is the only orchestral-capable neural realiser with a
VERIFIED clean licence. Symphony Rendering is a second orchestral-capable candidate with released
weights whose licence is UNSTATED — an outstanding read, not a refusal. And its twelve-Western-
composer corpus with a composer-label style control is a Western-idiom prior sitting one layer below
a world-idiom score, which is the care exposure §6 names, appearing here in its least visible form.
Verified: GPLv3; hosts VST instruments and effects; *"Rendering and saving multiple processors
simultaneously"* for faster-than-realtime batch from Python; **"MIDI playback in absolute time and
PPQN time" and "Parameter automation at audio-rate and at pulses-per-quarter-note"** — which is
precisely the expression layer §5.1 requires; Windows x86_64 with 64-bit Python 3.11-3.14.
GPLv3 binds redistribution of the SOFTWARE; we redistribute rendered audio, which is the same
reasoning sfz_palette.py already records for sfizz's BSD-2 notice condition. **Face-value: clean
as a build-time tool.**
from the repository: *"DawDreamer must be imported before JAX or other LLVM-based libraries"* to
avoid segmentation faults from LLVM initialisation conflicts. **Option A's render host and Option
B's neural realiser are therefore an import-order or separate-process problem by construction.**
It is a half-hour of engineering discipline, and it is the kind of thing that eats an evening when
it is discovered at 2 a.m. instead of read in a brief.
independently confirms it (*"NO — it is a MATERIAL source, not a realiser"*, conditioning is
period-composer-instrumentation only). Both docs agree. It writes its own tunes and is out under any
reading of composed-never-prompted.
— CC-BY-NC on the weights, and JASCO is the best melody-conditioning surface ever built for exactly
our problem. The open audio field's quality frontier and its usable frontier are not the same
frontier.
init_noise_levelranges 0.0-1.0, default 1.0, at which *"the init audio is fully replaced by noise and has no effect
(pure generation)"*; 0.1 produces close variations, 0.5 a halfway blend. Multi-region inpainting
is real — lists of (start, end) pairs, *"everything else is preserved."* Both models are already
installed and licence-pinned on this box.
---
STEPBACK_OPEN_STACK.md §0 states the composed-never-prompted ruling as an engineering requirement,
and it is the right frame: **a model that cannot be handed a specific sequence of specific pitches at
specific times is not a candidate at any quality level, because adopting it means Josh stops owning
the themes.** This table is the first filter every option below passes through.
| Component | Can it be GIVEN Josh's theme? | The mechanism | Fidelity | Licence for a shipped game | Verified this pass |
|---|---|---|---|---|---|
| Composer's Assistant 2 | YES — the cleanest instance found | empty MIDI items are targets; unselected items are fixed encoder context. The tune is not the model's to change, enforced by which items are empty rather than by a policy | EXACT — the theme is never regenerated | MIT code; weights trained only on PD and permissively-licensed MIDI | YES |
| Anticipatory Music Transformer | YES | extract_instruments() + supplied control events; generate() writes accompaniment around them | EXACT | Apache-2.0 code; weights licence unstated; Lakh MIDI provenance is an open clearance question | YES |
| Magenta RealTime 2 | YES — the only AUDIO model with a symbolic input port | 128-dim multihot MIDI pitch state per frame, 25 Hz | HIGH by construction; the 3 = model decides state is latitude we set per note | Apache-2.0 + CC-BY-4.0; "Google claims no rights in outputs" | YES |
| Symphony Rendering | YES | time-aligned piano-roll features at 100 Hz + composer embedding | HIGH | UNSTATED — weights released, no licence terms on the page | YES (that it is unstated) |
| Stable Audio 3 | INDIRECTLY — audio-to-audio | init_audio + swept init_noise_level; multi-region inpainting | TUNABLE — a band must be found by sweep, not guessed | Stability Community License, free under $1M revenue; Section III registration triggered and unsatisfied; Gemma Terms second chain | YES (semantics) |
| ACE-Step 1.5 | INDIRECTLY — audio-to-audio | Reference Audio, Cover, Repaint & Edit, Vocal2BGM. No MIDI input anywhere | TUNABLE | MIT | carried |
| BBC SO Discover via a host | YES — it is a sampler; it plays what it is given | MIDI + CC automation | EXACT | Grant fires on a face-value read; Section 10 forbids AI training on it | YES |
| NotaGen | NO | period-composer-instrumentation prompts only | — | MIT, but the constraint that binds it is composed-never-prompted, not the licence | carried |
| MusicGen / JASCO / MIDI-GPT / Khala | yes, and irrelevant | chromagram / rich control surfaces | — | REFUSED — CC-BY-NC-4.0 weights | carried |
The finding the table produces: every component in the recommended architecture sits in the top
half of it. Nothing in this brief asks a model to invent a theme.
---
All four preserve, without exception, everything in STEPBACK_POSTMORTEM.md §8.1: the leitmotif
architecture and theme registry, the care mechanism with its declared margins, the exemplar corpus and
its per-class bands as admission floors, the random-melody-null method, the mix and loudness policy,
the licence discipline, and the record-integrity method.
A hardware fact that applies to all four, verified in the open-stack leg: VRAM is no longer the
constraint. A symbolic model, an audio model and a sampler run concurrently on one 5090 with room
left. No option below is forced to choose one family, and no option needs a second box.
---
What it is. Josh's leitmotif stays locked and authored. A music-trained model writes the material
AROUND it that four rounds have not produced. A professional orchestral library, driven by a real
expression model, plays the result. A selection layer publishes survivors rather than the only
artifact made. Four stages, each standing alone and each provable separately.
sfizz/CC0 chain, unchanged. This is STEPBACK_POSTMORTEM.md §8.3 item 4 and it has never been run.
For four rounds we have never separated "the writing is bad" from "the players are bad."
orchestrate.py emits velocity shaping,CC1 and CC11 curves and articulation selection per phrase; then DawDreamer hosts BBC SO Discover
and plays it. Both rungs, and the order is load-bearing — see §5.1.
slot. Rank. Publish the top few. Josh hears three to five, each carrying its scores.
inner voices, accompaniment and per-section variation around it; our grammar, harmony rules and care
mechanism filter, develop and gate what returns. **AMT is the fallback if CA2's classical-choral-folk
prior proves too narrow, at the cost of the Lakh provenance question.**
Pros against the binding rulings.
| Ruling | How A stands |
|---|---|
| Melody-first, hummable thirty-year leitmotifs | The theme is locked material no model may overwrite. The piano-reduction thirty-year test stays as an accept gate on the SCORE, where it belongs, because a score survives every stage |
| Composed, never prompted | Mechanically enforced, not policed. Unselected MIDI items are encoder context; the model cannot write into them. This is a stronger guarantee than a rule anyone has to remember, and stronger than any conditioning-discipline promise Option B can make |
| Hummable | Unchanged — the hummable line is Josh's and is never regenerated. What changes is what surrounds it |
| Local-first on the 5090 | CA2 runs on CPU or a modest GPU and states it sends nothing over the internet. DawDreamer is local Python. The ranking models are small. The whole stack co-resides with headroom |
| No human composer hire | Nobody is hired. The pass ladder is code and local models — the factory's own ladder, exactly as ruled |
| Proof before spend | Stages 0, 1 and 2 cost nothing. The free host (DawDreamer) and the free library (Discover) make the entire rendering upgrade a zero-capex experiment, which revision 1 could not say because it assumed REAPER |
| The care line (CVD §17.1) | Allow and deny registers stay enforceable because a score survives every stage. The exposure is real and it is §6 |
Cons, stated without softening.
risk that §6 names as the strongest objection.
purchase** — it becomes both CA2's host and, conveniently, a second render path. Stages 0-2 keep the
spend at zero; stage 3 is where money enters, and by then three free proofs will have earned it.
name.
Cost.
are small. DawDreamer uses none. Everything fits simultaneously.
the weights licence line in the repo). DawDreamer GPLv3 as a build-time tool. BBC SO Discover:
grant fires on a face-value read; Section 10 forbids training AI on it.** ReaRender MIT if REAPER
enters.
verified this pass**). BBCSO Core is the optional paid library step-up; SWAM the optional modelling
path for exposed solo lines (neither price verified this pass).
Risk. Stage 1's null is the near-term risk: if the expression rung is skipped, a world-class
library plays a dead performance and we conclude the library did not help — a wrong conclusion bought
at the cost of a wave. STEPBACK_OPEN_STACK.md §8.2 makes that prediction on the record and this
brief adopts it. Stage 3's risk is research-shaped and care-shaped. Throughput is the unpriced number:
Josh's floor is 20-plus tracks per node across 76 region rows, on the order of 1,500 tracks, and
SMART's published bottleneck is render cost per candidate, so N-candidates × render-time is what
decides whether stage 2 survives contact with the real corpus.
FIRST PROOF — one wave, zero spend, zero purchase. Five arms on one cue slot, all rendered, all
scored by timbre_critic.py blind plus the existing battery as a floor:
1. The staged round-4 master, as control.
2. A known-strong public-domain melody through our current CC0 chain — the arm that has never
existed, and the only one that separates the writing from the players.
3. Round-4 ROAD, current library, plus the expression layer — isolating rung one alone.
4. Round-4 ROAD, expression layer plus BBC SO Discover through DawDreamer — isolating rung two.
5. Best-of-forty from the existing planner, through whichever renderer arms 3 and 4 rank higher.
Josh hears five things where he has heard one, four rounds running. **Every arm is a single-variable
change from the arm above it**, which is what revision 1's proof did not guarantee.
---
What it is. The sampler comes out of the chain entirely, or sits in it only as a scaffold. Our
composed note stream drives a model trained on recorded performance, and the model produces the audio.
Three arms, in descending order of licence cleanliness:
orchestrate.py's note stream becomes the 128-dim multihot pitch stateper frame; style arrives as 12 MusicCoCa tokens from a text prompt or from audio we own; our
loop_seams.py owns long-form continuity because the model's effective receptive field is 20
seconds. The only arm whose licence is verified clean end to end.
composer embedding for style. 150M params. **Licence unstated — a gating read before a byte is
published.**
Stable Audio 3's init_audio at a swept init_noise_level, or ACE-Step's Repaint. Structure
survives because it is physically present in the input. **Zero install — both are already on this
box, pinned.**
Pros against the binding rulings. Melody-first is untouched: the notes do not change, they are
realised. Composed-never-prompted holds under B1 and B2 mechanically, since the conditioning is
symbolic; **under B3 it holds only if the conditioning stays our audio or our MIDI and never a text
description of a style, which is a line to write into the lane contract rather than trust.** Local-first
holds on all three — B1 is ~5 GB, B2 is 150M params, B3's models are installed. No hire. Zero money.
And B1's licence is genuinely the best available: no revenue threshold, no registration, no second
chain, no rights claimed in outputs.
Cons, and they are structural rather than incidental.
the lane census and a meaningful part of the measurement rig operate on lanes. Either we lose those
instruments or we render per-lane and re-mix, which is untested and may not sum coherently.
sounding like production library music — the precise aesthetic he has rejected four times.
forthcoming technical report."* We would be adopting on architecture, not on evidence.
world-idiom score — the care exposure at its least visible depth. And it is trained on
YouTube-sourced recordings, which is a provenance question the authors deserve credit for publishing
and which is still a question.
the concatenative stage bounds everything downstream — poor samples directly degrade the initial
synthesis and limit what refinement can recover. Our sampler is the weakest in the survey.
Refining on top of the known-worst acoustic prior produces a result that cannot be read as an upper
bound on anything.
silence-budget dimensions our cards measure.
Cost. Zero money on all three arms. VRAM: B1 ~5 GB, B2 trivial, B3 1.69-6.52 GB for SA3 or ≥24 GB
for ACE-Step XL-SFT. Licences: B1 verified clean; B2 unstated and gating; B3's SA3 carries the live
unsatisfied Section III registration obligation, ACE-Step carries none.
Risk. Moderate to high, concentrated in the unknowns rather than in the engineering. B1's quality
is genuinely unknown. B2's licence is genuinely unknown. B3's usable band may not exist.
FIRST PROOF — one wave, zero spend. Two nights, run in parallel and independent of everything
else. Night one: install JAX on sm_120, render ONE 30-second leitmotif statement through Magenta
RT 2 base with a text style prompt, score it blind on timbre_critic.py plus a motif-preservation
check against the same statement rendered through Option A's stage-1 chain. Night two, zero install:
sweep SA3 init_noise_level across 0.1 / 0.2 / 0.35 / 0.5 on the round-4 Flores render under one
caption, and plot melody preservation against timbre-critic distance. **The output is a curve, and the
curve either has a usable band or it eliminates a family.** Both are cheap enough that their results
are worth having regardless of which option wins.
---
What it is. Option A's stage 1, standing alone. Keep the 88,000-line composition engine exactly as
it is. Change only the sound: emit real expression from orchestrate.py, then host BBC SO Discover in
DawDreamer and render offline from Python. No models anywhere. No selection. No new note comes from
anything but our own code.
Pros against the binding rulings. Every ruling holds trivially and unarguably, because nothing
about how the notes are chosen changes. It is the option with the least distance from where we stand,
the fewest new questions, and no fork for Josh to rule on. It has zero research risk. It is the only
option whose licence surface is now fully read. And it answers, precisely and directly, the most
concrete sentence in the round-4 grade.
Cons. It answers one of Josh's four clauses. The melody, the slowness and the absence of
movement all live upstream, and STEPBACK_OPEN_STACK.md §9 states the consequence as an on-record
prediction: *"If round 5 ships better instruments playing the same melody, it will be graded down
again."* Two further cons that must not hide behind a general upgrade:
change from community samples, not a ceiling.
instruments in 12-TET, and a baroque recorder standing in for a bamboo flute, are **not fixed by an
orchestral library.** The non-Western palette is a separate sourcing decision the care doctrine makes
non-optional, and an orchestral upgrade would let it hide.
Cost. VRAM zero — it is a CPU sampler chain. Money zero: DawDreamer GPLv3, Discover free, no
REAPER purchase needed. Licences: the Spitfire grant fires on a face-value read, Section 10 forbids
training on it, GPLv3 binds software redistribution and we redistribute audio. **The one integration
unknown to size early rather than assume is whether DawDreamer hosts an authorised commercial Spitfire
plugin headlessly**; if it does not, REAPER CLI is the fallback and it costs money — which is exactly
the spend proof-before-spend permits, because C's own proof would have earned it.
Risk. Low, and the low risk is the point. The one real risk is the ordering: **skip the expression
rung and this option produces a null result and a wrong conclusion.**
FIRST PROOF — one wave, zero spend. The cleanest single-variable ladder available to this lane,
and it is three arms rather than two so that no arm changes two things at once: the round-4 ROAD score
rendered (i) through the current CC0 sfizz palette exactly as staged, (ii) through the same palette
with the expression layer added, (iii) through Discover in DawDreamer with the same expression layer.
Blind timbre critic on all three. **If the critic does not move between (i) and (ii), the expression
hypothesis is wrong. If it moves between (ii) and (iii) but not between (i) and (ii), the library is
the whole story. We learn which in a night either way.**
---
What it is. Option A's stage 2, standing alone, and the pattern the commercial-engines survey
found in every product that satisfies listeners. Change no generator and no renderer. Produce a large
candidate pool per cue slot, score every candidate, publish only survivors. Generator-agnostic, so it
multiplies whatever composer wins later.
What makes it nearly free. The machinery is already written and bypassed. emit_review_picks.py
carries a generic ranked path that sorts by rank_in_theme and counts kept_in_theme; the three round
emitters hardcode both to 1 because the composer produced one cue. **The field exists and the pool
does not.** Every stochastic choice in the current planner — seeds, form templates, development
operation sequences, tempi, orchestration assignments — is already a parameter the engine commits to
one draw of.
The scoring stack. Audiobox Aesthetics on rendered audio for four axes including the closest
published proxy for "does a person like this"; CLaMP 3 for cross-modal similarity against Josh's
approved corpus, trained across global traditions, which matters for a 79-chapter world score; our own
battery as the craft floor and never as the verdict; and timbre_critic.py, which we own, which is
anchored on 1,182 measured cards of real released game music, and **which no battery or realiser
currently imports.**
The calibration precondition that makes it honest rather than another self-graded battery. We hold
a graded corpus: eighteen round-1 tracks, two round-2 cues, three round-3 and three round-4, all
turned down, with verbatim reasons. **Any proposed ranker must reproduce Josh's ordering of the rounds
he has already graded — in particular it must rank round 2 above round 4, the ordering that broke our
current battery.** A ranker that cannot recover a known verdict is not qualified to select an unknown
one.
Pros against the binding rulings. Every ruling holds unchanged. No new generator, no new corpus,
no capex, no provenance question, no research risk, no fork for Josh to rule. **It is the only option
in the brief with none of those.** And the open-stack leg makes it cheaper still: SA3 and ACE-Step are
already installed, so the pool can be diversified at the audio tier as well as the plan tier.
Cons. Selection cannot exceed the generator's ceiling. If all forty candidates share the same
two-idea form, the same slow melodic surface and the same CC0 instruments, the best of forty is still
a rejected cue and Josh will say so in one sentence. And optimising against an automatic aesthetic
score produces music that scores well and sounds worse — the ranker becomes a target and stops being a
measure. Both are why the human stays at the end of the loop and why the calibration check is a
precondition rather than a nicety.
Cost. Compute we already own. The real price is render time per candidate, which is SMART's
published bottleneck and the number that must be measured in the proof rather than assumed — because
it is also the number that decides whether this layer survives 1,500 tracks.
FIRST PROOF — one wave, zero spend. Forty candidates for one slot from the existing planner,
ranked, top three staged against the round-4 master, all through the current renderer so selection is
the only variable. Run the calibration check before anything is staged, and report render-seconds
per candidate as a first-class number.
---
| A — hybrid: locked theme + infill + pro rendering + selection | B — neural audio steered by the composed melody | C — rendering upgrade only | D — selection engine only | |
|---|---|---|---|---|
| Layers it fixes | generation, rendering, curation | rendering | rendering | curation |
| Clauses it answers | all four, staged | instruments, and possibly liveness | instruments | movement and melody, partially and indirectly |
| Melody conditioning | EXACT — theme is encoder context, never a target | HIGH (B1/B2, frame-wise symbolic) / TUNABLE (B3, audio-to-audio) | EXACT — a sampler plays what it is given | untouched |
| Composed-never-prompted | mechanically enforced by which MIDI items are empty | mechanical under B1/B2; B3 needs a written conditioning rule | untouched | untouched |
| Keeps stems and the measurement rig | yes | no (B1/B2 emit one stream; B3 keeps pre-refinement stems as fallback) | yes | yes |
| Care exposure | CA2's classical-choral-folk prior at stage 3; non-Western palette unsolved | B1's 71k h stock prior; B2's 12-Western-composer prior, one layer down | palette substitutions persist and could hide | none |
| Provenance | best in survey — PD and permissive MIDI only, MIT code | B1 Apache + CC-BY, outputs unclaimed; B2 YouTube-sourced, licence unstated | Spitfire EULA now read; Section 10 forbids AI training | none |
| VRAM on the 5090 | comfortable | comfortable | zero | comfortable |
| Money before proof | zero | zero | zero | zero |
| Research risk | real, at stage 3 only | moderate to high, concentrated in unknowns | low | none |
| Needs Josh's §7 ruling | yes, for stage 3 only | yes for B1/B2 (a model is in the note-to-audio path) | no | no |
| Proof cost | one wave, five arms, compound | two nights, parallel | one night, three arms | one wave |
---
**Adopt Option A, run in proof-cost order, with stages 0-1, 2 and 3's proof running CONCURRENTLY in
one wave rather than in series. Run Option B's two cheap probes in parallel as an independent track.
Do not start at stage 3's adoption, and do not ship a renderer-only round 5.**
The reasoning is the ordering and the concurrency, not the ambition.
sub-systems, and running them first is free, fast, and the only way to learn whether stage 3 is
needed at all. Three of the four commercial architectures surveyed get their quality from a
selection loop and a recorded sound source. We have neither.
instruments, and has never heard the best of forty. Proof-before-spend applies to our own diagnosis,
not only to purchases — the expensive hypothesis is the one that needs the proof.
STEPBACK_OPEN_STACK.md §9closes with the objection that a renderer-only round 5 gets graded down again, and its own
recommendation puts AMT *"in the round-5 architecture alongside whichever renderer wins, not after
it."* STEPBACK_POSTMORTEM.md §6 records that "never any harmonies added on" identified a
missing architectural layer — not thin: absent. A symbolic infill model is the direct answer to
that clause and it is small, local and cheap to prove. **Holding it back one more wave to preserve a
clean experimental ordering would be buying methodological tidiness with the thing Josh actually
asked for.** The stage-3 PROOF runs in wave one; stage-3 ADOPTION still waits for the stage-1 and
stage-2 results.
NotaGen would write its own tunes from a period-composer prompt. AMT does the right thing with a
Lakh-derived provenance question attached. CA2 does the right thing, in the DAW that is already the
rendering host, with 34 named controls and the cleanest training provenance of any generative
component in any of the three step-back docs.
purchase. DawDreamer is GPLv3, hosts VST3, renders offline in batch from Python, and does MIDI at
PPQN with audio-rate parameter automation — which is the expression layer itself. **The rendering
upgrade is now a zero-capex experiment.** REAPER is bought if and when CA2 is adopted, or if
DawDreamer will not host the Spitfire plugin headlessly.
timbre_critic.py into the battery before round 5 stages anything. We own an instrumentbuilt precisely to catch a fake-sounding render, and no battery or realiser imports it. An ear must
enter the loop before Josh's — it is STEPBACK_POSTMORTEM.md §8.3 requirement 3 and it costs an
import.
WAV four times. In shipped practice a three-minute cue is a set of stems and segments that recombine
for as long as the player stays. Grading linear WAVs will keep producing the no-movement complaint
no matter how good the cue gets.
It is fast, free, licence-unconditional and good enough to iterate against. It is not what ships.
| Named defect in the current chain, from our own honest-limits list | What the recommendation changes it to |
|---|---|
| "One dynamic layer a note" | This is a property of OUR REALISER, not of the library, and it is rung one. VSCO 2 CE ships multiple velocity layers on many instruments and we play one. orchestrate.py starts emitting velocity shaping, CC1 and CC11 curves and per-phrase articulation selection, driven into the host at PPQN and audio rate. Zero cost, zero VRAM, no new licence — and §5.1's prediction is that everything else in the rendering upgrade depends on it |
| Free CC0 community sample sets | BBC SO Discover — a professional orchestra recorded at Air Studios, free, and the EULA grant now verified to fire for a mixed orchestral cue. A categorical step, not an incremental one. Discover is a single mix signal, so this is a class change and not a ceiling; Core is the paid step-up if the free arm proves the direction |
| Round robins on almost nothing, producing the machine-gun artifact | Recorded articulations and repetition handling that exist in the library or do not. An asset property, not an effort property — no amount of further code closes it |
| Scripted rather than recorded transitions; no true legato | Recorded legato transitions between specific note pairs — the single most-cited realism factor in the mockup literature |
| CC1 measured as a no-op on our palette; zero continuous controller data reaching the sampler | The controller work becomes meaningful because there are layers to crossfade between. SWAM physical modelling stays the option for exposed solo winds, brass and strings, where the honest trade is that modelling puts the life on the driver rather than on the take |
| One global synthetic diffuse field for a mix | A real recorded hall in the samples, with our mix and loudness policy unchanged on top and the synthetic IR retained so no third-party IR licence can ever be withdrawn from under a shipped game |
| The tuned-bronze half of a gong-waning ensemble played by concert instruments in 12-TET; a baroque recorder standing in for a bamboo flute | NOT fixed by an orchestral library, and named as its own work item so it cannot hide behind a general upgrade. Discover has no gamelan. The non-Western palette is a separate sourcing decision that the care doctrine makes non-optional |
| Everything above is a SAMPLER answer | And the parallel track tests a different answer entirely. Magenta RT 2 realises our notes from a model trained on recorded performance, under the cleanest licence in the survey. If a sampled orchestra still reads as a mockup, that is the arm that would not |
This complaint has four distinct mechanisms. Answering one leaves it standing, and revision 1's
apportionment survives the open-stack leg unchanged. **The second row is the one to read: it is the
defect we were scoring as a pass.**
| Mechanism | The evidence | What the recommendation changes |
|---|---|---|
| Composition — two distinct musical ideas in a 223-second cue | round-4 measurement; adding notes to two ideas does not make three. STEPBACK_POSTMORTEM.md §6 records "never any harmonies added on" as a missing architectural layer, absent rather than thin | Stage 3's whole task: multi-track infilling that writes ideas three, four and five as counter-lines, inner voices and re-orchestrations around the locked theme. CA2's DNOC token exists specifically to prevent octave-shifted copies of what is already there — the exact failure mode a naive layer-adder would produce. Its proof runs in wave one, not after |
| The objective function rewarded stasis | measured boundaries on the ROAD lineage fell 22 → 10 → 8; the longest hold on ROUNDS went from none → 34.92 s → 53.50 s at the tier we published; dynamic_arc scored it a pass every round. PASS4_VERDICT.md §1a documents a "hold spine" being ADDED to make the row pass, and §3 documents the pre-render gate being rewritten when the spine tripped it | hold_share is retired as an optimisation target and kept only as a detector. And Round 1's own floor — *"I dont like repetitious loops lasting more than 3-5 cycles without introducing a new instrument or catch or melody or beat drop"* — becomes an armed rule with a must-fire control. It is his own sentence, it has been on the record since round 1, and we never encoded it. instr_dynamic_arc encoded half of his round-2 sentence and none of "that still never lose the listeners," and was then optimised against faithfully all the way past what he meant |
| No selection | emit_review_picks.py hardcodes rank 1 of 1 on all three rounds; every round shipped the only artifact the generator made | Stage 2 samples forty draws across form templates, operation sequences, tempi and orchestration. Variance between takes is where movement is found rather than engineered, and a pool is where a two-idea cue loses to a four-idea one |
| Delivery — he has been graded on a linear WAV, four times | shipped scores manufacture a large share of perceived movement at runtime through vertical layering and horizontal resequencing | Stems and segments, built regardless of which generator wins. Part of this complaint is about a layer that was never in the artifact he was given — which does not excuse the cue, and does mean some of the complaint is cheap to answer |
slow, for two measured reasons: the axis pools eleven melodic layers, so a busy accompaniment reads
as a fast tune; and the development grammar's augmentation operation literally slows the theme and
is scored as a spine-changing operation. **A cue can be dense and its tune still crawl, and no axis
we own distinguishes those.** Stage 3 replaces the guess with a dial — horizontal onset density in
six bins from half notes to 4.5-plus onsets per quarter note, set per track. Before stage 3, stage 2
answers it by sampling tempo and surface density across forty candidates instead of committing to
one draw.
claiming otherwise would be the same mistake made four times.** It claims that those stages are the
experiment that finally tells us whether the defect is in the writing or in the delivery of the
writing, and that stage 3's concurrent proof means round 5 is not a renderer-only round. **If the
best of forty through a real orchestra with real expression still draws that sentence, the writing
is convicted on evidence rather than on argument** — and §7 becomes the only lever left.
---
**Every option in this brief locks Josh's theme by construction — and the theme may be the defect.
"The melody sucks" names the one artifact the entire architecture is built to protect, and
protecting it harder is the one move that cannot help.**
This is sharper than revision 1's care-idiom objection, and it is the objection that should be read
first, because it is the one that could invalidate the whole option space rather than one stage of it.
The argument runs as follows, and each step is on the record:
STEPBACK_POSTMORTEM.md §3 establishes that the melodic source of rounds 3 and 4 is three Python literals at pass3_flores.py:114-144 — six notes, five notes, seven notes — and that
theme_compositions.py holds the thirteen canonical theme heads the same way. **The tune Josh has
rejected three consecutive times is one authoring act, performed once, in a code editor.**
ours. They are our composition, not his — which means "Josh owns the themes" is currently true
as an authority claim and false as an authorship fact. **Composed-never-prompted is protecting our
typed integers, not his melodies.**
and samplers to surround it better. **If the head cells are the defect, all four options are scoped
to miss it**, and the fifth turndown will read exactly like the fourth.
thematic material, NotaGen-class, which the survey found is the strongest classical-craft symbolic
model available — is excluded by composed-never-prompted under a strict reading.**
STEPBACK_OPEN_STACK.md §8.5 names the cost of that exclusion and declines to absorb it silently.
This brief agrees: the cost is real and it is possibly the whole problem.
Why the recommendation survives the objection anyway, stated so it can be attacked:
under every reading, because a theme cannot be judged through a broken window: no one can say
whether the ROAD cell is a bad tune when it has only ever been heard as one dynamic layer per note
on community samples with no expression data. Stage 0 exists precisely to answer this — a
known-strong public-domain melody through our own chain tells us how much of "the melody sucks" our
own renderer manufactures.
by SELECTION is authorship: Josh picks the head cell from N machine-proposed candidates, exactly as
he would pick from N candidates a hired composer offered — except no one is hired, which keeps the
2026-08-08 ruling intact. That is a different act from prompt-and-pray, and it is the one lever that
reaches the clause he put first.
the melodic source stays authored, the brief's stages still run, and the program accepts that its
ceiling on melodic quality is whatever a code model typing integers can reach. That is a legitimate
choice — it is his product — but it should be made knowingly, not inherited from a rule written for
a different failure.
**A Western-corpus prior asked to develop a non-Western theme produces a plausible Western
approximation of it, and that is a care defect, not a quality defect — which means our instruments
will not catch it and Josh may not either.**
Verified this pass: CA2's own paper states its models are *"most proficient at infilling classical,
choral, and folk music."* Symphony Rendering's corpus is twelve Western composers. Magenta RT 2's is
71,000 hours of stock music. **All three candidate priors are the same class of exposure at different
depths**, and the deepest one is the hardest to see: unlike "the melody sucks," which Josh names in
one sentence, an idiom defect passes as competent to a listener who is not from that tradition. The
palette substitutions in §6.1 are the same class of problem already live in the rendering layer today.
Why it does not overturn the recommendation:
a further argument for the ordering rather than against it.
and deny registers, which read off the region page's own Section 10, stay enforceable because a
score survives; and the care mechanism's declared margins gate what returns.
must be run on a non-Western cue and the result read by the care mechanism with its margins
published. If it fails there, the honest outcome is that the learned layer is scoped to
Western-idiom chapters while world-idiom chapters keep an authored melodic source — a real and
acceptable answer, not a failure of the program.
"You are recommending four stages plus a parallel track, and calling it one architecture." Partly
true, and it is the real cost. The mitigations: every stage stands alone and is separately vetoable;
stages 0, 1 and 2 are nights rather than weeks; the Option B probes are two nights and independent of
everything; and each puts something new in front of Josh. **What makes the sequencing defensible
rather than evasive is that no amount of composition work is interpretable until the window is clean —
four rounds of composition work have already been graded through a broken window and we cannot read
the results.**
True, and it is the strongest thing anyone can say against any learned-model direction. **The answer
is that the arrangements are not comparable.** Round 1 used a learned model as a one-shot full-mix
generator with a prose caption as its only control, and the measured failure was the control surface,
not the musicality — roughly half the card dimensions did not respond to the generator's conditioning.
Nobody has tested a learned model as a constrained infiller around locked material inside our
care-enforcing architecture. The objection correctly refutes "learned models are automatically
better." It does not support "typed pitch literals are the right melodic source," which has now been
graded down three consecutive times.
---
STEPBACK_POSTMORTEM.md §8.3 ruled that the comparator lane must put this to him as a full brief
before any round-5 architecture is chosen. This is that section, and it is the one place in this
document where the recommendation is advisory rather than delegated.
The question. composed-never-prompted is Josh's standing ruling, and the AST-verified generator
firewall at pass2_realise.py:1205-1224 is that ruling in code. **Does it ban prompt-and-pray full-mix
generation, or does it ban any learned model touching the notes?**
Reading 1 — it bans prompt-and-pray. The ruling was made in response to round 1, which was
caption-in, mix-out, with nothing controlling what entered when. Under this reading a model that
writes an inner voice around a theme Josh owns, in measures he designated, under density and register
controls he set, is composition — he is the composer and the model is the section-writer, which is
what an orchestrator has always been in the chain the AAA industry actually uses.
Reading 2 — it bans any model touching the notes. Under this reading, authorship means every pitch
traces to a human decision or to a rule a human can read, and a learned prior anywhere in the note path
is the thing the firewall exists to prevent, regardless of how it is conditioned.
What the CA2 arrangement does to the fork. It narrows it considerably. Josh's theme is encoder
context, never a target: the model cannot alter a note of it, and that is enforced by which MIDI items
are empty rather than by a policy anyone has to remember. So the fork is not "may a model write Josh's
melodies." It is **"may a model write the counter-lines and accompaniment around Josh's locked
melodies, under his density, register and idiom controls."**
§7 establishes that the locked themes are typed literals authored by a code model under Josh's own
delegation, not melodies Josh wrote. That makes a second question unavoidable, and it is the one that
reaches the clause he put first:
**May a music-trained model PROPOSE thematic candidates that Josh SELECTS from — authorship by
selection rather than authorship by typing?**
occurs; and the selected head cell then flows through the identical locked-theme architecture every
option in this brief describes. It is exactly the posture Suno, AIVA and Soundraw users occupy, and
the posture STEPBACK_COMMERCIAL_ENGINES.md §7.2 identifies as the largest gap-to-cost ratio in the
entire survey. **It is also the only proposal in this document that touches "the melody sucks"
directly.**
not feel like writing, and a theme he picked but did not author may not carry thirty years. Only he
can say. And the strict reading has a real principle behind it: a leitmotif that must survive three
decades arguably should not begin life as a sample from a distribution.
Recommendation on both forks: Reading 1 on the first, narrowed to the sentence above and written
into the lane contract as a hard line — conditioning is symbolic and structural only, never a text
description of a style, and the theme tracks are never targets. On the second fork, **a bounded
experiment rather than a policy**: propose theme candidates for ONE cue, put them in front of him
beside the current head cell, and let the comparison rule rather than the argument. **Strongest
objection to that recommendation:** the distinction between "the model developed my theme" and "the
model wrote most of what you hear" is a matter of degree, and in a dense cue the infilled material is
the majority of the notes. A technically-locked melody inside a model-written arrangement may still
not feel like his.
**Stages 0, 1 and 2 need no ruling on either fork — they introduce no model into the note path at
all, and they should start regardless.** Stage 3's proof and Option B's probes need the first ruling.
The theme-candidate arm needs the second.
---
layer is inference from measurements plus his words. Where they disagree — his "how slow the melody
is" against an IN-BAND density reading — his words are treated as ground truth and the number as the
suspect. That is the correct ordering and it is also an assumption.
about Composer's Assistant 2, Magenta RealTime 2, Symphony Rendering, BBC SO Discover, DawDreamer,
the Anticipatory Music Transformer, Stable Audio 3, ACE-Step, Audiobox Aesthetics, CLaMP 3,
ReaRender or SWAM is the paper's or the vendor's claim, **verified as a claim, not as a
measurement.** That is what the proofs in §4 are for.
disk under docs/licence_records/ with commit-SHA pinning, and nothing in §4 enters a pipeline
before it goes through that door. The Spitfire EULA was READ this pass and must now be PINNED;
CA2's weights-licence line, Symphony Rendering's licence and AMT's weights licence are three
outstanding reads, and the AMT one carries the Lakh provenance question behind it.
them:** STEPBACK_OPEN_STACK.md §3.1's "exactly one symbolic model" finding (CA2 is a second and
better one); its §5.3/§5.4 omission of Symphony Rendering and the resulting overstatement of "one
model wide"; its VERIFIED-tier claim that Magenta RT 2 emits a mixdown (the model card does not say);
and the EULA URL both step-back docs cite, which served a privacy policy on fetch this pass.
sibling documents that this pass did not re-read. No spend decision should quote them without a
fresh read.
paper, ACE-Step and AIVA are carried, not re-fetched.** They were verified in revision 1 and are
cited here at that tier.
this product, he has been consistent across six recorded verdicts, and three of his four round grades
located a defect our own instruments later confirmed in our own files. Treating that record as
reliable is a judgement, and it is the judgement this brief makes.
---
DERIVED FROM:
docs/review_candidates.json — PASS7_FLORES_ROAD.grade_verbatim, read at HEAD and quoted verbatimin §0; the round-1 through round-4 takedown notes.
docs/proposals/music/STEPBACK_POSTMORTEM.md §2, §3, §5, §6, §8.1, §8.2, §8.3, §9, §10 — theanti-correlation and its three self-corrections, the typed pitch literals, the four structural
limits, the grade apportionment, the survives and retires lists, the six requirements on any round-5
architecture, and the fork carried into §8 here.
docs/proposals/music/STEPBACK_COMMERCIAL_ENGINES.md §2, §3, §4, §5, §6, §7, §9 — the four-layerframe, the AIVA and Suno reads, the delivery-layer practice, the rule-interaction literature, and the
three patterns.
docs/proposals/music/STEPBACK_OPEN_STACK.md §0, §2, §3, §4, §5, §6, §7, §8, §9 — **the legrevision 1 declared missing.** Its melody-conditioning frame, its licence board, its render ladder,
its seven conditioning primitives, its expression-rung prediction, and its three candidate pipelines.
Four corrections to it are recorded in §2 and §9.
docs/pipeline_review/tech_research/PIPE_MUSIC_MODELS_2026-08-06.md — the owning PIPE dossier thatbinds this lane.
harness/music_gen/emit_review_picks.py, the timbre_critic import census, sfz_palette.py, sa3_generate.py, fetch_sfz_stack.py — read at HEAD.
music-direction-melody-first-30-year-bar (melody-first, thirty-year bar,human composer removed permanently, "composer_in_loop" = the factory's own pass ladder); memory
factory-first-resolution-mandate (proof before spend); memory 5090-box-live-remote-stack
(local-first); memory license-reads-not-conservative-solo-dev-tooling (face-value licence reads,
flag only an explicitly triggered prohibition — §2.1 flags Section 10 under exactly that rule);
memory pipe-dossiers-bind-generation-lanes; CLAUDE.md content bar and the CVD §17 floor with the
§17.1 care doctrine; the standing decision protocol (options, pros and cons, recommendation,
strongest objection).
NOT DERIVED (authored judgment, and why it had no canon home):
this lane's synthesis of the three step-back docs plus §2's verification.
stages 1 and 2. It is an inference from the two sibling docs' own on-record prediction that a
renderer-only round 5 gets graded down, weighed against the proof-before-spend ordering. Revision 1
ruled the other way; this revision names the change so it can be vetoed back.
ACE-Step. Grounded in the verified licence and conditioning reads at the point of use.
reasoning that the locked head cells are our authorship rather than Josh's. It follows from
STEPBACK_POSTMORTEM.md §3 plus his own 2026-08-05 delegation, but no document draws the
conclusion, and the conclusion is uncomfortable enough that it is named as authored judgment.
STEPBACK_OPEN_STACK.md §8.5names the cost of excluding NotaGen-class proposal and declines to decide; this brief converts that
into a bounded, ruleable question rather than leaving it as a cost. It is advisory only and
explicitly Josh's.
records the clause; nothing in canon splits it. Each row cites the measurement it rests on.
dissimilar" grant**, and that Section 10 reaches training but not inference-time conditioning. Both
are readings of verbatim text under the standing licence-read memory, and both are the director's to
veto because Option A stage 1 and Option C depend on the first one.