music/STEPBACK_COMMERCIAL_ENGINES.md
Tier: PROPOSAL. Nothing here is canon, nothing here is ratified, and nothing here changes a locked
pillar. It is an architecture read of the music-generation products that people pay for and keep
using, written because four rounds of our own music have been turned down and Josh ordered the step
back before any round 5.
Written: 2026-08-09. Lane: music architecture review. Consumer: the round-5 architecture ruling.
The grade that ordered it, verbatim from docs/review_candidates.json, row PASS7_FLORES_ROAD,
field grade_verbatim, 2026-08-09:
you are using. I hate how slow the melody is and there is no variation or changes or movement
throughout the songs. Everything sucks."
Four rounds, four turndowns, and the diagnosis has moved down a layer every time. Round 1 was
rejected for its ARCHITECTURE — ACE-Step full renders, single-shot, placed nothing
(round_1_takedown_note). Round 2 fixed the architecture and was rejected for its MELODIES —
"the melodies are way too simple... put some intelligence into the melodies and how everything
orchestrates" (round_2_takedown_note). Round 3 was rejected in one sentence — "it sounds just like
a bunch of noise and these main melodies really suck. They clash and have no rythm or complexity.
Just a few notes slowly played back to back. Never any harmonies added on" (round_3_takedown_note).
Round 4 answered that sentence clause by clause, moved the measured numbers, and was graded WORSE
than round 2.
That last fact is the one this doc is built on. When a system's own instruments say it improved on
every axis and the listener says it got worse, the instruments are measuring inside an architecture
that cannot produce the thing being asked for. So the question stops being "which axis do we add"
and becomes "what shape do the systems have that DO produce it."
The honest naming of our own stack, from the round-4 card's honest_limits_short:
layer a note. This is not a scoring session, no human has touched a note of it, and the honest tier
is STRUCTURE plus a measured mix. A real player's phrasing, breath and bow are absent by
construction, and they are a large part of what separates this from a recording."
We wrote that limit down before Josh heard round 4 and then were surprised when he named the
instruments. The limit was correct. This doc is about what the products that clear that bar do
instead.
These bind any architecture proposed at the end. They are not negotiable by this lane.
(memory music-direction-melody-first-30-year-bar).
traditions that sfz_palette.py already applies by declining VCSL's living-tradition instruments.
docs/proposals/music/craft_research/VIRTUAL_ORCHESTRATION_AND_MOCKUP_REALISM.md already contains a
measured, source-grounded diagnosis of why OUR renders sound fake — zero continuous controller data
reaching the sampler, mathematically perfect simultaneity, a per-section step function for dynamics,
CC1 a no-op on our palette, one global diffuse field for a mix. That work stands and this doc does
not repeat it. docs/proposals/music/MUSIC_PRACTICE_GAP_AUDIT.md §4 carries the ten craft changes
round 4 was built to land. Both are craft-layer documents. This one is an ARCHITECTURE-layer
document: not "what should the composer do differently" but "what are the four layers a commercial
engine has, and which of them do we have at all."
---
Every claim in this doc traces to one of these. Confidence is marked where a vendor has published no
technical paper and the public account is inference or secondary reporting.
architecture, corpus-based learning from Bach/Beethoven/Mozart, Music Engine product from January
2019, SACEM registration, Genesis (2016) and Among the Stars (2018) albums, Avignon Symphonic
Orchestra performance April 2017.
https://datainnovation.org/2019/05/5qs-for-pierre-barreau-ceo-of-aiva/ — the founder's own account
of reading scores to infer compositional rules.
https://stayrelevant.globant.com/en/meet-pierre-barreau-expert-behind-algorithm-creates-music-artificial-intelligence-ai/
— "looks at large amounts of scores... to infer rules about how music is composed," analysing
"patterns in melody, harmony, structure, instrumentation," then "converts these pieces of written
scores into audio."
https://soundcloud.com/theaipodcast/ep-34 — the reading-30,000-scores account.
https://aiva.crisp.help/en/article/general-user-manual-44klp4/ — the Influence feature, Style
Designer, preset styles, piano-roll editor.
https://aisongcreator.pro/blog/aiva-ai-review — the practitioner account of preview renders vs
export, and the MIDI-into-a-real-library workflow.
influence upload behaviour.
CONFIDENCE NOTE: AIVA has published no technical paper. The corpus figure is quoted as 15,000
digitised partitions in one account and 30,000 scores in another; both are founder-adjacent
secondary reporting and the discrepancy is unresolved. The architectural SHAPE — symbolic corpus in,
score out, audio rendered afterwards — is consistent across every source and is what this doc relies
on. The exact model family is not public.
https://www.prnewswire.com/news-releases/former-google-deepmind-researchers-assemble-luminaries-across-music-and-tech-to-launch-udio-a-new-ai-powered-app-that-allows-anyone-to-create-extraordinary-music-in-an-instant-302113166.html
— founding team, Uncharted Labs.
https://venturebeat.com/ai/former-google-deepmind-researchers-launch-ai-powered-music-creation-app-udio
latent-codec + neural-vocoder reading, and the explicit statement that Udio has published no
detailed architecture paper.
https://musicgeneratorai.io/posts/how-does-suno-ai-create-music — the transformer-generates-tokens,
latent-diffusion-refines-spectrogram, EnCodec-family-vocoder reading; the Bark lineage.
https://vi-control.net/community/threads/how-exactly-do-suno-ai-and-udio-com-work-technical-view.151041/
in — AudioLDM, https://proceedings.mlr.press/v202/liu23f.html
and the AI music lawsuits timeline — https://dynamoi.com/learn/ai-music-distribution/ai-music-copyright-cases-timeline
— the 60,000+ fingerprinted training recordings, the Warner settlement and licensing deal, the
continuing Sony/UMG litigation.
https://www.forbes.com/sites/virginieberger/2025/12/18/launch-train-settle-how-suno-and-udios-licensing-deals-made-copyright-infringement-profitable/
CONFIDENCE NOTE: neither Suno nor Udio has published an architecture paper. Everything in §4 about
their internals is reconstruction from founder statements, the open-source sibling Bark, the
published academic family the founders came from, and practitioner analysis. It is directionally
reliable and specifically unreliable. The PRODUCT behaviour in §4.3 is directly observable and is
the part that matters most for us.
https://soundraw.io/blog/post/soundraw-revealed-how-our-ai-generates-music — "all the samples and
sounds used by our AI are created by our internal team of talented music producers"; no text
prompts; tag-driven; stems.
https://docs.channel.io/soundraw-faq/en/articles/how-soundraw-ai-works-698bcb78
by human musicians and sound designers, prompt-to-tag vector matching, arrangement assembly.
https://github.com/fantasy209jk/Mubert
https://support.boomy.com/hc/en-us/articles/17795213541773-How-does-Boomy-use-AI — the
statistics-based starting point plus user customisation account. (Fetch returned 403 at time of
writing; the quoted phrasing is from the indexed search summary of that page and is marked
SECONDARY. Aggregator claims that Boomy uses GANs are unverified and are not relied on here.)
https://www.audiokinetic.com/en/courses/wwise201/?id=lesson_5_creating_interaction_understanding_stingers/
https://www.audiokinetic.com/en/courses/wwise201/?id=smoothing_transition_decisions_transitioning_to_specific_playlist_items/
https://www.audiokinetic.com/en/blog/making-interactive-music-in-real-life-with-wwise/
https://alessandrofama.com/tutorials/fmod/fmod-studio/vertical-reorchestration
https://www.thegameaudioco.com/making-your-game-s-music-more-dynamic-vertical-layering-vs-horizontal-resequencing
https://www.arcellasound.com/post/interactive-music-in-aaa-games-vertical-layering-vs-horizontal-re-sequencing
https://www.resolutiongames.com/blog/behind-the-scenes-recording-demeos-soundtrack-with-a-live-orchestra
— the score-to-orchestrator-to-Prague-session chain as ordinary practice.
https://arxiv.org/pdf/1402.0585
https://arxiv.org/pdf/1712.04371
https://arxiv.org/html/2403.07995v1 — long-term structure through motives, patterns and variations
as the requirement, not an extra.
https://arxiv.org/pdf/1812.04832
https://www.mdpi.com/2078-2489/16/8/656 — the finding that hand-crafted rules interact and the
interaction is what destroys musicality.
Reports — https://www.nature.com/articles/s41598-025-13064-6
— https://arxiv.org/abs/2502.18008 and https://github.com/ElectricAlexis/NotaGen — 1.6M-piece ABC
pre-training, ~9K classical fine-tune, period-composer-instrumentation conditioning, CLaMP-DPO.
control by interleaving events and controls; trained on Lakh MIDI.
https://www.metacreation.net/projects/mmm-multi-track-music-machine — bar-level and track-level
inpainting with instrument and density control.
a DAW, which is the existence proof that this tier of model is a desktop citizen.
contrastive alignment of symbolic music, audio and multilingual text; open weights.
https://github.com/facebookresearch/audiobox-aesthetics — four-axis automatic aesthetic scoring
(Production Quality, Production Complexity, Content Enjoyment, Content Usefulness), open weights.
https://arxiv.org/pdf/2504.16839 — the closed loop: generate symbolic, render with a soundfont,
score the AUDIO, optimise the SYMBOLIC policy.
~150M parameters, conditioned on time-aligned piano-roll features at 100 Hz plus a composer
embedding, 12 composers / 216 recordings / ~62 hours; MIDI-to-symphony and audio-to-symphony
modes; code, weights and preprocessing scripts released.
https://arxiv.org/pdf/2410.16785 — render with a sampler first, then let a diffusion model polish
it; outperforms end-to-end because the sampler supplies the acoustic prior.
https://arxiv.org/html/2309.12283 and https://benadar293.github.io/midipm/
https://arxiv.org/pdf/2512.02652 — current expressive-performance-rendering models.
https://github.com/ace-step/ACE-Step-1.5 — LM planner plus DiT decoder, sub-4GB VRAM, LoRA from a
few songs, repainting and audio2audio, and a published limitations list.
driven by continuous MIDI expression rather than sample selection.
https://www.soundonsound.com/reviews/audio-modeling-swam-string-sections — the honest trade-off.
— free, Maida Vale recorded; and BBCSO Core — https://www.spitfireaudio.com/bbc-symphony-orchestra-core/
https://github.com/YatingMusic/ReaRender ; REAPER technical page —
https://www.reaper.fm/about.php#technical — VST/VST3/CLAP hosting, batch and queued rendering.
https://andrewfeazelle.com/how-to-make-an-orchestral-sample-library-sound-real/
https://www.soundonsound.com/techniques/sampled-orchestra-part3
---
Read across all of them and the same four layers appear. The products differ in which layer they own
and which they buy, but none of them is missing a layer.
| Layer | What it decides | The failure it prevents |
|---|---|---|
| GENERATION | the notes: melody, harmony, rhythm, form | "the melody sucks", "no variation" |
| RENDERING | the sound: timbre, articulation, phrasing, room, mix | "I hate the instruments you are using" |
| CURATION | which of many candidates is published | shipping the first thing the generator made |
| DELIVERY | how the music behaves during play | a track that starts and never changes |
Our stack, honestly placed against that table:
Josh has now rejected four times.
dynamic layer a note, no continuous controller data, a synthetic impulse response. This is not a
layer we chose over an alternative; it is the free option.
generator produced. Rounds 2, 3 and 4 each staged one cue per slot from one plan. There has never
been a candidate pool and there has never been a selection step.
been graded on a linear WAV, and "no variation or changes or movement" is partly a complaint about
a layer that is not in the artifact he was given.
That table is the whole finding of this document. Three of our four layers are either absent or
at their floor, and all four rounds of work went into deepening the one layer we already owned.
---
AIVA is the closest commercial analogue to what we are trying to be, and the most instructive,
because its architecture is the one Josh's rulings actually permit: it composes a SCORE, the score
is editable, and the human owns the direction.
AIVA learns from a symbolic corpus, not from recordings. Barreau's own account is that the system
"looks at large amounts of scores... to infer rules about how music is composed," analysing
"patterns in melody, harmony, structure, instrumentation and more"
(https://stayrelevant.globant.com/en/meet-pierre-barreau-expert-behind-algorithm-creates-music-artificial-intelligence-ai/).
Wikipedia records deep learning plus reinforcement learning architectures and corpus-based learning
from the classical masters (https://en.wikipedia.org/wiki/AIVA). The corpus size is variously
reported as 15,000 partitions or 30,000 scores; treat the number as unconfirmed and the method as
confirmed.
The load-bearing distinction is this: AIVA INFERRED its rules from a corpus rather than having them
written down. The rules it uses are not smaller than ours or better argued than ours — they are of a
different kind. A learned prior encodes the tens of thousands of soft constraints that a composer
obeys without being able to state, and a hand-written rule system encodes only the constraints
somebody thought to state. §7.1 develops why that gap is the melody gap.
AIVA is NOT primarily a text-prompt product, and this is the single most useful fact in this doc.
Its conditioning surface, per the official user manual
(https://aiva.crisp.help/en/article/general-user-manual-44klp4/) and feature documentation
(https://siteefy.com/tools/aiva):
a free-text description.
harmonic structure and writes a new piece over that structure with its own melody. Practitioner
accounts are consistent that MIDI influences work well and audio influences work badly, which
tells you the conditioning is genuinely SYMBOLIC and the audio path is a lossy front end onto it.
That is the shape of "composed, never prompted" as a product feature. The human supplies structural
musical material; the machine develops it. If we ever adopt a learned generator, the Influence
pattern is the precedent that keeps Josh's authorship intact: his themes go in as symbolic
conditioning, and what comes back is a development of HIS material rather than a sample from a
distribution.
AIVA's product loop is generate, listen, edit, regenerate. It presents the full score as a piano
roll where a user can change individual notes, adjust velocities, reassign instruments and
restructure sections before export
(https://aisongcreator.pro/blog/aiva-ai-review). The human is inside the loop on every track, and
the loop is cheap enough to run many times.
This is the layer we do not have. AIVA users do not publish AIVA's first output; they generate,
reject, regenerate, then edit. The published artifact is the survivor of a selection process. Our
published artifact is the only thing that was made.
This is where the read overturned an assumption.
AIVA's own audio does not carry AIVA's reputation. Its in-app previews render through stock samples
and are widely described as flat and MIDI-esque. The professional workflow that produces the results
people cite is: compose in AIVA, EXPORT MIDI, load it in a DAW, and trigger a real orchestral sample
library — Spitfire, EastWest, Kontakt instruments — then mix it
(https://aisongcreator.pro/blog/aiva-ai-review, https://aibuilderhub.dev/en/use-ai/aiva). AIVA
supports this directly: MIDI export, individual instrument stems, and chord-progression export are
first-class outputs.
So AIVA does not solve the rendering problem. It DECLINES to solve the rendering problem, and hands
a score to a rendering stack that was solved by the sample-library industry over twenty-five years of
recording real players in real halls with articulation trees, velocity layers, round robins and
recorded legato transitions.
The implication for us is direct and uncomfortable. We have been trying to get a shippable orchestral
sound out of the free tier of a solved industry. Josh's "I hate the instruments you are using" is not
an aesthetic quibble that better composition will overcome. It is an accurate report that we are
using the wrong instruments, and no amount of work in the generation layer will change it.
stated ones.
developed applies to it unchanged.
synthesise a violin, because a violin was already recorded.
Secondary reviews repeatedly say the Pro plan grants full copyright. AIVA's own legal text
(https://www.aiva.ai/legal/1) does not say that. Reading the terms directly:
composition in content the licensee holds rights over.
Twitch, TikTok, Instagram.
or similar service; and USING THE AUDIO OR MIDI AS PART OF A TRAINING DATASET FOR ANY MACHINE
LEARNING.
Consequences for us, stated plainly:
grant is scoped to four social platforms and a game is not one of them.
AIVA-generated material is closed by their terms.
supply route.
---
These products have no score. There is no MIDI inside them and no orchestration decision that could
be inspected. A text prompt and lyrics go in; a mixed, mastered, performed-sounding stereo recording
comes out.
The public reconstruction of Suno's pipeline — inference, not documentation — is a prompt-parsing
language model, a transformer that predicts sequences of audio tokens carrying structure, melody,
harmony and lyrics, a latent diffusion stage that refines toward a spectrogram, and an EnCodec-family
neural vocoder that produces the waveform
(https://musicgeneratorai.io/posts/how-does-suno-ai-create-music). Udio is described in the same
family: transformer backbone over neural-codec latents with a neural vocoder, the technique family
established by AudioLM, MusicLM and Lyria — which is where its five DeepMind founders came from
(https://www.emergentmind.com/topics/udio,
https://www.prnewswire.com/news-releases/former-google-deepmind-researchers-assemble-luminaries-across-music-and-tech-to-launch-udio-a-new-ai-powered-app-that-allows-anyone-to-create-extraordinary-music-in-an-instant-302113166.html).
Because it never renders anything. The timbre, the bow noise, the breath, the room, the compression,
the mastering chain and the human phrasing were all present in the training recordings, and the model
reproduces the joint distribution of all of them. There is no articulation-switching problem because
nobody switched an articulation; the model learned what a violin sounds like WHEN PLAYING THAT
PHRASE, in a room, through a mix.
That is the deepest structural reason Suno beats every mockup pipeline on raw sonic believability,
and it is also the reason it cannot take direction at the note level. The two facts are the same
fact.
Product behaviour is observable and is the most transferable finding in this section:
the interface itself.
many generations, a listening pass, and a pick.
Suno's quality, as experienced by a listener, is a joint product of a strong generator AND a heavy
selection process. We have been comparing our single output against their selected output. That is
not a fair comparison and, more importantly, it is a fixable one — selection is the cheapest layer to
add and we have never had it.
Suno's later versions honour structural meta-tags such as verse and chorus markers, with community
analysis attributing that to reinforcement learning from human feedback that rewarded outputs which
followed the requested song map (https://musicgeneratorai.io/posts/how-does-suno-ai-create-music).
This is worth naming precisely: even a company with a frontier audio model had to add a separate
mechanism to make structure obey instruction, because structure does not emerge reliably from
next-token prediction over audio. Our round-1 rejection — "single-shot generation placed nothing" —
was the same finding arrived at independently, and it remains true.
Sony and Universal identified 60,000+ of their copyrighted recordings in Suno's training data by
audio fingerprinting, with an amended complaint alleging acquisition by stream-ripping around
YouTube's DRM. Warner settled and entered a licensing partnership; Sony and UMG continue to litigate
and are moving to expand the case to 61,026 recordings
(https://www.aimusicpreneur.com/knowledge-base/legal/riaa-suno-copyright-case/,
https://dynamoi.com/learn/ai-music-distribution/ai-music-copyright-cases-timeline,
https://www.forbes.com/sites/virginieberger/2025/12/18/launch-train-settle-how-suno-and-udios-licensing-deals-made-copyright-infringement-profitable/).
For this project the reading is not moralistic and it is not conservative either — it is a supply
risk read. A shipped 30M-word RPG carries its soundtrack for its whole commercial life. The
end-to-end audio route's quality advantage comes from exactly the corpus that is under active
litigation, and any pipeline of ours built on a model of that lineage inherits the provenance
question. Our own licence discipline already has teeth — fetch_sfz_stack.py stores the verbatim
licence body of every component under docs/licence_records/ and pins by commit SHA rather than
branch. Whatever we adopt has to survive that same read.
---
Mubert states it directly: all sounds — separate loops for bass, leads and the rest — are created by
musicians and sound designers and are NOT synthesised by neural networks; the proprietary technology
analyses and selects relevant sounds and builds arrangements from them
(https://mubert.com/api, https://landing.mubert.com/). A prompt is encoded to a latent vector, matched
against tag vectors, and the matched tags drive library retrieval and arrangement.
Soundraw says the same in its own words: "all the samples and sounds used by our AI are created by
our internal team of talented music producers," and its generation is tag-driven rather than
prompt-driven — "SOUNDRAW's music generation is driven by a unique AI system that doesn't rely on
text prompts or mimicking existing songs"
(https://soundraw.io/blog/post/soundraw-revealed-how-our-ai-generates-music). The user sets mood,
genre, instruments and length; the system returns a list of candidate tracks; the user picks one and
then edits it, with stems available for a DAW.
Boomy's own support material describes a statistics-based starting point that the user then
customises with accessible editing tools (SECONDARY, see §1.3).
Because a human played every sound in it. The machine's entire job is combinatorial: which loop,
which key, which section order, which density. It never has to make a violin sound like a violin,
never has to phrase a line, never has to decide a bow direction — those decisions are frozen into the
assets. What it can get wrong is limited to arrangement, and arrangement errors are far less
offensive to a listener than timbre and phrasing errors.
This architecture cannot serve a leitmotif. It has no way to state a specific melody and develop it
across seventy-nine chapters, because it does not compose melodies — it retrieves phrases. It is
excellent for background beds and structurally incapable of the thing Josh's melody-first ruling
requires.
The lesson is the SOURCING PRINCIPLE, and it is the same one AIVA's users demonstrate from the other
direction: every commercial engine whose output sounds like music got its sound from recorded human
performance. Soundraw and Mubert recorded it themselves. AIVA's users buy it from Spitfire. Suno
learned it from records. Three completely different architectures, one universal fact.
We are the only architecture in this survey that tries to source its sound from a free
community-contributed sample set with one dynamic layer per note. That is the outlier, and Josh
identified it by ear without seeing any of this.
---
The standard chain is unromantic and worth stating because it is the bar: a composer writes the cue,
an orchestrator produces parts, real players record it in a studio, and the result is delivered as
STEMS into middleware. Resolution Games documents exactly that chain for Demeo — score and audio
reference to an orchestrator, sheet music for all individual instruments, recording at the Czech
National Symphony Orchestra studio in Prague
(https://www.resolutiongames.com/blog/behind-the-scenes-recording-demeos-soundtrack-with-a-live-orchestra).
The Philharmonia describes game soundtracks as a significant and routine part of studio output
(https://philharmonia.co.uk/what-we-do/in-the-studio/game-soundtracks/).
Nothing is generated at runtime. Everything is composed and recorded, then reassembled.
several synchronised stems that all play at once, with individual volumes driven by game
parameters, so the arrangement thickens and thins with intensity
(https://alessandrofama.com/tutorials/fmod/fmod-studio/vertical-reorchestration).
state changes.
Practitioner guidance is that vertical layering serves real-time intensity shifts in combat,
exploration and open-world play, while horizontal resequencing serves story-driven and segmented
gameplay — boss fights, cutscenes, linear levels
(https://www.thegameaudioco.com/making-your-game-s-music-more-dynamic-vertical-layering-vs-horizontal-resequencing,
https://www.arcellasound.com/post/interactive-music-in-aaa-games-vertical-layering-vs-horizontal-re-sequencing).
Stems map to real-time parameter controls in Wwise or FMOD.
Segments hold the audio. Playlist containers order and repeat segments with weights and loop counts.
Switch containers with transition rules choose which segment plays for the current game state, and
any segment can serve as a transition segment inside a rule. Stingers are short motifs registered to
triggers, fired on punctual events and synchronised to the next beat, bar or cue so they land in
musical context
(https://www.audiokinetic.com/en/courses/wwise201/?id=lesson_5_creating_interaction_understanding_stingers/,
https://www.audiokinetic.com/en/courses/wwise201/?id=smoothing_transition_decisions_transitioning_to_specific_playlist_items/,
https://www.audiokinetic.com/en/blog/making-interactive-music-in-real-life-with-wwise/).
Part of that complaint belongs to the composition, and round 4's own numbers admit some of it — one
measured hold, two distinct musical ideas in a 223-second cue. But part of it belongs to a layer that
was never in the artifact. In shipped practice, a three-and-a-half-minute cue is not a
three-and-a-half-minute experience; it is a set of stems and segments that recombine for as long as
the player stays. The variation a player experiences is produced at RUNTIME by the delivery layer,
and we have been asking a linear WAV to carry a job that no commercial score carries alone.
This does not excuse the cue. It does mean that some of the complaint is cheap to answer, and that
we should stop grading linear WAVs as if they were the product.
---
Five mechanisms. Each is sourced, and each maps to something specific in our own returns.
The survey literature is consistent. Markov and rule-based systems capture local note-to-note
transitions and fail at long-range structure; for longer works Markov models over-reuse corpus
melodies and become monotonous; and — the finding that indicts our specific approach — the
INTERACTION among hand-crafted rules introduces consistency problems, so outputs lack musicality and
coherence even when each rule is individually correct
(https://www.mdpi.com/2078-2489/16/8/656, https://arxiv.org/pdf/1402.0585,
https://arxiv.org/pdf/1712.04371).
That is a precise description of our round 3 to round 4 transition. We added a functional harmonic
schedule with named cadence families, snapped structural melody tones onto it, added a second melodic
voice in thirds and sixths, replaced pads with voice-led figures, and built a measured groove with a
metrical hierarchy. Every one of those rules is defensible in isolation, all of them are real
practice, and the combined output was graded worse than round 2. The literature predicts exactly
that: rule interaction is the failure mode, not rule absence.
It also explains why our measurement battery kept saying yes. The battery measures the rules. If the
defect lives in the interaction between rules, an instrument per rule cannot see it — which is the
same class of blindness that MUSIC_PRACTICE_GAP_AUDIT already caught once, when nine instruments
read sections, boundaries, drops and holds and not one read harmony, rhythm or melodic substance.
Suno returns multiple candidates by default. AIVA users regenerate until something is worth editing.
Soundraw returns a LIST of tracks for a tag set. In all three the published artifact is a survivor.
Our pipeline publishes the only artifact it makes. emit_review_picks.py carries the vocabulary of
selection — it has rank_in_theme and kept_in_theme fields — and for rounds 2, 3 and 4 both are
hardcoded to 1, one cue per theme, because the composer produced one. The field exists and the pool
does not. Even holding the generator completely fixed, a generate-many-and-rank loop would have
raised what Josh heard, because the variance between takes of a stochastic composer is large and we
have been sampling it once.
This is the single largest gap-to-cost ratio in the whole survey.
Stated as a table, because the pattern is unanimous.
| Product | Where its SOUND comes from |
|---|---|
| AIVA (professional use) | commercial sample libraries the user owns — Spitfire, EastWest, Kontakt |
| Suno / Udio | learned from commercial recordings of real performances |
| Soundraw | recorded in-house by their own producers |
| Mubert | recorded by contracted musicians and sound designers |
| AAA game scores | recorded by a live orchestra in a studio |
| OUR STACK | free CC0 community sample sets, one dynamic layer a note, offline sampler |
Why free sample sets cannot close that gap is well documented and is not a matter of effort. What
makes a sampled orchestra believable is recorded legato transitions between specific note pairs,
multiple round robins per transition to defeat the machine-gun effect, many velocity and dynamic
layers crossfaded by a continuous controller, several note-length articulations to switch between,
phase-aligned transitions, and divisi so a chord splits a section instead of stacking copies of it
(https://andrewfeazelle.com/how-to-make-an-orchestral-sample-library-sound-real/,
https://www.soundonsound.com/techniques/sampled-orchestra-part3,
https://www.musicnation.co.nz/exploring-orchestral-articulations-in-sample-libraries-from-common-to-unusual-techniques/).
Those are RECORDINGS THAT EITHER EXIST IN THE LIBRARY OR DO NOT. Our palette has one dynamic layer,
so there is no crossfade to perform, and VIRTUAL_ORCHESTRATION_AND_MOCKUP_REALISM.md §1.6 already
measured that CC1 is a no-op on it.
The alternative to buying recordings is modelling instead of sampling. SWAM's physical and
behavioural modelling requires and rewards continuous expression input rather than sample selection,
and independent review is candid about the trade — sample libraries carry the inherent life of a real
take, while with modelling it is on the driver to create that life
(https://audiomodeling.com/, https://www.soundonsound.com/reviews/audio-modeling-swam-string-sections).
For a machine driver with a rich expression model that trade can favour modelling, especially for
solo woodwind, brass and strings where samples are weakest.
The gap between a score and a performance is a research field with current models, not a fudge
factor. Expressive performance rendering generates realistic performances from note sequences, in
two branches: modelling expressive parameters in MIDI, and synthesising performance directly in
audio (https://www.nature.com/articles/s41598-025-13064-6). Current systems include PianoKontext,
a flow-matching renderer in a pretrained latent space (https://arxiv.org/html/2606.12282), and
Pianist Transformer, self-supervised on a large unlabelled MIDI corpus and forced to internalise
harmonic function and melodic direction because those cues inform performance choices
(https://arxiv.org/pdf/2512.02652).
Our stack currently has no expression model at all. It has constants. That is a whole named layer
of the problem, staffed by zero code.
Round 4 raised density and lengthened the new-tune span from 6.03 to 12.9 bars, and the listener
still heard no movement. The literature is explicit that what a listener experiences as movement is
long-term structure — motives, patterns and variations across a span — rather than local activity
(https://arxiv.org/html/2403.07995v1, https://arxiv.org/pdf/1812.04832). And in shipped games, a
large share of perceived movement is manufactured by the delivery layer at runtime, per §6.
Two distinct ideas in a 223-second cue is the number to look at. Adding notes to two ideas does not
make three.
---
| His words | The layer that owns it | What every commercial engine does there | What we do |
|---|---|---|---|
| "The melody sucks" | GENERATION | a learned prior over a real corpus, then human selection | hand-written rules, no selection |
| "I hate the instruments you are using" | RENDERING | buy recordings, record them, or learn them from records | free CC0 sets, one dynamic layer, no CC |
| "no variation or changes or movement" | GENERATION form plus DELIVERY | multi-idea forms, plus stems recombined at runtime | one linear WAV, two ideas |
| "how slow the melody is" | GENERATION | tempo and rhythmic surface are style-conditioned by the corpus | derived from our own rules and plateau policy |
Three of the four clauses point at layers we either do not have or have at their floor. Round 5 spent
inside the generation layer would be the fifth attempt to fix a rendering complaint with composition
work.
This is not a proposal to throw the lane away. Specific assets survive any architecture change:
LEITMOTIF_ARCHITECTURE.md and theme_architecture_rows.json are Josh's authorship layer and are exactly the symbolic
conditioning that AIVA's Influence pattern consumes. In a learned-generator architecture they stop
being generation code and become the CONDITIONING, which is a promotion.
a RANKING FUNCTION over a candidate pool it is immediately useful, because ranking is robust to the
absolute miscalibration that made it say 57-of-57 on a cue Josh rejected. A ranker only has to be
right about which of two candidates is better.
whatever renders next, and several findings — the controller triad, perceptual attack time,
systematic microtiming, depth as four independent cues — are renderer-agnostic.
than followed. Any new component enters through that door.
phrasing, before Josh heard it. That was correct and it is why this step back is possible.
The procedural composition engine plus CC0 sampler architecture cannot reach the bar Josh is grading
against, and the reason is not that it is unfinished. Every commercial engine in this survey sources
its SOUND from recorded human performance and its NOTES from either a learned prior or a human, and
runs a selection loop over multiple candidates before anything reaches a listener. We do none of
those three things. Round 5 has to change the architecture, not deepen the rules.
---
Each pattern is stated with what it is, how it satisfies every binding ruling, what it costs, the
proof-before-spend first step, and the strongest objection to it. A recommendation follows.
What it is. Replace the hand-written rule composer with a learned symbolic model conditioned on
Josh's themes, and replace the CC0 sampler with a purchased, articulation-complete orchestral
library driven by a real expression model and rendered headlessly.
The three sub-systems, all local and all with existing open components:
NotaGen is the strongest published musicality result of the three, pre-trained on 1.6M pieces and
fine-tuned on ~9K classical works with period-composer-instrumentation conditioning, with weights
and code released (https://arxiv.org/abs/2502.18008, https://github.com/ElectricAlexis/NotaGen).
Anticipatory Music Transformer and MMM are the CONTROL-shaped members of the family: infilling,
accompaniment generation, bar-level and track-level inpainting with instrument and density control
(https://arxiv.org/abs/2306.08620, https://arxiv.org/pdf/2008.06048). Control is what makes
"composed, never prompted" mechanically true — Josh's theme is stated, the model develops,
reharmonises, counterpoints and orchestrates AROUND fixed material rather than inventing the tune.
Composer's Assistant 2 is the existence proof that this tier runs locally in a DAW
(https://arxiv.org/pdf/2407.14700). All of it fits a 5090 with enormous headroom.
legato, driven by our craft docs' controller work, rendered headlessly through REAPER, which hosts
VST/VST3/CLAP and supports queued batch rendering, with ReaRender as the existing Python harness
(https://www.reaper.fm/about.php#technical, https://github.com/YatingMusic/ReaRender). Free entry
point for the proof: Spitfire BBC Symphony Orchestra Discover, recorded at Maida Vale
(https://www.spitfireaudio.com/bbc-symphony-orchestra-discover). Paid step-up: BBCSO Core. For solo
winds, brass and exposed lines where samples are weakest, SWAM physical modelling, which consumes
exactly the continuous expression our craft docs specify (https://audiomodeling.com/).
Audiobox Aesthetics on the rendered audio for the four aesthetic axes, CLaMP 3 for symbolic and
cross-modal similarity to Josh's approved corpus, and our own battery as a third opinion
(https://arxiv.org/abs/2502.05139, https://arxiv.org/abs/2502.10362). SMART is the published
precedent for the whole closed loop — generate symbolic, render, score the audio, optimise the
symbolic policy — and its honest limitation is that rendering every candidate is the bottleneck
(https://arxiv.org/pdf/2504.16839).
Rulings check:
test stays as an accept gate on the SCORE, where it belongs.
text prompt anywhere in the chain. AIVA's Influence feature is the commercial precedent.
sfz_palette.py already applies; the corpus provenance question is named honestly below.
possible at zero capex.
First proof step, cheap. Take ONE existing round-4 cue's score. Render it twice — once through the
current CC0 sfizz palette, once through BBCSO Discover in REAPER with the craft docs' controller work
applied — and put both in front of Josh with nothing else changed. That isolates the rendering
variable completely and answers "I hate the instruments" as a yes-or-no question in one sitting,
before a single line of learned-generator code is written.
Cost and time: the render probe is days. The generator swap is the real project — weeks, and it
carries genuine research risk. Library capex is optional at first and low later.
Strongest objection. NotaGen and its siblings are trained on classical Western corpora. Our score
spans seventy-nine chapters of world idioms, and the care doctrine means those idioms must be
rendered with authenticity rather than as classical music wearing a costume. A classically-trained
prior conditioned toward, say, Flores gamelan-adjacent material may produce a plausible Western
approximation of it, which is precisely the failure the care doctrine exists to prevent. The
mitigation is that the conditioning and the palette stay ours and the model is used for DEVELOPMENT
of stated material rather than idiom invention — but this objection is real, it is a care-tier
question and not just a quality one, and it must be tested on a non-Western cue before this pattern
is adopted wholesale. The second objection is corpus provenance: Lakh MIDI and the large ABC corpora
have their own licence questions, and our licence discipline has to be run over any weights we adopt
before they enter a shipping pipeline.
What it is. Change nothing in the generation layer. Render the score through a sampler exactly as now
to establish the acoustic prior — correct notes, correct timing, correct instrument identity — and
then pass that render through a neural model that rewrites it into something that sounds recorded.
The published basis is strong and recent:
renders with a sampler first and polishes with a diffusion model, and BEATS end-to-end generation
because the concatenative stage supplies the acoustic prior — note timing, dynamics and instrument
identity — so the generative model only has to supply quality
(https://arxiv.org/pdf/2410.16785). Its stated limitation is that it depends on sampler
availability and quality, which is a caution for our palette specifically.
style, using a flow-matching DiT of about 150M parameters conditioned on time-aligned piano-roll
features at 100 Hz plus a composer embedding, trained on 216 recordings and about 62 hours, with
MIDI-to-symphony and audio-to-symphony modes and released code, weights and preprocessing scripts
(https://symphony-rendering.github.io/). 150M parameters is nothing on a 5090.
adjacent published approaches (https://arxiv.org/html/2309.12283,
https://ismir2025program.ismir.net/poster_208.html).
with LoRA trainable from a few songs (https://ace-step.github.io/ace-step-v1.5.github.io/). Used as
a GENERATOR it was rejected in round 1 and correctly so. Used as a RENDERER conditioned on our own
audio, it is a different tool doing a different job, and that distinction is the one that keeps
"composed, never prompted" intact.
Rulings check. Melody-first is untouched because the notes do not change. Composed never prompted
holds if and only if the neural stage is conditioned on OUR audio or OUR MIDI and never on a text
description of a style — a hard line worth writing into the lane contract. Care line needs a
provenance read on any weights adopted; Symphony Rendering's corpus is published and auditable, which
is a point in its favour. No hire. Fully local. Cheapest possible proof.
First proof step. Run one round-4 cue's existing DRY render through Symphony Rendering in
MIDI-to-symphony mode and through a repaint pass, and compare all three against the current master.
Nothing else changes. Days, not weeks, and it directly interrogates the instrument complaint.
Strongest objection. It answers exactly one of Josh's four clauses. The melody, the slowness and the
absence of variation all live upstream and this pattern does not touch them. There is also a real
risk it makes matters worse in a specific way: a neural renderer that beautifies a weak line makes
the weakness MORE audible, not less, because it removes the sampler artifacts that were previously
absorbing some of the listener's attention. And Symphony Rendering's corpus is twelve Western
composers, which imports the same care-tier idiom question as Pattern A, one layer lower and harder
to see.
What it is. Stop shipping cues and start shipping SURVIVORS. Make the factory produce a large
candidate pool per cue slot, score every candidate through a listener-proxy stack calibrated against
Josh's own graded history, and publish only the top few. Generator-agnostic: it multiplies whatever
composer we have, including the one we already built.
The components:
different form templates, different theme-development operation sequences, different tempi,
different orchestration assignments. The engine already has all of these as parameters; it simply
commits to one draw of each.
which is the closest published proxy for "does a person like this"
(https://arxiv.org/abs/2502.05139). CLaMP 3 gives cross-modal similarity so a candidate can be
scored against Josh's approved reference corpus rather than against an abstract ideal
(https://arxiv.org/abs/2502.10362). Our battery supplies the craft floors. NotaGen's CLaMP-DPO
demonstrates that a contrastive model can drive quality improvement WITHOUT human annotations or
hand-designed rewards, which is exactly our situation (https://arxiv.org/abs/2502.18008).
a graded corpus: eighteen round-1 tracks, two round-2 cues, three round-3 cues and three round-4
cues, all turned down, with verbatim reasons, and the round-4 comparison cards carry per-axis
numbers for rounds 2 through 4. Any proposed ranker must be able to reproduce Josh's ORDERING of
the rounds he has already graded — in particular it must rank round 2 above round 4, which is the
ordering that broke our current battery. A ranker that cannot recover a known verdict is not
qualified to select an unknown one.
Rulings check. Every ruling holds unchanged, because this pattern adds no generator and no corpus.
It is the only one of the three with no research risk, no capex and no provenance question.
First proof step. Take the round-4 planner exactly as it stands, generate forty candidates for ONE
cue slot by sampling its existing stochastic axes, render them all, and score them. Then check the
calibration claim against the graded history before anything is staged. If the top-ranked of forty is
plainly better than what was staged as round 4, the layer has proved itself for the cost of compute
we already own.
Strongest objection. Selection cannot exceed the generator's ceiling. If every one of forty
candidates has the same two-idea form, the same slow melodic surface and the same CC0 instruments,
the best of forty is still a rejected cue, and Josh will say so in one sentence. Selection multiplies
quality; it does not create it. There is also a real risk that optimising against an automatic
aesthetic score produces music that scores well and sounds worse — the ranker becomes a target and
stops being a measure — which is why the human stays at the end of the loop and why the calibration
check against his graded history is a precondition rather than a nicety.
Run them in the order of proof cost, not in the order of ambition.
code and hardware we already have, neither spends money, and between them they isolate the two
variables Josh named most sharply — the instruments and the absence of choice. Each is days.
a known quantity, and the A/B against our CC0 palette on an identical score is the single most
informative experiment available to this lane. It also de-risks Pattern B by telling us whether a
neural renderer is even needed, or whether the sampler was simply the wrong sampler.
and the honest position is that the four turndowns have not yet isolated it: Josh has never heard
our composer through good instruments, and he has never heard the best of forty. Those two probes
have to run before we can know whether the generation layer needs replacing or was merely being
heard through a broken window.
selection layers are fixed. If they do, Pattern A is the destination and the care-tier idiom
objection in §9.1 is the first thing that must be tested, on a non-Western cue.
of "no variation or changes or movement" is a stems-and-segments question, and grading a linear WAV
will keep producing that complaint no matter how good the cue gets.
The one-line version: we have been improving the only layer we own, and the layers we do not own are
where every commercial engine keeps its quality. Fix the cheap ones first and let them tell us
whether the expensive one is really the problem.
---
paper at all. Section 4's internals and section 3.1's corpus figures are reconstruction and
secondary reporting, flagged as such at the point of use. The PRODUCT behaviour and the LICENCE
terms are directly sourced and are the parts the recommendation actually rests on.
phrasing comes from the indexed summary of it. Aggregator claims that Boomy uses GANs are
unverified and are not used.
Symphony Rendering, Audiobox Aesthetics, CLaMP 3, ACE-Step 1.5, BBCSO Discover or SWAM is the
vendor's or the paper's claim, not our measurement. That is what the proof steps are for.
licence bodies on disk and commit-SHA pinning, and nothing in §9 enters a pipeline before it goes
through that door. Spitfire's EULA in particular was not retrievable at time of writing and must be
read directly before any Discover or Core render is published, with specific attention to
commercial game-soundtrack use and to any machine-learning restriction of the kind AIVA's terms
carry.
raises and does not answer.
Josh's verbatim grades or read out of the round cards.
DERIVED FROM:
docs/review_candidates.json — grade_verbatim and grade_reading on PASS7_FLORES_ROAD / _FALLS / _ROUNDS (round 4), round_1_takedown_note, round_2_takedown_note,
round_3_takedown_note, round_3_note, round_4_note, and the honest_limits_short and
comparison_card axis values quoted in §0 and §7.
harness/music_gen/pass7_realise.py (module docstring) — the renderer lineage, the synthetic-roommaster, the dry-and-mastered artifact pair.
harness/music_gen/fetch_sfz_stack.py (module docstring) — sfizz plus VSCO 2 CE plus VCSL, the commit-SHA pinning, the verbatim licence bodies under docs/licence_records/, and the care
decision to decline VCSL's living-tradition instruments.
docs/proposals/music/craft_research/VIRTUAL_ORCHESTRATION_AND_MOCKUP_REALISM.md §1.1-§1.8 and§2-§5 — the measured renderer diagnosis this doc cites rather than repeats.
docs/proposals/music/MUSIC_PRACTICE_GAP_AUDIT.md §4 and §4A — the ten round-4 changes.docs/proposals/music/LEITMOTIF_ARCHITECTURE.md and docs/proposals/music/theme_architecture_rows.json — named in §8.2 as the conditioning asset.
music-direction-melody-first-30-year-bar (melody-first, thirty-yearbar, no human composer hire, composer_in_loop is the factory's own pass ladder); CLAUDE.md
§"Content bar" and the CVD §17 floor; memory factory-first-resolution-mandate (proof before
spend); memory 5090-box-live-remote-stack (local-first target).
NOT DERIVED (authored judgment, and why it had no canon home):
separate layers; it is this doc's synthesis of the six product architectures surveyed, and it is
offered as an analytical tool rather than as a finding.
selects an architecture, and none of the three has been run.
emit_review_picks.py, where rank_in_theme and kept_in_theme are hardcoded to 1 on the round-2, round-3 and round-4 emitters
(lines 386, 504, 649), and off the absence of any generate-many step in the pass2-7 realisers. It
is an inference from the code's shape, not a statement any doc makes.