STEPBACK_COMMERCIAL_ENGINES.md

music/STEPBACK_COMMERCIAL_ENGINES.md

STEPBACK — how the commercial engines that actually satisfy listeners are built

Tier: PROPOSAL. Nothing here is canon, nothing here is ratified, and nothing here changes a locked

pillar. It is an architecture read of the music-generation products that people pay for and keep

using, written because four rounds of our own music have been turned down and Josh ordered the step

back before any round 5.

Written: 2026-08-09. Lane: music architecture review. Consumer: the round-5 architecture ruling.

0. Why this doc exists, in Josh's words

The grade that ordered it, verbatim from docs/review_candidates.json, row PASS7_FLORES_ROAD,

field grade_verbatim, 2026-08-09:

you are using. I hate how slow the melody is and there is no variation or changes or movement

throughout the songs. Everything sucks."

Four rounds, four turndowns, and the diagnosis has moved down a layer every time. Round 1 was

rejected for its ARCHITECTURE — ACE-Step full renders, single-shot, placed nothing

(round_1_takedown_note). Round 2 fixed the architecture and was rejected for its MELODIES —

"the melodies are way too simple... put some intelligence into the melodies and how everything

orchestrates" (round_2_takedown_note). Round 3 was rejected in one sentence — "it sounds just like

a bunch of noise and these main melodies really suck. They clash and have no rythm or complexity.

Just a few notes slowly played back to back. Never any harmonies added on" (round_3_takedown_note).

Round 4 answered that sentence clause by clause, moved the measured numbers, and was graded WORSE

than round 2.

That last fact is the one this doc is built on. When a system's own instruments say it improved on

every axis and the listener says it got worse, the instruments are measuring inside an architecture

that cannot produce the thing being asked for. So the question stops being "which axis do we add"

and becomes "what shape do the systems have that DO produce it."

The honest naming of our own stack, from the round-4 card's honest_limits_short:

layer a note. This is not a scoring session, no human has touched a note of it, and the honest tier

is STRUCTURE plus a measured mix. A real player's phrasing, breath and bow are absent by

construction, and they are a large part of what separates this from a recording."

We wrote that limit down before Josh heard round 4 and then were surprised when he named the

instruments. The limit was correct. This doc is about what the products that clear that bar do

instead.

0.1 The rulings this doc is written against

These bind any architecture proposed at the end. They are not negotiable by this lane.

(memory music-direction-melody-first-30-year-bar).

traditions that sfz_palette.py already applies by declining VCSL's living-tradition instruments.

0.2 What this doc deliberately does not redo

docs/proposals/music/craft_research/VIRTUAL_ORCHESTRATION_AND_MOCKUP_REALISM.md already contains a

measured, source-grounded diagnosis of why OUR renders sound fake — zero continuous controller data

reaching the sampler, mathematically perfect simultaneity, a per-section step function for dynamics,

CC1 a no-op on our palette, one global diffuse field for a mix. That work stands and this doc does

not repeat it. docs/proposals/music/MUSIC_PRACTICE_GAP_AUDIT.md §4 carries the ten craft changes

round 4 was built to land. Both are craft-layer documents. This one is an ARCHITECTURE-layer

document: not "what should the composer do differently" but "what are the four layers a commercial

engine has, and which of them do we have at all."

---

1. SOURCES

Every claim in this doc traces to one of these. Confidence is marked where a vendor has published no

technical paper and the public account is inference or secondary reporting.

1.1 AIVA

architecture, corpus-based learning from Bach/Beethoven/Mozart, Music Engine product from January

2019, SACEM registration, Genesis (2016) and Among the Stars (2018) albums, Avignon Symphonic

Orchestra performance April 2017.

https://datainnovation.org/2019/05/5qs-for-pierre-barreau-ceo-of-aiva/ — the founder's own account

of reading scores to infer compositional rules.

https://stayrelevant.globant.com/en/meet-pierre-barreau-expert-behind-algorithm-creates-music-artificial-intelligence-ai/

— "looks at large amounts of scores... to infer rules about how music is composed," analysing

"patterns in melody, harmony, structure, instrumentation," then "converts these pieces of written

scores into audio."

https://soundcloud.com/theaipodcast/ep-34 — the reading-30,000-scores account.

https://aiva.crisp.help/en/article/general-user-manual-44klp4/ — the Influence feature, Style

Designer, preset styles, piano-roll editor.

https://aisongcreator.pro/blog/aiva-ai-review — the practitioner account of preview renders vs

export, and the MIDI-into-a-real-library workflow.

influence upload behaviour.

CONFIDENCE NOTE: AIVA has published no technical paper. The corpus figure is quoted as 15,000

digitised partitions in one account and 30,000 scores in another; both are founder-adjacent

secondary reporting and the discrepancy is unresolved. The architectural SHAPE — symbolic corpus in,

score out, audio rendered afterwards — is consistent across every source and is what this doc relies

on. The exact model family is not public.

1.2 Suno and Udio

https://www.prnewswire.com/news-releases/former-google-deepmind-researchers-assemble-luminaries-across-music-and-tech-to-launch-udio-a-new-ai-powered-app-that-allows-anyone-to-create-extraordinary-music-in-an-instant-302113166.html

— founding team, Uncharted Labs.

https://venturebeat.com/ai/former-google-deepmind-researchers-launch-ai-powered-music-creation-app-udio

latent-codec + neural-vocoder reading, and the explicit statement that Udio has published no

detailed architecture paper.

https://musicgeneratorai.io/posts/how-does-suno-ai-create-music — the transformer-generates-tokens,

latent-diffusion-refines-spectrogram, EnCodec-family-vocoder reading; the Bark lineage.

https://vi-control.net/community/threads/how-exactly-do-suno-ai-and-udio-com-work-technical-view.151041/

in — AudioLDM, https://proceedings.mlr.press/v202/liu23f.html

and the AI music lawsuits timeline — https://dynamoi.com/learn/ai-music-distribution/ai-music-copyright-cases-timeline

— the 60,000+ fingerprinted training recordings, the Warner settlement and licensing deal, the

continuing Sony/UMG litigation.

https://www.forbes.com/sites/virginieberger/2025/12/18/launch-train-settle-how-suno-and-udios-licensing-deals-made-copyright-infringement-profitable/

CONFIDENCE NOTE: neither Suno nor Udio has published an architecture paper. Everything in §4 about

their internals is reconstruction from founder statements, the open-source sibling Bark, the

published academic family the founders came from, and practitioner analysis. It is directionally

reliable and specifically unreliable. The PRODUCT behaviour in §4.3 is directly observable and is

the part that matters most for us.

1.3 Loop and pattern assembly

https://soundraw.io/blog/post/soundraw-revealed-how-our-ai-generates-music — "all the samples and

sounds used by our AI are created by our internal team of talented music producers"; no text

prompts; tag-driven; stems.

https://docs.channel.io/soundraw-faq/en/articles/how-soundraw-ai-works-698bcb78

by human musicians and sound designers, prompt-to-tag vector matching, arrangement assembly.

https://github.com/fantasy209jk/Mubert

https://support.boomy.com/hc/en-us/articles/17795213541773-How-does-Boomy-use-AI — the

statistics-based starting point plus user customisation account. (Fetch returned 403 at time of

writing; the quoted phrasing is from the indexed search summary of that page and is marked

SECONDARY. Aggregator claims that Boomy uses GANs are unverified and are not relied on here.)

1.4 Game-audio delivery practice

https://www.audiokinetic.com/en/courses/wwise201/?id=lesson_5_creating_interaction_understanding_stingers/

https://www.audiokinetic.com/en/courses/wwise201/?id=smoothing_transition_decisions_transitioning_to_specific_playlist_items/

https://www.audiokinetic.com/en/blog/making-interactive-music-in-real-life-with-wwise/

https://alessandrofama.com/tutorials/fmod/fmod-studio/vertical-reorchestration

https://www.thegameaudioco.com/making-your-game-s-music-more-dynamic-vertical-layering-vs-horizontal-resequencing

https://www.arcellasound.com/post/interactive-music-in-aaa-games-vertical-layering-vs-horizontal-re-sequencing

https://www.resolutiongames.com/blog/behind-the-scenes-recording-demeos-soundtrack-with-a-live-orchestra

— the score-to-orchestrator-to-Prague-session chain as ordinary practice.

1.5 Why rule systems fail — the literature

https://arxiv.org/pdf/1402.0585

https://arxiv.org/pdf/1712.04371

https://arxiv.org/html/2403.07995v1 — long-term structure through motives, patterns and variations

as the requirement, not an extra.

https://arxiv.org/pdf/1812.04832

https://www.mdpi.com/2078-2489/16/8/656 — the finding that hand-crafted rules interact and the

interaction is what destroys musicality.

Reports — https://www.nature.com/articles/s41598-025-13064-6

1.6 The pieces a local rebuild would use

— https://arxiv.org/abs/2502.18008 and https://github.com/ElectricAlexis/NotaGen — 1.6M-piece ABC

pre-training, ~9K classical fine-tune, period-composer-instrumentation conditioning, CLaMP-DPO.

control by interleaving events and controls; trained on Lakh MIDI.

https://www.metacreation.net/projects/mmm-multi-track-music-machine — bar-level and track-level

inpainting with instrument and density control.

a DAW, which is the existence proof that this tier of model is a desktop citizen.

contrastive alignment of symbolic music, audio and multilingual text; open weights.

https://github.com/facebookresearch/audiobox-aesthetics — four-axis automatic aesthetic scoring

(Production Quality, Production Complexity, Content Enjoyment, Content Usefulness), open weights.

https://arxiv.org/pdf/2504.16839 — the closed loop: generate symbolic, render with a soundfont,

score the AUDIO, optimise the SYMBOLIC policy.

~150M parameters, conditioned on time-aligned piano-roll features at 100 Hz plus a composer

embedding, 12 composers / 216 recordings / ~62 hours; MIDI-to-symphony and audio-to-symphony

modes; code, weights and preprocessing scripts released.

https://arxiv.org/pdf/2410.16785 — render with a sampler first, then let a diffusion model polish

it; outperforms end-to-end because the sampler supplies the acoustic prior.

https://arxiv.org/html/2309.12283 and https://benadar293.github.io/midipm/

https://arxiv.org/pdf/2512.02652 — current expressive-performance-rendering models.

https://github.com/ace-step/ACE-Step-1.5 — LM planner plus DiT decoder, sub-4GB VRAM, LoRA from a

few songs, repainting and audio2audio, and a published limitations list.

driven by continuous MIDI expression rather than sample selection.

https://www.soundonsound.com/reviews/audio-modeling-swam-string-sections — the honest trade-off.

— free, Maida Vale recorded; and BBCSO Core — https://www.spitfireaudio.com/bbc-symphony-orchestra-core/

https://github.com/YatingMusic/ReaRender ; REAPER technical page —

https://www.reaper.fm/about.php#technical — VST/VST3/CLAP hosting, batch and queued rendering.

https://andrewfeazelle.com/how-to-make-an-orchestral-sample-library-sound-real/

https://www.soundonsound.com/techniques/sampled-orchestra-part3

---

2. THE FRAME — every engine that satisfies listeners has four layers

Read across all of them and the same four layers appear. The products differ in which layer they own

and which they buy, but none of them is missing a layer.

LayerWhat it decidesThe failure it prevents
GENERATIONthe notes: melody, harmony, rhythm, form"the melody sucks", "no variation"
RENDERINGthe sound: timbre, articulation, phrasing, room, mix"I hate the instruments you are using"
CURATIONwhich of many candidates is publishedshipping the first thing the generator made
DELIVERYhow the music behaves during playa track that starts and never changes

Our stack, honestly placed against that table:

Josh has now rejected four times.

dynamic layer a note, no continuous controller data, a synthetic impulse response. This is not a

layer we chose over an alternative; it is the free option.

generator produced. Rounds 2, 3 and 4 each staged one cue per slot from one plan. There has never

been a candidate pool and there has never been a selection step.

been graded on a linear WAV, and "no variation or changes or movement" is partly a complaint about

a layer that is not in the artifact he was given.

That table is the whole finding of this document. Three of our four layers are either absent or

at their floor, and all four rounds of work went into deepening the one layer we already owned.

---

3. AIVA — the symbolic composer with a bought orchestra

AIVA is the closest commercial analogue to what we are trying to be, and the most instructive,

because its architecture is the one Josh's rulings actually permit: it composes a SCORE, the score

is editable, and the human owns the direction.

3.1 Generation layer

AIVA learns from a symbolic corpus, not from recordings. Barreau's own account is that the system

"looks at large amounts of scores... to infer rules about how music is composed," analysing

"patterns in melody, harmony, structure, instrumentation and more"

(https://stayrelevant.globant.com/en/meet-pierre-barreau-expert-behind-algorithm-creates-music-artificial-intelligence-ai/).

Wikipedia records deep learning plus reinforcement learning architectures and corpus-based learning

from the classical masters (https://en.wikipedia.org/wiki/AIVA). The corpus size is variously

reported as 15,000 partitions or 30,000 scores; treat the number as unconfirmed and the method as

confirmed.

The load-bearing distinction is this: AIVA INFERRED its rules from a corpus rather than having them

written down. The rules it uses are not smaller than ours or better argued than ours — they are of a

different kind. A learned prior encodes the tens of thousands of soft constraints that a composer

obeys without being able to state, and a hand-written rule system encodes only the constraints

somebody thought to state. §7.1 develops why that gap is the melody gap.

3.2 What conditions it — the part that matters for "composed, never prompted"

AIVA is NOT primarily a text-prompt product, and this is the single most useful fact in this doc.

Its conditioning surface, per the official user manual

(https://aiva.crisp.help/en/article/general-user-manual-44klp4/) and feature documentation

(https://siteefy.com/tools/aiva):

a free-text description.

harmonic structure and writes a new piece over that structure with its own melody. Practitioner

accounts are consistent that MIDI influences work well and audio influences work badly, which

tells you the conditioning is genuinely SYMBOLIC and the audio path is a lossy front end onto it.

That is the shape of "composed, never prompted" as a product feature. The human supplies structural

musical material; the machine develops it. If we ever adopt a learned generator, the Influence

pattern is the precedent that keeps Josh's authorship intact: his themes go in as symbolic

conditioning, and what comes back is a development of HIS material rather than a sample from a

distribution.

3.3 The curation loop

AIVA's product loop is generate, listen, edit, regenerate. It presents the full score as a piano

roll where a user can change individual notes, adjust velocities, reassign instruments and

restructure sections before export

(https://aisongcreator.pro/blog/aiva-ai-review). The human is inside the loop on every track, and

the loop is cheap enough to run many times.

This is the layer we do not have. AIVA users do not publish AIVA's first output; they generate,

reject, regenerate, then edit. The published artifact is the survivor of a selection process. Our

published artifact is the only thing that was made.

3.4 The rendering layer — and the finding that changes our plan

This is where the read overturned an assumption.

AIVA's own audio does not carry AIVA's reputation. Its in-app previews render through stock samples

and are widely described as flat and MIDI-esque. The professional workflow that produces the results

people cite is: compose in AIVA, EXPORT MIDI, load it in a DAW, and trigger a real orchestral sample

library — Spitfire, EastWest, Kontakt instruments — then mix it

(https://aisongcreator.pro/blog/aiva-ai-review, https://aibuilderhub.dev/en/use-ai/aiva). AIVA

supports this directly: MIDI export, individual instrument stems, and chord-progression export are

first-class outputs.

So AIVA does not solve the rendering problem. It DECLINES to solve the rendering problem, and hands

a score to a rendering stack that was solved by the sample-library industry over twenty-five years of

recording real players in real halls with articulation trees, velocity layers, round robins and

recorded legato transitions.

The implication for us is direct and uncomfortable. We have been trying to get a shippable orchestral

sound out of the free tier of a solved industry. Josh's "I hate the instruments you are using" is not

an aesthetic quibble that better composition will overcome. It is an accurate report that we are

using the wrong instruments, and no amount of work in the generation layer will change it.

3.5 Why it sounds musical where rule systems fail

stated ones.

developed applies to it unchanged.

synthesise a violin, because a violin was already recorded.

3.6 Licence — and a correction to the marketing summaries

Secondary reviews repeatedly say the Pro plan grants full copyright. AIVA's own legal text

(https://www.aiva.ai/legal/1) does not say that. Reading the terms directly:

composition in content the licensee holds rights over.

Twitch, TikTok, Instagram.

or similar service; and USING THE AUDIO OR MIDI AS PART OF A TRAINING DATASET FOR ANY MACHINE

LEARNING.

Consequences for us, stated plainly:

grant is scoped to four social platforms and a game is not one of them.

AIVA-generated material is closed by their terms.

supply route.

---

4. Suno and Udio — end-to-end neural audio

4.1 Generation and rendering are the same layer

These products have no score. There is no MIDI inside them and no orchestration decision that could

be inspected. A text prompt and lyrics go in; a mixed, mastered, performed-sounding stereo recording

comes out.

The public reconstruction of Suno's pipeline — inference, not documentation — is a prompt-parsing

language model, a transformer that predicts sequences of audio tokens carrying structure, melody,

harmony and lyrics, a latent diffusion stage that refines toward a spectrogram, and an EnCodec-family

neural vocoder that produces the waveform

(https://musicgeneratorai.io/posts/how-does-suno-ai-create-music). Udio is described in the same

family: transformer backbone over neural-codec latents with a neural vocoder, the technique family

established by AudioLM, MusicLM and Lyria — which is where its five DeepMind founders came from

(https://www.emergentmind.com/topics/udio,

https://www.prnewswire.com/news-releases/former-google-deepmind-researchers-assemble-luminaries-across-music-and-tech-to-launch-udio-a-new-ai-powered-app-that-allows-anyone-to-create-extraordinary-music-in-an-instant-302113166.html).

4.2 Why it sounds musical — the mechanism, stated exactly

Because it never renders anything. The timbre, the bow noise, the breath, the room, the compression,

the mastering chain and the human phrasing were all present in the training recordings, and the model

reproduces the joint distribution of all of them. There is no articulation-switching problem because

nobody switched an articulation; the model learned what a violin sounds like WHEN PLAYING THAT

PHRASE, in a room, through a mix.

That is the deepest structural reason Suno beats every mockup pipeline on raw sonic believability,

and it is also the reason it cannot take direction at the note level. The two facts are the same

fact.

4.3 The curation loop, which is the transferable part

Product behaviour is observable and is the most transferable finding in this section:

the interface itself.

many generations, a listening pass, and a pick.

Suno's quality, as experienced by a listener, is a joint product of a strong generator AND a heavy

selection process. We have been comparing our single output against their selected output. That is

not a fair comparison and, more importantly, it is a fixable one — selection is the cheapest layer to

add and we have never had it.

4.4 Structural control, and its limits

Suno's later versions honour structural meta-tags such as verse and chorus markers, with community

analysis attributing that to reinforcement learning from human feedback that rewarded outputs which

followed the requested song map (https://musicgeneratorai.io/posts/how-does-suno-ai-create-music).

This is worth naming precisely: even a company with a frontier audio model had to add a separate

mechanism to make structure obey instruction, because structure does not emerge reliably from

next-token prediction over audio. Our round-1 rejection — "single-shot generation placed nothing" —

was the same finding arrived at independently, and it remains true.

4.5 The care and licence posture

Sony and Universal identified 60,000+ of their copyrighted recordings in Suno's training data by

audio fingerprinting, with an amended complaint alleging acquisition by stream-ripping around

YouTube's DRM. Warner settled and entered a licensing partnership; Sony and UMG continue to litigate

and are moving to expand the case to 61,026 recordings

(https://www.aimusicpreneur.com/knowledge-base/legal/riaa-suno-copyright-case/,

https://dynamoi.com/learn/ai-music-distribution/ai-music-copyright-cases-timeline,

https://www.forbes.com/sites/virginieberger/2025/12/18/launch-train-settle-how-suno-and-udios-licensing-deals-made-copyright-infringement-profitable/).

For this project the reading is not moralistic and it is not conservative either — it is a supply

risk read. A shipped 30M-word RPG carries its soundtrack for its whole commercial life. The

end-to-end audio route's quality advantage comes from exactly the corpus that is under active

litigation, and any pipeline of ours built on a model of that lineage inherits the provenance

question. Our own licence discipline already has teeth — fetch_sfz_stack.py stores the verbatim

licence body of every component under docs/licence_records/ and pins by commit SHA rather than

branch. Whatever we adopt has to survive that same read.

---

5. Soundraw, Mubert, Boomy — loop and pattern assembly

5.1 The architecture

Mubert states it directly: all sounds — separate loops for bass, leads and the rest — are created by

musicians and sound designers and are NOT synthesised by neural networks; the proprietary technology

analyses and selects relevant sounds and builds arrangements from them

(https://mubert.com/api, https://landing.mubert.com/). A prompt is encoded to a latent vector, matched

against tag vectors, and the matched tags drive library retrieval and arrangement.

Soundraw says the same in its own words: "all the samples and sounds used by our AI are created by

our internal team of talented music producers," and its generation is tag-driven rather than

prompt-driven — "SOUNDRAW's music generation is driven by a unique AI system that doesn't rely on

text prompts or mimicking existing songs"

(https://soundraw.io/blog/post/soundraw-revealed-how-our-ai-generates-music). The user sets mood,

genre, instruments and length; the system returns a list of candidate tracks; the user picks one and

then edits it, with stems available for a DAW.

Boomy's own support material describes a statistics-based starting point that the user then

customises with accessible editing tools (SECONDARY, see §1.3).

5.2 Why it sounds musical

Because a human played every sound in it. The machine's entire job is combinatorial: which loop,

which key, which section order, which density. It never has to make a violin sound like a violin,

never has to phrase a line, never has to decide a bow direction — those decisions are frozen into the

assets. What it can get wrong is limited to arrangement, and arrangement errors are far less

offensive to a listener than timbre and phrasing errors.

5.3 What it cannot do, and why that matters to us

This architecture cannot serve a leitmotif. It has no way to state a specific melody and develop it

across seventy-nine chapters, because it does not compose melodies — it retrieves phrases. It is

excellent for background beds and structurally incapable of the thing Josh's melody-first ruling

requires.

5.4 The transferable lesson, which is not the architecture

The lesson is the SOURCING PRINCIPLE, and it is the same one AIVA's users demonstrate from the other

direction: every commercial engine whose output sounds like music got its sound from recorded human

performance. Soundraw and Mubert recorded it themselves. AIVA's users buy it from Spitfire. Suno

learned it from records. Three completely different architectures, one universal fact.

We are the only architecture in this survey that tries to source its sound from a free

community-contributed sample set with one dynamic layer per note. That is the outlier, and Josh

identified it by ear without seeing any of this.

---

6. Game-audio middleware practice — the delivery layer

6.1 What AAA actually does

The standard chain is unromantic and worth stating because it is the bar: a composer writes the cue,

an orchestrator produces parts, real players record it in a studio, and the result is delivered as

STEMS into middleware. Resolution Games documents exactly that chain for Demeo — score and audio

reference to an orchestrator, sheet music for all individual instruments, recording at the Czech

National Symphony Orchestra studio in Prague

(https://www.resolutiongames.com/blog/behind-the-scenes-recording-demeos-soundtrack-with-a-live-orchestra).

The Philharmonia describes game soundtracks as a significant and routine part of studio output

(https://philharmonia.co.uk/what-we-do/in-the-studio/game-soundtracks/).

Nothing is generated at runtime. Everything is composed and recorded, then reassembled.

6.2 The two mechanisms that produce movement

several synchronised stems that all play at once, with individual volumes driven by game

parameters, so the arrangement thickens and thins with intensity

(https://alessandrofama.com/tutorials/fmod/fmod-studio/vertical-reorchestration).

state changes.

Practitioner guidance is that vertical layering serves real-time intensity shifts in combat,

exploration and open-world play, while horizontal resequencing serves story-driven and segmented

gameplay — boss fights, cutscenes, linear levels

(https://www.thegameaudioco.com/making-your-game-s-music-more-dynamic-vertical-layering-vs-horizontal-resequencing,

https://www.arcellasound.com/post/interactive-music-in-aaa-games-vertical-layering-vs-horizontal-re-sequencing).

Stems map to real-time parameter controls in Wwise or FMOD.

6.3 Wwise's interactive music hierarchy, as the concrete shape

Segments hold the audio. Playlist containers order and repeat segments with weights and loop counts.

Switch containers with transition rules choose which segment plays for the current game state, and

any segment can serve as a transition segment inside a rule. Stingers are short motifs registered to

triggers, fired on punctual events and synchronised to the next beat, bar or cue so they land in

musical context

(https://www.audiokinetic.com/en/courses/wwise201/?id=lesson_5_creating_interaction_understanding_stingers/,

https://www.audiokinetic.com/en/courses/wwise201/?id=smoothing_transition_decisions_transitioning_to_specific_playlist_items/,

https://www.audiokinetic.com/en/blog/making-interactive-music-in-real-life-with-wwise/).

6.4 What this means for Josh's "no variation or changes or movement"

Part of that complaint belongs to the composition, and round 4's own numbers admit some of it — one

measured hold, two distinct musical ideas in a 223-second cue. But part of it belongs to a layer that

was never in the artifact. In shipped practice, a three-and-a-half-minute cue is not a

three-and-a-half-minute experience; it is a set of stems and segments that recombine for as long as

the player stays. The variation a player experiences is produced at RUNTIME by the delivery layer,

and we have been asking a linear WAV to carry a job that no commercial score carries alone.

This does not excuse the cue. It does mean that some of the complaint is cheap to answer, and that

we should stop grading linear WAVs as if they were the product.

---

7. WHY RULE SYSTEMS FAIL WHERE THESE SUCCEED

Five mechanisms. Each is sourced, and each maps to something specific in our own returns.

7.1 A learned prior encodes what nobody can write down

The survey literature is consistent. Markov and rule-based systems capture local note-to-note

transitions and fail at long-range structure; for longer works Markov models over-reuse corpus

melodies and become monotonous; and — the finding that indicts our specific approach — the

INTERACTION among hand-crafted rules introduces consistency problems, so outputs lack musicality and

coherence even when each rule is individually correct

(https://www.mdpi.com/2078-2489/16/8/656, https://arxiv.org/pdf/1402.0585,

https://arxiv.org/pdf/1712.04371).

That is a precise description of our round 3 to round 4 transition. We added a functional harmonic

schedule with named cadence families, snapped structural melody tones onto it, added a second melodic

voice in thirds and sixths, replaced pads with voice-led figures, and built a measured groove with a

metrical hierarchy. Every one of those rules is defensible in isolation, all of them are real

practice, and the combined output was graded worse than round 2. The literature predicts exactly

that: rule interaction is the failure mode, not rule absence.

It also explains why our measurement battery kept saying yes. The battery measures the rules. If the

defect lives in the interaction between rules, an instrument per rule cannot see it — which is the

same class of blindness that MUSIC_PRACTICE_GAP_AUDIT already caught once, when nine instruments

read sections, boundaries, drops and holds and not one read harmony, rhythm or melodic substance.

7.2 Selection is where quality comes from, and we have none

Suno returns multiple candidates by default. AIVA users regenerate until something is worth editing.

Soundraw returns a LIST of tracks for a tag set. In all three the published artifact is a survivor.

Our pipeline publishes the only artifact it makes. emit_review_picks.py carries the vocabulary of

selection — it has rank_in_theme and kept_in_theme fields — and for rounds 2, 3 and 4 both are

hardcoded to 1, one cue per theme, because the composer produced one. The field exists and the pool

does not. Even holding the generator completely fixed, a generate-many-and-rank loop would have

raised what Josh heard, because the variance between takes of a stochastic composer is large and we

have been sampling it once.

This is the single largest gap-to-cost ratio in the whole survey.

7.3 The rendering layer is a separate industry and every winner pays for it

Stated as a table, because the pattern is unanimous.

ProductWhere its SOUND comes from
AIVA (professional use)commercial sample libraries the user owns — Spitfire, EastWest, Kontakt
Suno / Udiolearned from commercial recordings of real performances
Soundrawrecorded in-house by their own producers
Mubertrecorded by contracted musicians and sound designers
AAA game scoresrecorded by a live orchestra in a studio
OUR STACKfree CC0 community sample sets, one dynamic layer a note, offline sampler

Why free sample sets cannot close that gap is well documented and is not a matter of effort. What

makes a sampled orchestra believable is recorded legato transitions between specific note pairs,

multiple round robins per transition to defeat the machine-gun effect, many velocity and dynamic

layers crossfaded by a continuous controller, several note-length articulations to switch between,

phase-aligned transitions, and divisi so a chord splits a section instead of stacking copies of it

(https://andrewfeazelle.com/how-to-make-an-orchestral-sample-library-sound-real/,

https://www.soundonsound.com/techniques/sampled-orchestra-part3,

https://www.musicnation.co.nz/exploring-orchestral-articulations-in-sample-libraries-from-common-to-unusual-techniques/).

Those are RECORDINGS THAT EITHER EXIST IN THE LIBRARY OR DO NOT. Our palette has one dynamic layer,

so there is no crossfade to perform, and VIRTUAL_ORCHESTRATION_AND_MOCKUP_REALISM.md §1.6 already

measured that CC1 is a no-op on it.

The alternative to buying recordings is modelling instead of sampling. SWAM's physical and

behavioural modelling requires and rewards continuous expression input rather than sample selection,

and independent review is candid about the trade — sample libraries carry the inherent life of a real

take, while with modelling it is on the driver to create that life

(https://audiomodeling.com/, https://www.soundonsound.com/reviews/audio-modeling-swam-string-sections).

For a machine driver with a rich expression model that trade can favour modelling, especially for

solo woodwind, brass and strings where samples are weakest.

7.4 Expressive performance is a model, not a constant

The gap between a score and a performance is a research field with current models, not a fudge

factor. Expressive performance rendering generates realistic performances from note sequences, in

two branches: modelling expressive parameters in MIDI, and synthesising performance directly in

audio (https://www.nature.com/articles/s41598-025-13064-6). Current systems include PianoKontext,

a flow-matching renderer in a pretrained latent space (https://arxiv.org/html/2606.12282), and

Pianist Transformer, self-supervised on a large unlabelled MIDI corpus and forced to internalise

harmonic function and melodic direction because those cues inform performance choices

(https://arxiv.org/pdf/2512.02652).

Our stack currently has no expression model at all. It has constants. That is a whole named layer

of the problem, staffed by zero code.

7.5 Movement is a form problem and a delivery problem, not a density problem

Round 4 raised density and lengthened the new-tune span from 6.03 to 12.9 bars, and the listener

still heard no movement. The literature is explicit that what a listener experiences as movement is

long-term structure — motives, patterns and variations across a span — rather than local activity

(https://arxiv.org/html/2403.07995v1, https://arxiv.org/pdf/1812.04832). And in shipped games, a

large share of perceived movement is manufactured by the delivery layer at runtime, per §6.

Two distinct ideas in a 223-second cue is the number to look at. Adding notes to two ideas does not

make three.

---

8. WHAT IT MEANS FOR OUR STACK

8.1 Josh's four clauses, mapped to layers

His wordsThe layer that owns itWhat every commercial engine does thereWhat we do
"The melody sucks"GENERATIONa learned prior over a real corpus, then human selectionhand-written rules, no selection
"I hate the instruments you are using"RENDERINGbuy recordings, record them, or learn them from recordsfree CC0 sets, one dynamic layer, no CC
"no variation or changes or movement"GENERATION form plus DELIVERYmulti-idea forms, plus stems recombined at runtimeone linear WAV, two ideas
"how slow the melody is"GENERATIONtempo and rhythmic surface are style-conditioned by the corpusderived from our own rules and plateau policy

Three of the four clauses point at layers we either do not have or have at their floor. Round 5 spent

inside the generation layer would be the fifth attempt to fix a rendering complaint with composition

work.

8.2 What is genuinely worth keeping

This is not a proposal to throw the lane away. Specific assets survive any architecture change:

theme_architecture_rows.json are Josh's authorship layer and are exactly the symbolic

conditioning that AIVA's Influence pattern consumes. In a learned-generator architecture they stop

being generation code and become the CONDITIONING, which is a promotion.

a RANKING FUNCTION over a candidate pool it is immediately useful, because ranking is robust to the

absolute miscalibration that made it say 57-of-57 on a cue Josh rejected. A ranker only has to be

right about which of two candidates is better.

whatever renders next, and several findings — the controller triad, perceptual attack time,

systematic microtiming, depth as four independent cues — are renderer-agnostic.

than followed. Any new component enters through that door.

phrasing, before Josh heard it. That was correct and it is why this step back is possible.

8.3 The uncomfortable conclusion, stated plainly

The procedural composition engine plus CC0 sampler architecture cannot reach the bar Josh is grading

against, and the reason is not that it is unfinished. Every commercial engine in this survey sources

its SOUND from recorded human performance and its NOTES from either a learned prior or a human, and

runs a selection loop over multiple candidates before anything reaches a listener. We do none of

those three things. Round 5 has to change the architecture, not deepen the rules.

---

9. THE THREE ARCHITECTURE PATTERNS

Each pattern is stated with what it is, how it satisfies every binding ruling, what it costs, the

proof-before-spend first step, and the strongest objection to it. A recommendation follows.

9.1 Pattern A — LEARNED COMPOSER, BOUGHT ORCHESTRA (the AIVA shape, rebuilt local)

What it is. Replace the hand-written rule composer with a learned symbolic model conditioned on

Josh's themes, and replace the CC0 sampler with a purchased, articulation-complete orchestral

library driven by a real expression model and rendered headlessly.

The three sub-systems, all local and all with existing open components:

NotaGen is the strongest published musicality result of the three, pre-trained on 1.6M pieces and

fine-tuned on ~9K classical works with period-composer-instrumentation conditioning, with weights

and code released (https://arxiv.org/abs/2502.18008, https://github.com/ElectricAlexis/NotaGen).

Anticipatory Music Transformer and MMM are the CONTROL-shaped members of the family: infilling,

accompaniment generation, bar-level and track-level inpainting with instrument and density control

(https://arxiv.org/abs/2306.08620, https://arxiv.org/pdf/2008.06048). Control is what makes

"composed, never prompted" mechanically true — Josh's theme is stated, the model develops,

reharmonises, counterpoints and orchestrates AROUND fixed material rather than inventing the tune.

Composer's Assistant 2 is the existence proof that this tier runs locally in a DAW

(https://arxiv.org/pdf/2407.14700). All of it fits a 5090 with enormous headroom.

legato, driven by our craft docs' controller work, rendered headlessly through REAPER, which hosts

VST/VST3/CLAP and supports queued batch rendering, with ReaRender as the existing Python harness

(https://www.reaper.fm/about.php#technical, https://github.com/YatingMusic/ReaRender). Free entry

point for the proof: Spitfire BBC Symphony Orchestra Discover, recorded at Maida Vale

(https://www.spitfireaudio.com/bbc-symphony-orchestra-discover). Paid step-up: BBCSO Core. For solo

winds, brass and exposed lines where samples are weakest, SWAM physical modelling, which consumes

exactly the continuous expression our craft docs specify (https://audiomodeling.com/).

Audiobox Aesthetics on the rendered audio for the four aesthetic axes, CLaMP 3 for symbolic and

cross-modal similarity to Josh's approved corpus, and our own battery as a third opinion

(https://arxiv.org/abs/2502.05139, https://arxiv.org/abs/2502.10362). SMART is the published

precedent for the whole closed loop — generate symbolic, render, score the audio, optimise the

symbolic policy — and its honest limitation is that rendering every candidate is the bottleneck

(https://arxiv.org/pdf/2504.16839).

Rulings check:

test stays as an accept gate on the SCORE, where it belongs.

text prompt anywhere in the chain. AIVA's Influence feature is the commercial precedent.

sfz_palette.py already applies; the corpus provenance question is named honestly below.

possible at zero capex.

First proof step, cheap. Take ONE existing round-4 cue's score. Render it twice — once through the

current CC0 sfizz palette, once through BBCSO Discover in REAPER with the craft docs' controller work

applied — and put both in front of Josh with nothing else changed. That isolates the rendering

variable completely and answers "I hate the instruments" as a yes-or-no question in one sitting,

before a single line of learned-generator code is written.

Cost and time: the render probe is days. The generator swap is the real project — weeks, and it

carries genuine research risk. Library capex is optional at first and low later.

Strongest objection. NotaGen and its siblings are trained on classical Western corpora. Our score

spans seventy-nine chapters of world idioms, and the care doctrine means those idioms must be

rendered with authenticity rather than as classical music wearing a costume. A classically-trained

prior conditioned toward, say, Flores gamelan-adjacent material may produce a plausible Western

approximation of it, which is precisely the failure the care doctrine exists to prevent. The

mitigation is that the conditioning and the palette stay ours and the model is used for DEVELOPMENT

of stated material rather than idiom invention — but this objection is real, it is a care-tier

question and not just a quality one, and it must be tested on a non-Western cue before this pattern

is adopted wholesale. The second objection is corpus provenance: Lakh MIDI and the large ABC corpora

have their own licence questions, and our licence discipline has to be run over any weights we adopt

before they enter a shipping pipeline.

9.2 Pattern B — KEEP THE COMPOSER, ADD A NEURAL RENDERER (the two-stage refinement shape)

What it is. Change nothing in the generation layer. Render the score through a sampler exactly as now

to establish the acoustic prior — correct notes, correct timing, correct instrument identity — and

then pass that render through a neural model that rewrites it into something that sounds recorded.

The published basis is strong and recent:

renders with a sampler first and polishes with a diffusion model, and BEATS end-to-end generation

because the concatenative stage supplies the acoustic prior — note timing, dynamics and instrument

identity — so the generative model only has to supply quality

(https://arxiv.org/pdf/2410.16785). Its stated limitation is that it depends on sampler

availability and quality, which is a caution for our palette specifically.

style, using a flow-matching DiT of about 150M parameters conditioned on time-aligned piano-roll

features at 100 Hz plus a composer embedding, trained on 216 recordings and about 62 hours, with

MIDI-to-symphony and audio-to-symphony modes and released code, weights and preprocessing scripts

(https://symphony-rendering.github.io/). 150M parameters is nothing on a 5090.

adjacent published approaches (https://arxiv.org/html/2309.12283,

https://ismir2025program.ismir.net/poster_208.html).

with LoRA trainable from a few songs (https://ace-step.github.io/ace-step-v1.5.github.io/). Used as

a GENERATOR it was rejected in round 1 and correctly so. Used as a RENDERER conditioned on our own

audio, it is a different tool doing a different job, and that distinction is the one that keeps

"composed, never prompted" intact.

Rulings check. Melody-first is untouched because the notes do not change. Composed never prompted

holds if and only if the neural stage is conditioned on OUR audio or OUR MIDI and never on a text

description of a style — a hard line worth writing into the lane contract. Care line needs a

provenance read on any weights adopted; Symphony Rendering's corpus is published and auditable, which

is a point in its favour. No hire. Fully local. Cheapest possible proof.

First proof step. Run one round-4 cue's existing DRY render through Symphony Rendering in

MIDI-to-symphony mode and through a repaint pass, and compare all three against the current master.

Nothing else changes. Days, not weeks, and it directly interrogates the instrument complaint.

Strongest objection. It answers exactly one of Josh's four clauses. The melody, the slowness and the

absence of variation all live upstream and this pattern does not touch them. There is also a real

risk it makes matters worse in a specific way: a neural renderer that beautifies a weak line makes

the weakness MORE audible, not less, because it removes the sampler artifacts that were previously

absorbing some of the listener's attention. And Symphony Rendering's corpus is twelve Western

composers, which imports the same care-tier idiom question as Pattern A, one layer lower and harder

to see.

9.3 Pattern C — THE SELECTION ENGINE (the Suno lesson, without Suno)

What it is. Stop shipping cues and start shipping SURVIVORS. Make the factory produce a large

candidate pool per cue slot, score every candidate through a listener-proxy stack calibrated against

Josh's own graded history, and publish only the top few. Generator-agnostic: it multiplies whatever

composer we have, including the one we already built.

The components:

different form templates, different theme-development operation sequences, different tempi,

different orchestration assignments. The engine already has all of these as parameters; it simply

commits to one draw of each.

which is the closest published proxy for "does a person like this"

(https://arxiv.org/abs/2502.05139). CLaMP 3 gives cross-modal similarity so a candidate can be

scored against Josh's approved reference corpus rather than against an abstract ideal

(https://arxiv.org/abs/2502.10362). Our battery supplies the craft floors. NotaGen's CLaMP-DPO

demonstrates that a contrastive model can drive quality improvement WITHOUT human annotations or

hand-designed rewards, which is exactly our situation (https://arxiv.org/abs/2502.18008).

a graded corpus: eighteen round-1 tracks, two round-2 cues, three round-3 cues and three round-4

cues, all turned down, with verbatim reasons, and the round-4 comparison cards carry per-axis

numbers for rounds 2 through 4. Any proposed ranker must be able to reproduce Josh's ORDERING of

the rounds he has already graded — in particular it must rank round 2 above round 4, which is the

ordering that broke our current battery. A ranker that cannot recover a known verdict is not

qualified to select an unknown one.

Rulings check. Every ruling holds unchanged, because this pattern adds no generator and no corpus.

It is the only one of the three with no research risk, no capex and no provenance question.

First proof step. Take the round-4 planner exactly as it stands, generate forty candidates for ONE

cue slot by sampling its existing stochastic axes, render them all, and score them. Then check the

calibration claim against the graded history before anything is staged. If the top-ranked of forty is

plainly better than what was staged as round 4, the layer has proved itself for the cost of compute

we already own.

Strongest objection. Selection cannot exceed the generator's ceiling. If every one of forty

candidates has the same two-idea form, the same slow melodic surface and the same CC0 instruments,

the best of forty is still a rejected cue, and Josh will say so in one sentence. Selection multiplies

quality; it does not create it. There is also a real risk that optimising against an automatic

aesthetic score produces music that scores well and sounds worse — the ranker becomes a target and

stops being a measure — which is why the human stays at the end of the loop and why the calibration

check against his graded history is a precondition rather than a nicety.

9.4 Recommendation

Run them in the order of proof cost, not in the order of ambition.

code and hardware we already have, neither spends money, and between them they isolate the two

variables Josh named most sharply — the instruments and the absence of choice. Each is days.

a known quantity, and the A/B against our CC0 palette on an identical score is the single most

informative experiment available to this lane. It also de-risks Pattern B by telling us whether a

neural renderer is even needed, or whether the sampler was simply the wrong sampler.

and the honest position is that the four turndowns have not yet isolated it: Josh has never heard

our composer through good instruments, and he has never heard the best of forty. Those two probes

have to run before we can know whether the generation layer needs replacing or was merely being

heard through a broken window.

selection layers are fixed. If they do, Pattern A is the destination and the care-tier idiom

objection in §9.1 is the first thing that must be tested, on a non-Western cue.

of "no variation or changes or movement" is a stems-and-segments question, and grading a linear WAV

will keep producing that complaint no matter how good the cue gets.

The one-line version: we have been improving the only layer we own, and the layers we do not own are

where every commercial engine keeps its quality. Fix the cheap ones first and let them tell us

whether the expensive one is really the problem.

---

10. HONEST LIMITS OF THIS DOCUMENT

paper at all. Section 4's internals and section 3.1's corpus figures are reconstruction and

secondary reporting, flagged as such at the point of use. The PRODUCT behaviour and the LICENCE

terms are directly sourced and are the parts the recommendation actually rests on.

phrasing comes from the indexed summary of it. Aggregator claims that Boomy uses GANs are

unverified and are not used.

Symphony Rendering, Audiobox Aesthetics, CLaMP 3, ACE-Step 1.5, BBCSO Discover or SWAM is the

vendor's or the paper's claim, not our measurement. That is what the proof steps are for.

licence bodies on disk and commit-SHA pinning, and nothing in §9 enters a pipeline before it goes

through that door. Spitfire's EULA in particular was not retrievable at time of writing and must be

read directly before any Discover or Core render is published, with specific attention to

commercial game-soundtrack use and to any machine-learning restriction of the kind AIVA's terms

carry.

raises and does not answer.

Josh's verbatim grades or read out of the round cards.

11. DERIVATION

DERIVED FROM:

_FALLS / _ROUNDS (round 4), round_1_takedown_note, round_2_takedown_note,

round_3_takedown_note, round_3_note, round_4_note, and the honest_limits_short and

comparison_card axis values quoted in §0 and §7.

master, the dry-and-mastered artifact pair.

commit-SHA pinning, the verbatim licence bodies under docs/licence_records/, and the care

decision to decline VCSL's living-tradition instruments.

§2-§5 — the measured renderer diagnosis this doc cites rather than repeats.

docs/proposals/music/theme_architecture_rows.json — named in §8.2 as the conditioning asset.

bar, no human composer hire, composer_in_loop is the factory's own pass ladder); CLAUDE.md

§"Content bar" and the CVD §17 floor; memory factory-first-resolution-mandate (proof before

spend); memory 5090-box-live-remote-stack (local-first target).

NOT DERIVED (authored judgment, and why it had no canon home):

separate layers; it is this doc's synthesis of the six product architectures surveyed, and it is

offered as an analytical tool rather than as a finding.

selects an architecture, and none of the three has been run.

rank_in_theme and kept_in_theme are hardcoded to 1 on the round-2, round-3 and round-4 emitters

(lines 386, 504, 649), and off the absence of any generate-many step in the pass2-7 realisers. It

is an inference from the code's shape, not a statement any doc makes.

Generated by harness/site/structure_site.py — the URL path is the repo path. review root