ORCHESTRATION_TEXTURE_AND_DENSITY.md

music/ORCHESTRATION_TEXTURE_AND_DENSITY.md

Orchestration, Texture and Density — how many layers, and how they are staged across a cue

CANON SUBORDINATION — this document is PROPOSAL-TIER: it serves canon and never outranks it.
Canon served: the composition floor at docs/spine/DECISIONS_PENDING_JOSH.md (commit da627060) -- the
"dozens of layers" and "strategically place the sounds and layers" clauses.
If this document disagrees with canon, CANON WINS and this document is the defect.

TIER: RESEARCH SYNTHESIS, proposal-tier. Canon is authority; nothing here amends canon. Every

substantive claim carries a source — a URL, or book plus author plus chapter. Where sources disagree

the disagreement is stated. Where a figure is an estimate it is labelled an estimate and its basis is

given. Where something is craft consensus rather than evidence it is labelled consensus. Where our

own pipeline was measured to produce a number, the measurement command is named so it can be re-run.

1. What this document answers

Josh graded round 1 of generated music and turned down all eighteen candidates. His composition floor

(docs/spine/DECISIONS_PENDING_JOSH.md, tail) contains the sentence this lane exists to answer:

each track will be dozens of layers, some an orchestra's worth of instruments, some simpler with

nature and ambience and whistling or humming or harmonizing, no track boring or filler, with the

sounds and layers and harmonies and clashes and ambience and nature and voice and chants and humming

and clapping strategically placed. Two questions fall out of it and both need real numbers.

in discrete parts on the page, in tracks in the session, and in stems delivered to implementation.

entrance schedule of a great three-to-four minute cue actually looks like.

The document answers both, then converts every principle into a PIPELINE HOOK: the specific thing our

stack composes, constrains, or measures. Hooks are marked as CONSTRAINT (encodable before the render,

which is worth more to us) or MEASURE (only checkable after the render).

2. The orchestration canon — who is authoritative for what

Five treatises carry the field. They do not say the same thing and they are not interchangeable.

2.1 Rimsky-Korsakov, Principles of Orchestration (1912, ed. Steinberg)

The only classic treatise that states balance as arithmetic. From the Project Gutenberg text

(https://www.gutenberg.org/files/33900/33900-h/33900-h.htm), Chapter I, "Comparison of resonance in

orchestral groups", page 33: in loud passages the horns are half as strong as the heavy brass, giving

1 trumpet = 1 trombone = 1 tuba = 2 horns; woodwind in forte are twice as weak again, giving

1 horn = 2 clarinets = 2 oboes = 2 flutes = 2 bassoons. For strings against wind in an orchestra of

medium formation, one whole string department equals one wind instrument at piano and two at forte

(Violins I = 1 flute at piano). Chained, that is roughly 1 heavy brass = 2 horns = 4 woodwind, and

one string desk-group in the same weight class as a single wind line.

The same chapter carries the warning that matters more than the table: constant use of compound

timbres in pairs and threes eliminates the characteristics of tone and produces, in his words, a

"dull, neutral texture" (page 33). That is Josh's "stacked noise", diagnosed in 1912.

Steinberg's Editor's Preface (page X) reduces the whole book to one clause: good orchestration means

proper handling of parts. Chapter II, "Melody" (page 36), states the doctrine our lead-line rule

descends from — melody should always stand out in relief from the accompaniment, achieved by

accentuated dynamic shading, selection and contrast of timbres, and strengthening by doubling.

Honest limit on this source: the Gutenberg page is large enough that automated retrieval returned the

front matter and Chapter I reliably but would not return the bodies of Chapter III ("Number of

harmonic parts — Duplication", page 64) or Chapter IV's tutti sections ("Full Tutti" page 101, "Tutti

in the wind" page 103, "Tutti pizzicato" page 103, "Tutti in one, two and three parts" page 104).

Those section titles are confirmed present in the book's own contents; their text is not quoted here

and no claim is made from them. The existence of a chapter enumerating tutti in one, two and three

parts is itself the evidence for the principle below, and it is stated as inference, not as quotation.

to full tutti, tutti in the wind, tutti pizzicato, and tutti in one, two and three parts. A tutti

that meant "all instruments play" would need one section, not four. Craft consensus, in the same

direction and independently sourced, appears in section 5 below from game-scoring practice.

2.2 Walter Piston, Orchestration (Norton, 1955)

Piston is authoritative for the systematic analysis of orchestral texture — the decomposition of a

score into melody, secondary melody, harmonic support, rhythmic support and bass, and the study of

how composers distribute those functions. Mark DeVoto's biographical essay for Tufts calls Piston's

book the best text on the subject in English and says it set a standard for systematic analysis of

orchestral texture that its competitors have not approached

(https://sites.tufts.edu/markdevoto/files/2015/10/Piston.pdf). Use Piston for the question "what job

is this part doing", which is exactly the question our part-role field already asks.

2.3 Samuel Adler, The Study of Orchestration (Norton, 4th ed. 2016)

The modern pedagogical standard: instrument-by-instrument capability, range, articulation and

transposition, taught through score excerpts and listening (https://wwnorton.com/books/9780393920659).

Adler's balance guidance is deliberately less absolute than Rimsky-Korsakov's arithmetic; practitioner

discussion notes that Adler and Koechlin give notes similar in kind but not mathematical, reflecting

the nuance of real ensembles (https://vi-control.net/community/threads/are-rimsky-korsakovs-balance-ratios-still-right-nowadays.91013/).

Sources disagree here, and the disagreement is real rather than an error: Rimsky-Korsakov's ratios are

a usable first approximation for a synthetic mix where every part is a fader, and they overstate their

own precision for a live room. Our renderer is faders, so the ratios are more directly usable to us

than to a live orchestrator — which is a reason to adopt them as a default gain law and then measure.

2.4 Alfred Blatter, Instrumentation and Orchestration (2nd ed.)

The reference-desk book. Strongest on notation, transposition, extended and contemporary techniques,

percussion of American and African origin, electronic instruments and sound modification, with

appendices on MIDI and guitar (https://www.amazon.com/Instrumentation-Orchestration-Alfred-Blatter/dp/0534251870,

https://archive.org/details/instrumentationo0000blat_u5w6). Use Blatter when the question is "can this

instrument physically do this and how is it written", including for percussion and non-orchestral

colour — which is the family our region cells lean on hardest.

2.5 Ertugrul Sevsay, The Cambridge Guide to Orchestration (CUP, 2013)

The practical-exercise book: musical excerpts given in reduction for the reader to orchestrate and

then compare against the original, ordered through the choirs from strings alone toward complex

combinations, with systematic analysis of orchestration technique in original scores including

twentieth-century repertoire (https://www.cambridge.org/core, frontmatter at

https://assets.cambridge.org/97811070/25165/frontmatter/9781107025165_frontmatter.pdf;

https://archive.org/details/cambridgeguideto0000sevs). Sevsay is the closest published analogue to

what our pipeline needs: a reduction plus a target orchestration plus a comparison. That is a training

loop shape, and it is the reason this treatise is named first among the five for our purposes.

manifest rather than hand-set gain_db per part: heavy brass 0 dB reference, horns -6 dB per

instrument to match two-for-one, woodwind -12 dB per instrument, string ensemble sections treated as

one wind-equivalent at piano and two at forte. Our composed tracks currently carry hand-authored

gain_db values (measured: -3.0 lead clarinet, -12.0 counter viola in HF_MT_02). A declared law

makes an out-of-balance mix a rule violation instead of a taste argument.

lane_analysis.interlap_matrix are the after-the-fact check on that law. EX_124 reads low_mid 0.381,

bass 0.290, low_bass 0.151, mid 0.133 — a bass-weighted profile with a clear single peak, not a flat

smear. Flatness across bands is the measurable signature of the dull neutral texture Rimsky-Korsakov

named.

3. Real layer counts — the number Josh's floor is graded against

"Dozens of layers" is only a floor if it has a unit. Four different units are in circulation and they

differ by two orders of magnitude. Naming which one Josh's sentence is graded against is the single

most useful thing this document does.

3.1 Players on the floor (session size)

brass four horns, four trumpets, three trombones, two bass trombones, one contrabass trombone, one

tuba, one contrabass tuba; woodwind piccolo, two flutes, two oboes, cor anglais, two clarinets, bass

clarinet, two bassoons, contrabassoon; percussion three players

(https://www.spitfireaudio.com/en-us/products/abbey-road-one-orchestral-foundations,

https://www.soundonsound.com/reviews/spitfire-audio-abbey-road-one-orchestral-foundations).

(https://filmmusictheory.com/article/orchestrating-the-apocalypse-the-music-of-elden-ring/,

https://blog.playstation.com/2022/09/09/elden-ring-composer-tsukasa-saito-on-creating-the-games-score-and-his-favorite-track/).

Cantorum in Iceland and a 48-singer choir in Prague, plus Nordic folk instruments including

nyckelharpa, hurdy-gurdy and Hardanger fiddle, recorded across Los Angeles, Nashville, London,

Iceland, Germany and Prague

(https://www.billboard.com/music/music-news/bear-mccreary-god-of-war-video-game-score-interview-8256913/,

https://bearmccreary.com/god-of-war/).

Read as instrument LINES rather than bodies, a 90-piece orchestra is roughly 25 to 30 distinct written

parts, because sixteen first violins are one part. That is the number that matters for texture, and it

is already "dozens" — but only just, and only in the largest sessions.

3.2 Discrete parts on the page

An orchestral score of the Abbey Road formation runs about 25 to 32 staves. Extended-technique

repertoire goes far higher by writing every player a separate part: Ligeti's Atmospheres contains a

mirror canon in forty-eight parts, twenty-eight violins descending against twenty violas and cellos

ascending, producing a cluster spanning nearly five octaves

(https://americansymphony.org/concert-notes/atmosphres-1961/, https://en.wikipedia.org/wiki/Micropolyphony);

Penderecki's Threnody is written for 52 strings as 52 individual parts

(https://en.wikipedia.org/wiki/Threnody_to_the_Victims_of_Hiroshima). These are the ceiling cases and

they are deliberately not perceived as counted voices — see section 6.1.

3.3 Tracks in the session (the number closest to Josh's sentence)

This is where the real "dozens" lives, and it is far past dozens.

sometimes more than 2000 tracks, one session per cue

(https://film-mixing.com/2016/07/28/film-score-mixing-with-alan-meyerson/). The same interview

records the structural reason for that count: the orchestra is captured with layered microphone

arrays — Decca Trees at twelve and eight feet, wide cardioids ten to twelve feet to the sides,

hypercardioid surrounds in a star pattern, overheads above sixteen feet — so a single violin section

arrives as many tracks, and synths, percussion and guitars come in as separate prerecords.

composer on VI-Control describes moving from an ~800-track disabled Cubase template to ~3000 tracks

including outputs, about 1500 actual virtual instruments

(https://vi-control.net/community/threads/how-do-you-set-up-your-orchestral-template.68795/).

Labelled consensus-of-practitioners rather than published data.

The honest distinction: a mockup's track count and a live session's part count measure different

things. A 1500-track template is an instrument PALETTE, of which one cue may use thirty. A 2000-track

Meyerson mix is one cue, but most of those tracks are microphone perspectives and overdub passes on a

much smaller set of musical lines. Neither number is "the number of layers a listener hears".

3.4 Overdub multiplication — how a small ensemble becomes a large texture

thumb is at least three passes so that the small take-to-take differences read as ensemble size

(https://studiopros.com/overdub-an-orchestra-section/, https://en.wikipedia.org/wiki/Overdubbing).

Labelled craft consensus.

times, so ninety voices are heard on the final track

(https://en.wikipedia.org/wiki/Dragonborn_(song)).

was produced by five people — Marty O'Donnell, Michael Salvatori and three colleagues from jingle

sessions — layered into a monastic choir

(https://en.wikipedia.org/wiki/Halo_Original_Soundtrack, https://www.halopedia.org/Halo_Theme).

3.5 Stems delivered to implementation

Stem counts are an order of magnitude below track counts and are set per project, not by a standard.

a separate LFE track (https://film-mixing.com/2016/07/28/film-score-mixing-with-alan-meyerson/).

That is a handful of stem GROUPS, each multichannel.

(https://winifredphillips.wpcomstaging.com/2015/09/29/arrangement-for-vertical-layers-pt-1-a-game-composers-guide/,

https://www.gamedeveloper.com/game-platforms/pure-vertical-layering-for-game-music-composers-from-spyder-to-sackboy-gdc-2021-).

want stems rarely want more than about ten

(https://vi-control.net/community/threads/delivering-stems-for-production-music-which-groups.51512/).

Labelled consensus.

sparse bed plus acoustic guitar, percussion, then brass and aggressive strings

(https://www.arcellasound.com/post/interactive-music-in-aaa-games-vertical-layering-vs-horizontal-re-sequencing).

the implementation lead re-mixed and re-cut per scenario in Unreal MetaSounds; no stem count was

disclosed (https://www.asoundeffect.com/clair-obscur-expedition-33-game-audio/). Named here as a

source that is genuinely thin on the number.

3.6 The ruling this section produces

Josh's "dozens of layers" is best read against the PART count, not the track count and not the stem

count. Dozens of parts — call it 24 to 40 discrete musical lines for a full cue, thinning to six or

fewer for the sparse nature-and-humming class he also named — is exactly a real orchestral score, is

achievable by our renderer, and is the number a listener could in principle enumerate if the

orchestration let them. Track count is a recording artefact of microphone arrays we do not have; stem

count is an implementation choice downstream.

the orchestration.parts arrays hold a mean of 9.3 parts, minimum 4, maximum 19 (HF_MT_10_CREDITS).

Re-derivable by reading orchestration.parts from each tracks/*/*.json. Against a 24-to-40 floor

the pipeline is short by roughly a factor of three on the full-cue class, and is already correct for

the deliberately sparse class.

beside the existing voice_count_target, banded by cue class: title and boss-apex 28 to 40,

ordinary battle and region festival 18 to 28, exploration and sanctuary 8 to 16, ambience and

nature-voice cells 3 to 8. Make the composer battery refuse a track whose part manifest falls

outside its band. This is the single highest-leverage change in the document, because it is a

pre-render constraint on the artefact Josh graded.

4. Texture types and their density signatures

Texture is the relationship between simultaneous lines, and each type has a measurable fingerprint.

The type list below is standard music-theoretic taxonomy; the density signatures beside each are

stated as our own operational mapping onto formula-card fields, and are labelled as such.

melodic_salience, register_spacing_octaves short.

repertoire our region cells draw from. Signature: high pairwise band overlap with high melodic

salience — variants share a register by definition.

synchronised across parts, few independent entry and exit events, high interlap_matrix values.

a clearly subordinate register-separated support mass.

scoring because it survives looping. Signature: high motif_economy.interval_3gram_repetition in

the accompaniment band with lower repetition in the lead band.

independent entry and exit, moderate simultaneity, distinct rhythmic profiles per lane.

Signature: high onset rate with low per-lane active_fraction, and high silence budget per lane.

and deliberately inaudible as counterpoint; what is heard is a woven texture in which the individual

lines are hidden (https://en.wikipedia.org/wiki/Micropolyphony). Signature: maximum lane count with

near-total band overlap and a collapsed register-spacing figure.

organic material, and the actual idiom of AAA. Its density signature is a bimodal band profile:

orchestral energy in the low-mid through presence bands, synth and design occupying sub, air and the

spaces between.

from the list above, and let the score generator select its part-writing routine from it. A cue that

declares the same texture type for all its sections is declaring its own monotony before a note is

rendered, and the battery can reject that without listening.

classifier for the achieved type. It cannot distinguish heterophony from unison doubling, and should

not be asked to.

5. Staged entrance architecture — the core deliverable

This is the part amateurs miss, and it is the part Josh's "strategically place the sounds and layers"

sentence is about. Below are documented entrance schedules from four repertoires, then the general

shape they share.

5.1 The textbook additive case — Ravel, Bolero (1928)

Bolero is fifteen minutes long, is built on exactly two melodies, and states them alternately while

adding one new colour per statement over an unvarying snare-drum ostinato repeated 169 times, with

pizzicato strings strumming beneath

(https://thelistenersclub.com/2021/05/03/bolero-ravels-sublime-orchestration-exercise/). The melodic

carrier order runs solo flute, clarinet, bassoon, E-flat clarinet, oboe d'amore, then trumpet with

flute, saxophones, celesta with horn, a quartet of reeds, a portamento trombone, the highest woodwind,

and only then the strings

(https://thelistenersclub.com/2021/05/03/bolero-ravels-sublime-orchestration-exercise/,

https://www.clevelandorchestra.com/posts/ravels-bolero). Four structural facts are transferable:

texture thickens by promotion, not by piling on.

The biggest colour is the last one spent.

to produce artificial overtones

(https://thelistenersclub.com/2021/05/03/bolero-ravels-sublime-orchestration-exercise/). A composite

timbre is introduced as a NEW instrument, not as a doubling.

5.2 The arch case — Barber, Adagio for Strings (1936/38)

Barber opens pianissimo from a single melodic cell in first violin, hands the principal cell to the

cellos for the expansive middle, moves the string choir up the scale into its highest register, hits a

fortissimo climax, and follows it with SILENCE before the restatement

(https://www.parlancechamberconcerts.org/individual-program-notes/samuel-barber-(1910-1981)/adagio-from-string-quartet-no.-1,-op.-11,

https://classicalexburns.com/2022/07/21/samuel-barber-adagio-for-strings-diving-into-an-emotional-abyss/,

https://fcsymphony.org/program-notes/barber-adagio-for-strings/). The transferable law: the loudest

moment is followed by nothing, and the pause is a structural member rather than a gap. Subtraction to

zero is the strongest event available.

5.3 The game-cue cases

orchestra; Limgrave uses feathered bowing (players alternating bow strokes rather than bowing in

unison) to produce a diffuse atmospheric bed, with Scottish Highlands field recordings folded into

the background texture; Rennala's first phase layers music-box elements with processed female vocals

and the second phase brings in full choir and orchestra

(https://filmmusictheory.com/article/orchestrating-the-apocalypse-the-music-of-elden-ring/). Note

that the boss cue's layer escalation is bound to a PHASE change, which is our own boss-phase seam.

Uematsu has said he conceived it as a rock song rather than a choral-orchestral piece; it was

assembled by writing two to four measures a day until he had twenty to thirty fragments and then

rearranging them into an order that worked

(http://nobuouematsu-musicex.blogspot.com/2011/05/one-winged-angel.html,

https://marionmizuuki.wordpress.com/2017/03/06/1st-analysis-one-winged-angel-final-fantasy-vii/).

The transferable point is the assembly method: a MOTIF POOL first, an order second. Our arm-B

four-motif floor is already this method; the pool should be larger than the track needs.

know what the player was doing, with strings and harp added to build texture and variation subtly

(https://daily.redbullmusicacademy.com/2015/08/c418-interview/,

https://musescore.com/news/features/unlocking-the-secrets-of-c418s-sweden-the-minimalist-masterpiece-at-the-heart-of-minecrafts-soundtrack/).

The 2025 National Recording Registry essay on Minecraft: Volume Alpha is the scholarly citation

(https://www.loc.gov/static/programs/national-recording-preservation-board/documents/Minecraft-Volume-Alpha_Grosser.pdf).

This is the class Josh named as "some more simple with nature and ambience and whistling or humming".

5.4 The interactive-layering discipline, and the one prohibition

Winifred Phillips' vertical-layering guidance is the most directly applicable published craft doctrine

we have, because it governs layers that must work both alone and combined — which is our adaptive

requirement (https://winifredphillips.wpcomstaging.com/2015/09/29/arrangement-for-vertical-layers-pt-1-a-game-composers-guide/).

together they are the full composition.

perceived clearly in the full mix.

tutti. This is the independently-sourced companion to the Rimsky-Korsakov inference in section 2.1:

a texture where everyone plays the same rhythm cannot be decomposed into layers, and cannot be

subtracted from.

Corroborating measurement from our own corpus: EX_124, a title cue Josh's own exemplar set holds up as

a target, records event_grammar.counts.tuttis = 0 across 127 seconds, against 14 drops and 20 raises.

A beloved title cue spends its whole length on drops and raises and never once puts everyone in at

once. Three independent lines of evidence — treatise structure, game-audio craft doctrine, and our own

exemplar measurement — arrive at the same rule.

5.5 The general shape, stated as a schedule

Synthesising the four repertoires above into one transferable schedule for a 3:00 to 4:00 cue. This is

our synthesis, labelled as such, and each clause traces to a source above.

inside the first few seconds where the cue class allows one (our own EARWORM rule; EX_124 measures

time_to_hook_s = 0.104 on the Deus Ex title cue).

promotion law). Part count roughly doubles. Full texture is still far off; EX_124's

time_to_full_texture_s is 11.56 on a 127-second cue, which is early and is a property of that

title-cue class rather than a universal.

support. Harmonic rhythm may quicken. This is where the 3-to-5-cycle variation law bites: a

repetition must acquire a new instrument, hook, melody or beat event before the fourth pass.

part that has been present but masked. This is the move that makes the return feel large without

adding anything new.

where the withheld colour is spent — Bolero's strings, Rennala's phase-two choir, Elden Ring's full

orchestra after the sparse chorale.

a single sustained element. EX_124's event_grammar.silence_budget.fraction of 0.0232 across the

whole track shows how small the absolute quantity of silence is even when it is structurally load-

bearing; the value is in placement, not duration.

orchestration.parts gains entry_bar, exit_bar and optionally a list of rest_spans. The part

manifest stops being a roster and becomes a SCHEDULE. Once it is a schedule, three things become

checkable before rendering: no part is active for the whole cue except a declared ostinato; at least

one subtraction event of at least N parts occurs after the midpoint; the peak part count occurs in

the last third. Every one of those is a Josh-floor clause turned into a predicate.

a carrier that just finished — Bolero's law, made explicit and countable.

flow_curve.peak_position_normalised, event_grammar.drops/raises/breaks with their

lanes_added/lanes_removed deltas, and dynamics.peak_position_normalised verify the schedule

after the fact. EX_124's flow peak sits at 0.272 and its dynamic peak at 0.4595 — the energy peak

and the loudness peak are in different places, which is itself a finding worth carrying to rung 2.

6. Density management — why stacked noise happens, and what prevents it

6.1 The perceptual ceiling on simultaneous streams

The strongest evidence available, and it is directly on point. David Huron, "Voice Denumerability in

Polyphonic Music of Homogeneous Timbres", Music Perception 6(4), 1989, pages 361 to 382

(https://online.ucpress.edu/mp/article/6/4/361/62869/Voice-Denumerability-in-Polyphonic-Music-of):

expert musicians become slower to detect added voices and less accurate at counting them as the number

rises, and accuracy drops markedly at the step from three voices to four. Huron's conclusion is that

the auditory system follows a one-two-three-or-many rule and that it may be impossible to process more

than about four concurrent streams.

Two consequences, and they point in opposite directions on purpose.

perceptual STREAMS. Bolero's accompaniment mass is one stream regardless of how many players are in

it. This is the reconciliation of "dozens of layers" with the four-stream ceiling, and it is the

single most important idea in this document.

defect Josh named is a GROUPING failure, not a count failure.

The mechanism behind grouping is codified in Stephen McAdams, Meghan Goodchild and Kit Soden,

"A Taxonomy of Orchestral Grouping Effects Derived from Principles of Auditory Perception", Music

Theory Online 28.3, 2022 (https://mtosmt.org/issues/mto.22.28.3/mto.22.28.3.mcadams.html). Its three

classes and their operative subtypes:

timbral AUGMENTATION, where a dominant instrument is coloured by a subordinate one [4.5]; timbral

EMERGENCE, where the fusion produces a new timbre identifiable as none of its constituents [4.9];

timbral HETEROGENEITY, where parts group but do not fully blend and some instruments stay audible

as themselves [4.13]. Fusion is strengthened by onset synchrony, harmonicity, and parallel changes

in amplitude and frequency [3.7].

stream SEGREGATION, defined as two or more clearly distinguishable voices of near-equivalent

prominence [5.12], and STRATIFICATION, defined as layers separated into more and less prominent

strands [5.16], on the other. Stratification prominence is analysed via Koechlin's extensity

(auditory size) and intensite (inherent force) [5.18].

juxtapositions and sectional boundaries. Segmentation strength rises when several parameters change

together [6.1]. The taxonomy notes explicitly that none of these effects is all-or-nothing [7.4].

Stratification is the concept our pipeline is missing by name. A forty-part cue with three declared

strata — foreground carrier, middleground counter-material, background bed — satisfies both Josh's

floor and Huron's ceiling at once.

background, and enforce three to four active strata rather than three to four active parts. Enforce

that no stratum holds more than about 60 percent of the summed gain, and that the foreground stratum

is never the largest by part count.

5.455, max 9). Read as STRATA that is far too many; read as spectral bands it is unremarkable. The

card cannot currently tell the two readings apart, which is exactly the limit section 10 addresses.

6.2 Frequency-slot allocation

Every part needs a lane, and lanes are finite. The register-spacing figure on the formula card is the

existing proxy: EX_124's register_spacing_octaves reads 0.99, 0.99, 0.99, 0.99, 1.0, 1.0 — an almost

perfectly even one-octave spacing between its active bands, which is what a well-slotted arrangement

looks like. Practitioner reports on orchestral mixing converge on 250 to 500 Hz as the accumulation

zone where full-orchestra energy piles into boxiness, with cuts of 6 to 8 dB reported as necessary

(https://vi-control.net/community/threads/your-go-to-eq-tricks-and-tips-when-mixing-orchestral-music.40771/,

https://www.waves.com/tips-for-mixing-film-tv-scores-free-presets). Labelled craft consensus, and it

matches the low_mid band being EX_124's largest at 0.381 of total energy.

Meyerson's stated rule is spatial rather than spectral and is worth carrying: do not build mixes in

the middle, and use small time delays in the Haas range of roughly 150 to 250 samples to separate

elements that would otherwise mask each other

(https://film-mixing.com/2016/07/28/film-score-mixing-with-alan-meyerson/). Panning and micro-delay

are a third axis of separation beyond register and time, and our renderer already has pan per part.

6.3 Masking, stated precisely

Masking is the process by which the threshold of hearing of one sound is raised by the presence of

another, raising the masked detection threshold — the minimum intensity a sound needs to stay audible

(Amelie Bernier-Robert and Ben Duinker, "Masking", Timbre and Orchestration Resource, 13 November

2023, https://timbreandorchestration.org/writings/timbre-lingo/masking). Two kinds:

processed the masked sound, so it never reaches higher processing.

masker, by attention, and by auditory-scene-analysis grouping.

The same source states plainly that predicting masking is very complex, because a composer must

account for each masking type per component and for interactions between maskers, which may reduce

each other or combine to increase total masking through cochlear distortion. Critical-band theory

supplies the frequency geometry: the auditory system integrates energy over critical bands on the Bark

scale, band width grows with centre frequency while covering constant distance on the basilar

membrane, and simultaneous masking is strongest within the masker's own band

(https://support.ircam.fr/docs/AudioSculpt/3.0/co/Masking%20Effect%20Intro.html).

Stacked noise, defined acoustically: it is what happens when many parts share critical bands, share

onsets, and share amplitude envelopes. Shared onsets and parallel amplitude change are precisely the

cues that FUSE events (McAdams et al. [3.7]) — so a homophonic tutti of many instruments in one

register is maximally fused and maximally masked at the same time. Every part contributes energy and

almost none contributes information. That is one sentence with three independent citations behind it,

and it is the technical answer to Josh's "there can't just be stacked noise".

score in which more than two parts hold the same slot in the same bar unless one is explicitly

marked as a doubling of the other. Refuse a bar in which more than a declared fraction of parts share

an onset, outside sections whose texture_type is homophonic.

EX_124's low_mid|mid pair reads 0.969 and mid|air 0.907 — very high overlap that a band-based

measure cannot distinguish from good octave doubling. Treat high interlap as a FLAG requiring the

score-side slot declaration to justify it, never as a verdict on its own.

7. Orchestrational variety as the anti-boredom device

Josh's 3-to-5-cycle rule and the treatises agree: repetition is not the problem, unvaried repetition

is. Bolero repeats two melodies for fifteen minutes and is not boring because the COLOUR changes every

statement (section 5.1). The standard vocabulary of colour change, drawn from the treatise tradition

and from the McAdams segmental categories, is small enough to enumerate and therefore small enough to

encode.

timbre covaries with pitch, playing effort and articulation (McAdams et al. [5.1]).

extensity in Koechlin's sense [5.18].

Limgrave device).

because it changes meaning rather than surface.

horn plus celesta organ-mixture does. This is McAdams' timbral EMERGENCE [4.9] used deliberately.

the first [6.4, 6.5].

declared law to license one. Extend that: when a repeat IS licensed, require a

colour_change field naming which device from the list above is applied, and make the absence of a

named device a battery failure. That converts the 3-to-5-cycle rule from a hope into a predicate,

and it is enforceable entirely at composition time.

measure whether repetition happened; they cannot see whether it was re-coloured. The constraint is

the real control here and the measure is only corroboration.

8. The non-orchestral layers Josh named

Ambience, nature, voice, chant, humming, whistling, clapping. Treated as musical parts, not decoration.

8.1 Field recordings and nature as pitched material

Elden Ring's Limgrave folds Scottish Highlands field recordings into the background texture

(https://filmmusictheory.com/article/orchestrating-the-apocalypse-the-music-of-elden-ring/). The

technique tradition runs from Schaeffer's musique concrete through contemporary practice; the standard

production move is to pitch-shift a recording so it functions as a pad or lead rather than as ambience,

with even a few semitones of shift changing its character

(https://blog.landr.com/field-recording-tips/, https://en.wikipedia.org/wiki/Pitch_shifting). Real-time

ensemble performance of field-recorded environmental sound is an active research area, which is

evidence that treating recordings as instruments is a defined practice rather than a metaphor

(https://arxiv.org/pdf/2006.09645).

The doctrine our lane should adopt: a nature layer is TUNED to the cue's key or it is a sound effect.

A river, a wind, a cicada bed and a fire all have a spectral centre; shifting that centre onto the

tonic or fifth makes the ambience a drone part with a register slot, subject to the same masking rules

as any other part. Where an ambience is unpitched by nature — footsteps, rain on stone, a rope creak —

its musical function is RHYTHMIC and it belongs on the grid.

8.2 Voice, chant and humming as parts

designed to rhyme in both that language and in English

(https://en.wikipedia.org/wiki/Dragonborn_(song)).

(https://en.wikipedia.org/wiki/Halo_Original_Soundtrack).

(https://www.billboard.com/music/music-news/bear-mccreary-god-of-war-video-game-score-interview-8256913/).

arrives in phase two

(https://filmmusictheory.com/article/orchestrating-the-apocalypse-the-music-of-elden-ring/).

the part instead

(https://blog.playstation.com/2022/09/09/elden-ring-composer-tsukasa-saito-on-creating-the-games-score-and-his-favorite-track/).

That is carrier substitution at production time, and it is the exact device section 7 lists first.

A hum and a whistle are SOLO WIND parts with an unusual timbre: narrow dynamic range, limited range,

strong fundamental, weak upper partials. They sit cleanly in a register slot and mask almost nothing,

which is why they work as the top line over a sparse bed and why C418's class of cue can carry one.

8.3 Clapping and body percussion as rhythmic parts

Flamenco palmas is the most fully theorised body-percussion practice available and its distinctions are

directly encodable. Palmas sordas are cupped and muffled, sitting under the music without cutting

through; palmas claras are sharp and bright and mark accents unmistakably; palmeros reach for sordas

when they want to hold the compas present without competing with the singer, and for claras when an

accent must be unmissable. Contratiempo is clapping on the off-beat

(https://www.siudyflamenco.org/blog/palmas-and-compas-flamenco-rhythm,

https://www.grangalaflamenco.com/en/blog/types-of-palmas-in-flamenco-and-their-sounds/,

https://en.wikipedia.org/wiki/Palmas_(music)). Two parameters — timbre bright or dark, placement on or

off the beat — give four distinct rhythmic parts from a pair of hands, each with its own frequency

slot. That is a real orchestration resource, and it costs nothing to render.

register_slot, stratum, entry_bar, exit_bar, role. A nature bed with no register slot and

no entry schedule is decoration; with them it is a part. Add tuned_to on ambience parts naming the

scale degree its spectral centre is shifted onto.

because a family head realised on an mbira would be scoring a living tradition against a card that

has not authorised it. The same rule applies with equal force to chant, to language, and to

culturally-specific clapping traditions: the region page's own instrumentation section is the

licence, and palmas belongs to an Andalusian cell rather than to any cue that wants a clap. Where a

cue simply needs body percussion with no cultural claim attached, it is written as body percussion

and named as such.

9. What a doctorate drills that a self-taught pipeline misses

Stated as the gap list, each item with the section above that supplies the fix.

orchestrate, then compare against the original — is a supervised training loop, and it is the one

thing on this list our pipeline could literally automate (section 2.5).

reflexive; a self-taught mix reaches for a fader instead (section 2.1).

support and bass is applied to every bar of every score studied. A part with no declared function is

the commonest amateur defect (section 2.2).

cannot be reduced is usually a texture with no strata (section 6.1).

as on adding; the amateur instinct is monotone accumulation (section 5.2 and 5.5).

octaves appearing by accident between two instruments neither of which is doubling, and inner parts

that leap because nobody tracked them.

an instrument's weak register sounds wrong even when rendered by samples, because the sample set was

recorded from a real player (section 2.4).

Adagio's silence, Ligeti's hidden canon — these are moves in a shared vocabulary, and a composer who

does not know them re-invents the weaker version.

and they are the difference between "add more" and "group better" (section 6.1).

10. The honest measurement limit, and how we should actually count layers

10.1 The limit, stated plainly

Our formula card's lanes are TEN SPECTRAL BANDS — sub, low_bass, bass, low_mid, mid, upper_mid,

presence, brilliance, air, ultra — not instruments. The card says so itself in

lane_analysis.method_note: two instruments in one octave read as one lane, one instrument spanning

two octaves reads as two, and the card declares its confidence as LOW-FOR-ROSTER,

USABLE-FOR-ARRANGEMENT. The ceiling is therefore ten, the observed maximum on EX_124 is nine, and Josh

is asking for dozens. No amount of tuning makes a ten-band measure report a thirty-part texture. The

card must never be quoted as evidence about layer count, and any claim that it was is a defect.

What the band measure IS good for stands undiminished: entry and exit shapes, the simultaneity curve's

form over time, register spacing, interlap as a masking flag, and the event grammar's drops and

raises. Those are arrangement measures and they are trustworthy as such.

10.2 The four candidate ways to count instruments from audio, ranked

bass, other, vocals — with a six-source variant adding guitar and piano, and its own documentation

notes the piano source underperforms and that leakage and artefacts are normal in complex mixes

(https://github.com/adefossez/demucs, https://github.com/facebookresearch/demucs). Verdict: it counts

to four or six, not to thirty. It cannot answer the question.

spectra and time-varying gains and is a standard transcription and separation front end

(https://www.ee.columbia.edu/~dpwe/e6820/papers/SmarB03-nmf.pdf). The component count is a MODEL

ORDER chosen by the analyst, and model-order selection is itself an open research problem addressed

by model averaging over several orders (as in the ArXiv literature on multiple-order NMF,

https://arxiv.org/pdf/1605.07469 and related). Verdict: it returns whatever count you asked for. A

measurement whose answer is its own input parameter is not a measurement.

and multi-instrument multi-pitch estimation are active MIR fields; polyphony can be obtained

implicitly by counting active predicted pitches or modelled explicitly as local polyphony

(https://link.springer.com/article/10.1186/s13636-025-00398-2,

https://arxiv.org/pdf/1811.01143). Verdict: the most honest audio-side option, giving a defensible

estimate of concurrent PITCHED VOICES, which is closer to Huron's unit than to a part count. Worth

building as a corroborating measure and never as the primary one.

free, and already three-quarters built here.

10.3 The recommendation

Count layers at COMPOSITION time, from the part manifest, and treat every audio-side figure as

corroboration only.

Our pipeline already emits orchestration.parts per composed track, with part, instrument, role,

gain_db, pan, detune_cents and notes per entry, and the identity card already carries

instrumentation and voice_count_target. The measured state today is a mean of 9.3 parts per track

across the ten arm-B tracks, ranging 4 to 19. The gap to a real orchestral texture is therefore known

exactly, in the right unit, before anything is rendered — which is the whole argument for this

recommendation.

The proposed STEM MANIFEST schema, as the additions to the existing parts entries:

fieldtypepurpose
stratumforeground / middleground / backgroundHuron and McAdams grouping; three to four strata regardless of part count
register_slotoctave band identifierfrequency-lane allocation and the masking predicate
entry_bar / exit_barintthe staged entrance schedule; makes subtraction checkable
rest_spanslist of bar pairsrest as a first-class parameter per part
promoted_frompart idBolero's carrier-to-accompaniment promotion, made countable
doublespart idlicenses two parts in one slot; absence makes co-slotting a violation
colour_changedevice namenames which variety device licenses a repeat
texture_typeper form sectiondeclares the density regime the section is writing in
tuned_toscale degreefor ambience and nature parts, the pitch centre they are shifted onto
cultural_licenceregion page referencecare-line: which page's instrumentation section authorises this part

With that schema, layer_count is a field, not an inference — and every clause of Josh's floor

sentence becomes a predicate a battery can enforce before a single sample is rendered.

11. The eight actionable principles

1. Count layers as PARTS, not tracks and not stems: 24 to 40 for a full cue, 3 to 8 for the sparse

nature class. HOOK: part_count_floor/ceiling per cue class on the identity card; battery refuses

out-of-band. Our measured mean today is 9.3.

2. Dozens of parts must group into three or four STRATA, because listeners cannot count past about

four concurrent voices (Huron 1989). HOOK: stratum per part; enforce strata count, not part count.

3. A tutti is a balance construction, never everyone playing the same rhythm; the full homophonic tutti

is the one arrangement that cannot be layered or subtracted (Phillips). HOOK: onset-synchrony

fraction ceiling per bar outside declared homophonic sections.

4. Balance is arithmetic first — 1 heavy brass = 2 horns = 4 woodwind (Rimsky-Korsakov, p. 33). HOOK:

default gain_db law derived from the ratios; hand values become declared deviations.

5. The part manifest must be a SCHEDULE, not a roster: entry, exit and rest per part, with the peak in

the last third and at least one real subtraction after the midpoint (Bolero, Barber). HOOK:

entry_bar/exit_bar/rest_spans plus three schedule predicates.

6. Subtraction is the strongest event available, and the loudest moment is followed by the least

(Barber's climax into silence). HOOK: require at least one lanes_removed event of declared depth

after the flow peak; corroborate with event_grammar.drops and silence_budget.

7. Repetition is licensed only by a NAMED colour change from the enumerated device list — carrier,

family, register, solo/tutti, mute, attack mode, articulation, counter-line, reharmonisation,

composite timbre, timbral echo. HOOK: colour_change required on every licensed verbatim repeat;

absence is a battery failure. This is the 3-to-5-cycle law made mechanical.

8. Nature, voice, hum, whistle and clap are PARTS with register slots and entry schedules, and pitched

ambience is tuned to the cue's key. HOOK: identical manifest fields for non-orchestral parts, plus

tuned_to, plus cultural_licence naming the region page that authorises the material.

12. Source note

Every external claim above carries its URL or its book-and-chapter inline, so no separate bibliography

is repeated here. Source-strength tiers used throughout: PUBLISHED RESEARCH (Huron 1989; McAdams,

Goodchild and Soden 2022; Bernier-Robert and Duinker 2023; the NMF, Demucs and multi-pitch

literature); TREATISE (Rimsky-Korsakov, Piston, Adler, Blatter, Sevsay, Phillips); PRACTITIONER

TESTIMONY (Meyerson, McCreary, Saitoh, Joffres, O'Donnell, C418); and CRAFT CONSENSUS, which is the

VI-Control and mixing-tips material and is labelled as such at every point of use.

Repository artefacts read: build/audio/exemplars/formula_cards/EX_124.json (every card figure quoted);

build/audio/complete_tracks/tracks/*/*.json (the 9.3-part measurement);

build/audio/complete_tracks/cards/HF_MT_02_EXPLORATION_TRAVERSAL.json; harness/music_gen/orchestrate.py,

sfz_palette.py, theme_compositions.py, compose_arm_b.py, to_midi.py, stem_probe.py;

docs/spine/DECISIONS_PENDING_JOSH.md (the composition floor).

Thinnest evidence, named honestly: per-cue stem counts for shipped AAA titles are largely undisclosed

(the Expedition 33 interview is the clearest case of a detailed audio postmortem that gives no number);

Rimsky-Korsakov's tutti chapters could not be retrieved in full and no claim is quoted from them; and

no published source was found that states a track or part count for a single named game cue, so

section 3.3's figures come from film scoring and from composer templates rather than from game audio

directly.

Generated by harness/site/structure_site.py — the URL path is the repo path. review root