music/ORCHESTRATION_TEXTURE_AND_DENSITY.md
CANON SUBORDINATION — this document is PROPOSAL-TIER: it serves canon and never outranks it.
Canon served: the composition floor atdocs/spine/DECISIONS_PENDING_JOSH.md(commitda627060) -- the
"dozens of layers" and "strategically place the sounds and layers" clauses.
If this document disagrees with canon, CANON WINS and this document is the defect.
TIER: RESEARCH SYNTHESIS, proposal-tier. Canon is authority; nothing here amends canon. Every
substantive claim carries a source — a URL, or book plus author plus chapter. Where sources disagree
the disagreement is stated. Where a figure is an estimate it is labelled an estimate and its basis is
given. Where something is craft consensus rather than evidence it is labelled consensus. Where our
own pipeline was measured to produce a number, the measurement command is named so it can be re-run.
Josh graded round 1 of generated music and turned down all eighteen candidates. His composition floor
(docs/spine/DECISIONS_PENDING_JOSH.md, tail) contains the sentence this lane exists to answer:
each track will be dozens of layers, some an orchestra's worth of instruments, some simpler with
nature and ambience and whistling or humming or harmonizing, no track boring or filler, with the
sounds and layers and harmonies and clashes and ambience and nature and voice and chants and humming
and clapping strategically placed. Two questions fall out of it and both need real numbers.
in discrete parts on the page, in tracks in the session, and in stems delivered to implementation.
entrance schedule of a great three-to-four minute cue actually looks like.
The document answers both, then converts every principle into a PIPELINE HOOK: the specific thing our
stack composes, constrains, or measures. Hooks are marked as CONSTRAINT (encodable before the render,
which is worth more to us) or MEASURE (only checkable after the render).
Five treatises carry the field. They do not say the same thing and they are not interchangeable.
The only classic treatise that states balance as arithmetic. From the Project Gutenberg text
(https://www.gutenberg.org/files/33900/33900-h/33900-h.htm), Chapter I, "Comparison of resonance in
orchestral groups", page 33: in loud passages the horns are half as strong as the heavy brass, giving
1 trumpet = 1 trombone = 1 tuba = 2 horns; woodwind in forte are twice as weak again, giving
1 horn = 2 clarinets = 2 oboes = 2 flutes = 2 bassoons. For strings against wind in an orchestra of
medium formation, one whole string department equals one wind instrument at piano and two at forte
(Violins I = 1 flute at piano). Chained, that is roughly 1 heavy brass = 2 horns = 4 woodwind, and
one string desk-group in the same weight class as a single wind line.
The same chapter carries the warning that matters more than the table: constant use of compound
timbres in pairs and threes eliminates the characteristics of tone and produces, in his words, a
"dull, neutral texture" (page 33). That is Josh's "stacked noise", diagnosed in 1912.
Steinberg's Editor's Preface (page X) reduces the whole book to one clause: good orchestration means
proper handling of parts. Chapter II, "Melody" (page 36), states the doctrine our lead-line rule
descends from — melody should always stand out in relief from the accompaniment, achieved by
accentuated dynamic shading, selection and contrast of timbres, and strengthening by doubling.
Honest limit on this source: the Gutenberg page is large enough that automated retrieval returned the
front matter and Chapter I reliably but would not return the bodies of Chapter III ("Number of
harmonic parts — Duplication", page 64) or Chapter IV's tutti sections ("Full Tutti" page 101, "Tutti
in the wind" page 103, "Tutti pizzicato" page 103, "Tutti in one, two and three parts" page 104).
Those section titles are confirmed present in the book's own contents; their text is not quoted here
and no claim is made from them. The existence of a chapter enumerating tutti in one, two and three
parts is itself the evidence for the principle below, and it is stated as inference, not as quotation.
to full tutti, tutti in the wind, tutti pizzicato, and tutti in one, two and three parts. A tutti
that meant "all instruments play" would need one section, not four. Craft consensus, in the same
direction and independently sourced, appears in section 5 below from game-scoring practice.
Piston is authoritative for the systematic analysis of orchestral texture — the decomposition of a
score into melody, secondary melody, harmonic support, rhythmic support and bass, and the study of
how composers distribute those functions. Mark DeVoto's biographical essay for Tufts calls Piston's
book the best text on the subject in English and says it set a standard for systematic analysis of
orchestral texture that its competitors have not approached
(https://sites.tufts.edu/markdevoto/files/2015/10/Piston.pdf). Use Piston for the question "what job
is this part doing", which is exactly the question our part-role field already asks.
The modern pedagogical standard: instrument-by-instrument capability, range, articulation and
transposition, taught through score excerpts and listening (https://wwnorton.com/books/9780393920659).
Adler's balance guidance is deliberately less absolute than Rimsky-Korsakov's arithmetic; practitioner
discussion notes that Adler and Koechlin give notes similar in kind but not mathematical, reflecting
the nuance of real ensembles (https://vi-control.net/community/threads/are-rimsky-korsakovs-balance-ratios-still-right-nowadays.91013/).
Sources disagree here, and the disagreement is real rather than an error: Rimsky-Korsakov's ratios are
a usable first approximation for a synthetic mix where every part is a fader, and they overstate their
own precision for a live room. Our renderer is faders, so the ratios are more directly usable to us
than to a live orchestrator — which is a reason to adopt them as a default gain law and then measure.
The reference-desk book. Strongest on notation, transposition, extended and contemporary techniques,
percussion of American and African origin, electronic instruments and sound modification, with
appendices on MIDI and guitar (https://www.amazon.com/Instrumentation-Orchestration-Alfred-Blatter/dp/0534251870,
https://archive.org/details/instrumentationo0000blat_u5w6). Use Blatter when the question is "can this
instrument physically do this and how is it written", including for percussion and non-orchestral
colour — which is the family our region cells lean on hardest.
The practical-exercise book: musical excerpts given in reduction for the reader to orchestrate and
then compare against the original, ordered through the choirs from strings alone toward complex
combinations, with systematic analysis of orchestration technique in original scores including
twentieth-century repertoire (https://www.cambridge.org/core, frontmatter at
https://assets.cambridge.org/97811070/25165/frontmatter/9781107025165_frontmatter.pdf;
https://archive.org/details/cambridgeguideto0000sevs). Sevsay is the closest published analogue to
what our pipeline needs: a reduction plus a target orchestration plus a comparison. That is a training
loop shape, and it is the reason this treatise is named first among the five for our purposes.
manifest rather than hand-set gain_db per part: heavy brass 0 dB reference, horns -6 dB per
instrument to match two-for-one, woodwind -12 dB per instrument, string ensemble sections treated as
one wind-equivalent at piano and two at forte. Our composed tracks currently carry hand-authored
gain_db values (measured: -3.0 lead clarinet, -12.0 counter viola in HF_MT_02). A declared law
makes an out-of-balance mix a rule violation instead of a taste argument.
register.octave_band_profile and lane_analysis.interlap_matrix are the after-the-fact check on that law. EX_124 reads low_mid 0.381,
bass 0.290, low_bass 0.151, mid 0.133 — a bass-weighted profile with a clear single peak, not a flat
smear. Flatness across bands is the measurable signature of the dull neutral texture Rimsky-Korsakov
named.
"Dozens of layers" is only a floor if it has a unit. Four different units are in circulation and they
differ by two orders of magnitude. Naming which one Josh's sentence is graded against is the single
most useful thing this document does.
brass four horns, four trumpets, three trombones, two bass trombones, one contrabass trombone, one
tuba, one contrabass tuba; woodwind piccolo, two flutes, two oboes, cor anglais, two clarinets, bass
clarinet, two bassoons, contrabassoon; percussion three players
(https://www.spitfireaudio.com/en-us/products/abbey-road-one-orchestral-foundations,
https://www.soundonsound.com/reviews/spitfire-audio-abbey-road-one-orchestral-foundations).
(https://filmmusictheory.com/article/orchestrating-the-apocalypse-the-music-of-elden-ring/,
https://blog.playstation.com/2022/09/09/elden-ring-composer-tsukasa-saito-on-creating-the-games-score-and-his-favorite-track/).
Cantorum in Iceland and a 48-singer choir in Prague, plus Nordic folk instruments including
nyckelharpa, hurdy-gurdy and Hardanger fiddle, recorded across Los Angeles, Nashville, London,
Iceland, Germany and Prague
(https://www.billboard.com/music/music-news/bear-mccreary-god-of-war-video-game-score-interview-8256913/,
https://bearmccreary.com/god-of-war/).
Read as instrument LINES rather than bodies, a 90-piece orchestra is roughly 25 to 30 distinct written
parts, because sixteen first violins are one part. That is the number that matters for texture, and it
is already "dozens" — but only just, and only in the largest sessions.
An orchestral score of the Abbey Road formation runs about 25 to 32 staves. Extended-technique
repertoire goes far higher by writing every player a separate part: Ligeti's Atmospheres contains a
mirror canon in forty-eight parts, twenty-eight violins descending against twenty violas and cellos
ascending, producing a cluster spanning nearly five octaves
(https://americansymphony.org/concert-notes/atmosphres-1961/, https://en.wikipedia.org/wiki/Micropolyphony);
Penderecki's Threnody is written for 52 strings as 52 individual parts
(https://en.wikipedia.org/wiki/Threnody_to_the_Victims_of_Hiroshima). These are the ceiling cases and
they are deliberately not perceived as counted voices — see section 6.1.
This is where the real "dozens" lives, and it is far past dozens.
sometimes more than 2000 tracks, one session per cue
(https://film-mixing.com/2016/07/28/film-score-mixing-with-alan-meyerson/). The same interview
records the structural reason for that count: the orchestra is captured with layered microphone
arrays — Decca Trees at twelve and eight feet, wide cardioids ten to twelve feet to the sides,
hypercardioid surrounds in a star pattern, overheads above sixteen feet — so a single violin section
arrives as many tracks, and synths, percussion and guitars come in as separate prerecords.
composer on VI-Control describes moving from an ~800-track disabled Cubase template to ~3000 tracks
including outputs, about 1500 actual virtual instruments
(https://vi-control.net/community/threads/how-do-you-set-up-your-orchestral-template.68795/).
Labelled consensus-of-practitioners rather than published data.
The honest distinction: a mockup's track count and a live session's part count measure different
things. A 1500-track template is an instrument PALETTE, of which one cue may use thirty. A 2000-track
Meyerson mix is one cue, but most of those tracks are microphone perspectives and overdub passes on a
much smaller set of musical lines. Neither number is "the number of layers a listener hears".
thumb is at least three passes so that the small take-to-take differences read as ensemble size
(https://studiopros.com/overdub-an-orchestra-section/, https://en.wikipedia.org/wiki/Overdubbing).
Labelled craft consensus.
times, so ninety voices are heard on the final track
(https://en.wikipedia.org/wiki/Dragonborn_(song)).
was produced by five people — Marty O'Donnell, Michael Salvatori and three colleagues from jingle
sessions — layered into a monastic choir
(https://en.wikipedia.org/wiki/Halo_Original_Soundtrack, https://www.halopedia.org/Halo_Theme).
Stem counts are an order of magnitude below track counts and are set per project, not by a standard.
a separate LFE track (https://film-mixing.com/2016/07/28/film-score-mixing-with-alan-meyerson/).
That is a handful of stem GROUPS, each multichannel.
(https://winifredphillips.wpcomstaging.com/2015/09/29/arrangement-for-vertical-layers-pt-1-a-game-composers-guide/,
https://www.gamedeveloper.com/game-platforms/pure-vertical-layering-for-game-music-composers-from-spyder-to-sackboy-gdc-2021-).
want stems rarely want more than about ten
(https://vi-control.net/community/threads/delivering-stems-for-production-music-which-groups.51512/).
Labelled consensus.
sparse bed plus acoustic guitar, percussion, then brass and aggressive strings
(https://www.arcellasound.com/post/interactive-music-in-aaa-games-vertical-layering-vs-horizontal-re-sequencing).
the implementation lead re-mixed and re-cut per scenario in Unreal MetaSounds; no stem count was
disclosed (https://www.asoundeffect.com/clair-obscur-expedition-33-game-audio/). Named here as a
source that is genuinely thin on the number.
Josh's "dozens of layers" is best read against the PART count, not the track count and not the stem
count. Dozens of parts — call it 24 to 40 discrete musical lines for a full cue, thinning to six or
fewer for the sparse nature-and-humming class he also named — is exactly a real orchestral score, is
achievable by our renderer, and is the number a listener could in principle enumerate if the
orchestration let them. Track count is a recording artefact of microphone arrays we do not have; stem
count is an implementation choice downstream.
the orchestration.parts arrays hold a mean of 9.3 parts, minimum 4, maximum 19 (HF_MT_10_CREDITS).
Re-derivable by reading orchestration.parts from each tracks/*/*.json. Against a 24-to-40 floor
the pipeline is short by roughly a factor of three on the full-cue class, and is already correct for
the deliberately sparse class.
part_count_floor and part_count_ceiling on the identity card beside the existing voice_count_target, banded by cue class: title and boss-apex 28 to 40,
ordinary battle and region festival 18 to 28, exploration and sanctuary 8 to 16, ambience and
nature-voice cells 3 to 8. Make the composer battery refuse a track whose part manifest falls
outside its band. This is the single highest-leverage change in the document, because it is a
pre-render constraint on the artefact Josh graded.
Texture is the relationship between simultaneous lines, and each type has a measurable fingerprint.
The type list below is standard music-theoretic taxonomy; the density signatures beside each are
stated as our own operational mapping onto formula-card fields, and are labelled as such.
simultaneity_curve.mean, high melodic_salience, register_spacing_octaves short.
repertoire our region cells draw from. Signature: high pairwise band overlap with high melodic
salience — variants share a register by definition.
synchronised across parts, few independent entry and exit events, high interlap_matrix values.
a clearly subordinate register-separated support mass.
scoring because it survives looping. Signature: high motif_economy.interval_3gram_repetition in
the accompaniment band with lower repetition in the lead band.
independent entry and exit, moderate simultaneity, distinct rhythmic profiles per lane.
Signature: high onset rate with low per-lane active_fraction, and high silence budget per lane.
and deliberately inaudible as counterpoint; what is heard is a woven texture in which the individual
lines are hidden (https://en.wikipedia.org/wiki/Micropolyphony). Signature: maximum lane count with
near-total band overlap and a collapsed register-spacing figure.
organic material, and the actual idiom of AAA. Its density signature is a bimodal band profile:
orchestral energy in the low-mid through presence bands, synth and design occupying sub, air and the
spaces between.
texture_type field per FORM SECTION on the identity card, drawnfrom the list above, and let the score generator select its part-writing routine from it. A cue that
declares the same texture type for all its sections is declaring its own monotony before a note is
rendered, and the battery can reject that without listening.
texture_and_density block plus lane_analysis gives a coarseclassifier for the achieved type. It cannot distinguish heterophony from unison doubling, and should
not be asked to.
This is the part amateurs miss, and it is the part Josh's "strategically place the sounds and layers"
sentence is about. Below are documented entrance schedules from four repertoires, then the general
shape they share.
Bolero is fifteen minutes long, is built on exactly two melodies, and states them alternately while
adding one new colour per statement over an unvarying snare-drum ostinato repeated 169 times, with
pizzicato strings strumming beneath
(https://thelistenersclub.com/2021/05/03/bolero-ravels-sublime-orchestration-exercise/). The melodic
carrier order runs solo flute, clarinet, bassoon, E-flat clarinet, oboe d'amore, then trumpet with
flute, saxophones, celesta with horn, a quartet of reeds, a portamento trombone, the highest woodwind,
and only then the strings
(https://thelistenersclub.com/2021/05/03/bolero-ravels-sublime-orchestration-exercise/,
https://www.clevelandorchestra.com/posts/ravels-bolero). Four structural facts are transferable:
texture thickens by promotion, not by piling on.
The biggest colour is the last one spent.
to produce artificial overtones
(https://thelistenersclub.com/2021/05/03/bolero-ravels-sublime-orchestration-exercise/). A composite
timbre is introduced as a NEW instrument, not as a doubling.
Barber opens pianissimo from a single melodic cell in first violin, hands the principal cell to the
cellos for the expansive middle, moves the string choir up the scale into its highest register, hits a
fortissimo climax, and follows it with SILENCE before the restatement
(https://www.parlancechamberconcerts.org/individual-program-notes/samuel-barber-(1910-1981)/adagio-from-string-quartet-no.-1,-op.-11,
https://classicalexburns.com/2022/07/21/samuel-barber-adagio-for-strings-diving-into-an-emotional-abyss/,
https://fcsymphony.org/program-notes/barber-adagio-for-strings/). The transferable law: the loudest
moment is followed by nothing, and the pause is a structural member rather than a gap. Subtraction to
zero is the strongest event available.
orchestra; Limgrave uses feathered bowing (players alternating bow strokes rather than bowing in
unison) to produce a diffuse atmospheric bed, with Scottish Highlands field recordings folded into
the background texture; Rennala's first phase layers music-box elements with processed female vocals
and the second phase brings in full choir and orchestra
(https://filmmusictheory.com/article/orchestrating-the-apocalypse-the-music-of-elden-ring/). Note
that the boss cue's layer escalation is bound to a PHASE change, which is our own boss-phase seam.
Uematsu has said he conceived it as a rock song rather than a choral-orchestral piece; it was
assembled by writing two to four measures a day until he had twenty to thirty fragments and then
rearranging them into an order that worked
(http://nobuouematsu-musicex.blogspot.com/2011/05/one-winged-angel.html,
https://marionmizuuki.wordpress.com/2017/03/06/1st-analysis-one-winged-angel-final-fantasy-vii/).
The transferable point is the assembly method: a MOTIF POOL first, an order second. Our arm-B
four-motif floor is already this method; the pool should be larger than the track needs.
know what the player was doing, with strings and harp added to build texture and variation subtly
(https://daily.redbullmusicacademy.com/2015/08/c418-interview/,
https://musescore.com/news/features/unlocking-the-secrets-of-c418s-sweden-the-minimalist-masterpiece-at-the-heart-of-minecrafts-soundtrack/).
The 2025 National Recording Registry essay on Minecraft: Volume Alpha is the scholarly citation
(https://www.loc.gov/static/programs/national-recording-preservation-board/documents/Minecraft-Volume-Alpha_Grosser.pdf).
This is the class Josh named as "some more simple with nature and ambience and whistling or humming".
Winifred Phillips' vertical-layering guidance is the most directly applicable published craft doctrine
we have, because it governs layers that must work both alone and combined — which is our adaptive
requirement (https://winifredphillips.wpcomstaging.com/2015/09/29/arrangement-for-vertical-layers-pt-1-a-game-composers-guide/).
together they are the full composition.
perceived clearly in the full mix.
tutti. This is the independently-sourced companion to the Rimsky-Korsakov inference in section 2.1:
a texture where everyone plays the same rhythm cannot be decomposed into layers, and cannot be
subtracted from.
Corroborating measurement from our own corpus: EX_124, a title cue Josh's own exemplar set holds up as
a target, records event_grammar.counts.tuttis = 0 across 127 seconds, against 14 drops and 20 raises.
A beloved title cue spends its whole length on drops and raises and never once puts everyone in at
once. Three independent lines of evidence — treatise structure, game-audio craft doctrine, and our own
exemplar measurement — arrive at the same rule.
Synthesising the four repertoires above into one transferable schedule for a 3:00 to 4:00 cue. This is
our synthesis, labelled as such, and each clause traces to a source above.
inside the first few seconds where the cue class allows one (our own EARWORM rule; EX_124 measures
time_to_hook_s = 0.104 on the Deus Ex title cue).
promotion law). Part count roughly doubles. Full texture is still far off; EX_124's
time_to_full_texture_s is 11.56 on a 127-second cue, which is early and is a property of that
title-cue class rather than a universal.
support. Harmonic rhythm may quicken. This is where the 3-to-5-cycle variation law bites: a
repetition must acquire a new instrument, hook, melody or beat event before the fourth pass.
part that has been present but masked. This is the move that makes the return feel large without
adding anything new.
where the withheld colour is spent — Bolero's strings, Rennala's phase-two choir, Elden Ring's full
orchestra after the sparse chorale.
a single sustained element. EX_124's event_grammar.silence_budget.fraction of 0.0232 across the
whole track shows how small the absolute quantity of silence is even when it is structurally load-
bearing; the value is in placement, not duration.
orchestration.parts gains entry_bar, exit_bar and optionally a list of rest_spans. The part
manifest stops being a roster and becomes a SCHEDULE. Once it is a schedule, three things become
checkable before rendering: no part is active for the whole cue except a declared ostinato; at least
one subtraction event of at least N parts occurs after the midpoint; the peak part count occurs in
the last third. Every one of those is a Josh-floor clause turned into a predicate.
promoted_from to a part that takes over an accompaniment role froma carrier that just finished — Bolero's law, made explicit and countable.
opening_and_flow_law.time_to_full_texture_s, flow_curve.peak_position_normalised, event_grammar.drops/raises/breaks with their
lanes_added/lanes_removed deltas, and dynamics.peak_position_normalised verify the schedule
after the fact. EX_124's flow peak sits at 0.272 and its dynamic peak at 0.4595 — the energy peak
and the loudness peak are in different places, which is itself a finding worth carrying to rung 2.
The strongest evidence available, and it is directly on point. David Huron, "Voice Denumerability in
Polyphonic Music of Homogeneous Timbres", Music Perception 6(4), 1989, pages 361 to 382
(https://online.ucpress.edu/mp/article/6/4/361/62869/Voice-Denumerability-in-Polyphonic-Music-of):
expert musicians become slower to detect added voices and less accurate at counting them as the number
rises, and accuracy drops markedly at the step from three voices to four. Huron's conclusion is that
the auditory system follows a one-two-three-or-many rule and that it may be impossible to process more
than about four concurrent streams.
Two consequences, and they point in opposite directions on purpose.
perceptual STREAMS. Bolero's accompaniment mass is one stream regardless of how many players are in
it. This is the reconciliation of "dozens of layers" with the four-stream ceiling, and it is the
single most important idea in this document.
defect Josh named is a GROUPING failure, not a count failure.
The mechanism behind grouping is codified in Stephen McAdams, Meghan Goodchild and Kit Soden,
"A Taxonomy of Orchestral Grouping Effects Derived from Principles of Auditory Perception", Music
Theory Online 28.3, 2022 (https://mtosmt.org/issues/mto.22.28.3/mto.22.28.3.mcadams.html). Its three
classes and their operative subtypes:
timbral AUGMENTATION, where a dominant instrument is coloured by a subordinate one [4.5]; timbral
EMERGENCE, where the fusion produces a new timbre identifiable as none of its constituents [4.9];
timbral HETEROGENEITY, where parts group but do not fully blend and some instruments stay audible
as themselves [4.13]. Fusion is strengthened by onset synchrony, harmonicity, and parallel changes
in amplitude and frequency [3.7].
stream SEGREGATION, defined as two or more clearly distinguishable voices of near-equivalent
prominence [5.12], and STRATIFICATION, defined as layers separated into more and less prominent
strands [5.16], on the other. Stratification prominence is analysed via Koechlin's extensity
(auditory size) and intensite (inherent force) [5.18].
juxtapositions and sectional boundaries. Segmentation strength rises when several parameters change
together [6.1]. The taxonomy notes explicitly that none of these effects is all-or-nothing [7.4].
Stratification is the concept our pipeline is missing by name. A forty-part cue with three declared
strata — foreground carrier, middleground counter-material, background bed — satisfies both Josh's
floor and Huron's ceiling at once.
stratum to every part, valued foreground, middleground orbackground, and enforce three to four active strata rather than three to four active parts. Enforce
that no stratum holds more than about 60 percent of the summed gain, and that the foreground stratum
is never the largest by part count.
lane_analysis.simultaneity_curve gives mean and max lanes (EX_124: mean5.455, max 9). Read as STRATA that is far too many; read as spectral bands it is unremarkable. The
card cannot currently tell the two readings apart, which is exactly the limit section 10 addresses.
Every part needs a lane, and lanes are finite. The register-spacing figure on the formula card is the
existing proxy: EX_124's register_spacing_octaves reads 0.99, 0.99, 0.99, 0.99, 1.0, 1.0 — an almost
perfectly even one-octave spacing between its active bands, which is what a well-slotted arrangement
looks like. Practitioner reports on orchestral mixing converge on 250 to 500 Hz as the accumulation
zone where full-orchestra energy piles into boxiness, with cuts of 6 to 8 dB reported as necessary
(https://vi-control.net/community/threads/your-go-to-eq-tricks-and-tips-when-mixing-orchestral-music.40771/,
https://www.waves.com/tips-for-mixing-film-tv-scores-free-presets). Labelled craft consensus, and it
matches the low_mid band being EX_124's largest at 0.381 of total energy.
Meyerson's stated rule is spatial rather than spectral and is worth carrying: do not build mixes in
the middle, and use small time delays in the Haas range of roughly 150 to 250 samples to separate
elements that would otherwise mask each other
(https://film-mixing.com/2016/07/28/film-score-mixing-with-alan-meyerson/). Panning and micro-delay
are a third axis of separation beyond register and time, and our renderer already has pan per part.
Masking is the process by which the threshold of hearing of one sound is raised by the presence of
another, raising the masked detection threshold — the minimum intensity a sound needs to stay audible
(Amelie Bernier-Robert and Ben Duinker, "Masking", Timbre and Orchestration Resource, 13 November
2023, https://timbreandorchestration.org/writings/timbre-lingo/masking). Two kinds:
processed the masked sound, so it never reaches higher processing.
masker, by attention, and by auditory-scene-analysis grouping.
The same source states plainly that predicting masking is very complex, because a composer must
account for each masking type per component and for interactions between maskers, which may reduce
each other or combine to increase total masking through cochlear distortion. Critical-band theory
supplies the frequency geometry: the auditory system integrates energy over critical bands on the Bark
scale, band width grows with centre frequency while covering constant distance on the basilar
membrane, and simultaneous masking is strongest within the masker's own band
(https://support.ircam.fr/docs/AudioSculpt/3.0/co/Masking%20Effect%20Intro.html).
Stacked noise, defined acoustically: it is what happens when many parts share critical bands, share
onsets, and share amplitude envelopes. Shared onsets and parallel amplitude change are precisely the
cues that FUSE events (McAdams et al. [3.7]) — so a homophonic tutti of many instruments in one
register is maximally fused and maximally masked at the same time. Every part contributes energy and
almost none contributes information. That is one sentence with three independent citations behind it,
and it is the technical answer to Josh's "there can't just be stacked noise".
register_slot (an octave band) and refuse ascore in which more than two parts hold the same slot in the same bar unless one is explicitly
marked as a doubling of the other. Refuse a bar in which more than a declared fraction of parts share
an onset, outside sections whose texture_type is homophonic.
lane_analysis.interlap_matrix is the after-the-fact masking proxy.EX_124's low_mid|mid pair reads 0.969 and mid|air 0.907 — very high overlap that a band-based
measure cannot distinguish from good octave doubling. Treat high interlap as a FLAG requiring the
score-side slot declaration to justify it, never as a verdict on its own.
Josh's 3-to-5-cycle rule and the treatises agree: repetition is not the problem, unvaried repetition
is. Bolero repeats two melodies for fifteen minutes and is not boring because the COLOUR changes every
statement (section 5.1). The standard vocabulary of colour change, drawn from the treatise tradition
and from the McAdams segmental categories, is small enough to enumerate and therefore small enough to
encode.
timbre covaries with pitch, playing effort and articulation (McAdams et al. [5.1]).
extensity in Koechlin's sense [5.18].
Limgrave device).
because it changes meaning rather than surface.
horn plus celesta organ-mixture does. This is McAdams' timbral EMERGENCE [4.9] used deliberately.
the first [6.4, 6.5].
declared law to license one. Extend that: when a repeat IS licensed, require a
colour_change field naming which device from the list above is applied, and make the absence of a
named device a battery failure. That converts the 3-to-5-cycle rule from a hope into a predicate,
and it is enforceable entirely at composition time.
motif_economy.self_similarity_lift and the n-gram repetition figuresmeasure whether repetition happened; they cannot see whether it was re-coloured. The constraint is
the real control here and the measure is only corroboration.
Ambience, nature, voice, chant, humming, whistling, clapping. Treated as musical parts, not decoration.
Elden Ring's Limgrave folds Scottish Highlands field recordings into the background texture
(https://filmmusictheory.com/article/orchestrating-the-apocalypse-the-music-of-elden-ring/). The
technique tradition runs from Schaeffer's musique concrete through contemporary practice; the standard
production move is to pitch-shift a recording so it functions as a pad or lead rather than as ambience,
with even a few semitones of shift changing its character
(https://blog.landr.com/field-recording-tips/, https://en.wikipedia.org/wiki/Pitch_shifting). Real-time
ensemble performance of field-recorded environmental sound is an active research area, which is
evidence that treating recordings as instruments is a defined practice rather than a metaphor
(https://arxiv.org/pdf/2006.09645).
The doctrine our lane should adopt: a nature layer is TUNED to the cue's key or it is a sound effect.
A river, a wind, a cicada bed and a fire all have a spectral centre; shifting that centre onto the
tonic or fifth makes the ambience a drone part with a register slot, subject to the same masking rules
as any other part. Where an ambience is unpitched by nature — footsteps, rain on stone, a rope creak —
its musical function is RHYTHMIC and it belongs on the grid.
designed to rhyme in both that language and in English
(https://en.wikipedia.org/wiki/Dragonborn_(song)).
(https://en.wikipedia.org/wiki/Halo_Original_Soundtrack).
(https://www.billboard.com/music/music-news/bear-mccreary-god-of-war-video-game-score-interview-8256913/).
arrives in phase two
(https://filmmusictheory.com/article/orchestrating-the-apocalypse-the-music-of-elden-ring/).
the part instead
(https://blog.playstation.com/2022/09/09/elden-ring-composer-tsukasa-saito-on-creating-the-games-score-and-his-favorite-track/).
That is carrier substitution at production time, and it is the exact device section 7 lists first.
A hum and a whistle are SOLO WIND parts with an unusual timbre: narrow dynamic range, limited range,
strong fundamental, weak upper partials. They sit cleanly in a register slot and mask almost nothing,
which is why they work as the top line over a sparse bed and why C418's class of cue can carry one.
Flamenco palmas is the most fully theorised body-percussion practice available and its distinctions are
directly encodable. Palmas sordas are cupped and muffled, sitting under the music without cutting
through; palmas claras are sharp and bright and mark accents unmistakably; palmeros reach for sordas
when they want to hold the compas present without competing with the singer, and for claras when an
accent must be unmissable. Contratiempo is clapping on the off-beat
(https://www.siudyflamenco.org/blog/palmas-and-compas-flamenco-rhythm,
https://www.grangalaflamenco.com/en/blog/types-of-palmas-in-flamenco-and-their-sounds/,
https://en.wikipedia.org/wiki/Palmas_(music)). Two parameters — timbre bright or dark, placement on or
off the beat — give four distinct rhythmic parts from a pair of hands, each with its own frequency
slot. That is a real orchestration resource, and it costs nothing to render.
register_slot, stratum, entry_bar, exit_bar, role. A nature bed with no register slot and
no entry schedule is decoration; with them it is a part. Add tuned_to on ambience parts naming the
scale degree its spectral centre is shifted onto.
because a family head realised on an mbira would be scoring a living tradition against a card that
has not authorised it. The same rule applies with equal force to chant, to language, and to
culturally-specific clapping traditions: the region page's own instrumentation section is the
licence, and palmas belongs to an Andalusian cell rather than to any cue that wants a clap. Where a
cue simply needs body percussion with no cultural claim attached, it is written as body percussion
and named as such.
Stated as the gap list, each item with the section above that supplies the fix.
orchestrate, then compare against the original — is a supervised training loop, and it is the one
thing on this list our pipeline could literally automate (section 2.5).
reflexive; a self-taught mix reaches for a fader instead (section 2.1).
support and bass is applied to every bar of every score studied. A part with no declared function is
the commonest amateur defect (section 2.2).
cannot be reduced is usually a texture with no strata (section 6.1).
as on adding; the amateur instinct is monotone accumulation (section 5.2 and 5.5).
octaves appearing by accident between two instruments neither of which is doubling, and inner parts
that leap because nobody tracked them.
an instrument's weak register sounds wrong even when rendered by samples, because the sample set was
recorded from a real player (section 2.4).
Adagio's silence, Ligeti's hidden canon — these are moves in a shared vocabulary, and a composer who
does not know them re-invents the weaker version.
and they are the difference between "add more" and "group better" (section 6.1).
Our formula card's lanes are TEN SPECTRAL BANDS — sub, low_bass, bass, low_mid, mid, upper_mid,
presence, brilliance, air, ultra — not instruments. The card says so itself in
lane_analysis.method_note: two instruments in one octave read as one lane, one instrument spanning
two octaves reads as two, and the card declares its confidence as LOW-FOR-ROSTER,
USABLE-FOR-ARRANGEMENT. The ceiling is therefore ten, the observed maximum on EX_124 is nine, and Josh
is asking for dozens. No amount of tuning makes a ten-band measure report a thirty-part texture. The
card must never be quoted as evidence about layer count, and any claim that it was is a defect.
What the band measure IS good for stands undiminished: entry and exit shapes, the simultaneity curve's
form over time, register spacing, interlap as a masking flag, and the event grammar's drops and
raises. Those are arrangement measures and they are trustworthy as such.
bass, other, vocals — with a six-source variant adding guitar and piano, and its own documentation
notes the piano source underperforms and that leakage and artefacts are normal in complex mixes
(https://github.com/adefossez/demucs, https://github.com/facebookresearch/demucs). Verdict: it counts
to four or six, not to thirty. It cannot answer the question.
spectra and time-varying gains and is a standard transcription and separation front end
(https://www.ee.columbia.edu/~dpwe/e6820/papers/SmarB03-nmf.pdf). The component count is a MODEL
ORDER chosen by the analyst, and model-order selection is itself an open research problem addressed
by model averaging over several orders (as in the ArXiv literature on multiple-order NMF,
https://arxiv.org/pdf/1605.07469 and related). Verdict: it returns whatever count you asked for. A
measurement whose answer is its own input parameter is not a measurement.
and multi-instrument multi-pitch estimation are active MIR fields; polyphony can be obtained
implicitly by counting active predicted pitches or modelled explicitly as local polyphony
(https://link.springer.com/article/10.1186/s13636-025-00398-2,
https://arxiv.org/pdf/1811.01143). Verdict: the most honest audio-side option, giving a defensible
estimate of concurrent PITCHED VOICES, which is closer to Huron's unit than to a part count. Worth
building as a corroborating measure and never as the primary one.
free, and already three-quarters built here.
Count layers at COMPOSITION time, from the part manifest, and treat every audio-side figure as
corroboration only.
Our pipeline already emits orchestration.parts per composed track, with part, instrument, role,
gain_db, pan, detune_cents and notes per entry, and the identity card already carries
instrumentation and voice_count_target. The measured state today is a mean of 9.3 parts per track
across the ten arm-B tracks, ranging 4 to 19. The gap to a real orchestral texture is therefore known
exactly, in the right unit, before anything is rendered — which is the whole argument for this
recommendation.
The proposed STEM MANIFEST schema, as the additions to the existing parts entries:
| field | type | purpose |
|---|---|---|
stratum | foreground / middleground / background | Huron and McAdams grouping; three to four strata regardless of part count |
register_slot | octave band identifier | frequency-lane allocation and the masking predicate |
entry_bar / exit_bar | int | the staged entrance schedule; makes subtraction checkable |
rest_spans | list of bar pairs | rest as a first-class parameter per part |
promoted_from | part id | Bolero's carrier-to-accompaniment promotion, made countable |
doubles | part id | licenses two parts in one slot; absence makes co-slotting a violation |
colour_change | device name | names which variety device licenses a repeat |
texture_type | per form section | declares the density regime the section is writing in |
tuned_to | scale degree | for ambience and nature parts, the pitch centre they are shifted onto |
cultural_licence | region page reference | care-line: which page's instrumentation section authorises this part |
With that schema, layer_count is a field, not an inference — and every clause of Josh's floor
sentence becomes a predicate a battery can enforce before a single sample is rendered.
1. Count layers as PARTS, not tracks and not stems: 24 to 40 for a full cue, 3 to 8 for the sparse
nature class. HOOK: part_count_floor/ceiling per cue class on the identity card; battery refuses
out-of-band. Our measured mean today is 9.3.
2. Dozens of parts must group into three or four STRATA, because listeners cannot count past about
four concurrent voices (Huron 1989). HOOK: stratum per part; enforce strata count, not part count.
3. A tutti is a balance construction, never everyone playing the same rhythm; the full homophonic tutti
is the one arrangement that cannot be layered or subtracted (Phillips). HOOK: onset-synchrony
fraction ceiling per bar outside declared homophonic sections.
4. Balance is arithmetic first — 1 heavy brass = 2 horns = 4 woodwind (Rimsky-Korsakov, p. 33). HOOK:
default gain_db law derived from the ratios; hand values become declared deviations.
5. The part manifest must be a SCHEDULE, not a roster: entry, exit and rest per part, with the peak in
the last third and at least one real subtraction after the midpoint (Bolero, Barber). HOOK:
entry_bar/exit_bar/rest_spans plus three schedule predicates.
6. Subtraction is the strongest event available, and the loudest moment is followed by the least
(Barber's climax into silence). HOOK: require at least one lanes_removed event of declared depth
after the flow peak; corroborate with event_grammar.drops and silence_budget.
7. Repetition is licensed only by a NAMED colour change from the enumerated device list — carrier,
family, register, solo/tutti, mute, attack mode, articulation, counter-line, reharmonisation,
composite timbre, timbral echo. HOOK: colour_change required on every licensed verbatim repeat;
absence is a battery failure. This is the 3-to-5-cycle law made mechanical.
8. Nature, voice, hum, whistle and clap are PARTS with register slots and entry schedules, and pitched
ambience is tuned to the cue's key. HOOK: identical manifest fields for non-orchestral parts, plus
tuned_to, plus cultural_licence naming the region page that authorises the material.
Every external claim above carries its URL or its book-and-chapter inline, so no separate bibliography
is repeated here. Source-strength tiers used throughout: PUBLISHED RESEARCH (Huron 1989; McAdams,
Goodchild and Soden 2022; Bernier-Robert and Duinker 2023; the NMF, Demucs and multi-pitch
literature); TREATISE (Rimsky-Korsakov, Piston, Adler, Blatter, Sevsay, Phillips); PRACTITIONER
TESTIMONY (Meyerson, McCreary, Saitoh, Joffres, O'Donnell, C418); and CRAFT CONSENSUS, which is the
VI-Control and mixing-tips material and is labelled as such at every point of use.
Repository artefacts read: build/audio/exemplars/formula_cards/EX_124.json (every card figure quoted);
build/audio/complete_tracks/tracks/*/*.json (the 9.3-part measurement);
build/audio/complete_tracks/cards/HF_MT_02_EXPLORATION_TRAVERSAL.json; harness/music_gen/orchestrate.py,
sfz_palette.py, theme_compositions.py, compose_arm_b.py, to_midi.py, stem_probe.py;
docs/spine/DECISIONS_PENDING_JOSH.md (the composition floor).
Thinnest evidence, named honestly: per-cue stem counts for shipped AAA titles are largely undisclosed
(the Expedition 33 interview is the clearest case of a detailed audio postmortem that gives no number);
Rimsky-Korsakov's tutti chapters could not be retrieved in full and no claim is quoted from them; and
no published source was found that states a track or part count for a single named game cue, so
section 3.3's figures come from film scoring and from composer templates rather than from game audio
directly.