REPETITION_AND_VARIATION_FORM.md

music/REPETITION_AND_VARIATION_FORM.md

Repetition with Variation — the anti-boredom architecture

CANON SUBORDINATION — this document is PROPOSAL-TIER: it serves canon and never outranks it.
Canon served: the composition floor at docs/spine/DECISIONS_PENDING_JOSH.md (commit da627060) -- the
3-to-5-cycle clause and "no track will be boring or filler".
If this document disagrees with canon, CANON WINS and this document is the defect.

TIER: RESEARCH SYNTHESIS — PROPOSAL-TIER. This document is HOW, never WHAT. It sets no canon, names

no region content, and changes no spine or registry row. Every substantive claim carries a source (a

URL, or book plus author plus chapter); where sources disagree the disagreement is stated; where a

statement is craft consensus rather than evidence it is labelled CRAFT CONSENSUS. Every principle

carries a PIPELINE HOOK line declaring whether our stack can COMPOSE it as a constraint, MEASURE it

from the formula card, or neither — and where neither, the missing measurement is named rather than

proxied. Measurements labelled MEASURED HERE were derived in this lane from the cards already on

disk; their pool, their caveats and their honest register are stated in section 1.6 and nowhere

weakened afterwards.

WHAT THIS LANE ANSWERS. Josh graded round 1 of the generated music, turned down all eighteen

candidates, and wrote a composition floor that contains one measurable law

(docs/spine/DECISIONS_PENDING_JOSH.md, tail, commit da627060): "I dont like repetitious loops

lasting more than 3-5 cycles without introducing a new instrument or catch or melody or beat drop",

and beside it "no track will be boring or filler." The law is being instrumented as a gate in a

parallel lane. This lane supplies the craft that makes a track pass it musically rather than merely

satisfy a counter — and, because the measured corpus appears at first reading to say the opposite of

the law, it does the reconciliation first.

WHERE IT SITS AMONG ITS SIBLINGS. COUNTERPOINT_AND_VOICE_LEADING.md in this directory owns the

vertical axis — how many simultaneous layers stay separable. This document owns the horizontal axis —

what changes over time and when. school_grammars/KONDO_SCHOOL.md and

school_grammars/SNES_JRPG_SCHOOL.md already state the rules of two dialects

(FORM-01..05, LOOP-01..10); this document supplies the reasoning underneath them, the traditions

they descend from, and the schedule they leave unstated.

---

1. The reconciliation — why "loved tracks repeat more" and the 3-5-cycle law are the same fact

1.1 The apparent contradiction

build/audio/exemplars/PATTERN_FINDINGS_V2.md is the measured derivation over the tracks Josh named

as lifelong favourites, scored against their own album siblings. Six axes clear a calibrated

multiplicity bar at N = 28, and they are all measures of structural scale: section_count (AUC

0.750, hit median 17 against a sibling median 9), total_length_s (0.732), bars (0.718),

repeat_sim_mean (0.714, hit median 0.902), n_drops (0.710) and novelty_mean — which clears the

bar in the negative direction, hit median 0.556 against a sibling median 0.678 (§4 and §10 of that

document).

Read naively, the last two say the loved tracks repeat themselves more than the filler beside them

and change less. Josh's floor says change every three to five cycles. Those pull opposite ways, and

the resolution is the most valuable paragraph this lane can write. So it is derived rather than

asserted.

1.2 What novelty_mean actually measures

The card's duration_and_structure.novelty_peak_strength is a checkerboard-kernel novelty score at

each section boundary of a chroma-plus-MFCC self-similarity matrix, and the values are normalised to

each track's own maximum — EX_124.json carries a 1.0 in its list, which is the normaliser showing

itself. novelty_mean is therefore not "how much this track changes." It is the SHAPE of the

boundary-strength distribution: how equal the track's structural events are to one another. A low

mean with a high boundary count describes a track with a few large events standing above many small

ones. A high mean describes a track whose boundaries are all about equally strong.

That is a testable difference, so this lane tested it.

1.3 The measurement

MEASURED HERE, over the formula cards at HEAD via harness/music_gen/derive_patterns.py's own

load_cardsflattenlabel_rows path, restricted to tracks of 45 seconds or more (28 corpus

rows against 630–927 siblings depending on which card fields are present):

what was measuredcorpus rowsalbum siblingsAUC
interior section boundaries per track16110.698
median section length, in the card's own bars5.79 bars5.34 bars0.540
share of boundaries at or above 0.7 of the track's own maximum17.9%36.4%0.313
share of boundaries below 0.428.6%17.6%0.609
strong boundaries per minute0.881.890.308
strong boundaries, absolute count340.422
strongest boundary divided by the median boundary1.911.650.646
raises that add at least one lane, per minute4.553.100.568
mean magnitude of a raise7.49 dB10.09 dB0.425
raises of 10 dB or more, per minute0.380.520.442
complexity steps per track640.622
share of complexity steps landing at a section boundary0.3580.5000.463

Six independent readings, one story. A loved track carries about half again as many structural

boundaries as its own album's filler, and it makes each one smaller. It brings a new voice in twice

as often and raises the level less when it does. It has fewer big moments in absolute terms — three,

against four — across twice the running time, so the rate of big moments is less than half. And the

ratio between its largest event and its typical one is steeper, which is what a hierarchy looks like

when it is real rather than flat.

1.4 The resolution, stated

The loved tracks do not repeat more and change less. They repeat their MATERIAL and change their

TREATMENT, continuously, in small increments, under a sparse ridge of three or so genuine structural

events. repeat_sim_mean sits high because the material returns; novelty_mean sits low because

most of the returns are re-treatments rather than ruptures; section_count sits high because a

re-treatment still registers as a boundary. Josh's law and the measurement are the same fact read at

two scales. His floor governs the dense tier — something must change every few cycles, and in the

corpus something changes every 5.79 bars. The measurement governs the sparse tier — most of those

changes must be small, because a track that makes every change a big one spends its hierarchy and

has nothing left to arrive at.

That reading also explains why the corpus puts a SMALLER share of its complexity steps at section

boundaries than the filler does (0.358 against 0.500). If change were carried by the form, the steps

would land on the seams. They do not. They land inside the sections, which is what continuous

re-treatment over a returning cycle looks like from the outside. This directly contradicts the

expectation written into docs/proposals/music/MASTERPIECE_PROGRAM.md §11.6 — "masterpieces step

complexity at STRUCTURAL boundaries and hold it flat inside sections" — and that section's own

instruction is "If the corpus says otherwise, the corpus wins." The corpus says otherwise. This is

reported as a finding for the program to rule on, not applied as a change.

event_grammar and complexity_law on the existing card. None of it needs a new extractor. The

gate lane can adopt strong_boundary_rate_per_min, boundary_strength_max_over_median and

boundary_step_share as reported diagnostics immediately.

1.5 What would test it

The claim above is inferred from proxies. The direct test needs a measurement the card cannot make.

Take every pair of sections whose repeat_map similarity is 0.95 or higher — sections that are, in

chroma and MFCC terms, the same music — and ask whether the SET OF SOUNDING INSTRUMENTS differs

between them. If the corpus rows change instrumentation across near-identical sections at a higher

rate than their siblings do, the hypothesis is confirmed directly rather than by inference. The card

cannot do this: lane_analysis declares its own confidence as "LOW-FOR-ROSTER" and

PATTERN_FINDINGS_V2 §10 records that max_lanes is pinned at 9 for hits and siblings alike, which

is a ceiling in the spectral_band_proxy method rather than a fact about the music. The measurement

that closes it is a real instrument roster — source separation, or a trained multi-instrument

recogniser — reporting a per-section instrument set. That is the single highest-value addition to the

measurement rig this lane can name.

1.6 Honest register on the numbers in 1.3

seconds or more, and they have no calibrated multiplicity null of their own. PATTERN_FINDINGS_V2

§3 set its bar at AUC 0.695 for a 103-feature bank at N = 28; nothing in the table above may be

called "clearing" anything.

percentile. The rest are descriptive reads that agree with it.

and rate-normalised figures (per minute) are the ones to trust over absolute counts.

novelty_mean are related measures over the same matrix), which the source document already says.

---

2. The variation canon — what actually changes, and in what order

2.1 Decide the invariant before deciding the variations

The first analytical move in a variation set is not what changes but what is held. Grove's first

edition already separates the cases: some sets are bound to the theme "mainly through the melody," in

others "the succession of the harmonies is the chief bond of connection," a third class repeats a

bass figure under new material, a fourth recasts the theme in an old contrapuntal form

(https://en.wikisource.org/wiki/A_Dictionary_of_Music_and_Musicians/Variations). Grove also states

what must survive in every case: harmonic succession, phrase and period lengths, and the bass

outline.

The Goldberg Variations are the canonical demonstration that choosing the bass buys the most freedom.

The invariant is the aria's 32-bar ground and its harmony; the aria's melody is never varied at all,

and Kirkpatrick's performing edition notates the frame as figured bass, which makes the work

functionally a very large passacaglia (https://en.wikipedia.org/wiki/Goldberg_Variations;

https://www.laphil.com/musicdb/pieces/4840/goldberg-variations-bwv-988).

list and a form map. An explicit invariant field per track — naming which of bass, harmony,

melody or formal outline is held across sections — would let the verbatim-run detector

(S_verbatim) distinguish a legitimate restatement of the invariant from a lazy copy of a bar run,

which is the defect RECOMPOSE_RUNG_TWO_FINDINGS §1.3 records.

2.2 The taxonomy of variation types

TypeWhat changesCanonical instance
Melodic ornamentation, figuralsurface decorated, frame audible throughoutGoldberg arabesque variations 5, 14, 20, 23
Harmonicnew progression under the same outlinethe subject of a CWU thesis on Brahms Op. 24, https://digitalcommons.cwu.edu/etd/1340/
Rhythmic, metricpulse broken or subdivision shrunkBeethoven Op. 111/ii, 9/16 to 6/16 to 12/32 to triplet-32nds
Texturalvoice count, spacing, register redistributedHerrmann's Psycho "Cellar", fugal texture over 8-bar units
Timbral, orchestrationalthe invariant re-scoredBrahms Op. 56a against the two-piano 56b
Contrapuntalthe invariant put under imitative lawGoldberg canons 3 through 27; Op. 24's fugue
Modalparallel-mode shift, often with added suspension and imitationGoldberg 15, 21, 25; Brahms Op. 56 variations 2, 4, 8
Charactergenre or affect recast — march, siciliana, musette, overtureBrahms Op. 24 variations 13, 19, 22; Goldberg 16
Double variationtwo alternating themes, usually minor and majorHaydn's speciality

Types and definitions from https://en.wikipedia.org/wiki/Variation_(music) and Grove as above;

Robert U. Nelson's The Technique of Variation (UC Press, 1949) is the source monograph and carries

exactly these categories in its index (https://archive.org/details/in.ernet.dli.2015.177964). Elaine

Sisman's Haydn and the Classical Variation (Harvard, 1993) and her revised New Grove article are the

modern authorities; the Grove text is paywalled and her specific category names are not verified

here and should not be quoted as hers.

The same article draws the distinction this project needs most: "Variation depends upon one type of

presentation at a time, while development is carried out upon portions of material treated in many

different presentations simultaneously." Variation is the modular form. Development is the

through-composed form. A loop is a variation set that happens to have no last variation.

2.3 The ordering logic — where a set escalates and where it relieves

through, if the theme be in the major, there will be a minor variation, and vice versâ" (Wikisource,

as above). It holds across the repertoire: Goldberg's Var. 25 at 25 of 30, Brahms Op. 24's Hungarian

funeral march at 13 of 25, and the minore at 21 of 23 in Beethoven's original 1819 Diabelli draft.

unbarred cadenza" — the set is braked immediately before the finale rather than run straight into it.

the exact midpoint and bisects the work, following Bach's habit in the other Clavier-Übung volumes.

Op. 24 variations 14 to 18 — "a wonderful decrescendo of tone and crescendo of Romantic beauty" —

and John Rink tracks the same set's dynamics as flux early, intensity deliberately suppressed after

variations 13 to 15, then a massive build from variation 23

(https://en.wikipedia.org/wiki/Variations_and_Fugue_on_a_Theme_by_Handel).

"edge-related" rather than concentric, "each variation being lent significance by its relationship

with what comes before and after," producing family resemblances rather than derivation from the

theme (ibid.). A variation set is a sequence graph, not a star graph.

1819 Diabelli draft had 23 variations, and that variations 1, 15 and 25 are late 1822–23 insertions,

each making "pointed and humorous reference to the original theme in its original register"

(William Kinderman, "Beethoven's Diabelli Variations", Arietta 2013, pp. 6–7,

https://williamkinderman.com/wp-content/uploads/2014/01/Arietta-Diabelli-2013.pdf). The final

compositional act was to re-seed recognisability at intervals across a set that already escalated.

dynamics.arc_curve_db and complexity_law.index_per_bar — a local minimum in both, sited between

0.6 and 0.75 of the running time. Josh's floor names "no track will be filler"; a relief section is

the one place a thin texture is correct, and declaring it on the card is what distinguishes composed

relief from an empty stretch.

2.4 A diminution ladder is finite, and the top rung must change kind

Op. 111's Arietta is the paradigm escalation engine: the pulse is fixed and only the subdivision

shrinks, 9/16 to a notated 6/16 (for 18/32) to 12/32 (for 36/64) and then back to 9/16 filled with

triplet 32nds, an effective 27/32

(https://en.wikipedia.org/wiki/Piano_Sonata_No._32_(Beethoven)). Perceived tempo accelerates while

actual tempo does not. The jazz and boogie-woogie reading of the third variation is DISPUTED —

Uchida, Stravinsky and Denk hear it, Schiff explicitly rejects the comparison — and both should be

reported. What matters structurally is what Beethoven does when the ladder runs out: the movement

substitutes sustained shimmer, the long trills over the returning theme, for further rhythmic

activity. CRAFT CONSENSUS, from the shape of the repertoire: a set that escalates only by subdivision

hits a wall, and the change of kind at the top rung has to be planned rather than discovered.

a diminution ladder will show as a monotone staircase in the index. A track whose index rises

monotonically to its own maximum and then stops is exactly the failure mode; the constraint is that

the last complexity step before the final section must come from a different mechanism than the

three before it, which step_map can be extended to record.

2.5 The return of the theme is never neutral

Goldberg ends with the aria da capo, note for note. Peter Williams: "no such return can have a

neutral effect" — the melody now reads wistful, or nostalgic, or resigned

(https://en.wikipedia.org/wiki/Goldberg_Variations). The material is literally unchanged and the

listener has been changed. For a looping cue this is the cheapest strong device in the entire

tradition, because the loop point already guarantees the return; what it does not guarantee is that

the return arrives after something has happened.

---

3. Ostinato — what develops over a figure that does not

3.1 The fixed floor licenses maximal freedom

The Baroque doctrine is explicit in the sources and in the manuscripts: early dance manuscripts "give

the bass line ostinato and nothing else; all the rest of it would have been improvised over the

ostinato line" (San Francisco Conservatory, Ostinato and Variation analysis lecture,

https://sfcm.edu/study/majors/academics/music-theory-and-musicianship/sfcm-theory/online-materials/analysis-lectures/ostinato-and-variation).

The repeating bass supplies constant harmonic drive so an elaborate melody cannot get lost, and that

stability is what licensed improvisation in the first place (English Touring Opera,

https://englishtouringopera.wordpress.com/2018/08/24/get-a-good-grounding-the-importance-of-ground-bass-in-baroque-music/).

Ostinato is originally a scaffolding for real-time invention, which is precisely the game-audio case.

The passacaglia-versus-chaconne distinction is DISPUTED and the dispute is itself the finding.

Goetschius defined chaconne by a recurring soprano over a harmonic sequence and passacaglia by a

ground bass; Clarence Lucas defined the pair in precisely the opposite way; the SFCM lecture gives a

third account by bass motion. Silbiger in Grove settles it by dissolving it — composers "often used

the terms chaconne and passacaglia indiscriminately," and "modern attempts to arrive at a clear

distinction are arbitrary and historically unfounded" (https://en.wikipedia.org/wiki/Chaconne;

https://en.wikipedia.org/wiki/Passacaglia). Do not build a taxonomy on the distinction.

3.2 The nine parameters that develop over an unchanging figure

ground "twelve times in the bass, three statements in the treble, and one more statement in the

bass, upon which the theme returns for full orchestra"

(https://www.hollywoodbowl.com/musicdb/pieces/4466/variations-on-a-theme-by-haydn). Note that

several programme notes give a bare count of 17 instead; the registral plan is the interesting fact

and the counts DISAGREE.

1 to 6 admit many harmonisations and notes 7 to 10 do not.

shorten while the surface accelerates and the perceived event rate rises twice over

(https://www.corilon.com/us/library/practical-issues/johann-sebastian-bach-chaconne). Segmentation

of that work is DISPUTED — most analyses assume 64 four-bar variations, others count 34.

3.3 The failure mode, and the tradition's fixes

No scholarly source located uses the phrase "ostinato trap," so the term is CRAFT CONSENSUS. The

nearest sourced diagnosis is local rather than global, and it is better: David Feurzeig identifies "the

fourth-measure doldrums in the ground that would otherwise sap the momentum of the song" in Purcell's

"Ah! Belinda," and shows Purcell answering it by concentrating melodic entries precisely where the

ground goes slack (David Feurzeig, "On Shifting Grounds: Meandering, Modulating, and Möbius

Passacaglias", SMT 2010, pp. 2–3, https://www.uvm.edu/~dfeurzei/PassacagliaDF.pdf). The transferable

technique is exact: find the dead beat in your loop and put the foreground event on it.

The rest of the tradition's answers, all from Feurzeig unless noted:

upper voices and modulates briefly for colour while staying anchored in the tonic

(https://en.wikipedia.org/wiki/Passacaglia_and_Fugue_in_C_minor,_BWV_582).

tonic," or a theme hovering between two equally plausible tonics, which "transform[s] the passacaglia

loop into a Möbius strip, whose tonally contrasting sides follow one another with no discontinuity"

(Feurzeig, pp. 1, 8–12).

initial direction even as the theme repeats" (Feurzeig, p. 1).

becomes progressively buried until it is "effectively inaudible," and Ligeti keeps it strict except

for a single discontinuity placed just after the climax (Feurzeig, pp. 7–8). One rule-break at the

point of maximum attention is worth more than continuous variation.

statement of the theme, wrapping rather than cadencing (Feurzeig, p. 11).

whole to half to quarter notes across choruses — the same loop at three temporal scales (Feurzeig,

pp. 9–11). This is the ostinato's version of the Op. 111 diminution ladder.

cadence at the wrap are already law and already instrumented — harness/music_gen/loop_cadence.py

reads the final and penultimate harmonies off the shipped MIDI and refuses a perfect authentic

cadence, with a positive control that overwrites a real track's last two bars with a textbook PAC to

prove the reading path can see one. harness/music_gen/loop_seams.py composes the turnaround under

five named rules rather than fading. The unimplemented ones are ground migration between intensity

states and reharmonisation of a fixed bass, both of which are score-side constraints

(arm_b_scores.py) with no measurement needed.

3.4 The phase parameter, which is the tradition's best trick and nobody's default

Purcell's grounds are 3, 5 and 6 bars precisely so the upper line cannot lock to them. Dido's Lament

runs a five-bar ground eleven times, which Feurzeig calls "highly irregular," forcing nine-bar vocal

groups of 4+5 and cadences that overlap the ground; Dido enters on the last note of the first

statement rather than at the head of the second

(https://en.wikipedia.org/wiki/Dido%27s_Lament). In the B section the melody of "Remember me" lags

behind its earlier position relative to the ground, which Feurzeig notes "hints at 20th-C

compositional procedure" (p. 3). Britten weaponises the same independence: in the Peter Grimes

passacaglia, a seven-note bass figure repeats 39 times and "the beginnings and ends of the variations

don't synchronize with the repetitions of the ground bass; indeed, they go their tumultuous way almost

in competition with it" (https://www.laphil.com/musicdb/pieces/2673/passacaglia-from-peter-grimes).

path. An odd-length ostinato — five or seven bars under a four-bar melodic phrase — guarantees a

changing alignment for twenty bars from a single decision, with no extra material. It also

guarantees that a naive verbatim-run detector will not see repetition where a listener does not hear

it. The measurement that would confirm it is motif_economy.recurrence_lag_s against the section

grid; note that PATTERN_FINDINGS_V2 §6 retired recurrence_lag_norm as a discriminator, so use it

as a diagnostic and never as a band.

3.5 Why the modern action ostinato reads generic

The scholarly framing of the Remote Control school is composer-as-engineer: Zimmer prioritises sonic

texture over melodic content, and in The Dark Knight "it is the timbre of the cello, not its melody,

that carries its identifying features"; action sequences are structured by "string ostinato and

taikos"; the same article carries the criticism that in places "the orchestration is mushy and sounds

overly processed" (Sounding Out!, "Sculptural Dissonance: Hans Zimmer and the Composer as Engineer",

https://soundstudiesblog.com/2014/07/10/sculptural-dissonance-hans-zimmer-and-the-composer-as-engineer/).

Practitioner criticism that the spiccato patterns "aren't saying anything harmonically" is forum-tier

and the page 403'd on direct fetch, so treat it as low-confidence (VI-Control thread 28230).

CRAFT CONSENSUS, and it is the diagnosis this project needs: the grammar is criticised for exactly the

failure mode 3.3 describes. It escalates by layer count and dynamics — one of the nine parameters in

3.2, applied monotonically. The Baroque ground-bass repertoire varies six to nine of them over the same

fixed floor. That is the whole difference, and it is a difference of discipline rather than of budget.

---

4. Minimalism — the tradition that wrote down the rules for literal repetition

4.1 The audible-process doctrine, and where the interest actually lives

Reich's 1968 essay is the founding statement and it is short enough to read whole (full text at

https://www.bussigel.com/systemsforplay/wp-content/uploads/2014/02/Reich_Gradual-Process.pdf;

reprinted in Writings on Music 1965–2000, ed. Hillier, OUP 2002). The load-bearing sentences: "The

distinctive thing about musical processes is that they determine all the note-to-note (sound-to-sound)

details and the over all form simultaneously"; "I am interested in perceptible processes. I want to be

able to hear the process happening throughout the sounding music"; "To facilitate closely detailed

listening a musical process should happen extremely gradually."

And then the passage that matters most for craft, because it locates the content of process music

somewhere other than the process: "Even when all the cards are on the table and everyone hears what is

gradually happening in a musical process, there are still enough mysteries to satisfy all. These

mysteries are the impersonal, unattended, psycho-acoustic by-products of the intended process. These

might include sub-melodies heard within repeated melodic patterns, stereophonic effects due to listener

location, slight irregularities in performance, harmonics, difference tones, etc." The process is a

machine for manufacturing epiphenomena, and the epiphenomena are the music.

4.2 Phasing is ramps and stations, and the stations are the point

In Piano Phase both players play a twelve-semiquaver figure of only five distinct pitch classes; one

accelerates imperceptibly until exactly one semiquaver ahead, then locks and the pair hold that offset

while the composite is absorbed, twelve times, until unison returns

(https://en.wikipedia.org/wiki/Piano_Phase). The drift is a transition; the locked offset is a station,

held long enough for the ear to build a new gestalt. Reich confirmed the priority by abandoning the

ramp entirely in Clapping Music, where one player jumps a beat ahead after twelve repetitions — all

stations, no ramps.

Reich also stopped leaving the resulting patterns to chance. In Music for Mallet Instruments, Voices

and Organ, once the marimbas and glockenspiels reach maximum activity, a third woman's voice doubles

some of the short melodic patterns emerging from the combination of the four marimba players

(https://stevereich.com/composition/music-for-mallet-instruments-voices-and-organ/). The emergent

illusion is promoted to a real voice, which teaches the listener to hear it, after which they hear it

everywhere. That is the highest-leverage single trick in this section.

4.3 Density as a continuously variable parameter over an invariant cycle

Drumming runs one 12/8 cycle under a work of 55 to 85 minutes. Reich names three techniques as new

there: rhythmic construction and reduction, which is "the process of gradually substituting beats for

rests (or rests for beats)"; the simultaneous combination of different timbre families; and the human

voice imitating the instruments (https://en.wikipedia.org/wiki/Drumming_(Reich)). The cycle is

invariant. What rotates is which beats of it are sounding and which instrument family is sounding it.

The four-section form is a timbral modulation scheme standing in for a harmonic one.

4.4 Reich's own thesis sentence

From the Music for 18 Musicians programme note, read in full

(https://www.vapmedia.com/uploads/8/1/6/4/81640608/reichmusicfor18programnotes.pdf): "A melodic

pattern may be repeated over and over again, but by introducing a two- or four-chord cadence

underneath it, first beginning on one beat of the pattern, and then beginning on a different beat, a

sense of changing accent in the melody will be heard… Its effect, by change of accent, is to vary that

which is in fact unchanging."

That last clause is the thesis of this entire document, written from the inside by the composer who

had the most to lose from admitting that repetition alone is not enough.

The same note supplies four more directly portable facts. The whole work is an 11-chord cycle played

fast at the opening and close, then augmented roughly twentyfold so each chord becomes a five-minute

section — "a kind of pulsing cantus for the entire piece," so the introduction is not an introduction

but the piece at another scale. Every section is one of exactly two archetypes: an arch,

A-B-C-D-C-B-A, or "a musical process, like that of substituting beats for rests, working itself out

from beginning to end." Material recurs "surrounded by different harmony and instrumentation," and

Reich says the relationship between sections "is thus best understood in terms of resemblances between

members of a family." And the cueing is diegetic: section changes are called by the vibraphone,

explicitly "much as in a Balinese Gamelan a drummer will audibly call for changes of pattern, or as the

master drummer will call for changes of pattern in West African music," because "audible cues become

part of the music and allow the musicians to keep listening."

can carry per section. Harmonic-rhythm scaling — the same chord succession stated fast as an

exposition and slow as the form — is a compositional decision with a measurable signature in

tempo_and_time_grammar.dominant_onset_period_s against section length. Family resemblance rather

than literal reprise is exactly what the S_verbatim detector should be rewarding, and the missing

measurement is the instrument-roster one named in 1.5.

4.5 The rest of the mechanism catalogue

repetition while deleting the previous repetition's last note, so the phrase boundary keeps sliding

against the meter (https://musictheory.pugetsound.edu/mt21c/AdditiveMinimalism.html). Reich changes

the relationship between two copies of an invariant; Glass changes the length of the invariant.

single dominant-eleventh chord over a constant maracas grid.

other," never to "race too far ahead or to lag too far behind," and to "stay on a pattern long enough

to interlock with other patterns being played" (quoted at

https://teropa.info/blog/2017/01/23/terry-rileys-in-c.html; the authoritative text is the published

score). Enough desynchronisation to guarantee unrepeatable heterophony, not enough to lose the frame.

was added in 1964 rehearsals, attributed to Reich, when the ensemble needed an anchor. Pure freedom

was unlistenable and the fix was a constant.

electronics and runs a modulating square wave, one state Lydian and one Phrygian, through half the

keys by the circle of fifths (https://en.wikipedia.org/wiki/Phrygian_Gates). The local surface may

repeat freely because the arc is what is moving. For long-form scoring this is the most useful single

structural idea in the section. Note that "a Minimalist bored with Minimalism" is a writer's phrase

about Adams and not his own words; do not quote it as his.

foreign object — the hymn "Be Thou My Vision" surfacing near the end of Eastman's Femenine

(https://en.wikipedia.org/wiki/Julius_Eastman); Andriessen's De Staat, where "any slight variation

comes with total emphasis, so drastic that you can hardly discern anything that remains of its past

form" (https://www.boosey.com/cr/music/Louis-Andriessen-De-Staat/1425).

---

5. Cyclic traditions as engineering, not colour

These are living traditions and they solved the long-cycle problem centuries before Western concert

music posed it. They are read here for mechanism, in their own terms. The vertical and textural

aspects of kotekan and forest polyphony are treated in the sibling counterpoint document; what

follows is the time axis.

5.1 Javanese irama — the density-level shift, which is the mechanism this project needs most

Javanese karawitan nests time inside time. The punctuating gongs are the clock and their pattern is

the form's identity: gong ageng marks the gongan, kenong divides it into two or four kenongan, kempul

usually halves a kenongan, kethuk and kempyang "mark the pulse of the structure in between stronger

beats" (Barry Drummond, Javanese Gamelan Terminology, https://gamelanbvg.com/gendhing/gamelanGlossary.pdf,

p. 2). A ladrang is 32 balungan beats with kenong on 8, 16, 24 and 32, kempul on 12, 20 and 28, and a

wela — a structural silence where a kempul is expected — at beat 4 (Drummond pp. 11–13). The larger

gendhing forms are classified by kethuk density and run to 64, 128 or 256 beats per gongan.

Irama is a level of density measured against that fixed cycle, and it is not a tempo. Drummond fixes

five levels by the ratio of peking strokes to balungan beats: lancar 1, tanggung 2, dadi 4, wilet 8,

rangkep 16 — each step a doubling (p. 5). Sumarsam describes what happens at a shift: "First, the

drum leads the ensemble to gradually speed up or slow down the tempo… when the elaborating instruments

reach a point where playing their instruments is uncomfortably too slow, then they have to make an

upward adjustment of their tempo, accompanied by expanding the number of pulses within the gongan

structure of the piece. In essence, when an irama changes, the tempo returns back to the same tempo

before the change, but the piece becomes more expansive since the gongan structure is expanded"

(Sumarsam, "Temporal and Density Flow in Javanese Gamelan", in Thought and Play in Musical Rhythm, OUP

2020, pp. 126–127, https://sumarsam.faculty.wesleyan.edu/files/2023/01/4_Temporal_and_Density_Flow.pdf).

Restated as an algorithm: hold the fast layer's absolute note-rate roughly constant, halve the rate of

the skeleton and the punctuation, and the same cycle now takes twice the wall-clock time at twice the

density. The pattern never changed. Benamou's Western analogy, quoted by Sumarsam at p. 124, is exact —

a 2/2 variation with eighth-note figuration moving to a 4/4 variation with sixteenths, where the theme

is twice as long but the figuration goes by at the same speed.

Three details make it implementable rather than decorative:

Pangkur (pp. 128–129): the drum cues wilet by slowing while switching to the livelier kendhang

ciblon, whereupon the gendèr switches from lomba to rangkep technique and the bonang switches from

pipilan to imbal. One cue reconfigures three instrument idioms at once.

LAND on a structural point — Sumarsam's listening guide to gendhing Jaladara has the change back to

irama dadi landing on the gong and the shift to tanggung cued three gatra before it (pp. 136–138).

irama because it doubles density by repeating sections of an existing pattern rather than splitting

one pattern into two, and notes that "any piece in whatever irama can be performed in rangkep"

(Sumarsam p. 132, citing Supanggah 2011, 295). Sources also disagree on whether lancar counts, and

count a lancaran at 8 or at 16 pulses per gongan depending on the counting unit — a unit artefact,

not a factual conflict, but any implementation must fix its unit explicitly.

density-level shift is a novelty event that requires no new material, no new harmony and no new

motif — the same cycle, re-filled at twice the rate, with the elaborating idiom switched. It is

exactly what Josh's "introducing a new instrument or catch" asks for, delivered from material the

track already contains. It is measurable after the fact as a step in texture_and_density.note_rate_per_s

and complexity_law.index_per_bar with no corresponding change in duration_and_structure.repeat_map

similarity — high similarity across a complexity step is the signature, and it is checkable today.

5.2 Bali — the pyramid sounded all at once, and the cued break

Where Java changes which density level is active over time, Bali sounds them simultaneously. Tenzer

gives the strata as simple duple ratios: kotekan composite at 4 (sometimes 8) tones per beat, neliti at

1, pokok at 1 per 2 beats, jegogan at 1 per 4, 8, 16 or 32

(https://hugoribeiro.com.br/biblioteca-digital/Tenzer-Theory_Analysis_Melody_Balinese_Gamelan.pdf, §2.2

and Fig. 2). Kotekan is the interlocking pair, polos and sangsih, producing a composite line neither

plays and neither could play alone.

Two things port directly. First, norot is a fully deterministic generative rule: the gangsa polos

oscillates between the pokok tone and the tone above it at eight notes per pokok tone, the sangsih plays

a fixed kempyung four keys above, a fixed transition figure precedes each pokok change, and where the

register ceiling makes the kempyung impossible the sangsih collapses to unison

(Gong Workshop, Explanation of the notation of Balinese music, p. 2,

https://gongworkshop.nl/wp-content/uploads/docs/explanation-notation-balinese-music.pdf). Given a

skeleton, the elaboration writes itself, including its fallback.

Second, the angsel taxonomy is cycle-positional. Research from ISI Denpasar identifies nine main types

in gong kebyar — kempli, kempul, kemong, gong, tugak, sigug, bawak, lantang and suwud — and most are

named for the colotomic stroke the break resolves against, the rest for length or for function. The

taxonomy is literally where in the cycle the break lands, crossed with how long it is, and the kendang

pair cues it. Short cycles are what make an interruption legible: gilak is eight beats.

Tenzer adds the caveat that kotekan style is a salience dial rather than a fixture, ranging from

patterns that "follow the contour of the pokok exactly and [are] not articulated as an independent

stratum" to patterns standing distinctly apart from it (§3.8). The definition of ubit-ubitan is

CONTESTED — McGraw's glossary has the parts coinciding at regular intervals, a widely circulated

alternative has them coinciding at irregular ones. Treat it as an umbrella term.

arm_b_scores.py as a typed event with a declared landing point, and it is a far better model for

Josh's "beat drop" than an undifferentiated dB dip. It is MEASURABLE as a event_grammar.drops entry

whose at_s lands on a section boundary or a named subdivision of one.

5.3 West Africa — the fixed timeline and the indexed lead library

The reference is a bell pattern; Nketia's adopted term is "time line." The commonest is seven strokes

over twelve pulses, durations 2 2 1 2 2 2 1, placing onsets on pulses 1, 3, 5, 6, 8, 10 and 12 (Kofi

Agawu, "Structural Analysis or Cultural Analysis? Competing Perspectives on the 'Standard Pattern' of

West African Rhythm", JAMS 59/1, 2006, https://academicworks.cuny.edu/gc_pubs/935/). The asymmetry is

what makes cycle position identifiable at all.

David Locke's account is the most implementable. Performers "set up dynamic steady states"; the metric

matrix is "an unsounded temporal structure… a matrix of beats of different duration and position within

an isochronous time span that recycles repeatedly" (David Locke, "The Metric Matrix", Analytical

Approaches to World Music 1/1, 2011, p. 1,

https://iftawm.org/journal/oldsite/articles/2011a/Locke_AAWM_Vol_1_1.pdf). Beats are found through the

bell, not abstractly — Locke's teacher Godwin Agbeli told him "You should use the bell to find the

beats. Then feel the drum phrase in relation to the beats" (p. 7 n. 8). Locke also grades positional

stability: within a four-beat set, beats run 1-3-4-2 from most stable to most motile, so an onset on

beat 2 is inherently more in-motion than one on beat 1.

The architecture that matters for a loop: fixed timeline, fixed supporting parts, and a lead drummer

who selects from a finite indexed library of cycle-length pre-composed phrases. Locke's Dagomba example

gives the lead's material as "talks", each "a drummed theme consisting of one or more short phrases,

typically not longer than eight ternary beats," shaped by the Dagbani text it sets; after the lead

"calls in" the ensemble the talks are ordered live across a performance "that might last for several

hours" (pp. 61–62). The generation vocabulary is a language, which is what bounds it.

The analytic frame is DISPUTED and must be reported as open. A. M. Jones's "three against two" is

reframed by Locke as "three WITH two — the simultaneous presence of both ways to organize perception of

musical time" (p. 8), and Locke rejects the essentialising that came with the original. Agawu goes

further and disputes the apparatus itself, arguing that African music operates within a single meter

with "a firm and stable background" and a "fluid foreground," and that the scholarly fixation on

polyrhythm and additive rhythm has functioned to exoticise the music (Representing African Music,

Routledge, 2003; summary at https://en.wikipedia.org/wiki/Cross-rhythm). Anku takes a third position,

that time-points serve a structural rather than a metric purpose. Chernoff supplies the aesthetic

argument rather than the technical one: repetition's function is participatory — the cycle persists so

that other people can act inside it (African Rhythm and African Sensibility, Chicago, 1979,

https://archive.org/details/africanrhythmsens00cher; characterised from the book's argument, not

quoted).

5.4 Central African forest polyphony — a constrained variation grammar

Arom's model, as summarised in the survey literature, has at most four parts, is "an ostinato with

variations" comparable to a passacaglia, consists of "the repetition of periods of equal length, which

each singer divides using different rhythmic figures," and rests on the idea that "patterns are based

on a super-pattern which is never heard" (https://en.wikipedia.org/wiki/Music_of_the_Central_African_Republic).

The same source is careful that the performers "do not learn or think of their music in this

theoretical framework," which is load-bearing: the periodic model is an analyst's construct that

predicts the surface, and Arom's validation method — recording parts separately and playing the results

back to the musicians — exists precisely to test the gap.

Susanne Fürniss gives the concrete grammar for Aka polyphony (Analytical Approaches to World Music

Monograph Series vol. 1, 2021, CC-BY, https://zenodo.org/records/5608397, pp. 12–13). Four named parts

enter at fixed staggered positions — mòtángòlè on beat 1, dìyèí and òsêsê on beat 5, ngúé wà lémbò on

beat 6 — so no two parts articulate the cycle boundary together. Variation runs under two named

techniques: kɛtɛ bányɛ, "take a shortcut," which is mainly substitution of degrees at fixed intervals,

a fifth below or a fourth above or a neighbouring degree, plus embroidery around a sustained tone; and

kùká ngó dìkùkɛ, "simply cut it," which splits the cycle into short cells and varies those. Yodel is

itself listed as a variation process because switching laryngeal mechanism changes register and vowel

colour over the same pitch content. Singers change which part they hold mid-performance, without cue —

continuous re-voicing at zero structural cost. Fürniss is explicit that her barlines "are in no way

measure bars," so no downbeat accent may be imported with the transcription.

score-side constraints and both are checkable by a predicate over the notes, which is the shape

theme_compositions.py already uses for its one-deviant-per-head rule.

5.5 India — the density ladder, the grid re-tiling, and the computable cadence

Carnatic laya degrees run chauka 1, vilamba 2, madhyama 4, drut 8, adi-drut 16 strokes per beat

(https://en.wikipedia.org/wiki/Tala_(music)) — the same 1:2:4:8:16 ladder as Javanese irama, reached

independently. Hindustani layakari adds odd multipliers: barabar ×1, dugun ×2, tigun ×3, chaugun ×4.

Carnatic gati or nadai goes further and re-tiles the grid itself, swapping the beat's subdivision from

four to three, five, seven or nine mid-piece over the same tala — the most aggressive density device in

any tradition surveyed here, because it changes the grid rather than the fill rate.

The development architecture is a graded onboarding ramp rather than a switch: alap is unmetered, jor

begins "when a steady pulse is introduced," jhala arrives "when the tempo has been greatly increased,

or when the rhythmic element overtakes the melodic" (https://en.wikipedia.org/wiki/Alap). In khyal the

mechanism keeping a very long slow cycle alive is stated directly: "In each of these two songs, the

rate of the tala counts gradually increases during the course of their performance"

(https://en.wikipedia.org/wiki/Khyal).

The tihai is the tradition's computable closure device: a phrase repeated three times, calculated so

the third repetition's final stroke lands exactly on the target. With P the phrase length in matras, G

the gap and D the matras remaining to the landing, bedam solves 3P = D and damdar solves 3P + 2G = D

(https://en.wikipedia.org/wiki/Tihai). The landing target is not necessarily beat 1 — Carnatic practice

resolves onto the composition's eduppu, which may be offset (Glossary of Carnatic music). Tihais

"distort the listeners' perception of time, only to reveal the consistent underlying cycle at the sam."

That is a precise description of what a good pre-loop turnaround should do.

a cadential figure that is guaranteed to land on the loop point while obscuring the meter for the

three or four bars before it. harness/music_gen/loop_seams.py currently composes turnarounds under

harmonic rules only; a rhythmic-cadence rule of this kind is the natural S6.

5.6 Sri Lanka and Tamil Nadu — what is documented, and what is not

This is the thinnest section in the research and the thinness is real rather than an artefact of one

search. What is retrievable: the Kandyan ensemble uses the geta beraya with thammattama twin drums and

thalampota cymbals, whose stated function is "to help dancers maintain rhythm" — a dedicated timekeeping

idiophone separate from the drums, the same functional slot as the West African bell and the Balinese

kempli (https://en.wikipedia.org/wiki/Kandyan_dance). The vannam repertoire binds a specific drum part,

a specific sung text and specific choreography into one named item, eighteen classical vannam with a

seven-component internal form — per-item co-composition rather than a generic cycle. In Tamil Nadu the

periya melam works within Carnatic tala frameworks, so 5.5 applies, but no source retrieved gives

cycle-specific detail (https://en.wikipedia.org/wiki/Nadaswaram).

The most important retrievable finding is negative and must not be smoothed over. The University of

Canterbury thesis on low-country Raigama ritual drumming characterises the tradition by "an expressive

and illusive sense of timing which makes it appear to be free of beat, pulse and metre"

(https://ir.canterbury.ac.nz/items/fd1d277d-a46f-49d4-b6a2-fe51c6eac3da; full text inaccessible). Low

country drumming is not a fixed-cycle tradition and forcing it into the colotomic frame used above

would misrepresent it. Named gaps left open: the processional mallari repertoire, and parai and parai

attam, the latter carrying substantial recent scholarship on its status as a Dalit art form and its

contested caste politics — a gap that should be filled from dedicated Tamil-language and Indian

university sources rather than inferred.

5.7 The convergence worth noticing

Two traditions with no contact arrived at the same ladder. Javanese irama runs 1:2:4:8:16 peking

strokes per balungan beat; Carnatic laya degrees run 1:2:4:8:16 strokes per beat. India adds odd

ratios and grid re-tiling that Java and Bali do not use. Bali sounds the whole pyramid at once rather

than moving through it. Four independent answers to one question — how do you change the amount of

music without changing the music — and all four keep the cycle and move the density.

---

6. Game loop form, specifically

6.1 The channel budget was a form constraint, and looping was an aesthetic

With three to five voices you cannot state bass, chords, melody and drums at once, so every classic

technique is a compression. NESdev defines arpeggio precisely as rapidly alternating "a tone

generator's period among two to four pitches to create a warbly approximation of a chord," typically

updated at the 50 or 60 Hz vblank rate (https://www.nesdev.org/wiki/Arpeggio); the ludomusicology

analysis names the purpose as "harmonic compression; reducing all explicitly stated harmonic content to

the fewest voices possible," which "liberates channels for other purposes"

(https://www.ludomusicology.org/2015/07/16/compositional-strategies-for-programmable-sound-generators-with-limited-polyphony/).

The famichord drops the fifth from seventh chords so a four-note harmony fits three channels. The Follin

brothers folded echo into a volume envelope where DuckTales spent a second channel on it

(https://retrogameaudio.tumblr.com/post/18825096545/nes-audio-single-channel-echo-by-the-follin-bros).

The folk history says loops existed because memory was small. Two sources say otherwise and this is

the most important DISAGREEMENT in the section. Karen Collins argues looping was at least in part an

aesthetic that grew as games became more complex, "rather than being the consequence of the limited

memory available" ("In the Loop: Creativity and Constraint in 8-bit Video Game Audio",

Twentieth-Century Music 4(2), 2007). Samuel Hunt measured it: 21,391 MIDI pieces across 11 console

platforms, segmented at bar level, mean repetition score 62.79% with SD 3.49 — "on average 62.79% of

the musical clips in the music are a repeat of a previous clip" — and the figure barely moves across

hardware generations despite roughly a tenfold storage jump from N64 to PlayStation

(https://boblsturm.github.io/aimusic2020/papers/CSMC__MuMe_2020_paper_10.pdf). Storage did not buy less

repetition. It bought longer tracks and more voices. That figure is already carried as FORM-03 in

school_grammars/KONDO_SCHOOL.md.

6.2 The tracker order list is a form editor, and Deus Ex is our own exemplar

A tracker module separates content — patterns — from arrangement, the order list of pattern indices,

with position-jump and pattern-break as runtime control flow

(https://en.wikipedia.org/wiki/Module_file). Three consequences: reuse is free, so long forms are

cheap and variation is what costs; the order list is rewritable at runtime, which is the primitive

every dynamic re-sequencing system descends from; and per-pattern variation within a repeated group is

the anti-fatigue lever, which is widely practised and, as far as this research could establish,

nowhere formally codified — so FOLKLORE.

Unreal Engine 1 shipped this as adaptive music. A single UMX held multiple named sections of one piece

— Action, Ambient, Level Intro, Pre-Victory, Suspense, Tension, Victory — and the engine pattern-jumped

between them, with a level property setting SongSection and a MusicEvent trigger changing it at runtime

(BeyondUnreal Legacy:Music and OldUnreal community documentation; the wiki 403'd on direct fetch, so

this is at one remove). Deus Ex's music "changes to a different iteration of the currently playing song

based on the player's actions" (https://deusex.fandom.com/wiki/Deus_Ex_Soundtrack). The architecture is

efficient because one module is one instrument bank plus one pattern pool plus N arrangements — the

ambient and combat versions share their samples and most of their patterns, and vertical layering and

horizontal re-sequencing become the same operation at pattern level.

This matters to us specifically: EX_124 is the Deus Ex UNATCO theme, it is a corpus row, and its

card is the worked example this lane and the gate lane were both pointed at — a track whose

repeat_map marks ten of its eleven sections as repeats at similarities from 0.9623 to 0.9994, whose

boundaries fall every 8.6 to 15.6 seconds, and which carries exactly two boundaries above 0.9 in a

list of ten. It is section 8.2's two-tier schedule, measured on a track Josh named.

Alexander Brandon's own anti-fatigue statement is blunt: "repeat a theme

too much and you get players that start listening to their own music collection and turning the game

music off" (https://ocremix.org/info/Composer_Interview:_Alexander_Brandon).

6.3 The JRPG loop, with numbers

A survey of ten canonical tracks by Uematsu and Kikuta found 6 of 10 in ternary form, 2 of 10 in four

equal sections, 8 of 10 employing ostinatos throughout, all but two carrying "a tag or a breakdown to

rhythmic hits" for relief, section entries "almost all square and on the beat," and — the number worth

carrying — "most of the tunes centre around a continuous chunk of thematic melody of around 30-50

seconds' length… often that's 16 bars long." Loops ran from under a minute to 2:30

(https://drumchant.wordpress.com/2021/02/14/jrpg-song-forms/). These figures already ground LOOP-01

through LOOP-06 in school_grammars/SNES_JRPG_SCHOOL.md.

Sugiyama supplies the counter-doctrine, and it is a real fork. He wrote around 60 pieces for Dragon

Quest VI and cut to 30, reasoning that "the more songs there are, the less of an impression each song

leaves" (https://shmuplations.com/dragonquestvi/). Yuzo Koshiro supplies the best single framing

sentence in the whole game-music literature for our purposes: he likens house music to game music

because of "the way the phrases loop" (https://shmuplations.com/sormusic/). Loop-native dance music is

the correct model for a cue, not song form.

6.4 The two primitives, and what the middleware actually exposes

Michael Sweet's Writing Interactive Music for Video Games (Addison-Wesley, 2014) is the standard

reference and its chapter structure is verifiable from the publisher's sample: Ch. 7 "Composing and

Editing Music Loops" (including Reverb Tails and Long Decays at p. 136), Ch. 8 "Horizontal

Resequencing" (Crossfading, Transitional and Branching Scores), Ch. 9 "Vertical Remixing" (Deciding

How Many Layers to Use at p. 158; Nonsynchronization of Layers at p. 162), Ch. 10 "Writing Transitions

and Stingers" (Transition Matrixes at p. 167), and Ch. 18's "Variation and Randomization" and "Random

Playlists, Track Variation, and Alternative Start Points" at p. 255

(https://ptgmedia.pearsoncmg.com/images/9780321961587/samplepages/9780321961587.pdf). Sweet defines a

stinger as "a short musical phrase (3–12 seconds) that acts like a musical exclamation point," and

attributes to George Sanger the framing "Repetition is the problem," observing that "the play

experience is typically far longer than the music can support."

Bear McCreary's sidebar in that book (ch. 1, pp. 22–23) supplies the mechanism and it is habituation,

not boredom: "Our brains have evolved to filter out information that has no meaning… Before long, the

subconscious makes a connection between that music and that event and filters out the music, because

the information no longer carries meaning. Music you wrote to be as ominous as a lion in the reeds is

now no more effective than a babbling brook."

Winifred Phillips (A Composer's Guide to Game Music, MIT Press, 2014) supplies the term repetition

fatigue and prescribes instrumentation change, theme transformation and fragmentation, and for loops

specifically slow textures, perpetual development and variations. She also treats repetition as an

asset first, citing the four-chord motif of Assassin's Creed Liberation recurring fourteen times in the

main theme alone before resurfacing in menus and stealth sequences

(https://www.gamedeveloper.com/audio/game-composers-and-the-importance-of-themes-repetition-in-game-music-pt-2-).

Phillips and McCreary DISAGREE, and the resolution is the useful part: recurrence across contexts

accrues meaning, recurrence within one continuous listening session decays it. The book text itself was

not fetchable, so treat the terminology as reliable and the nuance as second-hand.

Wwise's object model runs Music Track to Music Segment to Playlist or Switch Container. A segment

carries an Entry Cue and an Exit Cue, with pre-entry and post-exit audio that can sound simultaneously

across a change — which is the engine-level answer to the reverb-tail problem, where a naive loop cuts

the last bar's decay dead. Transition rules are defined per source-to-destination pair with an exit

condition on the source (Immediate, Next Grid, Next Beat, Next Bar, Next Cue, Next Custom Cue, Exit

Cue), an entry condition on the destination, fade curves, and an optional transition segment; zero-length

transition segments are the standard way to fire a stinger across a transition; and placing custom cues

inside long segments is how transition latency is tuned, so cue density is the latency dial

(http://gamesounddesign.com/in-depth-creation-of-dynamic-music-in-a-video-game-page-four.html; the

wwiser reverse-engineering documentation at https://github.com/bnnm/wwiser/blob/master/doc/WWISER.md).

FMOD expresses the same primitives on a timeline: loop regions, transition markers and regions with

quantization "to the pulse of the music," transition timelines, marker conditions, and sustain points

that halt the cursor until released (https://limulo.github.io/game-sound-sae2017/fmod.html). Both

descend from iMUSE, whose branch and loop markers in a MIDI sequence let the score function "as a pit

orchestra, playing back shorter or longer sections of music while waiting for key events"

(https://en.wikipedia.org/wiki/IMUSE). Audiokinetic's and FMOD's own doc domains returned 403 or empty

bodies to automated fetch, so the enum names above are cited to doc-derived secondary sources and are

labelled as such.

6.5 Why the famous loops do not wear out, structurally

non-committal material — "music that doesn't really explain anything. It doesn't say if it's battle

or if it's night" — because the composer could not know the player's context

(https://daily.redbullmusicacademy.com/2015/08/c418-interview/). Systemically the game plays music

rarely: gameplay music resumes 10 to 20 minutes after the previous track finishes

(https://minecraft.wiki/w/Music, which flags the figure for verification). C418: "You should probably

play the music as few times as possible… the pause in between actually helps." The silence is the

anti-fatigue mechanism.

'Environmental BGM'… the environmental noise becoming the BGM"; Wakai judged that "even if we had a

piece of music in there, it wouldn't be able to match that sense of inspiration the player already

finds in that world"

(https://nintendoeverything.com/breath-of-the-wild-composers-on-changing-up-zeldas-music-formula-and-more/).

Structurally: randomised solo-piano fragments have no loop period to detect.

shifts or subtle harmonic change rather than exact repeat, with a modulation the melody never needs to

adapt to

(https://usutheoryiv.wordpress.com/2016/10/27/skyrim-and-immersion-jeremy-soules-use-of-ostinatos-and-instrumentation/).

A harmonically static bed is what makes a varied surface legible as variation rather than as a new

event. Note: no primary Soule quotation using the word "wallpaper" could be verified; do not attribute

it to him.

the sparse version, and instrumentation is the area's identity

(https://indiegamefans.com/soundtrack-spotlight-hollow-knight/). Vertical layering used as a spatial

map, so the player's own movement continuously re-orchestrates the loop.

(https://medium.com/game-audio-lookout/coherent-situational-music-in-undertale-e9d682c366ff).

in an exploration game like Outer Wilds, the music can lose importance and become wallpaper"

(https://www.gamedeveloper.com/audio/behind-the-hauntingly-beautiful-music-of-i-outer-wilds-i-). His

solution is fewer cues, not better loops — music marks discovery, not duration.

structural claim is made about them.

6.6 The independently-derived 3-5 number

Marcelo Martins, writing in Game Developer in 2012, gives the fatigue thresholds directly: loops

repeated one to two times remain tolerable, three to five approach problematic, and more than five

create significant listening fatigue; the target is that "music length should match or exceed average

gameplay time"

(https://www.gamedeveloper.com/audio/rethinking-the-audio-loop-in-games). He also names the structural

failure modes precisely — a key change that builds intensity collapses at the loop point, and

orchestral growth reverses there, "creating an 'empty' feeling" — and gives five cheap fixes: part

inversion (A-B-C to B-C-A), melody muting on alternate passes, strategic silence (he cites Halo muting

after three repetitions), random playlists, and an expandable roughly 30-second tail.

Josh's floor and Martins's threshold are the same number, arrived at independently, and neither cites

the other. That is the strongest external corroboration the composition floor has, and it is worth

recording as such. Note carefully that Martins is counting repeats of a WHOLE LOOP, while Josh's

sentence counts cycles of a repeating figure inside a track — the same number at two different scales,

which is exactly the two-tier structure section 1 derived from the corpus.

6.7 The two other real disagreements

6.4.

five minutes, briefly ran about 15, and shipped about 6, keeping narrative music at a single tempo and

key (https://en.wikipedia.org/wiki/Music_of_Red_Dead_Redemption_2); Rainbow Six 3 scored only a few

dramatic contexts to dodge the transition problem entirely. Against that, Sweet and Whitmore's

variation-bank doctrine says more variants is the fix. Both camps agree on the underlying principle —

the enemy is predictability, not repetition — and attack it from opposite ends.

---

7. What the empirical literature can and cannot license

the pooled exposure-affect effect at about r = .26 (Psychological Bulletin 106(2), 1989, DOI

10.1037/0033-2909.106.2.265; the effect size is second-hand from summaries). Modest is the operative

word: repetition nudges liking, it does not manufacture it.

review categorises 57 studies as 50 compatible, 5 mixed and 2 incompatible, and endorses the

descriptive adequacy of the curve while conceding inconsistencies with Berlyne's arousal mechanism

(Psychology of Music 45(6), 2017, DOI 10.1177/0305735617697507,

https://anthonychmiel.com/wp-content/uploads/2025/08/Chmiel2017_BackToTheInvertedU.pdf). It is a

vote-count, not a pooled effect size, and the model is rescued by being piecewise, which weakens its

bite.

and Pliner tested 0, 2, 8 and 32 exposures and found liking peaking at 8 and falling back to baseline

by 32 under FOCUSED listening, but rising linearly all the way to 32 under INCIDENTAL listening (JEP:

LMC 30(2), 2004, DOI 10.1037/0278-7393.30.2.370). Madison and Schiölde ran 28 spaced presentations of

real music, one per day over about four weeks, and found liking increasing monotonically at every

complexity level with "no indication of a U-shaped relation" (Frontiers in Neuroscience 11:147, 2017,

https://pmc.ncbi.nlm.nih.gov/articles/PMC5374342/). Their own reconciliation is methodological:

studies finding decline typically used stimuli that were "monophonic, synthetic, or short… repeated

within the same listening session." Decline under repetition is real but conditional — it belongs to

the listening condition, not to repetition itself.

one-minute excerpts of Berio and Elliott Carter either unaltered or artificially edited so segments

repeated; both repetition conditions were rated more enjoyable, more interesting and more artistic

than the originals, and a majority reported the spliced version sounded more likely composed by a

human artist (Empirical Studies of the Arts 31(1), 2013, 45–57, DOI 10.2190/EM.31.1.c). Correction to

a widely circulated misquote: the composers were Berio and Carter, not Boulez. She also showed that

attention migrates across time-scales with re-hearing — listeners initially detect roughly 1-second

repetitions more readily and only detect roughly 8-second ones after repeated exposure, with the

improvement on long units coming at the cost of the short ones (Music Perception 29(4), 2012). Her

reviewers note explicitly that On Repeat does not resolve when repetition ceases to be pleasurable, so

do not cite her for a satiation threshold (https://mtosmt.org/issues/mto.14.20.4/mto.14.20.4.albrecht.html).

literature's. ITPRA splits the response into Imagination and Tension before the outcome and

Prediction, Reaction and Appraisal after it; the prediction response rewards accurate expectation with

positive valence independently of whether the predicted event is good; and contrastive valence

explains why a braced-for outcome that resolves better than expected produces surplus pleasure (David

Huron, Sweet Anticipation, MIT Press, 2006; definitions cross-checked against two secondary sources,

the book and press page having returned 403). Meyer's earlier formulation is that affect arises when a

tendency to respond is inhibited (Emotion and Meaning in Music, Chicago, 1956). Repetition is how the

schema gets installed; the schema is what makes a well-timed violation pay.

Four things practitioners routinely over-extrapolate, stated so this project does not:

excerpt in one lab session, immediately contradicted by 28 spaced exposures with no decline. Any

"loop it four times" rule is CRAFT CONSENSUS.

for background use and weakest for the attentive listening most composers design for. Practitioners

usually invert this.

time inside a single listening. Applying it to "the third minute of the track" is a category error.

exists to stop the gate lane from writing an evidence-flavoured threshold that the evidence does not

support. Josh's 3-5 is a craft floor and Martins's 3-5 is a craft threshold; both are legitimate as

craft and neither is an empirical constant.

---

8. The novelty schedule doctrine

8.1 What makes a novelty event count

The distinction Josh's floor implies but does not spell out is between an event that changes what is

sounding and an event that changes only how it sounds. Both are useful; only the first can carry the

3-5-cycle law.

TierExamplesWhy it rates there
REAL — carries the lawa new instrument entering with its own line; a counter-melody; a density-level shift; a mode change; a metric or subdivision change; a texture inversion where lead and accompaniment swap roles; a genuine drop that removes lanes and re-entersadds or removes an independently followable voice, or changes the rule the music is running under
REAL — conditionala register migration of an existing line; a reharmonisation under an unchanged melody; a phase displacement of the melody against the ostinatocounts only if the moved element was already salient enough for the move to be heard as an event
CHEAP — never counts alonea filter sweep; a reverb change; a volume ramp with no lane change; a pan move; a one-shot riser; a delay throwchanges the treatment of an unchanged texture, which the ear absorbs as production rather than as music

The measured warrant for the cheap category is in section 1.3: corpus rows raise by 7.49 dB on

average against their siblings' 10.09 dB and fire fewer raises of 10 dB or more, while firing

lane-adding raises at 4.55 per minute against 3.10. Loved tracks do not get their novelty from

loudness. They get it from arrivals.

drops[].lanes_removed separate lane-changing events from level-only events, and the ten spectral

bands are a usable proxy for arrangement even though the card correctly refuses to call them a

roster. A cheap event is a raise or drop whose lanes_added or lanes_removed rounds to zero. That

test can be armed now and should be, because it is the one that stops a filter sweep from satisfying

a cycle counter.

8.2 The two-tier schedule

Derived from section 1.3 and from every tradition in sections 3 through 5, and stated as a target band

rather than a constant:

boundary every 5.79 bars, p10 to p90 of 3.06 to 15.81. This is the tier Josh's 3-5-cycle law governs:

at a two-bar figure, six bars is three cycles; at a four-bar phrase it is one and a half. Most of

these must be small — the corpus makes 28.6% of its boundaries weak and only 17.9% strong.

three-to-four-minute track. The corpus median is three strong boundaries at 0.88 per minute, against

siblings' 1.89 per minute. The ratio of the largest to the typical event should be about 1.9 or

steeper; flat is the filler signature.

measures as a sibling; a track that fires only tide has nothing to arrive at. The 3-5-cycle law is

satisfied by the tide. The thirty-year bar is carried by the landmarks.

8.3 What kind of novelty, in what order

CRAFT CONSENSUS, synthesised from sections 2, 3 and 5, stated as an ordering rather than a formula:

boundary at 5.1% of the running time — about 10 seconds into a three-minute cue. The first change

should be small, because its job is to teach the listener that this track changes.

cost nothing structurally and can be spent repeatedly; mode change, metric change and the arrival of

a genuinely new melodic idea each work once.

(section 2.3). A texture that thins after having thickened is relief; a texture that thickens more

slowly is nothing.

and note that the corpus's own loudest moment sits at 0.526 of the running time, so relief after the

peak is where the tradition and the measurement agree.

discontinuity at the point of maximum attention outperforms continuous variation.

theme-recalling variations into an already-escalating order. On a loop this means: once the schedule

is built, go back and add a plain statement of the head somewhere in the last third, so the wrap

arrives after a recognition rather than after a stranger.

8.4 The escalation question, answered by measurement rather than instinct

The natural instinct is that the last minute must be bigger than the first. The corpus does not support

it, and the honest result is more useful than the instinct.

MEASURED HERE. The mean strength of second-half boundaries minus the mean strength of first-half

boundaries is −0.014 for corpus rows against +0.006 for siblings, AUC 0.445, with only 42.9% of corpus

rows positive. Pooling all 106 strong boundaries across the 28 corpus rows, their median position is

0.477 of the running time and 47.2% fall in the second half — no back-loading at all. The single

strongest boundary sits at a median position of 0.547, and the dynamic peak at 0.526. Loved tracks do

not save their biggest event for the end. They put it just past the middle.

What they do instead is refuse to decay. Complexity in the last third minus complexity in the first

third is −0.036 for corpus rows and −0.056 for siblings across the whole pool. Excluding the eleven

corpus rows with a baked fade-out — a mastering artefact that depresses every late-track measure — the

seventeen that end by looping or through-composing read −0.020 against their siblings' −0.094, AUC

0.725. Siblings run out of ideas and coast to the end. Corpus rows arrive at the loop point carrying

the same complexity they had at the start.

So the doctrine is NON-DECAY, not escalation. The last minute must not be the first minute, and the

way it differs is by being differently constituted at equal weight — different instruments, different

register, different treatment of the same material — rather than by being louder or busier. That is

also the only reading compatible with a cue that loops, because a track that escalates to its end and

then restarts is exactly Martins's documented failure: "orchestral growth reverses at the loop point,

creating an 'empty' feeling."

read directly, and cadence_and_ending.ending supplies the fade-out exclusion the read requires. A

band of roughly −0.05 to +0.10 on that difference, computed only on non-fade endings, is defensible as

a diagnostic. It must not be stated as a gate until it has a calibrated null of its own.

8.5 The worked schedule

For a three-to-four-minute cue built on a repeating cycle, as a starting target to be varied by purpose

class, not as a template to be filled:

PositionWhat happensTier
0 to 5%the material states plainly; the head lands inside the first phrase
~5%first change, deliberately small — one voice added or one register openedtide
5 to 35%tide every 4 to 8 bars: counter-melody in and out, octave shifts, instrumentation swaps, the cheap-to-generate kindstide
~30 to 40%LANDMARK 1 — a density-level shift or a new voice with its own line; the texture changes rule, not just colourlandmark
40 to 55%tide continues at the new density; the dynamic peak lands around 0.53tide
~55 to 65%LANDMARK 2 — the largest event; mode change, metric change, or full tutti arrivallandmark
~65 to 75%the deep relief; the thinnest texture in the track, and the one place a sparse bar is correctlandmark by absence
~75 to 85%LANDMARK 3 — the return, re-orchestrated; a plain restatement of the head so the wrap follows a recognitionlandmark
85 to 97%tide at held complexity; no decay, no escalationtide
final 1 to 2 barsthe composed turnaround: weakened cadence, no perfect authentic cadence, ideally a three-fold cadential figure landing exactly on the wrap

8.6 What the doctrine forbids

(strong-boundary share 36.4% against the corpus's 17.9%).

binds the thin sections hardest; a relief section is a scored event with a declared function, not a

gap.

---

9. The measurement ledger

9.1 Measurable from the card as it stands

novelty_peak_strength, remembering it is normalised to each track's own maximum and is therefore a

shape measure, not an absolute one.

drops[].lanes_removed, db_delta.

boundary_step_share.

harness/music_gen/loop_cadence.py, both of which already carry positive controls.

9.2 Not measurable today, with the measurement that would close each

candidate proxies both fail for stated reasons. complexity_law.bar_length_s inherits whichever

metrical level the beat tracker chose, and the card's own octave_ambiguity_note says bar figures

must be compared within a card and never across cards. tempo_and_time_grammar.dominant_onset_period_s

finds a sub-phrase period: MEASURED HERE, its median across the 28 corpus rows is 1.74 seconds, and

dividing section length by it puts only 36% of corpus rows' medians inside a strict 3-to-5 band

against 28% of siblings, which is no signal. The measurement that closes it is a per-section CYCLE

ESTIMATOR — the lag maximising the diagonal autocorrelation of the self-similarity matrix within a

section, reported with a confidence — which would give the gate a real denominator instead of a

borrowed one. This is the single most important addition for the parallel gate lane.

event added a VOICE rather than a BAND. lane_analysis declares its own confidence as

LOW-FOR-ROSTER and max_lanes is pinned at 9 for hits and siblings alike, which two full re-runs have

confirmed is a method ceiling. Source separation or a trained multi-instrument recogniser closes it.

transcription; the card says so itself in its cadence note. Until then, a mode or harmony change cannot

be verified from audio and must be declared on the score side.

parameter. It needs the cycle estimator above plus a melodic onset track, and it is measurable in

principle once both exist.

9.3 Compositional constraint against after-the-fact measurement

PrincipleWhere it lives
declared invariant per track; odd-length ostinato; staggered entries; bounded substitution grammar; tihai cadence; two section archetypes; density-level shift with declared idiom switchCONSTRAINT — harness/music_gen/arm_b_scores.py, theme_compositions.py, loop_seams.py
boundary spacing; boundary-strength hierarchy; lane-adding event rate; complexity non-decay; cheap-versus-real event classification; peak position; silence budgetMEASUREMENT — the formula card, computable today
cycles-since-last-real-event; instrumentation change across a repeated section; reharmonisation detection; phase displacementNEITHER YET — see 9.2

---

10. Honest register

spine and the region pages outrank every line of it.

calibrated multiplicity null of their own. Only section_count and its restatements carry an

out-of-sample claim, and that claim belongs to PATTERN_FINDINGS_V2, not to this document.

as a finding for that program to rule on. Nothing is changed here.

and FMOD official documentation wording (both blocked automated fetch, so middleware enum names are at

one remove); Sisman's Grove article; Kinderman's OUP monograph; the full texts of Sweet, Phillips and

Collins; Riley's published performing directions; radio-play saturation research, on which nothing is

asserted; Stardew Valley and Celeste primary interviews; and the Sri Lankan and Tamil material in 5.6,

including the politically sensitive parai scholarship, which should be filled from dedicated

Tamil-language and Indian university sources rather than inferred.

substrate for any given chapter is settled by its region page and by MUSIC_COMPOSITION_DOCTRINE.md

§4, not by this document, which names no chapter and no region.

Generated by harness/site/structure_site.py — the URL path is the repo path. review root