HOOK_AND_TENSION_ARCHITECTURE.md

music/HOOK_AND_TENSION_ARCHITECTURE.md

HOOK CONSTRUCTION AND TENSION/RELEASE ARCHITECTURE

CANON SUBORDINATION — this document is PROPOSAL-TIER: it serves canon and never outranks it.
Canon served: the composition floor at docs/spine/DECISIONS_PENDING_JOSH.md (commit da627060), and the
per-chapter music_mood anchors of docs/spine/CH_03.md, docs/spine/CH_04.md and
docs/spine/CH_10.md.
If this document disagrees with canon, CANON WINS and this document is the defect.
Tier: RESEARCH SYNTHESIS, proposal-tier. This document is subordinate to canon: the CVD, the T1
foundation docs, the spine (docs/spine/CH_*.md), the Tier-2 region pages
(_source/02_Tier_2_Region_Pages/) and the T0 registries. Where it disagrees with canon, canon
wins and this document is the defect. Nothing here is applied until it is ratified.
Every substantive claim carries a source. Claims that are craft consensus rather than evidence are
labelled CONSENSUS. Claims measured from our own formula cards for this document are labelled
MEASURED-HERE and carry no calibrated multiplicity bar — they are descriptive reads at N = 28,
in the same honest register as PATTERN_FINDINGS_V2.md section 10.
Commissioned by: the composition floor at docs/spine/DECISIONS_PENDING_JOSH.md (MUSIC ROUND 1
GRADED + THE COMPOSITION FLOOR, commit da627060) — "its got to have the hooks and complexities.
There cant just be stacked noise", "no track will be boring or filler", and the requirement to set
"the next bar for video game sound tracks".

0. DERIVATION — what canon was read before this was written

docs/spine/CH_NN.md (Asset anchors, the music_mood bullet), the region page Section 10

(the Music Sound and Composition Brief), registries/T0_Theme_Registry,

docs/translation/T99_Translation_Audio.md Music Steps 1-7, and

docs/proposals/MUSIC_COMPOSITION_DOCTRINE.md.

docs/spine/CH_10.md L224 (the music_mood asset anchors); and Section 10 of

_source/02_Tier_2_Region_Pages/ for flores_island (L844), bali (L1894),

central_africa_congo (L1017), ethiopia (L713), south_africa (L599), south_india (L787),

sri_lanka (L908), sumatra_java (L2229), swahili_coast (L838), rift_valley (L624) and

west_africa (L2398).

formula-card/v1), build/audio/exemplars/PATTERN_FINDINGS_V2.md,

harness/music_gen/exemplar_hook_signatures.json (13 signatures, 14 declared gaps),

harness/music_gen/nostalgia_score.py and docs/proposals/music/NOSTALGIA_RUBRIC.md.

identity a player carries, and the region pages' Section 10 briefs, which already specify the

cultural substrate. This document supplies the craft layer between those two and never re-decides

either. It changes no canon; every canon-facing consequence below is stated as a finding.

1. What a hook is

1.1 Three traditions, one object

6/1, 1987, 1-20, https://www.tagg.org/xpdfs/burns87.pdf) a hook is the part of a record that

stands out and recurs, and Burns's contribution is that the hook need not be melodic at all. He

sorts hooks into musical (rhythm, melody, harmony, instrumentation), textual (lyric), performance

(tempo, dynamics, improvisation) and production/technological (editing, mix, channel balance,

signal distortion, sound effects) families. The practical consequence for us is direct: a

gong-waning cycle, a whip-crack placed on a beat, a single detuned bronze pair, or a sudden

register drop are each admissible hooks by Burns's taxonomy, and a cue with no singable tune is

not automatically hookless.

rest of the movement is heard against. Caplin's account (Classical Form, Oxford, 1998) makes this

operational rather than impressionistic — see section 2.

memory claim, not a structural one, and no purely structural measurement can certify it. That gap

is the honest centre of this document.

1.2 Hook against riff and ostinato

summarized at https://mtosmt.org/issues/mto.02.8.4/leydon_text.html). Musematic repetition is

near-unvaried repetition of a museme — the smallest meaningful unit — and its paradigm case is the

riff or ostinato; it is circular, synchronic and open. Discursive repetition is repetition of

longer, syntactically complex units — phrases, strophes, sections — and it is linear and

self-sufficient.

saturation and groove, a hook by shape and return. A cue that has only musematic repetition is the

thing Josh named as stacked noise: material that recurs without ever being answered.

begins, goes somewhere and comes back — even when its surface is an ostinato culture.

musematic density. motif_economy.self_similarity_lift (peak minus baseline) is the closest thing

the card has to a discursive-return measure. Corpus medians, MEASURED-HERE at N = 28: 3-gram

repetition 0.432 (p10 0.157, p90 0.791), self-similarity lift 0.181 (p10 0.058, p90 0.328). A

candidate with high 3-gram repetition and near-zero lift is a riff with no theme.

2. Phrase structure — the engineering under a memorable tune

2.1 The sentence

(being in the middle) and cadential (ending). The presentation is a two-bar basic idea plus its

near-literal repetition, usually over a tonic prolongation. The continuation fragments the basic

idea, accelerates harmonic rhythm and increases surface activity. The cadential function

liquidates the idea's characteristic features and closes

(https://mtosmt.org/issues/mto.14.20.1/mto.14.20.1.aziz.html;

http://shanahdt.github.io/MUSI4331/lessons/phrases1.html).

unit and immediately proves it is a unit by repeating it; from bar 3 the listener can predict bar

4. The continuation then rewards that acquired competence by breaking the unit into pieces the

listener already owns. Fragmentation is only intelligible because the whole was stated twice

first.

similarity, hard-gated at 0.80; P2 continuation fragmentation; P3/P4 the 8-bar grid). The

census reports P1 at 3 of 14 — our landed cohort answers its head with a mirror instead of

repeating it (docs/proposals/music/NOSTALGIA_RUBRIC.md section 6). That is the single most

diagnosable phrase-level defect we have measured.

2.2 The period, elision and hypermeter

idea and closes strongly. Caplin treats antecedent/consequent as functionally comparable to

presentation/continuation, which is why a tune can feel classically shaped without being a

textbook example of either.

removing the seam between phrases without shortening either. CONSENSUS.

(Rothstein, Phrase Rhythm in Tonal Music, 1989, cited in

https://mtosmt.org/issues/mto.22.28.4/mto.22.28.4.temperley.pdf). Hypermetric regularity is what

makes a listener feel where a section will end before it ends, and hypermetric irregularity

(a five-bar or six-bar group) is a tension device precisely because that expectation exists.

bar arithmetic at an assumed 4/4 (tempo_and_time_grammar.metre.measured is null on EX_124), so

hypermeter has no representation. Closing measurement named in section 12.

3. Melodic anatomy

3.1 Contour archetypes

contours over ascending and descending shapes ("The Melodic Arch in Western Folksongs", Computing

in Musicology 10, 1996; https://www.academia.edu/50292501/The_Melodic_Arch_in_Western_Folksongs).

The arch is the default phrase shape of a large, mostly European folk corpus — which is a fact

about that corpus, and the Essen database's over-use as a proxy for music in general is a known

methodological complaint (https://dl.acm.org/doi/10.1145/3469013.3469016).

filling the gap) archetypes come from Meyer and Narmour. Gap-fill originates with Meyer (1973):

large intervals imply smaller intervals in the opposite direction

(https://mutor-2.github.io/ScienceOfMusic/units/08/).

Reanalyses find that gap-fill does poorly as a way listeners classify melodies

(https://www.researchgate.net/publication/271681259_Questioning_a_Melodic_Archetype_Do_Listeners_Use_Gap-Fill_to_Classify_Melodies).

is axial with repeated pitches, not an arch. Our own card carries a gap-fill-after-leap axis whose

hit median and sibling median are both 0.000 with AUC 0.533 (PATTERN_FINDINGS_V2.md section 10)

— gap-fill does not separate our corpus rows from their album siblings, and should be used as a

craft option, never as a gate.

3.2 Post-skip reversal, and the deflationary explanation

two notes: a small interval implies continuation in the same direction, a large interval implies

reversal. Huron reduced it to five governing principles — registral direction, intervallic

difference, registral return, proximity, closure (Huron 1997; summarized at

https://mutor-2.github.io/ScienceOfMusic/units/08/).

to the mean: a leap tends to land far from the tessitura's centre, and the next note statistically

returns toward it, which looks like a reversal principle without being one ("Why Do Skips Precede

Reversals? The Effect of Tessitura on Melodic Structure", Music Perception;

https://www.researchgate.net/publication/224982434).

then a step back has done the ordinary thing, and the ordinary thing is not a fingerprint. If a

leap is to be an identity, it has to be the same leap every time, in the same metric position.

and interval_vocabulary_semitones. EX_124 carries a single -14 semitone leap at note index 6 out

of an interval vocabulary of exactly four values {0, 10, 12, 14}. That is the shape to aim for:

a tiny vocabulary with one deliberate outlier at a fixed index.

3.3 Tessitura, span and rhythmic distinctiveness

can reproduce without notation must fit an untrained voice. Our rubric encodes it as C5

(range <= 17 semitones, alphabet <= 8 pitches, tessitura MIDI 55-84), hard-gated

(NOSTALGIA_RUBRIC.md axis 2).

pitches — Burns's rhythm-hook category exists for exactly this reason, and Kondo's own

three-note head cells are described by their half-quarter-half template as much as by their

intervals (NOSTALGIA_RUBRIC.md H7, citing ECON-07 rhythmic economy). An upbeat entry changes

which note lands on the downbeat and therefore which note is heard as the identity.

carries two unisons in six intervals — a fact our own rubric had to write into feature I0 as a

declared derived floor after a texture control exposed the naive reading

(NOSTALGIA_RUBRIC.md axis 3).

the_hook.range_semitones. Corpus hook range median 22.0 semitones, MEASURED-HERE, range

0.67-52.0 — far wider than the rubric's 7-17 singability band, because the card's hook extractor

reads the audio's melodic band and the rubric reads an authored head cell. These two numbers are

not comparable and must never be quoted against each other.

4. The empirical memorability literature, and its limits

4.1 Earworms

imagery tunes against 100 never-named tunes controlled for popularity and style, and compared them

on 83 melodic features (Psychology of Aesthetics, Creativity, and the Arts 11/2, 2017, 122-135;

https://www.apa.org/pubs/journals/releases/aca-aca0000090.pdf). The reported discriminating

features are faster tempo, more common global melodic contour patterns (the plain rise-then-fall

arch being a positive predictor), and unusual interval structure — larger or less expected leaps

than the contour alone would imply.

recency were themselves strong predictors; the musical effects were modest; the design is

correlational and the sample is UK self-report. The authors' own framing is that these properties

interact rather than that any single feature determines stickiness.

detail. Put the novelty in the interval and the harmony; keep the contour common. Our own rubric

already says this at axis 2 and cites the same paper.

4.2 Repetition and mere exposure

participant — repeated material is heard as something you are inside rather than something you are

watching (https://mtosmt.org/issues/mto.14.20.4/mto.14.20.4.albrecht.html).

model adds a tedium factor, producing the inverted-U: liking rises with exposure to a peak and

then falls. The peak's location is stimulus-dependent and is not a constant anyone has pinned

down (https://arxiv.org/pdf/2210.16226).

cue heard in a loop, the tedium factor overtakes the habituation factor at roughly the fourth

cycle unless something new enters. That number is a ruling, not a measurement, and it is

consistent with the shape of the literature without being derivable from it.

median is 17 against a sibling median of 9 (PATTERN_FINDINGS_V2.md section 10), the strongest

measured axis in our whole bank. A 17-section track that runs 196 s changes something roughly

every 11 s. That is the measured form of the 3-to-5-cycle law.

4.3 Expectancy models

per event, a probability distribution over the next event, and from it an information content

(surprise) and entropy (uncertainty). It outperforms static rule-based models such as I-R for

pitch expectation in many contexts (https://onlinelibrary.wiley.com/doi/full/10.1111/j.1756-8765.2012.01214.x;

https://nyaspubs.onlinelibrary.wiley.com/doi/10.1111/nyas.13654).

(Music Perception 37/2, 165; https://online.ucpress.edu/mp/article-abstract/37/2/165/109530), and

the predictability-liking relation recovered from it is the same inverted-U shape as section 4.2.

candidates, so no expectancy model can be run in our pipeline today. Named as a gap in section 12.

5. Hook placement and economy

to 2015 and found time-before-voice-entry and time-before-first-title-mention both fell

significantly across the period (Musicae Scientiae 22/3, 2018, 291-304;

https://journals.sagepub.com/doi/abs/10.1177/1029864917698010), with the popular summary of the

same work putting average intros at over 20 s in the mid-1980s against about 5 s today

(https://news.osu.edu/has-music-streaming-killed-the-instrumental-intro/). This is a fact about a

distribution channel, not about memorability, and it should be read as one.

across the 28 hits has median 0.575 s (p10 0.104, p90 5.12, max 61.6). Half of the tracks Josh

named put identifiable melodic material inside the first six-tenths of a second. But

time_to_hook_frac sits at AUC 0.344 against siblings (PATTERN_FINDINGS_V2.md section 4) —

the siblings do it too, so early arrival is a floor and not a differentiator.

MEASURED-HERE, time_to_full_texture_s / total_length_s has median 0.032 across the 28: the

arrangement is essentially open within the first three percent of the runtime. The corpus does not

do long ramps into the material. It states the thing, then varies it for three minutes.

should appear once in its plainest form before it is transformed, because every later variation is

parsed against that statement. Caplin's presentation function is the classical form of this rule;

the film-scoring version is the restraint argument — the theme lands at the climax because the

film withheld it (https://www.cinemagicscoring.com/post/behind-the-music-how-film-scores-work).

bullet: state a fragment plainly early, withhold the complete harmonized statement for a later

structural payoff. This is the design the leitmotif architecture already assumes.

time_to_hook_bars, time_to_full_texture_s) plus motif_economy.recurrence_lag_s. Note that

recurrence_lag_norm was V1's best hypothesis and is now the second-worst axis in the bank at

AUC 0.505 (PATTERN_FINDINGS_V2.md section 6) — do not build a delayed-return gate on it.

6. Hook in a non-vocal, non-Western, period-authentic frame

The slice's traditions are not tune cultures with exotic timbres. In several of them the

identity-bearing unit is not a melody at all, and the region pages already say so. The craft answer

per tradition:

6.1 The cycle is the hook (Bali, Java, Sumatra)

punctuated by a colotomic structure of gongs marking a cycle of fixed length, with elaborating

parts above (https://ecampusontario.pressbooks.pub/beyondtheclassroom/chapter/balinese-music-gamelan/;

https://grokipedia.com/page/Colotomy). Kotekan is the interlocking of paired parts into a single

line faster than either player produces (https://en.wikipedia.org/wiki/Kotekan).

cycle closes and hears everything against that. The identity a player carries out of the room is

a period and a punctuation pattern, not a tune.

(L1894) says Bali is "the arc's first region where music is not scoring but structure" and that

the music "is built to be READ"; it also documents the ombak beat-note engineered by deliberately

mistuned pairs [SRC_00157, p. 316-317, p. 319-320]. sumatra_java.md Section 10 (L2229) states

plainly that its processed sources carry almost no instrumental ethnography and that the region's

sound is built around the voice; and the Minangkabau talempong pacik tradition is itself

interlocking, three players on six gong-chimes playing anak, tengah and peningkah.

i.e. 20.84 cycles per minute) read together with section_boundaries_s. A colotomic cue should

show a strong, stable dominant onset period and section boundaries that are integer multiples of

it. Nothing in the card checks that multiple relation today; it is a cheap addition.

6.2 The phrase is speech (Ile-Ife/Yoruba, and the talking drum)

lexical tone levels of Yoruba, and the repertoire consists of real texts — oriki praise poetry and

owe proverbs (https://www.frontiersin.org/journals/communication/articles/10.3389/fcomm.2021.652542/full).

Villepastour's study of the bata shows the encoding is no less precise for being less iconic

(https://ethnomusicologyreview.ucla.edu/journal/volume/15/piece/481).

the tradition's own scholar that the dundun "can use codes to produce a type of language which

nearly all Yoruba people can decode" [SRC_00260, p. 135], and records the Ifa corpus imitating the

drums by name [SRC_00229, p. 112].

verbatim because it is a sentence, and it is memorable for the same reason a catchphrase is. Write

a drum phrase of fixed length and fixed contour, use it as a call, and let the ensemble answer it.

Do not develop it — a proverb that is developed stops being a proverb.

melodic band and returns available: false or noise on a drum-led cue. Named in section 12.

6.3 The ostinato is a body (Ethiopian highlands, Sri Lanka, Tamil South India)

anchihoye — carried on krar, begena, masenqo and washint

(https://en.wikipedia.org/wiki/Qenet; https://www.ncbi.nlm.nih.gov/pmc/articles/PMC10118148/).

Our own ethiopia.md Section 10 (L713) supplies the vernacular roster with descriptions attached

— karar the six-string lyre, bagana the eight-or-ten-string floor harp, masanqo the one-string

fiddle, kabaro, nagarit, sanasal the sistrum, maqwamiya the prayer-staff [SRC_00254, p. 104-105].

over a fixed cyclic accompaniment. The listener's memory hook is the interval that identifies the

qenet, not a phrase.

runs pallavi, anupallavi, charanam; the pallavi establishes the raga's identity and returns as a

refrain after each later section; the anupallavi rises in register and introduces contrasting

material; each line is then varied through composed sangati

(https://artiumacademy.com/blogs/what-is-kriti-in-carnatic-music/;

https://www.mtosmt.org/issues/mto.15.21.4/mto.15.21.4.schachter.html). That is a refrain-and-

graded-variation architecture that solves the loop-fatigue problem natively, and it is a far

better template for a 3-5 minute looping cue than a Western AABA.

all-night preaching hall and the protective chant with its procedure [SRC_00256, p. 138, p. 152;

SRC_00237, pp. 62-66]. The hook there is a recitation cadence and a room.

a raga is invisible to the card. nostalgia_score.py Y12 (characteristic-degree rate) is the

right shape of feature and exists only for authored symbolic material.

6.4 The texture is the hook (Baka/Mbuti forest, /Xam San)

hocketed motives and yodel register-shifts over interlocking ostinatos (Arom, African Polyphony

and Polyrhythm, Cambridge, 1991; https://en.wikipedia.org/wiki/Simha_Arom;

https://www.melodigging.com/genre/mbenga-mbuti-music). Among the Mbuti the singing uses short

hocket-like motives connected to the voco-instrumental hocket of the small flute.

and is already a complete composition brief: overlapping entries, each singer holding her note as

the next comes in, a chorus cascading downward like a waterfall, swelling until breath runs out,

then all stopping together sharply and cocking their heads to listen for the forest's echo

[SRC_00253, p. 171-172]. It names that listened-for reply as the region's musical signature.

not his melody family. The memorable object is the collective attack, the swell and the hard

unison stop.

organizing principle is that a song belongs to a specific creature, person or state, and names the

catalogue — the Cat's Song, the Song of the Caama Fox, the Broken String and the rest [SRC_00265].

Our canon's own design line is a catalogue of short named songs discovered rather than triggered,

sparse, vocal-first, with a deliberate absence of harmony. Khoisan practice supplies the technique:

cyclical forms, interlocking parts, handclap polyrhythm, yodel-like register shifts, hocketing,

and a gradual upward drift over the course of a song

(https://www.melodigging.com/genre/khoisan-folk-music).

for entry-pattern hooks, but max_lanes is pinned at 9 for hits and siblings alike — a

measurement ceiling in the spectral_band_proxy method, confirmed across two full re-runs

(PATTERN_FINDINGS_V2.md section 10). No arrangement-width claim is admissible until a real

instrument roster exists.

6.5 Olduvai, and the honest case of no record

and that no rhythmic structure, scale, call pattern or instrument is described anywhere on the

page. The chapter's musical content is the Ju|'hoansi healing-dance clapping carried from ratified

canon at reference register, plus "a deep-time silence at Olduvai".

Hohle Fels vulture-radius flute with five finger holes, roughly 35-40 thousand years old, plus

ivory flute fragments (https://www.sciencedaily.com/releases/2009/06/090624213346.htm). The

defensible register is therefore body percussion, voice, struck stone and end-blown pipe, and the

honest composition is near-absence with a floor of wind and stone, which is exactly what the

region page already specifies.

("no track will be boring or filler") and the canon brief ("the arc's quietest bed") pull against

each other. The resolution available inside both is that the Olduvai bed be short, be event-bearing

rather than static — a single struck-stone figure that recurs at a long, irregular period — and

never be the chapter's only cue. This is a craft recommendation to the music lane, not a canon

change.

7. The tension parameters, one by one

Farbood's model is the best single empirical map. Both her experiments manipulated harmony, pitch

height, melodic expectation, dynamics, onset frequency, tempo, meter, rhythmic regularity and

syncopation, and her model is built on a moving perceptual window and trend salience — tension at a

moment is a function of the recent trajectory, not the instantaneous state (Music Perception 29/4,

2012, 387-428; http://mp.ucpress.edu/content/29/4/387). Lerdahl and Krumhansl's hierarchical model

combines prolongational structure, pitch-space distance, surface dissonance and attraction, and

implementations of it correlate strongly with listener tension ratings (Music Perception 24/4, 2007,

329-366; https://www.fredlerdahl.com/s/Modeling-Tonal-Tension.pdf).

7.1 Harmonic

under changing harmony, and modal mixture all raise tension. Lerdahl and Krumhansl's pitch-space

distance is the formal version of tonal distance; surface dissonance is a separate additive term

in the same model.

and the card itself declares that a pivot-chord map needs symbolic transcription. Closing

measurement in section 12.

7.2 Melodic

strain — a line sustained near the top of its range — is the extreme case. An unresolved leading

tone is the melodic form of an unresolved dominant.

melodic-register curve on the card, only a global span; a time-series version is the cheap fix.

7.3 Rhythmic

dissonance into grouping dissonance (pulse layers at different speeds, of which hemiola is the

commonest case) and displacement dissonance (same speed, misaligned)

(https://viva.pressbooks.pub/openmusictheory/chapter/metrical-dissonance/;

https://mtosmt.org/issues/mto.14.20.2/mto.14.20.2.biamonte.html). Hemiola clusters at cadences in

triple metre precisely because it heightens the tension the cadence then resolves.

because it raises perceived intensity without breaking a loop's grid. CONSENSUS.

groove.onset_rate_per_s, windowed_tempo_stability. Corpus tempo_adjust_n median 13.5 against

a sibling median of 6 (PATTERN_FINDINGS_V2.md section 10) — the tracks Josh named bend tempo

more than twice as often as their album neighbours. Metric dissonance itself is unmeasured.

7.4 Dynamic

powerfully: sudden loudness is one of the acoustic events that trips the brain-stem reflex, which

has no opt-out, and cross-cultural work correlates chill incidence with sudden peaks in loudness,

brightness and roughness

(https://www.frontiersin.org/journals/psychology/articles/10.3389/fpsyg.2017.02044/full;

https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4756174/).

that has spent its range on a long ramp has none left for the surprise. This is the craft argument

against front-loading loudness.

event_grammar.raises[].slope_db_per_s. Corpus dynamic-range median 11.7 dB against a sibling

median of 13.7 (AUC 0.369) — the corpus rows are slightly narrower, not wider, which is a warning

against reading dynamic range as a quality axis.

7.5 Textural

is the strongest release gesture available without harmony. Sloboda's structural inventory found

shivers most reliably evoked by new or unexpected harmony and tears by appoggiatura and sequence

(Psychology of Music 19, 1991, 110-120;

https://journals.sagepub.com/doi/10.1177/0305735691192002), which locates the strongest affective

events in harmony and voice-leading rather than in texture — texture is the amplifier.

event_grammar.drops[].lanes_removed, complexity_law.index_per_bar.

7.6 Timbral

best-grounded psychoacoustic account: Plomp and Levelt showed maximum sensory dissonance at

roughly a quarter of the critical bandwidth, with consonance recovering as the separation exceeds

it, and the complex-tone case computed as the summed roughness of all partial pairs (JASA 38,

1965, 548-560; https://www.semanticscholar.org/paper/1d3ccc073b1b3f13f95e2392ff81a9eff0e7a4d2).

Vassilakis's model is the standard modern implementation

(https://www.mat.ucsb.edu/Masters/Brian_Hansen_Masters.pdf).

spectral_flatness_mean, signature_colour. Brightness is measured; roughness is not. Adding a

Plomp-Levelt/Vassilakis roughness curve is the single cheapest high-value feature we could add,

because it is computed from the spectrum we already extract.

8. Arch form, climax placement and release

8.1 Where the climax actually sits — our own answer

case that made it famous is contested on its own arithmetic: Lendvai's reading of the first

movement of Music for Strings, Percussion and Celesta requires an added bar of rest to reach 89,

the dynamic climax is at bar 55 but the tonal climax is at 44, and Somfai's work on the sketches

found no proportional calculation of any kind

(https://blogs.ams.org/jmm2021/2021/01/08/did-bartok-use-fibonacci-numbers-in-his-music/).

opening_and_flow_law.flow_curve.peak_position_normalised:

statisticdynamic peak positionflow-curve peak position
median0.5750.606
p100.1640.210
p900.8690.901
min / max0.000 / 0.9930.178 / 0.993
within 0.55-0.707 of 285 of 28
in the first half13 of 2813 of 28
after 0.757 of 2810 of 28

(0.618), and the distribution is nearly flat. Thirteen of twenty-eight tracks peak before the

midpoint on both measures. Golden

section describes the median of this pool and is not a rule any of these composers appear to have

followed. A generator that pins every peak at 0.618 would be more uniform than the corpus it is

imitating, which is itself a defect. Use 0.5-0.75 as a soft target and treat anything outside

0.16-0.87 as worth a second look, not a rejection.

BATTLE 0.672 dynamic / 0.282 flow; TENSION 0.704 / 0.809; CREDITS_TRIUMPH 0.543 / 0.493;

EXPLORATION 0.462 / 0.763; MELANCHOLIC 0.459 / 0.640; TITLE_CHARACTER 0.460 / 0.346;

BOSS 0.291 / 0.805. No per-class claim is a test; five classes sit below the five-row floor.

8.2 Single climax against multi-peak, and the release

strongest shape for a 90-second cue. A multi-peak design is the shape a 3-5 minute cue needs,

because a single peak leaves too much runtime as approach and aftermath.

CONSENSUS, and it follows from Farbood's moving-window result: tension is read from trajectory, so

a peak with no descent registers as a plateau. After a peak the arrangement must lose lanes, lose

register width, or lose harmonic motion — and it must do so for long enough to be heard as a

descent, not as a gap.

either withhold the expected full statement or resolve deceptively, so the listener's spent

expectation is available to spend again.

trajectory, so peak count, peak prominence and post-peak descent depth are all derivable today and

none of them is currently extracted. Adding peak-count and post-peak-descent-dB to the card is

low-cost and directly serves this section.

9. The drop

mechanics: extensive uplifters and risers, the drum-roll effect, large frequency changes and

filter sweeps, the removal and reintroduction of bass and bass drum, and a contrasting breakdown —

all of which create tension and anticipation

(https://dj.dancecult.net/index.php/dancecult/article/view/451).

containing break routines and one without, with skin conductance, self-reported affect and

embodied experience measured. The emotional-intensity peak starts at or immediately after the drop

and continues into the core section, accompanied by greater motion and increased skin-conductance

response; dancers' overall activity level shows a sudden decrease then increase matching the

music's structure (Music Perception 36/4, 2019, 371-389;

https://online.ucpress.edu/mp/article/36/4/371/62966).

breakdown removes the low end and the pulse, so the listener's reference level falls; when the

same bass returns it is heard against that lowered reference, not against the track's own opening.

This is the textural form of the subito argument in section 7.4, and it is why a drop preceded by

a full-texture build lands harder than a drop preceded by a louder build.

is guaranteed across the gap. A drop lands when the grid survives the removal — the listener can

still count, so the return is arrival rather than restart. A drop merely stops when tempo,

harmony and grid all vanish together, because then there is nothing to be early or late against.

CONSENSUS, and it is the operational content of Medina-Gray's meter parameter: strong smoothness

is all pulse streams continuing across the seam

(https://www.mtosmt.org/issues/mto.19.25.3/mto.19.25.3.medina.gray.html).

melody. The rule that follows is that a drop is a structural event and gets structural preparation:

it needs a build it terminates and a re-entry it justifies. A drop with neither is the over-drop

defect measured in section 10.

and re_enters_at_s, and raises[] carries slope_db_per_s and lanes_added. A real drop is

identifiable today as a raise whose end is within a bar of a drop's start, with the drop's

bars_held_s at least a bar and a re-entry that restores at least the removed lanes. Nothing

computes that pairing. It is the highest-value derived feature in this document.

10. Expectation mechanics, and the 2.68-17.43 band

10.1 ITPRA

Pre-outcome: imagination (contemplating the future state) and tension (arousal and attention

raised immediately before the outcome, scaled by uncertainty, stakes and time-to-outcome).

Post-outcome: prediction (reward or punishment for the accuracy of the expectation itself),

reaction (fast, worst-case) and appraisal (slow, considered)

(Sweet Anticipation, MIT Press, 2006; https://mtosmt.org/issues/mto.09.15.3/mto.09.15.3.aversa.html).

independently of whether the outcome was good. This is why a listener enjoys a resolution they saw

coming, and why a composer must establish a pattern before violating it — the violation is only

worth anything if the prediction response was engaged.

response scales with estimated time to outcome, so holding the dominant longer literally raises

arousal before the resolution arrives.

10.2 Silence

they are acoustically identical apart from duration: silences following tonal closure are

identified faster and rated less tense than silences following unclosed music, and the preceding

material and the expectation of what follows both "seep into the gap" (Music Perception 24/5,

2007, 485-506; https://online.ucpress.edu/mp/article-abstract/24/5/485/95245).

context. A general pause after an unresolved dominant is tension; the same pause after a cadence

is punctuation. Our own card already declares this posture — the silence_budget field carries

the note that section 11.2 "treats rest as a FIRST-CLASS parameter, not an absence."

10.3 What our narrow band is really measuring

silence_frac * 100 + drops_per_min. It is a composite of how much of the track is below the

silence threshold and how often the texture is cut.

one of three carried hypotheses that died at N = 28 (PATTERN_FINDINGS_V2.md section 6). It

SUCCEEDS as a containment axis: the hit envelope 2.68-17.43 holds 26 of 28 held-out hits and is

one of the three axes in the joint box (section 8 of the same document). Corpus rows occupy a

narrow band of composed rest without sitting high or low on it.

28 corpus rows32 round-1 candidates
X_rest_grammar median8.4515.92
X_rest_grammar range2.68 - 17.436.14 - 42.55
drops per minute, median5.5614.08
drops per minute, range1.46 - 14.02
silence fraction, median2.12%2.75%
outside the 2.68-17.43 band0 of 2814 of 32, all on the high side

the worst candidate sits at 42.55 against a corpus ceiling of 17.43. Every one of the fourteen

rejections is an over-drop, not an under-drop.

an expectation, and the expectation is built by the texture staying up. Cut every eight bars and

the listener stops predicting continuation, so removal stops being an event and becomes the

texture's normal behaviour — at which point it costs a lane and buys nothing. Farbood's

moving-window model says the same thing formally: tension is read from trend salience, and a

signal that reverses every few seconds has no trend to be salient. Over-dropping does not merely

fail to build tension; it destroys the mechanism by which tension could be built.

(5.56 per minute) with a hard ceiling near fourteen per minute, and a silence budget near two

percent. That is a descriptive band from a containment axis, and it is necessary-not-sufficient:

it rejects, it never certifies.

11. Long-form architecture for a 3-5 minute cue that loops

11.1 The two requirements are in tension, and the resolution is architectural

listener cannot hear. A cue that peaks at 0.6 and releases to nothing by 1.0 has a large

discontinuity at the wrap; a cue that is flat enough to wrap invisibly has no arc.

return to the pallavi after every section, so the loop point is a return the listener has already

heard four times rather than a splice. The Western equivalent is a rondo. Either way the wrap

falls on material that has already been established as a place the music comes back to.

close the loop with a perfect authentic cadence (NOSTALGIA_RUBRIC.md Y1, from Gervais on

Zelda's Theme's avoidance of the tonic and of PACs). A loop that closes fully asks to stop; a loop

that closes on an open half cadence asks to continue.

(do the pulse streams continue), timbre (is the instrumentation shared), pitch (are the new pitch

classes inside the previous five seconds' macroharmony), volume (is loudness consistent across the

seam) and abruptness (is the entry a cut or a decay). She explicitly refuses to collapse them into

one number, because the correct weighting is context-dependent, and she notes that disjunction is

a legitimate design choice when the game wants the player to notice a state change

(https://www.mtosmt.org/issues/mto.19.25.3/mto.19.25.3.medina.gray.html).

cadence_and_ending.loop_seam.seam_ratio has median 24.17 (p10 5.34, p90 63.98), and EX_124 reads

seam_clean: false with baked_fade_out: true. These are album masters with fades, not game

loops. Any loop-seam standard must be derived from authored material or from extracted game loops,

never from this pool.

11.2 Repetition density has a measured, era-invariant home

repetition structure essentially consistent across platforms and generations, disproving his own

hypothesis that early hardware limits produced above-average repetition

(https://boblsturm.github.io/aimusic2020/papers/CSMC__MuMe_2020_paper_10.pdf;

https://uwe-repository.worktribe.com/output/6818560/). Our rubric's L4 band (0.55-0.70) is taken

from that survey's mean of 62.79% with SD 3.49.

novelty_mean is below it, both clearing the calibrated bar, and the six clearing axes hold the

same sign and magnitude in the 16-bit half and the modern half of the pool

(PATTERN_FINDINGS_V2.md sections 4 and 9). Structural self-return is era-invariant across the

two eras we hold.

11.3 Per-purpose-class tension arcs

Recommendations below are craft syntheses grounded in the parameter evidence of section 7, the

class-descriptive medians of section 8.1, and the region-page briefs. They are proposal-tier.

dynamic peak around 0.45-0.50, no single dominating climax, wrap on an open cadence over a

returning refrain. Novelty enters as instrument swaps rather than as dynamic events.

time_to_full_texture median is 3% of runtime), flow peak early at about 0.28 and the dynamic

peak late near 0.67, so the arrangement is wide throughout while loudness still has somewhere to

go. One drop, held one to two bars, in the final third.

earlier. Each phase gets its own build and its own release; the last phase withholds the full

harmonized statement of the boss motif until the final section.

budget at or slightly above the corpus median. Tension is carried by harmony and appoggiatura

rather than by density; Sloboda's tears-inducing structures are the target.

flow). Sustained unresolved harmony, one long delayed resolution, drops used sparingly and always

with a preceding build.

corpus median time_to_hook_s is 0.575 s), dynamic peak near 0.46, and the definitive full

statement of the theme reserved for the last third.

central peak, release, then a second smaller peak that quotes the arc's earlier themes. This is

the one class where a perfect authentic cadence is correct, because it does not loop.

12. The nostalgia rubric — what it gets right, and what it is missing

12.1 What it gets right

fact that our own composite is 0.744 with all axes and 0.497 with the length-independent ones —

a single number would hide exactly that.

NOT_COMPUTABLE features excluded from the denominator). Section 11 of this document reaches the

same conclusion about seam measurement from the corpus side.

two mechanisms with the most direct bearing on a cue surviving a hundred hearings, and it grounds

L2 in Kondo's own stated method.

signatures with 14 declared gaps, and CLEAR explicitly meaning "no seed signature matched" rather

than "clean".

after it.

12.2 What this research says it is missing

climax placement or arc shape. Every principle in sections 7 through 10 above is outside its

scope. This is the largest structural gap.

cannot see one, because it scores a symbolic head cell rather than an arrangement.

Margulis's result is that rest VALUE comes from context — the same silence after a cadence and

after an unresolved dominant are different objects. A density gate cannot distinguish them.

says nothing about their scheduling, and Josh's floor is explicitly about scheduling: no more than

three to five cycles without a new instrument, catch, melody or drop.

at 7-17 semitones from a symbolic head cell; the formula card's the_hook.range_semitones reads

the audio's melodic band and has a corpus median of 22. Neither is wrong; they measure different

objects, and nothing in the stack says so. FINDING: the two should be named distinctly

(head_cell_range against measured_melodic_band_range) before either is quoted in a verdict.

invented from memory is a false negative wearing a green tick. With the Nintendo discs arriving,

direct transcription of the named exemplars is now unblocked as a polish pass.

through 6.5 above describe five identity mechanisms — cycle, speech-phrase, modal ostinato,

entry-pattern, named-song catalogue — that the rubric would score as failures for not being tunes.

FINDING: the rubric needs a declared dialect parameter, or its verdicts will push every region

toward the same Western dialect the region pages explicitly refuse.

13. The measurement gaps, named

Principles that can be enforced as COMPOSITIONAL CONSTRAINTS in the authored path today:

axis 4, pre-render.

tappable rhythm) — axes 1 through 3, pre-render.

prompt, since the generator responds on only 4 of 8 probed dimensions and section_count is not

among them.

1.5-14) — enforceable as an arrangement rule in the authored path.

Principles measurable only after the fact, from the rendered audio:

Principles NOT measurable from the current card, with the measurement that would close each:

array gives three computable quantities — cloud diameter, cloud momentum and tensile strain (the

distance between local and global tonal context) — with a Python implementation in midi-miner

(http://dorienherremans.com/sites/default/files/paper_tenor_dh_preprint_small.pdf). Running it on

audio requires a chroma-to-pitch-set step we do not have; running it on the authored symbolic path

requires nothing new.

from the spectrum already extracted for timbre. This is the cheapest closure in this list.

(https://www.sciencedirect.com/science/article/abs/pii/S0165027024002929 for the current Python

implementation). Blocked on melodic transcription, which is also blocking cadence census, mode

identification and the hook-signature seed set — one blocker, four consumers.

and bar arithmetic assumes 4/4.

(max_lanes = 9 for hits and siblings alike across two re-runs). Source separation or a real

roster is the only route.

repertoire is invisible to the_hook.

THE SINGLE MOST IMPORTANT GAP: melodic transcription. Without a symbolic melody line, we cannot run

an expectancy model, cannot compute a cadence census, cannot identify a mode or a raga or a qenet,

cannot expand the anti-plagiarism seed set beyond its 13 signatures, and cannot connect the

rubric's authored-side features to the card's rendered-side features at all. Every other gap in this

list is a feature; this one is the bridge between our two measurement worlds.

14. Sources

https://www.tagg.org/xpdfs/burns87.pdf

https://mtosmt.org/issues/mto.14.20.1/mto.14.20.1.aziz.html ·

http://shanahdt.github.io/MUSI4331/lessons/phrases1.html

repetition. https://mtosmt.org/issues/mto.02.8.4/leydon_text.html

https://viva.pressbooks.pub/openmusictheory/chapter/metrical-dissonance/ ·

https://mtosmt.org/issues/mto.14.20.2/mto.14.20.2.biamonte.html

https://www.academia.edu/50292501/The_Melodic_Arch_in_Western_Folksongs

https://mtosmt.org/issues/mto.09.15.3/mto.09.15.3.aversa.html

Complexity (1992); Meyer 1973 on gap-fill. https://mutor-2.github.io/ScienceOfMusic/units/08/

Melodic Structure." Music Perception. https://www.researchgate.net/publication/224982434

Melodies?" https://www.researchgate.net/publication/271681259

Creativity, and the Arts 11/2 (2017), 122-135.

https://www.apa.org/pubs/journals/releases/aca-aca0000090.pdf

https://mtosmt.org/issues/mto.14.20.4/mto.14.20.4.albrecht.html

(2007), 485-506. https://online.ucpress.edu/mp/article-abstract/24/5/485/95245

learning and probabilistic prediction in music cognition," Annals NYAS (2018).

https://onlinelibrary.wiley.com/doi/full/10.1111/j.1756-8765.2012.01214.x ·

https://nyaspubs.onlinelibrary.wiley.com/doi/10.1111/nyas.13654 ·

https://www.sciencedirect.com/science/article/abs/pii/S0165027024002929

(2012), 387-428. http://mp.ucpress.edu/content/29/4/387

329-366. https://www.fredlerdahl.com/s/Modeling-Tonal-Tension.pdf

TENOR 2016. http://dorienherremans.com/sites/default/files/paper_tenor_dh_preprint_small.pdf

548-560. https://www.semanticscholar.org/paper/1d3ccc073b1b3f13f95e2392ff81a9eff0e7a4d2 ·

Vassilakis roughness: https://www.mat.ucsb.edu/Masters/Brian_Hansen_Masters.pdf

Music 19 (1991), 110-120. https://journals.sagepub.com/doi/10.1177/0305735691192002

https://dj.dancecult.net/index.php/dancecult/article/view/451

Perception 36/4 (2019), 371-389. https://online.ucpress.edu/mp/article/36/4/371/62966

https://www.mtosmt.org/issues/mto.19.25.3/mto.19.25.3.medina.gray.html

https://boblsturm.github.io/aimusic2020/papers/CSMC__MuMe_2020_paper_10.pdf

(2018), 291-304. https://journals.sagepub.com/doi/abs/10.1177/1029864917698010 ·

https://news.osu.edu/has-music-streaming-killed-the-instrumental-intro/

https://www.frontiersin.org/journals/psychology/articles/10.3389/fpsyg.2017.02044/full ·

https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4756174/

https://blogs.ams.org/jmm2021/2021/01/08/did-bartok-use-fibonacci-numbers-in-his-music/

https://en.wikipedia.org/wiki/Simha_Arom · https://www.melodigging.com/genre/mbenga-mbuti-music

https://www.frontiersin.org/journals/communication/articles/10.3389/fcomm.2021.652542/full ·

Villepastour, Ancient Text Messages of the Yorùbá Bàtá Drum:

https://ethnomusicologyreview.ucla.edu/journal/volume/15/piece/481

https://en.wikipedia.org/wiki/Kotekan ·

https://ecampusontario.pressbooks.pub/beyondtheclassroom/chapter/balinese-music-gamelan/

https://www.mtosmt.org/issues/mto.15.21.4/mto.15.21.4.schachter.html

https://www.ncbi.nlm.nih.gov/pmc/articles/PMC10118148/

https://www.auralarchipelago.com/auralarchipelago/talempongbotuang

_source/02_Tier_2_Region_Pages/*.md Section 10; build/audio/exemplars/PATTERN_FINDINGS_V2.md;

build/audio/exemplars/formula_cards/EX_124.json; harness/music_gen/derive_patterns.py L368;

harness/music_gen/nostalgia_score.py; docs/proposals/music/NOSTALGIA_RUBRIC.md;

harness/music_gen/exemplar_hook_signatures.json.

15. Reproduction of the MEASURED-HERE numbers

Every MEASURED-HERE figure in this document is read from the formula cards on disk through the same

fields derive_patterns.flatten uses. The 28 corpus rows are the hit list in

PATTERN_FINDINGS_V2.md section 2; the 32 candidates are

build/audio/generated/slice_v1/formula_cards/*.json. Fields used: dynamics.peak_position_normalised,

opening_and_flow_law.flow_curve.peak_position_normalised, opening_and_flow_law.time_to_hook_s,

opening_and_flow_law.time_to_full_texture_s, event_grammar.counts.drops,

event_grammar.silence_budget.fraction, duration_and_structure.total_length_s,

duration_and_structure.section_count, complexity_law.boundary_step_share,

cadence_and_ending.loop_seam.seam_ratio, motif_economy.self_similarity_lift,

motif_economy.interval_3gram_repetition, the_hook.range_semitones. No audio was decoded, no

exemplar bytes moved, and no calibrated multiplicity bar was computed for any of them — they are

descriptive reads, and the section 8.1 and section 10.3 tables are the two that matter.

Generated by harness/site/structure_site.py — the URL path is the repo path. review root