THE EARWORM AND LOOP SCIENCE — what makes a tune stick, measured
CANON SUBORDINATION — this document is PROPOSAL-TIER: it serves canon and never outranks it.
The canon: the CVD · the T1 foundation docs · the T0 registries (registries/) · the spine
(docs/spine/CH_*.md) · the region pages (_source/02_Tier_2_Region_Pages/). Authority order: docs/DOC_MAP.md § 0.
Canon served: NONE FOUND — this document cites no canon anchor anywhere in its body.
Do not build from it alone: open the spine node (docs/spine/CH_NN.md), the region page
(_source/02_Tier_2_Region_Pages/), and the T0 registry your work targets, and derive from those.
READ THAT CANON FIRST — open it and derive from it before you build anything from this document.
If this document disagrees with canon, CANON WINS and this document is the defect — fix the
document, never the canon. Nothing here is applied until it is ratified into canon.
The empirical layer under both schools: the involuntary-musical-imagery (INMI) literature, the Essen corpus contour base rates, the melodic-arch and post-skip-reversal universals, and the loop-fatigue findings. It is the source of the rules that are numbers rather than judgements, and of the CURRENT STATE measurements against the fourteen authored themes that were on the site when Josh ruled.
---
0. Provenance — where this text came from, and what was done to it
TIER: PROPOSAL. This is research, not canon. It binds nothing until ratified; the CVD and the
T1/T0 canon outrank every line of it.
SOURCE: the research return of workflow run wf_d1ba57a8-e47, agent a183e334faacf81d9, 2026-08-05. The return was a
structured object with three keys — rules, evidence, measurable_features — and all three are
reproduced below.
EXTRACTION: VERBATIM and mechanical. The strings were read out of the run journal by script and
written here unedited — not re-worded, not re-ordered, not summarised. Only the headings, the
bullet markers and this provenance section are new. That matters because NOSTALGIA_RUBRIC.md
cites these rules BY ID, and a citation that resolves to a paraphrase of its rule is the same
defect as one that resolves to nothing.
RULE IDS DEFINED HERE: R1..R14.
WHY IT EXISTS AS A FILE: the returns lived only in the run journal, which is agent scratch and
is not canon-readable. The rubric quoted them anyway and said the mechanism was "stated verbatim
from the school grammars" — which named documents that did not exist. That is finding F3 of the
cold verification of 2026-08-05, and this file is half of its fix (the other half is the rubric's own
provenance line).
WHAT IS NOT CLAIMED: that every source below was independently re-verified. The evidence list
carries the citations the research agent gave, with its own declared gaps and negative controls
left in. Verifier finding F4 stands open against this: several named scholarly sources are bare
author-year strings, and a reader cannot resolve them from this repo alone. Treat a citation here
as a pointer to check, never as a checked fact.
---
1. The rules
R1 TEMPO BAND — a flagship theme targets 110-150 BPM, hard floor 100 BPM. INMI tunes averaged 124.10 BPM (SD 28.73) against matched non-INMI controls at 115.79 (SD 25.39); tempo was one of only 3 features surviving from 83. Slower is legal ONLY for a row declared lament/atmosphere, never for the node's flagship. CURRENT STATE: 13/13 authored themes sit below the INMI mean and 11/13 below even the CONTROL mean (corpus mean 93.4 BPM) — the corpus is tempo-tuned to the class of tunes that do NOT stick.
R2 GLOBAL CONTOUR MUST BE THE ORDINARY ONE — reduce the primary phrase to Huron's 3-point sketch (first note, mean of interior notes, last note) and it must classify CONVEX/arch. Base rates over 35,793 Essen phrases: convex 28.59%, descending 27.14%, ascending 22.38%, concave 21.89%. Monotone-ascending and concave/zigzag are the two least common shapes and are BANNED for a flagship head. CURRENT STATE: only 3/13 head cells trace up-then-down; 3 are monotone ascending, 2 are zigzag, 1 is a repeated-note recitation.
R3 THE DEVIATION IS LOCAL, NEVER GLOBAL — the two contour features the INMI random forest selected point in OPPOSITE directions: common global shape AND uncommon average gradient between turning points. Measurable form: exactly ONE interval in the head phrase exceeds the phrase's median absolute interval by >= 4 semitones; every other interval is <= 4 semitones. Ordinary shape, one outlier gradient inside it. Two or more outliers fails; zero outliers fails.
R4 POST-SKIP REVERSAL IS MANDATORY ON THAT DEVIATION — any interval >= 6 semitones must be followed by a direction change (~72% of large leaps reverse across cultures; regression-to-the-mean/tessitura, not gap-fill). Never two same-direction intervals >= 5 semitones in succession. CURRENT STATE VIOLATIONS: FAMILIAR_BOND head [7, 9] (perfect fifth then major sixth, same direction), FAM_THE_SUBSTRATE [7, 5, 2] (three ascents), JOURNEY_WORLD [4, 3, 9] (terminal +9 with no reversal in the recorded head). The 'unusual_leap' deviant class is being applied without the reversal that makes a leap singable.
R5 THE HEAD IS A PHRASE, NOT A CELL — minimum 7 pitches spanning >= 2 bars for the identity statement. A 3-5 note cell has no global contour, so R2 and R3 are literally uncomputable on it: you cannot be 'the common shape' with three notes. CURRENT STATE: 13/13 authored head cells are 3-5 notes (2-4 intervals) — the corpus is measuring first-order interval CONTENT (triad_outlining, quartal_outlining, open_fifth_outlining) where the literature measures second-order corpus-relative SHAPE. The rig cannot currently see the single strongest melodic discriminator in the INMI evidence.
R6 THE HOOK LEADS — time from cue start to first complete head statement <= 5 seconds; intro before the head <= 4 bars; the head states >= 3 times per loop body. Sections of the SAME song differ significantly in how fast listeners recognise them, and the hook is by definition the fastest-recognised fragment; a loop restarts, so anything behind the first statement is heard fewer times per playthrough than the first statement is. A theme that arrives at 0:40 is a theme the player hears once per dwell.
R7 LOOP BUDGET — theme loop body 60-120 s (industry convention is 1-2 minutes), and the head must recur at gaps <= 30 s. Repetition becomes actively repelling at 3-5 loop repetitions. CURRENT STATE: 13/13 authored cues run 19.2-40.0 s (mean 27.3 s), so the documented annoyance threshold arrives after roughly 1.5-3 minutes of dwell — shorter than a single traversal segment.
R8 EXPOSURE CONVERTS A TUNE INTO NOSTALGIA, COMPOSITION ONLY QUALIFIES IT — in the INMI hurdle model, melodic features predicted WHETHER a tune is named at all, while popularity and recency predicted HOW MANY TIMES; familiarity drives the limbic and reward response; autobiographical salience is the strongest predictor of music-evoked nostalgia. Measurable: every flagship theme must accumulate >= 20 minutes of expected player exposure inside the slice, computed from cue-table dwell, and must recur in >= 3 separate chapters. A theme placed only in cutscenes cannot meet the 30-year bar by construction, however well composed.
R9 PLACE THEMES CARRY THE NOSTALGIA — the exemplar class Josh named (OoT, SM64, the SNES lineage) is ZONE-LOOP-anchored, not leitmotif-anchored: the tunes people hum for 30 years are the places they spent 40 minutes in, not the character motifs they heard for 20 seconds. Measurable: >= 1 full-phrase, hummable PLACE theme per playable region of the slice, held to R1-R7 identically. CURRENT STATE: 0 of 15 authored picks is a place theme (all are protagonist/companion/familiar/thread motifs), and region material is specified as 'one-to-two bars, none of them required to be hummable' — the architecture explicitly EXEMPTS from the hummability bar exactly the class that carries the exemplars' nostalgia load. This inversion is the single largest structural cause of the complaint.
R10 TRAVERSAL TEMPO ENTRAINS TO PLAYER CADENCE — INMI tempo is recalled near-veridically from long-term memory and correlates with arousal, and INMI episodes cluster in repetitive-movement and low-attention states. Measurable: traversal-cue BPM within +/-10% of the character's step cadence in steps/min, or an exact integer ratio (1:1, 1:2) of it. This is the empirical form of Kondo's 'recognise the intrinsic rhythms of the game' principle.
R11 VARIATION INSIDE THE LOOP, NEVER A SHORTER LOOP — the OoT field theme resequences twelve interchangeable phrases so travel does not tire. Measurable: >= 4 interchangeable phrase variants per traversal theme, every variant quoting the same head; melody-present and melody-absent renders alternate so head-statements-per-minute stays in 1-3. Shortening the cue to hide fatigue (the current 19-40 s shape) accelerates it instead.
R12 THE FLAGSHIP LEADS ON EVERY PLAYER-FACING SURFACE; BEDS STAY BEDS — LIVE DEFECT, and the site is the only surface anyone has actually heard. harness/site/generated_music.py:62 orders tiles by PRODUCER FILE (PICKS -> AUTHORED_PICKS -> BED_PICKS) and harness/site/build_progress_site.py:1072 renders generated -> own -> shared, so on ALL 14 slice nodes the first five audio tiles are five renders of the same JOURNEY_WORLD material — including one row whose own role is superseded_statement, one whose role is sibling_render, and one whose role is literally zone_bed. Every authored orchestral theme sits 6th or later, inside a stack of 18-24 tiles. REQUIRED: sort key = role rank (flagship theme > theme variant > bed), then node ownership, then tier; exactly one flagship per node in first position; rows whose role is superseded_statement or sibling_render never reach a player-facing page at all.
R13 ONE FLAGSHIP PER NODE, NOT THIRTEEN — a large catalogue is fine, a large RECOGNITION load is not (the Star Wars corpus documents dozens of motifs while audience recognition is carried by about five). Measurable: <= 3 distinct themes served per chapter page, exactly 1 designated flagship. CURRENT STATE: 11 of 15 authored picks serve >= 9 of the 14 slice nodes each (4 serve all 14), so nearly every chapter serves nearly the same undifferentiated pile and no node has an identity tune.
R14 THE GATE IS DISTRIBUTION OVERLAP WITH THE EXEMPLAR CLASS, NOT A PROXY PERCENTAGE — every feature in R1-R11 is computable on the exemplar set as well as on ours, and the pass condition is that our per-feature distribution overlaps the exemplar distribution (median inside the exemplar interquartile range, on tempo, contour class share, deviation index, post-skip reversal rate, head length, time-to-hook, loop length, head-recurrence gap). An absolute self-referential score with no comparator anchor is exactly the defect Josh named: the 14 themes passed proxy metrics that were never once pointed at the class they were supposed to join.
---
2. The evidence
Jakubowski, Finkel, Stewart & Mullensiefen (2017), 'Dissecting an Earworm: Melodic Features and Song Popularity Predict Involuntary Musical Imagery', Psychology of Aesthetics, Creativity, and the Arts 11(2), 122-135. 3,000 survey participants named INMI tunes; 100 frequently-named INMI tunes matched to 100 never-named controls on popularity and style; 83 statistical-summary and corpus-based melodic features compared via random forest. https://www.apa.org/pubs/journals/releases/aca-aca0000090.pdf
SAME PAPER, exact numbers (verified from the PDF text, not a summary): final model retained only THREE features — tempo, dens.int.cont.grad.mean, dens.step.cont.glob.dir (12 selected in the initial forest, from 83). Binary logistic regression: INMI tunes M = 124.10 bpm, SD = 28.73; non-INMI M = 115.79 bpm, SD = 25.39. Only dens.step.cont.glob.dir reached p < .05 (estimate 7.08, SE 3.31, z = 2.14, p = .03).
SAME PAPER, the directional rule in the authors' own words: 'If the melodic contour shape of a melody is highly congruent with established norms, then it is more likely for the tune to become INMI. If the melodic contour does not conform with norms, then it should have a highly unusual pattern of contour rises and falls to become an INMI tune.' Common-contour exemplars were arch-shaped phrases; uncommon-contour exemplars were ascending lines that never descend. Uncommon average gradient between turning points = many leaps or unusually large leaps.
SAME PAPER, the exposure mechanism: a hurdle model showed melodic-feature predictions significant in the HURDLE component (zero vs nonzero INMI counts) while chart popularity and recency were significant in the COUNT component (how many times named). Adding the melodic predictor to a popularity+recency-only model improved fit significantly, chi-sq(2) = 55.52, p < .001. Melody decides IF a tune can stick; exposure decides HOW MUCH.
SAME PAPER, methodology detail that fixes what 'the theme' means: where no section was specified, the CHORUS was extracted for analysis, 'as this is the section of a song that is most commonly reported to be experienced' as INMI. The hook section, not the whole work, is the unit the findings describe.
Huron (1996), 'The Melodic Arch in Western Folksongs', Computing in Musicology — Essen folksong database, 35,793 phrases reduced to a 3-point sketch (first note, mean of interior, last note): convex/arch 10,233 (28.59%), descending 9,714 (27.14%), ascending 8,012 (22.38%), concave 7,834 (21.89%). Arch is the single most common phrase archetype and short complete melodies tend to arch overall regardless of internal phrase shapes.
Von Hippel & Huron (2000), 'Why Do Skips Precede Reversals? The Effect of Tessitura on Melodic Structure', Music Perception — approximately 72% of large leaps are followed by a direction reversal; the effect holds across a wide variety of cultures and is explained by regression to the mean / tessitura constraints rather than by Meyer's gap-fill. Gives the post-skip-reversal test a cross-cultural warrant, which matters under the per-culture scoring law.
Burgoyne, Bountouridis, Van Balen & Wiering (2013), 'Hooked: A Game for Discovering What Makes Music Catchy', ISMIR 2013 — defines catchiness as long-term musical salience and a hook as 'the most salient, easiest-to-recall fragment of a piece of music'; found that different sections WITHIN the same song differ significantly in recognition time, i.e. hook placement is measurable, not a matter of taste. Follow-on analysis (Van Balen et al., 2015) found melodic repetitiveness, corpus-relative melodic conventionality, and vocal-line prominence predicted catchiness. https://archives.ismir.net/ismir2013/paper/000101.pdf
Williamson, Jilka, Fry, Finkel, Mullensiefen & Stewart (2012), 'How do earworms start? Classifying the everyday circumstances of Involuntary Musical Imagery', Psychology of Music — four onset categories from grounded-theory analysis: music exposure, memory triggers, affective states, and LOW ATTENTION STATES. The last two are the normal condition of a player traversing a zone, which is why the traversal loop is the highest-yield earworm slot in a game.
Jakubowski, Farrugia, Halpern, Sankarpandi & Stewart (2015), 'The speed of our mental soundtracks: Tracking the tempo of involuntary musical imagery in everyday life', Memory & Cognition 43(8), 1229-1242 — 4-day accelerometer tapping study: INMI tempo is recalled from long-term memory in highly veridical form (with slight regression to the mean) and correlates positively with subjective arousal; INMI episodes cluster during repetitive movement such as walking and running. Tempo is encoded with the tune, so the tempo choice is load-bearing, not a mix decision.
Pereira, Teixeira, Figueiredo, Xavier, Castro & Brattico (2011), 'Music and Emotions in the Brain: Familiarity Matters', PLoS ONE 6(11): e27241 — fMRI with familiarity and preference controlled: limbic/paralimbic emotion regions AND the reward circuitry were significantly more active for FAMILIAR than unfamiliar music. The affective payload the 30-year bar is chasing is unlocked by repeat exposure, not by the first hearing.
Barrett, Grimm, Robins, Wildschut, Sedikides & Janata (2010), 'Music-evoked nostalgia: affect, memory, and personality', Emotion — autobiographical salience was the strongest predictor of music-evoked nostalgia, with song familiarity also predicting nostalgia intensity; context-level constructs (salience, arousal, familiarity) outweighed person-level traits. Nostalgia is manufactured by hours-in-the-place, which is a level-design and cue-placement decision before it is a composition decision.
Koji Kondo, GDC 2007 keynote 'Painting an Interactive Musical Landscape' — three named essentials: rhythm, balance, interactivity; the composer must play the game to find its intrinsic rhythms (character movement, button presses). Concrete exemplar mechanism: the Ocarina of Time field theme was built from TWELVE phrases resequenced randomly so that travelling would not feel tiresome — variation inside the loop rather than a shorter loop. https://www.gamedeveloper.com/game-platforms/gdc-koji-kondo-s-interactive-musical-landscapes and https://archive.org/details/GDC2007Kondo
Game-industry loop practice, 'Rethinking the audio loop in games' (gamedeveloper.com): loops are typically one to two minutes; two repetitions inside a play session is tolerable, 3-5 repetitions 'can become too repetitive, repelling the player', more than five is the worst case; experienced players quickly identify the loop point and do not accept the illusion; recommended mitigation is alternating melody-present and melody-absent versions rather than shortening the cue. https://www.gamedeveloper.com/audio/rethinking-the-audio-loop-in-games
Phillips, 'A Composer's Guide to Game Music' (MIT Press) and Collins, 'Game Sound' (MIT Press) — repeating music must be constructed to avoid landmarks that alert listeners to the looping nature of the work; the historical response to loop fatigue was interactive/adaptive scoring, and unmitigated repetition drove players to mute game audio outright. Establishes the theme-vs-bed division of labour as a fatigue-management architecture, not a taste preference.
MEASURED IN-REPO, tempo: the 13 'orchestral rung one' authored themes in build/audio/generated/AUTHORED_PICKS.json run 68/76/80/84/84/88/90/96/96/100/112/120/120 BPM — mean 93.4, median 90. 13 of 13 fall below the INMI mean of 124.10; 11 of 13 fall below the non-INMI CONTROL mean of 115.79. The corpus is not merely off-target, it is tuned into the control distribution.
MEASURED IN-REPO, head length and contour: head_intervals are 2-4 intervals (3-5 pitches) on 13/13 rows. Sign patterns: monotone ascending on FAM_THE_SUBSTRATE [7,5,2], JOURNEY_WORLD [4,3,9], PROTAGONIST_THEME [2,1,2], FAMILIAR_BOND [7,9], THE_RECURRENCE [5,4]; zigzag on FAM_THE_LIVING_VOICE [4,-2,3] and FAM_THE_OTHERWORLD [7,-3,2]; only FAM_THE_MADE_PLACE [5,5,-3], FAM_THE_OTHER_SELF [3,4,-4] and HOME_AND_LOSS [2,2,-2,-2] trace up-then-down. Arch conformance is 3/13 (23%), BELOW the 28.59% base rate — the selection process was not pulling toward the common contour at all.
MEASURED IN-REPO, the taxonomy blind spot: contour_class values across the 13 rows are triad_outlining (6), purely_stepwise (3), interlocking, quartal_outlining, open_fifth_outlining, recitational. Not one class names a GLOBAL rise-fall shape. The rig classifies first-order interval content while the INMI evidence selected second-order corpus-relative contour commonness — the strongest single melodic discriminator in the literature is currently unmeasured by the pipeline.
MEASURED IN-REPO, duration: authored cue lengths are 19.2 to 40.0 seconds, mean 27.3; 13 of 13 are under 60 s, i.e. every one is under half the low end of the 1-2 minute loop convention, so the documented 3-5-repetition annoyance threshold is reached inside about 1.5-3 minutes of dwell.
MEASURED IN-REPO, site serving order (the surface Josh actually heard): harness/site/generated_music.py:62 iterates producers in the order (PICKS, AUTHORED_PICKS, BED_PICKS) and harness/site/build_progress_site.py:1072 renders generated -> own -> shared. Consequence, computed over all served nodes: on 14 of 14 slice nodes (CH_PROLOGUE, CH_01-CH_13) the FIRST tile is from PICKS.json. Those five PICKS rows are JOURNEY_WORLD_AUTHORED_HEAD_CELL, JOURNEY_WORLD_A1_R1_HOLD24, JOURNEY_WORLD_C04 (role: superseded_statement), JOURNEY_WORLD_VARIATION_REPAINT (role: sibling_render), JOURNEY_WORLD_ZONE_BED_FLOWERING_PLAIN (role: zone_bed). The 13 authored orchestral themes begin at tile 6, inside stacks of 18-24 tiles per page.
MEASURED IN-REPO, node differentiation: of the 15 authored picks, 4 serve all 14 slice nodes (FAM_THE_SUBSTRATE, FAM_THE_WRITTEN, JOURNEY_WORLD, PROTAGONIST_THEME) and 11 serve 9 or more; only COMPANION_HEALERS_CHILD (CH_13) and FAMILIAR_BOND (3 nodes) are node-specific, and 2 rows serve zero nodes. No chapter in the Josh-gate slice has a flagship of its own.
MEASURED IN-REPO, the place-theme gap: 0 of 15 authored picks is a place/zone theme — the roster is protagonist, companion, familiar-bond, home-and-loss, the-recurrence and eight FAM_* thread heads. docs/proposals/music/LEITMOTIF_ARCHITECTURE.md section 3.6 specifies the twelve region cells as 'one-to-two bars, none of them required to be hummable — only coherent with the parent', and harness/music_gen/theme_specs.json contains exactly one theme spec (JOURNEY_WORLD). The class that carries the named exemplars' nostalgia is both unbuilt and formally exempted from the bar.
---
3. The measurable features
tempo_bpm — cue tempo; gate band 110-150, hard floor 100 for a flagship; compare against INMI mean 124.10 / control mean 115.79
global_contour_class_huron3 — primary phrase reduced to (first pitch, mean of interior pitches, last pitch) and classified convex | descending | ascending | concave; flagship must be convex
contour_commonness_percentile — second-order corpus-relative commonness of the global rise-fall pattern (the dens.step.cont.glob.dir analogue), scored against BOTH a Western reference corpus and the register card's own declared corpus; flagship target >= 60th percentile
turning_point_gradient_percentile — corpus-relative mean absolute gradient of interpolation lines between contour turning points (the dens.int.cont.grad.mean analogue); flagship target <= 25th percentile (uncommon), i.e. deliberately opposed in direction to contour_commonness_percentile
head_phrase_pitch_count — pitches in the identity statement; minimum 7
head_phrase_bars — bars spanned by the identity statement; minimum 2
median_abs_interval_semitones — median |interval| across the head phrase; target <= 4
max_abs_interval_semitones — largest |interval| in the head phrase
deviation_index — max_abs_interval minus median_abs_interval; exactly one interval must clear +4, all others must not
deviation_count — number of intervals exceeding median+4 semitones; must equal 1
post_skip_reversal_rate — fraction of intervals >= 6 semitones followed by a direction change; must equal 1.0 (benchmark: ~0.72 observed in natural corpora)
consecutive_same_direction_large_leaps — count of adjacent interval pairs both >= 5 semitones in the same direction; must equal 0
repeated_note_fraction — proportion of zero-semitone intervals in the head; flags recitational heads that have no contour to be common
pitch_range_semitones — head-phrase ambitus; sanity band against the register card's declared tessitura
note_density_notes_per_second — melodic event rate of the head statement
time_to_first_head_statement_s — seconds from cue start to the first complete head statement; <= 5
intro_bars_before_head — bars of material before the head arrives; <= 4
head_statements_per_loop — complete head statements inside the loop body; >= 3
head_recurrence_gap_max_s — largest gap between consecutive head statements; <= 30
head_statements_per_minute — head density; band 1-3 across the melody-present / melody-absent alternation
loop_body_seconds — length of the looping body; band 60-120 for a theme, unbounded for a bed
loop_seam_discontinuity — spectral/level delta across the loop point, so an audible seam is a measured defect rather than an ear call
phrase_variant_count — interchangeable phrase variants quoting the same head, per traversal theme; >= 4
traversal_bpm_to_step_cadence_ratio — cue BPM divided by character step cadence in steps/min; must be within 10% of 1.0, 0.5 or 2.0
expected_player_exposure_minutes — dwell-weighted playtime a theme is expected to sound for across the slice, derived from build/audio/music_cue_table.csv; flagship >= 20
theme_chapter_recurrence_count — distinct chapters in which a theme sounds; flagship >= 3
place_theme_coverage — playable slice regions carrying a full-phrase hummable place theme, divided by playable slice regions; target 1.0, currently 0.0
role_rank — ordinal serving rank derived from role (flagship_theme > theme_variant > bed), the sort key that must replace producer-file order
flagship_rank_on_page — ordinal position of the node's flagship tile in the served list; must equal 1 on every node
tiles_per_node — audio tiles served on a chapter page; currently 18-24, target <= 5
distinct_themes_per_node — distinct themes served per page; <= 3
withdrawn_role_served_count — tiles whose role is superseded_statement or sibling_render reaching a player-facing page; must equal 0
exemplar_overlap_score — per feature, whether our corpus median falls inside the exemplar set's interquartile range for the same feature; the comparator anchor that was missing, and the actual pass condition
Generated by harness/site/structure_site.py — the URL path is the repo path. review root