music/HOOK_AND_TENSION_ARCHITECTURE.md
CANON SUBORDINATION — this document is PROPOSAL-TIER: it serves canon and never outranks it.
Canon served: the composition floor atdocs/spine/DECISIONS_PENDING_JOSH.md(commitda627060), and the
per-chaptermusic_moodanchors ofdocs/spine/CH_03.md,docs/spine/CH_04.mdand
docs/spine/CH_10.md.
If this document disagrees with canon, CANON WINS and this document is the defect.
Tier: RESEARCH SYNTHESIS, proposal-tier. This document is subordinate to canon: the CVD, the T1
foundation docs, the spine (docs/spine/CH_*.md), the Tier-2 region pages
(_source/02_Tier_2_Region_Pages/) and the T0 registries. Where it disagrees with canon, canon
wins and this document is the defect. Nothing here is applied until it is ratified.
Every substantive claim carries a source. Claims that are craft consensus rather than evidence are
labelled CONSENSUS. Claims measured from our own formula cards for this document are labelled
MEASURED-HERE and carry no calibrated multiplicity bar — they are descriptive reads at N = 28,
in the same honest register as PATTERN_FINDINGS_V2.md section 10.
Commissioned by: the composition floor at docs/spine/DECISIONS_PENDING_JOSH.md (MUSIC ROUND 1
GRADED + THE COMPOSITION FLOOR, commit da627060) — "its got to have the hooks and complexities.
There cant just be stacked noise", "no track will be boring or filler", and the requirement to set
"the next bar for video game sound tracks".
python harness/route.py music. It names, in order: docs/spine/CH_NN.md (Asset anchors, the music_mood bullet), the region page Section 10
(the Music Sound and Composition Brief), registries/T0_Theme_Registry,
docs/translation/T99_Translation_Audio.md Music Steps 1-7, and
docs/proposals/MUSIC_COMPOSITION_DOCTRINE.md.
docs/spine/CH_03.md L245, docs/spine/CH_04.md L242 and docs/spine/CH_10.md L224 (the music_mood asset anchors); and Section 10 of
_source/02_Tier_2_Region_Pages/ for flores_island (L844), bali (L1894),
central_africa_congo (L1017), ethiopia (L713), south_africa (L599), south_india (L787),
sri_lanka (L908), sumatra_java (L2229), swahili_coast (L838), rift_valley (L624) and
west_africa (L2398).
build/audio/exemplars/formula_cards/EX_124.json (schema formula-card/v1), build/audio/exemplars/PATTERN_FINDINGS_V2.md,
harness/music_gen/exemplar_hook_signatures.json (13 signatures, 14 declared gaps),
harness/music_gen/nostalgia_score.py and docs/proposals/music/NOSTALGIA_RUBRIC.md.
music_mood anchor's demand that a chapter's music be anidentity a player carries, and the region pages' Section 10 briefs, which already specify the
cultural substrate. This document supplies the craft layer between those two and never re-decides
either. It changes no canon; every canon-facing consequence below is stated as a finding.
6/1, 1987, 1-20, https://www.tagg.org/xpdfs/burns87.pdf) a hook is the part of a record that
stands out and recurs, and Burns's contribution is that the hook need not be melodic at all. He
sorts hooks into musical (rhythm, melody, harmony, instrumentation), textual (lyric), performance
(tempo, dynamics, improvisation) and production/technological (editing, mix, channel balance,
signal distortion, sound effects) families. The practical consequence for us is direct: a
gong-waning cycle, a whip-crack placed on a beat, a single detuned bronze pair, or a sudden
register drop are each admissible hooks by Burns's taxonomy, and a cue with no singable tune is
not automatically hookless.
rest of the movement is heard against. Caplin's account (Classical Form, Oxford, 1998) makes this
operational rather than impressionistic — see section 2.
memory claim, not a structural one, and no purely structural measurement can certify it. That gap
is the honest centre of this document.
summarized at https://mtosmt.org/issues/mto.02.8.4/leydon_text.html). Musematic repetition is
near-unvaried repetition of a museme — the smallest meaningful unit — and its paradigm case is the
riff or ostinato; it is circular, synchronic and open. Discursive repetition is repetition of
longer, syntactically complex units — phrases, strophes, sections — and it is linear and
self-sufficient.
saturation and groove, a hook by shape and return. A cue that has only musematic repetition is the
thing Josh named as stacked noise: material that recurs without ever being answered.
begins, goes somewhere and comes back — even when its surface is an ostinato culture.
motif_economy.interval_3gram_repetition and interval_4gram_repetition measure musematic density. motif_economy.self_similarity_lift (peak minus baseline) is the closest thing
the card has to a discursive-return measure. Corpus medians, MEASURED-HERE at N = 28: 3-gram
repetition 0.432 (p10 0.157, p90 0.791), self-similarity lift 0.181 (p10 0.058, p90 0.328). A
candidate with high 3-gram repetition and near-zero lift is a riff with no theme.
(being in the middle) and cadential (ending). The presentation is a two-bar basic idea plus its
near-literal repetition, usually over a tonic prolongation. The continuation fragments the basic
idea, accelerates harmonic rhythm and increases surface activity. The cadential function
liquidates the idea's characteristic features and closes
(https://mtosmt.org/issues/mto.14.20.1/mto.14.20.1.aziz.html;
http://shanahdt.github.io/MUSI4331/lessons/phrases1.html).
unit and immediately proves it is a unit by repeating it; from bar 3 the listener can predict bar
4. The continuation then rewards that acquired competence by breaking the unit into pieces the
listener already owns. Fragmentation is only intelligible because the whole was stated twice
first.
nostalgia_score.py axis 4 already encodes this (P1 basic-idea repeat similarity, hard-gated at 0.80; P2 continuation fragmentation; P3/P4 the 8-bar grid). The
census reports P1 at 3 of 14 — our landed cohort answers its head with a mirror instead of
repeating it (docs/proposals/music/NOSTALGIA_RUBRIC.md section 6). That is the single most
diagnosable phrase-level defect we have measured.
idea and closes strongly. Caplin treats antecedent/consequent as functionally comparable to
presentation/continuation, which is why a tune can feel classically shaped without being a
textbook example of either.
removing the seam between phrases without shortening either. CONSENSUS.
(Rothstein, Phrase Rhythm in Tonal Music, 1989, cited in
https://mtosmt.org/issues/mto.22.28.4/mto.22.28.4.temperley.pdf). Hypermetric regularity is what
makes a listener feel where a section will end before it ends, and hypermetric irregularity
(a five-bar or six-bar group) is a tension device precisely because that expectation exists.
complexity_law.bar_length_s and bars give bar arithmetic at an assumed 4/4 (tempo_and_time_grammar.metre.measured is null on EX_124), so
hypermeter has no representation. Closing measurement named in section 12.
contours over ascending and descending shapes ("The Melodic Arch in Western Folksongs", Computing
in Musicology 10, 1996; https://www.academia.edu/50292501/The_Melodic_Arch_in_Western_Folksongs).
The arch is the default phrase shape of a large, mostly European folk corpus — which is a fact
about that corpus, and the Essen database's over-use as a proxy for music in general is a known
methodological complaint (https://dl.acm.org/doi/10.1145/3469013.3469016).
filling the gap) archetypes come from Meyer and Narmour. Gap-fill originates with Meyer (1973):
large intervals imply smaller intervals in the opposite direction
(https://mutor-2.github.io/ScienceOfMusic/units/08/).
Reanalyses find that gap-fill does poorly as a way listeners classify melodies
(https://www.researchgate.net/publication/271681259_Questioning_a_Melodic_Archetype_Do_Listeners_Use_Gap-Fill_to_Classify_Melodies).
the_hook.pitch_contour is a U/D/- string; EX_124 reads ----UDUD-UDUD--, whichis axial with repeated pitches, not an arch. Our own card carries a gap-fill-after-leap axis whose
hit median and sibling median are both 0.000 with AUC 0.533 (PATTERN_FINDINGS_V2.md section 10)
— gap-fill does not separate our corpus rows from their album siblings, and should be used as a
craft option, never as a gate.
two notes: a small interval implies continuation in the same direction, a large interval implies
reversal. Huron reduced it to five governing principles — registral direction, intervallic
difference, registral return, proximity, closure (Huron 1997; summarized at
https://mutor-2.github.io/ScienceOfMusic/units/08/).
to the mean: a leap tends to land far from the tessitura's centre, and the next note statistically
returns toward it, which looks like a reversal principle without being one ("Why Do Skips Precede
Reversals? The Effect of Tessitura on Melodic Structure", Music Perception;
https://www.researchgate.net/publication/224982434).
then a step back has done the ordinary thing, and the ordinary thing is not a fingerprint. If a
leap is to be an identity, it has to be the same leap every time, in the same metric position.
the_hook.distinctive_deviation (largest interval, note index, is_licensed_leap) and interval_vocabulary_semitones. EX_124 carries a single -14 semitone leap at note index 6 out
of an interval vocabulary of exactly four values {0, 10, 12, 14}. That is the shape to aim for:
a tiny vocabulary with one deliberate outlier at a fixed index.
can reproduce without notation must fit an untrained voice. Our rubric encodes it as C5
(range <= 17 semitones, alphabet <= 8 pitches, tessitura MIDI 55-84), hard-gated
(NOSTALGIA_RUBRIC.md axis 2).
pitches — Burns's rhythm-hook category exists for exactly this reason, and Kondo's own
three-note head cells are described by their half-quarter-half template as much as by their
intervals (NOSTALGIA_RUBRIC.md H7, citing ECON-07 rhythmic economy). An upbeat entry changes
which note lands on the downbeat and therefore which note is heard as the identity.
carries two unisons in six intervals — a fact our own rubric had to write into feature I0 as a
declared derived floor after a texture control exposed the naive reading
(NOSTALGIA_RUBRIC.md axis 3).
register.melody_span_semitones and melody_tessitura_midi; the_hook.range_semitones. Corpus hook range median 22.0 semitones, MEASURED-HERE, range
0.67-52.0 — far wider than the rubric's 7-17 singability band, because the card's hook extractor
reads the audio's melodic band and the rubric reads an authored head cell. These two numbers are
not comparable and must never be quoted against each other.
imagery tunes against 100 never-named tunes controlled for popularity and style, and compared them
on 83 melodic features (Psychology of Aesthetics, Creativity, and the Arts 11/2, 2017, 122-135;
https://www.apa.org/pubs/journals/releases/aca-aca0000090.pdf). The reported discriminating
features are faster tempo, more common global melodic contour patterns (the plain rise-then-fall
arch being a positive predictor), and unusual interval structure — larger or less expected leaps
than the contour alone would imply.
recency were themselves strong predictors; the musical effects were modest; the design is
correlational and the sample is UK self-report. The authors' own framing is that these properties
interact rather than that any single feature determines stickiness.
detail. Put the novelty in the interval and the harmony; keep the contour common. Our own rubric
already says this at axis 2 and cites the same paper.
participant — repeated material is heard as something you are inside rather than something you are
watching (https://mtosmt.org/issues/mto.14.20.4/mto.14.20.4.albrecht.html).
model adds a tedium factor, producing the inverted-U: liking rises with exposure to a peak and
then falls. The peak's location is stimulus-dependent and is not a constant anyone has pinned
down (https://arxiv.org/pdf/2210.16226).
cue heard in a loop, the tedium factor overtakes the habituation factor at roughly the fourth
cycle unless something new enters. That number is a ruling, not a measurement, and it is
consistent with the shape of the literature without being derivable from it.
duration_and_structure.repeat_map plus section_count. Corpus section-count median is 17 against a sibling median of 9 (PATTERN_FINDINGS_V2.md section 10), the strongest
measured axis in our whole bank. A 17-section track that runs 196 s changes something roughly
every 11 s. That is the measured form of the 3-to-5-cycle law.
per event, a probability distribution over the next event, and from it an information content
(surprise) and entropy (uncertainty). It outperforms static rule-based models such as I-R for
pitch expectation in many contexts (https://onlinelibrary.wiley.com/doi/full/10.1111/j.1756-8765.2012.01214.x;
https://nyaspubs.onlinelibrary.wiley.com/doi/10.1111/nyas.13654).
(Music Perception 37/2, 165; https://online.ucpress.edu/mp/article-abstract/37/2/165/109530), and
the predictability-liking relation recovered from it is the same inverted-U shape as section 4.2.
candidates, so no expectancy model can be run in our pipeline today. Named as a gap in section 12.
to 2015 and found time-before-voice-entry and time-before-first-title-mention both fell
significantly across the period (Musicae Scientiae 22/3, 2018, 291-304;
https://journals.sagepub.com/doi/abs/10.1177/1029864917698010), with the popular summary of the
same work putting average intros at over 20 s in the mid-1980s against about 5 s today
(https://news.osu.edu/has-music-streaming-killed-the-instrumental-intro/). This is a fact about a
distribution channel, not about memorability, and it should be read as one.
opening_and_flow_law.time_to_hook_sacross the 28 hits has median 0.575 s (p10 0.104, p90 5.12, max 61.6). Half of the tracks Josh
named put identifiable melodic material inside the first six-tenths of a second. But
time_to_hook_frac sits at AUC 0.344 against siblings (PATTERN_FINDINGS_V2.md section 4) —
the siblings do it too, so early arrival is a floor and not a differentiator.
MEASURED-HERE, time_to_full_texture_s / total_length_s has median 0.032 across the 28: the
arrangement is essentially open within the first three percent of the runtime. The corpus does not
do long ramps into the material. It states the thing, then varies it for three minutes.
should appear once in its plainest form before it is transformed, because every later variation is
parsed against that statement. Caplin's presentation function is the classical form of this rule;
the film-scoring version is the restraint argument — the theme lands at the climax because the
film withheld it (https://www.cinemagicscoring.com/post/behind-the-music-how-film-scores-work).
bullet: state a fragment plainly early, withhold the complete harmonized statement for a later
structural payoff. This is the design the leitmotif architecture already assumes.
opening_and_flow_law (onset_type, first_sound_at_s, time_to_hook_s, time_to_hook_bars, time_to_full_texture_s) plus motif_economy.recurrence_lag_s. Note that
recurrence_lag_norm was V1's best hypothesis and is now the second-worst axis in the bank at
AUC 0.505 (PATTERN_FINDINGS_V2.md section 6) — do not build a delayed-return gate on it.
The slice's traditions are not tune cultures with exotic timbres. In several of them the
identity-bearing unit is not a melody at all, and the region pages already say so. The craft answer
per tradition:
punctuated by a colotomic structure of gongs marking a cycle of fixed length, with elaborating
parts above (https://ecampusontario.pressbooks.pub/beyondtheclassroom/chapter/balinese-music-gamelan/;
https://grokipedia.com/page/Colotomy). Kotekan is the interlocking of paired parts into a single
line faster than either player produces (https://en.wikipedia.org/wiki/Kotekan).
cycle closes and hears everything against that. The identity a player carries out of the room is
a period and a punctuation pattern, not a tune.
bali.md Section 10(L1894) says Bali is "the arc's first region where music is not scoring but structure" and that
the music "is built to be READ"; it also documents the ombak beat-note engineered by deliberately
mistuned pairs [SRC_00157, p. 316-317, p. 319-320]. sumatra_java.md Section 10 (L2229) states
plainly that its processed sources carry almost no instrumental ethnography and that the region's
sound is built around the voice; and the Minangkabau talempong pacik tradition is itself
interlocking, three players on six gong-chimes playing anak, tengah and peningkah.
tempo_and_time_grammar.dominant_onset_period_s (EX_124: 2.879 s, i.e. 20.84 cycles per minute) read together with section_boundaries_s. A colotomic cue should
show a strong, stable dominant onset period and section boundaries that are integer multiples of
it. Nothing in the card checks that multiple relation today; it is a cheap addition.
lexical tone levels of Yoruba, and the repertoire consists of real texts — oriki praise poetry and
owe proverbs (https://www.frontiersin.org/journals/communication/articles/10.3389/fcomm.2021.652542/full).
Villepastour's study of the bata shows the encoding is no less precise for being less iconic
(https://ethnomusicologyreview.ucla.edu/journal/volume/15/piece/481).
west_africa.md Section 10 (L2398) quotesthe tradition's own scholar that the dundun "can use codes to produce a type of language which
nearly all Yoruba people can decode" [SRC_00260, p. 135], and records the Ifa corpus imitating the
drums by name [SRC_00229, p. 112].
verbatim because it is a sentence, and it is memorable for the same reason a catchphrase is. Write
a drum phrase of fixed length and fixed contour, use it as a call, and let the ensemble answer it.
Do not develop it — a proverb that is developed stops being a proverb.
the_hook reads a pitched melodic band and returns available: false or noise on a drum-led cue. Named in section 12.
anchihoye — carried on krar, begena, masenqo and washint
(https://en.wikipedia.org/wiki/Qenet; https://www.ncbi.nlm.nih.gov/pmc/articles/PMC10118148/).
Our own ethiopia.md Section 10 (L713) supplies the vernacular roster with descriptions attached
— karar the six-string lyre, bagana the eight-or-ten-string floor harp, masanqo the one-string
fiddle, kabaro, nagarit, sanasal the sistrum, maqwamiya the prayer-staff [SRC_00254, p. 104-105].
over a fixed cyclic accompaniment. The listener's memory hook is the interval that identifies the
qenet, not a phrase.
runs pallavi, anupallavi, charanam; the pallavi establishes the raga's identity and returns as a
refrain after each later section; the anupallavi rises in register and introduces contrasting
material; each line is then varied through composed sangati
(https://artiumacademy.com/blogs/what-is-kriti-in-carnatic-music/;
https://www.mtosmt.org/issues/mto.15.21.4/mto.15.21.4.schachter.html). That is a refrain-and-
graded-variation architecture that solves the loop-fatigue problem natively, and it is a far
better template for a 3-5 minute looping cue than a Western AABA.
sri_lanka.md Section 10 (L908) documents theall-night preaching hall and the protective chant with its procedure [SRC_00256, p. 138, p. 152;
SRC_00237, pp. 62-66]. The hook there is a recitation cadence and a room.
tonality.key / mode / windowed_key_agreement see major/minor only. A qenet or a raga is invisible to the card. nostalgia_score.py Y12 (characteristic-degree rate) is the
right shape of feature and exists only for authored symbolic material.
hocketed motives and yodel register-shifts over interlocking ostinatos (Arom, African Polyphony
and Polyrhythm, Cambridge, 1991; https://en.wikipedia.org/wiki/Simha_Arom;
https://www.melodigging.com/genre/mbenga-mbuti-music). Among the Mbuti the singing uses short
hocket-like motives connected to the voco-instrumental hocket of the small flute.
central_africa_congo.md Section 10 (L1017) is more specific than any secondary sourceand is already a complete composition brief: overlapping entries, each singer holding her note as
the next comes in, a chorus cascading downward like a waterfall, swelling until breath runs out,
then all stopping together sharply and cocking their heads to listen for the forest's echo
[SRC_00253, p. 171-172]. It names that listened-for reply as the region's musical signature.
not his melody family. The memorable object is the collective attack, the swell and the hard
unison stop.
south_africa.md Section 10 (L599) records that the archive'sorganizing principle is that a song belongs to a specific creature, person or state, and names the
catalogue — the Cat's Song, the Song of the Caama Fox, the Broken String and the rest [SRC_00265].
Our canon's own design line is a catalogue of short named songs discovered rather than triggered,
sparse, vocal-first, with a deliberate absence of harmony. Khoisan practice supplies the technique:
cyclical forms, interlocking parts, handclap polyrhythm, yodel-like register shifts, hocketing,
and a gradual upward drift over the course of a song
(https://www.melodigging.com/genre/khoisan-folk-music).
lane_analysis.entry_exit_map and simultaneity_curve are the right instruments for entry-pattern hooks, but max_lanes is pinned at 9 for hits and siblings alike — a
measurement ceiling in the spectral_band_proxy method, confirmed across two full re-runs
(PATTERN_FINDINGS_V2.md section 10). No arrangement-width claim is admissible until a real
instrument roster exists.
rift_valley.md Section 10 (L624) opens by stating that no instruments are attested in its corpusand that no rhythmic structure, scale, call pattern or instrument is described anywhere on the
page. The chapter's musical content is the Ju|'hoansi healing-dance clapping carried from ratified
canon at reference register, plus "a deep-time silence at Olduvai".
Hohle Fels vulture-radius flute with five finger holes, roughly 35-40 thousand years old, plus
ivory flute fragments (https://www.sciencedaily.com/releases/2009/06/090624213346.htm). The
defensible register is therefore body percussion, voice, struck stone and end-blown pipe, and the
honest composition is near-absence with a floor of wind and stone, which is exactly what the
region page already specifies.
("no track will be boring or filler") and the canon brief ("the arc's quietest bed") pull against
each other. The resolution available inside both is that the Olduvai bed be short, be event-bearing
rather than static — a single struck-stone figure that recurs at a long, irregular period — and
never be the chapter's only cue. This is a craft recommendation to the music lane, not a canon
change.
Farbood's model is the best single empirical map. Both her experiments manipulated harmony, pitch
height, melodic expectation, dynamics, onset frequency, tempo, meter, rhythmic regularity and
syncopation, and her model is built on a moving perceptual window and trend salience — tension at a
moment is a function of the recent trajectory, not the instantaneous state (Music Perception 29/4,
2012, 387-428; http://mp.ucpress.edu/content/29/4/387). Lerdahl and Krumhansl's hierarchical model
combines prolongational structure, pitch-space distance, surface dissonance and attraction, and
implementations of it correlate strongly with listener tension ratings (Music Perception 24/4, 2007,
329-366; https://www.fredlerdahl.com/s/Modeling-Tonal-Tension.pdf).
under changing harmony, and modal mixture all raise tension. Lerdahl and Krumhansl's pitch-space
distance is the formal version of tonal distance; surface dissonance is a separate additive term
in the same model.
tonality.windowed_key_agreement is a modulation proxy only,and the card itself declares that a pivot-chord map needs symbolic transcription. Closing
measurement in section 12.
strain — a line sustained near the top of its range — is the extreme case. An unresolved leading
tone is the melodic form of an unresolved dominant.
register.melody_span_semitones, melody_tessitura_midi. There is no per-timemelodic-register curve on the card, only a global span; a time-series version is the cheap fix.
dissonance into grouping dissonance (pulse layers at different speeds, of which hemiola is the
commonest case) and displacement dissonance (same speed, misaligned)
(https://viva.pressbooks.pub/openmusictheory/chapter/metrical-dissonance/;
https://mtosmt.org/issues/mto.14.20.2/mto.14.20.2.biamonte.html). Hemiola clusters at cadences in
triple metre precisely because it heightens the tension the cadence then resolves.
because it raises perceived intensity without breaking a loop's grid. CONSENSUS.
tempo_function.adjustment_events (count, direction), groove.swing_ratio, groove.onset_rate_per_s, windowed_tempo_stability. Corpus tempo_adjust_n median 13.5 against
a sibling median of 6 (PATTERN_FINDINGS_V2.md section 10) — the tracks Josh named bend tempo
more than twice as often as their album neighbours. Metric dissonance itself is unmeasured.
powerfully: sudden loudness is one of the acoustic events that trips the brain-stem reflex, which
has no opt-out, and cross-cultural work correlates chill incidence with sudden peaks in loudness,
brightness and roughness
(https://www.frontiersin.org/journals/psychology/articles/10.3389/fpsyg.2017.02044/full;
https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4756174/).
that has spent its range on a long ramp has none left for the surprise. This is the craft argument
against front-loading loudness.
dynamics.dynamic_range_db_p95_minus_p10, arc_curve_db, and event_grammar.raises[].slope_db_per_s. Corpus dynamic-range median 11.7 dB against a sibling
median of 13.7 (AUC 0.369) — the corpus rows are slightly narrower, not wider, which is a warning
against reading dynamic range as a quality axis.
is the strongest release gesture available without harmony. Sloboda's structural inventory found
shivers most reliably evoked by new or unexpected harmony and tears by appoggiatura and sequence
(Psychology of Music 19, 1991, 110-120;
https://journals.sagepub.com/doi/10.1177/0305735691192002), which locates the strongest affective
events in harmony and voice-leading rather than in texture — texture is the amplifier.
lane_analysis.simultaneity_curve, register_spacing_octaves, event_grammar.drops[].lanes_removed, complexity_law.index_per_bar.
best-grounded psychoacoustic account: Plomp and Levelt showed maximum sensory dissonance at
roughly a quarter of the critical bandwidth, with consonance recovering as the separation exceeds
it, and the complex-tone case computed as the summed roughness of all partial pairs (JASA 38,
1965, 548-560; https://www.semanticscholar.org/paper/1d3ccc073b1b3f13f95e2392ff81a9eff0e7a4d2).
Vassilakis's model is the standard modern implementation
(https://www.mat.ucsb.edu/Masters/Brian_Hansen_Masters.pdf).
timbre.spectral_centroid_curve_hz, spectral_rolloff85_hz_mean, spectral_flatness_mean, signature_colour. Brightness is measured; roughness is not. Adding a
Plomp-Levelt/Vassilakis roughness curve is the single cheapest high-value feature we could add,
because it is computed from the spectrum we already extract.
case that made it famous is contested on its own arithmetic: Lendvai's reading of the first
movement of Music for Strings, Percussion and Celesta requires an added bar of rest to reach 89,
the dynamic climax is at bar 55 but the tonal climax is at 44, and Somfai's work on the sketches
found no proportional calculation of any kind
(https://blogs.ams.org/jmm2021/2021/01/08/did-bartok-use-fibonacci-numbers-in-his-music/).
dynamics.peak_position_normalised and opening_and_flow_law.flow_curve.peak_position_normalised:
| statistic | dynamic peak position | flow-curve peak position |
|---|---|---|
| median | 0.575 | 0.606 |
| p10 | 0.164 | 0.210 |
| p90 | 0.869 | 0.901 |
| min / max | 0.000 / 0.993 | 0.178 / 0.993 |
| within 0.55-0.70 | 7 of 28 | 5 of 28 |
| in the first half | 13 of 28 | 13 of 28 |
| after 0.75 | 7 of 28 | 10 of 28 |
(0.618), and the distribution is nearly flat. Thirteen of twenty-eight tracks peak before the
midpoint on both measures. Golden
section describes the median of this pool and is not a rule any of these composers appear to have
followed. A generator that pins every peak at 0.618 would be more uniform than the corpus it is
imitating, which is itself a defect. Use 0.5-0.75 as a soft target and treat anything outside
0.16-0.87 as worth a second look, not a rejection.
BATTLE 0.672 dynamic / 0.282 flow; TENSION 0.704 / 0.809; CREDITS_TRIUMPH 0.543 / 0.493;
EXPLORATION 0.462 / 0.763; MELANCHOLIC 0.459 / 0.640; TITLE_CHARACTER 0.460 / 0.346;
BOSS 0.291 / 0.805. No per-class claim is a test; five classes sit below the five-row floor.
strongest shape for a 90-second cue. A multi-peak design is the shape a 3-5 minute cue needs,
because a single peak leaves too much runtime as approach and aftermath.
CONSENSUS, and it follows from Farbood's moving-window result: tension is read from trajectory, so
a peak with no descent registers as a plateau. After a peak the arrangement must lose lanes, lose
register width, or lose harmonic motion — and it must do so for long enough to be heard as a
descent, not as a gap.
either withhold the expected full statement or resolve deceptively, so the listener's spent
expectation is available to spend again.
dynamics.arc_curve_db and opening_and_flow_law.flow_curve.values carry the fulltrajectory, so peak count, peak prominence and post-peak descent depth are all derivable today and
none of them is currently extracted. Adding peak-count and post-peak-descent-dB to the card is
low-cost and directly serves this section.
mechanics: extensive uplifters and risers, the drum-roll effect, large frequency changes and
filter sweeps, the removal and reintroduction of bass and bass drum, and a contrasting breakdown —
all of which create tension and anticipation
(https://dj.dancecult.net/index.php/dancecult/article/view/451).
containing break routines and one without, with skin conductance, self-reported affect and
embodied experience measured. The emotional-intensity peak starts at or immediately after the drop
and continues into the core section, accompanied by greater motion and increased skin-conductance
response; dancers' overall activity level shows a sudden decrease then increase matching the
music's structure (Music Perception 36/4, 2019, 371-389;
https://online.ucpress.edu/mp/article/36/4/371/62966).
breakdown removes the low end and the pulse, so the listener's reference level falls; when the
same bass returns it is heard against that lowered reference, not against the track's own opening.
This is the textural form of the subito argument in section 7.4, and it is why a drop preceded by
a full-texture build lands harder than a drop preceded by a louder build.
is guaranteed across the gap. A drop lands when the grid survives the removal — the listener can
still count, so the return is arrival rather than restart. A drop merely stops when tempo,
harmony and grid all vanish together, because then there is nothing to be early or late against.
CONSENSUS, and it is the operational content of Medina-Gray's meter parameter: strong smoothness
is all pulse streams continuing across the seam
(https://www.mtosmt.org/issues/mto.19.25.3/mto.19.25.3.medina.gray.html).
melody. The rule that follows is that a drop is a structural event and gets structural preparation:
it needs a build it terminates and a re-entry it justifies. A drop with neither is the over-drop
defect measured in section 10.
event_grammar.drops[] already carries db_delta, lanes_removed, bars_held_s and re_enters_at_s, and raises[] carries slope_db_per_s and lanes_added. A real drop is
identifiable today as a raise whose end is within a bar of a drop's start, with the drop's
bars_held_s at least a bar and a re-entry that restores at least the removed lanes. Nothing
computes that pairing. It is the highest-value derived feature in this document.
Pre-outcome: imagination (contemplating the future state) and tension (arousal and attention
raised immediately before the outcome, scaled by uncertainty, stakes and time-to-outcome).
Post-outcome: prediction (reward or punishment for the accuracy of the expectation itself),
reaction (fast, worst-case) and appraisal (slow, considered)
(Sweet Anticipation, MIT Press, 2006; https://mtosmt.org/issues/mto.09.15.3/mto.09.15.3.aversa.html).
independently of whether the outcome was good. This is why a listener enjoys a resolution they saw
coming, and why a composer must establish a pattern before violating it — the violation is only
worth anything if the prediction response was engaged.
response scales with estimated time to outcome, so holding the dominant longer literally raises
arousal before the resolution arrives.
they are acoustically identical apart from duration: silences following tonal closure are
identified faster and rated less tense than silences following unclosed music, and the preceding
material and the expectation of what follows both "seep into the gap" (Music Perception 24/5,
2007, 485-506; https://online.ucpress.edu/mp/article-abstract/24/5/485/95245).
context. A general pause after an unresolved dominant is tension; the same pause after a cadence
is punctuation. Our own card already declares this posture — the silence_budget field carries
the note that section 11.2 "treats rest as a FIRST-CLASS parameter, not an absence."
X_rest_grammar is defined in harness/music_gen/derive_patterns.py L368 as silence_frac * 100 + drops_per_min. It is a composite of how much of the track is below the
silence threshold and how often the texture is cut.
one of three carried hypotheses that died at N = 28 (PATTERN_FINDINGS_V2.md section 6). It
SUCCEEDS as a containment axis: the hit envelope 2.68-17.43 holds 26 of 28 held-out hits and is
one of the three axes in the joint box (section 8 of the same document). Corpus rows occupy a
narrow band of composed rest without sitting high or low on it.
| 28 corpus rows | 32 round-1 candidates | |
|---|---|---|
X_rest_grammar median | 8.45 | 15.92 |
X_rest_grammar range | 2.68 - 17.43 | 6.14 - 42.55 |
| drops per minute, median | 5.56 | 14.08 |
| drops per minute, range | 1.46 - 14.02 | — |
| silence fraction, median | 2.12% | 2.75% |
| outside the 2.68-17.43 band | 0 of 28 | 14 of 32, all on the high side |
the worst candidate sits at 42.55 against a corpus ceiling of 17.43. Every one of the fourteen
rejections is an over-drop, not an under-drop.
an expectation, and the expectation is built by the texture staying up. Cut every eight bars and
the listener stops predicting continuation, so removal stops being an event and becomes the
texture's normal behaviour — at which point it costs a lane and buys nothing. Farbood's
moving-window model says the same thing formally: tension is read from trend salience, and a
signal that reverses every few seconds has no trend to be salient. Over-dropping does not merely
fail to build tension; it destroys the mechanism by which tension could be built.
(5.56 per minute) with a hard ceiling near fourteen per minute, and a silence budget near two
percent. That is a descriptive band from a containment axis, and it is necessary-not-sufficient:
it rejects, it never certifies.
listener cannot hear. A cue that peaks at 0.6 and releases to nothing by 1.0 has a large
discontinuity at the wrap; a cue that is flat enough to wrap invisibly has no arc.
return to the pallavi after every section, so the loop point is a return the listener has already
heard four times rather than a splice. The Western equivalent is a rondo. Either way the wrap
falls on material that has already been established as a place the music comes back to.
close the loop with a perfect authentic cadence (NOSTALGIA_RUBRIC.md Y1, from Gervais on
Zelda's Theme's avoidance of the tonic and of PACs). A loop that closes fully asks to stop; a loop
that closes on an open half cadence asks to continue.
(do the pulse streams continue), timbre (is the instrumentation shared), pitch (are the new pitch
classes inside the previous five seconds' macroharmony), volume (is loudness consistent across the
seam) and abruptness (is the entry a cut or a decay). She explicitly refuses to collapse them into
one number, because the correct weighting is context-dependent, and she notes that disjunction is
a legitimate design choice when the game wants the player to notice a state change
(https://www.mtosmt.org/issues/mto.19.25.3/mto.19.25.3.medina.gray.html).
cadence_and_ending.loop_seam.seam_ratio has median 24.17 (p10 5.34, p90 63.98), and EX_124 reads
seam_clean: false with baked_fade_out: true. These are album masters with fades, not game
loops. Any loop-seam standard must be derived from authored material or from extracted game loops,
never from this pool.
repetition structure essentially consistent across platforms and generations, disproving his own
hypothesis that early hardware limits produced above-average repetition
(https://boblsturm.github.io/aimusic2020/papers/CSMC__MuMe_2020_paper_10.pdf;
https://uwe-repository.worktribe.com/output/6818560/). Our rubric's L4 band (0.55-0.70) is taken
from that survey's mean of 62.79% with SD 3.49.
repeat_sim_mean is above sibling median and novelty_mean is below it, both clearing the calibrated bar, and the six clearing axes hold the
same sign and magnitude in the 16-bit half and the modern half of the pool
(PATTERN_FINDINGS_V2.md sections 4 and 9). Structural self-return is era-invariant across the
two eras we hold.
Recommendations below are craft syntheses grounded in the parameter evidence of section 7, the
class-descriptive medians of section 8.1, and the region-page briefs. They are proposal-tier.
dynamic peak around 0.45-0.50, no single dominating climax, wrap on an open cadence over a
returning refrain. Novelty enters as instrument swaps rather than as dynamic events.
time_to_full_texture median is 3% of runtime), flow peak early at about 0.28 and the dynamic
peak late near 0.67, so the arrangement is wide throughout while loudness still has somewhere to
go. One drop, held one to two bars, in the final third.
earlier. Each phase gets its own build and its own release; the last phase withholds the full
harmonized statement of the boss motif until the final section.
budget at or slightly above the corpus median. Tension is carried by harmony and appoggiatura
rather than by density; Sloboda's tears-inducing structures are the target.
flow). Sustained unresolved harmony, one long delayed resolution, drops used sparingly and always
with a preceding build.
corpus median time_to_hook_s is 0.575 s), dynamic peak near 0.46, and the definitive full
statement of the theme reserved for the last third.
central peak, release, then a second smaller peak that quotes the arc's earlier themes. This is
the one class where a perfect authentic cadence is correct, because it does not loop.
fact that our own composite is 0.744 with all axes and 0.497 with the length-independent ones —
a single number would hide exactly that.
NOT_COMPUTABLE features excluded from the denominator). Section 11 of this document reaches the
same conclusion about seam measurement from the corpus side.
Y1) and the multi-motif loop (L1, L2), which are thetwo mechanisms with the most direct bearing on a cue surviving a hundred hearings, and it grounds
L2 in Kondo's own stated method.
signatures with 14 declared gaps, and CLEAR explicitly meaning "no seed signature matched" rather
than "clean".
after it.
climax placement or arc shape. Every principle in sections 7 through 10 above is outside its
scope. This is the largest structural gap.
cannot see one, because it scores a symbolic head cell rather than an arrangement.
C6 gates rest_density_melody >= 0.15, which is a budget.Margulis's result is that rest VALUE comes from context — the same silence after a cadence and
after an unresolved dominant are different objects. A density gate cannot distinguish them.
L2 requires three or four motifs per loop butsays nothing about their scheduling, and Josh's floor is explicitly about scheduling: no more than
three to five cycles without a new instrument, catch, melody or drop.
C3 bands melodic range at 7-17 semitones from a symbolic head cell; the formula card's the_hook.range_semitones reads
the audio's melodic band and has a corpus median of 22. Neither is wrong; they measure different
objects, and nothing in the stack says so. FINDING: the two should be named distinctly
(head_cell_range against measured_melodic_band_range) before either is quoted in a verdict.
invented from memory is a false negative wearing a green tick. With the Nintendo discs arriving,
direct transcription of the named exemplars is now unblocked as a polish pass.
through 6.5 above describe five identity mechanisms — cycle, speech-phrase, modal ostinato,
entry-pattern, named-song catalogue — that the rubric would score as failures for not being tunes.
FINDING: the rubric needs a declared dialect parameter, or its verdicts will push every region
toward the same Western dialect the region pages explicitly refuse.
Principles that can be enforced as COMPOSITIONAL CONSTRAINTS in the authored path today:
nostalgia_score.pyaxis 4, pre-render.
tappable rhythm) — axes 1 through 3, pre-render.
Y1, L6.L2, L3. prompt, since the generator responds on only 4 of 8 probed dimensions and section_count is not
among them.
1.5-14) — enforceable as an arrangement rule in the authored path.
Principles measurable only after the fact, from the rendered audio:
dynamics.arc_curve_db, flow_curve.silence_budget plus event_grammar.drops.timbre.spectral_centroid_curve_hz.cadence_and_ending.loop_seam.Principles NOT measurable from the current card, with the measurement that would close each:
array gives three computable quantities — cloud diameter, cloud momentum and tensile strain (the
distance between local and global tonal context) — with a Python implementation in midi-miner
(http://dorienherremans.com/sites/default/files/paper_tenor_dh_preprint_small.pdf). Running it on
audio requires a chroma-to-pitch-set step we do not have; running it on the authored symbolic path
requires nothing new.
from the spectrum already extracted for timbre. This is the cheapest closure in this list.
(https://www.sciencedirect.com/science/article/abs/pii/S0165027024002929 for the current Python
implementation). Blocked on melodic transcription, which is also blocking cadence census, mode
identification and the hook-signature seed set — one blocker, four consumers.
metre.measured is nulland bar arithmetic assumes 4/4.
spectral_band_proxy ceiling (max_lanes = 9 for hits and siblings alike across two re-runs). Source separation or a real
roster is the only route.
repertoire is invisible to the_hook.
THE SINGLE MOST IMPORTANT GAP: melodic transcription. Without a symbolic melody line, we cannot run
an expectancy model, cannot compute a cadence census, cannot identify a mode or a raga or a qenet,
cannot expand the anti-plagiarism seed set beyond its 13 signatures, and cannot connect the
rubric's authored-side features to the card's rendered-side features at all. Every other gap in this
list is a feature; this one is the bridge between our two measurement worlds.
https://www.tagg.org/xpdfs/burns87.pdf
https://mtosmt.org/issues/mto.14.20.1/mto.14.20.1.aziz.html ·
http://shanahdt.github.io/MUSI4331/lessons/phrases1.html
repetition. https://mtosmt.org/issues/mto.02.8.4/leydon_text.html
https://viva.pressbooks.pub/openmusictheory/chapter/metrical-dissonance/ ·
https://mtosmt.org/issues/mto.14.20.2/mto.14.20.2.biamonte.html
https://www.academia.edu/50292501/The_Melodic_Arch_in_Western_Folksongs
https://mtosmt.org/issues/mto.09.15.3/mto.09.15.3.aversa.html
Complexity (1992); Meyer 1973 on gap-fill. https://mutor-2.github.io/ScienceOfMusic/units/08/
Melodic Structure." Music Perception. https://www.researchgate.net/publication/224982434
Melodies?" https://www.researchgate.net/publication/271681259
Creativity, and the Arts 11/2 (2017), 122-135.
https://www.apa.org/pubs/journals/releases/aca-aca0000090.pdf
https://mtosmt.org/issues/mto.14.20.4/mto.14.20.4.albrecht.html
(2007), 485-506. https://online.ucpress.edu/mp/article-abstract/24/5/485/95245
learning and probabilistic prediction in music cognition," Annals NYAS (2018).
https://onlinelibrary.wiley.com/doi/full/10.1111/j.1756-8765.2012.01214.x ·
https://nyaspubs.onlinelibrary.wiley.com/doi/10.1111/nyas.13654 ·
https://www.sciencedirect.com/science/article/abs/pii/S0165027024002929
(2012), 387-428. http://mp.ucpress.edu/content/29/4/387
329-366. https://www.fredlerdahl.com/s/Modeling-Tonal-Tension.pdf
TENOR 2016. http://dorienherremans.com/sites/default/files/paper_tenor_dh_preprint_small.pdf
548-560. https://www.semanticscholar.org/paper/1d3ccc073b1b3f13f95e2392ff81a9eff0e7a4d2 ·
Vassilakis roughness: https://www.mat.ucsb.edu/Masters/Brian_Hansen_Masters.pdf
Music 19 (1991), 110-120. https://journals.sagepub.com/doi/10.1177/0305735691192002
https://dj.dancecult.net/index.php/dancecult/article/view/451
Perception 36/4 (2019), 371-389. https://online.ucpress.edu/mp/article/36/4/371/62966
https://www.mtosmt.org/issues/mto.19.25.3/mto.19.25.3.medina.gray.html
https://boblsturm.github.io/aimusic2020/papers/CSMC__MuMe_2020_paper_10.pdf
(2018), 291-304. https://journals.sagepub.com/doi/abs/10.1177/1029864917698010 ·
https://news.osu.edu/has-music-streaming-killed-the-instrumental-intro/
https://www.frontiersin.org/journals/psychology/articles/10.3389/fpsyg.2017.02044/full ·
https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4756174/
https://blogs.ams.org/jmm2021/2021/01/08/did-bartok-use-fibonacci-numbers-in-his-music/
https://en.wikipedia.org/wiki/Simha_Arom · https://www.melodigging.com/genre/mbenga-mbuti-music
https://www.frontiersin.org/journals/communication/articles/10.3389/fcomm.2021.652542/full ·
Villepastour, Ancient Text Messages of the Yorùbá Bàtá Drum:
https://ethnomusicologyreview.ucla.edu/journal/volume/15/piece/481
https://en.wikipedia.org/wiki/Kotekan ·
https://ecampusontario.pressbooks.pub/beyondtheclassroom/chapter/balinese-music-gamelan/
https://www.mtosmt.org/issues/mto.15.21.4/mto.15.21.4.schachter.html
https://www.ncbi.nlm.nih.gov/pmc/articles/PMC10118148/
https://www.auralarchipelago.com/auralarchipelago/talempongbotuang
docs/spine/CH_03.md L245, CH_04.md L242, CH_10.md L224; _source/02_Tier_2_Region_Pages/*.md Section 10; build/audio/exemplars/PATTERN_FINDINGS_V2.md;
build/audio/exemplars/formula_cards/EX_124.json; harness/music_gen/derive_patterns.py L368;
harness/music_gen/nostalgia_score.py; docs/proposals/music/NOSTALGIA_RUBRIC.md;
harness/music_gen/exemplar_hook_signatures.json.
Every MEASURED-HERE figure in this document is read from the formula cards on disk through the same
fields derive_patterns.flatten uses. The 28 corpus rows are the hit list in
PATTERN_FINDINGS_V2.md section 2; the 32 candidates are
build/audio/generated/slice_v1/formula_cards/*.json. Fields used: dynamics.peak_position_normalised,
opening_and_flow_law.flow_curve.peak_position_normalised, opening_and_flow_law.time_to_hook_s,
opening_and_flow_law.time_to_full_texture_s, event_grammar.counts.drops,
event_grammar.silence_budget.fraction, duration_and_structure.total_length_s,
duration_and_structure.section_count, complexity_law.boundary_step_share,
cadence_and_ending.loop_seam.seam_ratio, motif_economy.self_similarity_lift,
motif_economy.interval_3gram_repetition, the_hook.range_semitones. No audio was decoded, no
exemplar bytes moved, and no calibrated multiplicity bar was computed for any of them — they are
descriptive reads, and the section 8.1 and section 10.3 tables are the two that matter.