music/MELODIC_CRAFT_PRACTICE.md
CANON SUBORDINATION — this document is PROPOSAL-TIER: it serves canon and never outranks it.
Canon served: the composition floor at docs/spine/DECISIONS_PENDING_JOSH.md — MUSIC ROUND 1
GRADED + THE COMPOSITION FLOOR and its ROUND 2 GRADED successor; the melody-first / 30-year-bar
music direction (memory: music-direction-melody-first-30-year-bar); the theme rows of
registries/T0_Theme_Registry [DRAFT v0.1].
If this document disagrees with canon, CANON WINS and this document is the defect.
Tier: RESEARCH SYNTHESIS, proposal-tier. It contains no canon and rules on nothing.
Every claim about published craft carries a URL or a book, author and chapter. Every claim about
our own stack carries a file and a function, field or line. Claims are labelled **CRAFT
CONSENSUS** where practitioners and theorists agree without having measured anything, and
EVIDENCE where somebody measured something and published the number. The two are not
interchangeable and this document does not dress one as the other. A third label, **MEASURED
HERE**, marks numbers this document produced by running our own scorers — those are reproducible
from §5 and should be re-derived, not trusted.
CITATION AUDIT PASSED WITH REPAIRS — 2026-08-08. See §11 before citing anything from §4.
An independent pass fetched and re-read the source set. The published-corpus numbers in §3 came
through clean, cell for cell. Two fabricated claims and three misattributions were found — all of
them in the citation apparatus or in §4's game-practice prose — and all are repaired in place with
the original error stated rather than quietly overwritten.
Sibling volumes. MOTIVIC_DEVELOPMENT_AND_PHRASE_GRAMMAR.md covers how a tune develops;
HOOK_AND_TENSION_ARCHITECTURE.md covers where the hook sits in a cue;
ORNAMENT_AND_MELODIC_SURFACE.md covers the divided surface. This volume covers the one question
those three assume answered: what does a good melodic line actually look like, in numbers.
---
Josh graded our melodies as *a few notes slowly played back to back*. That is not a vague
complaint — it is a precise one, and it names four measurable properties at once:
Our two named melodic files are harness/music_gen/pass3_grammar.py (which builds the phrase
objects) and harness/music_gen/instr_melodic_intelligence.py (which scores the result). Between
them they measure thirteen things — METRIC_KEYS at instr_melodic_intelligence.py:1767-1772 —
and not one of them is note density, pitch range, or climax placement. The grade Josh gave
names three properties our instrument cannot see. That is the gap this document exists to close,
and §8 states it as findings the audit phase can consume.
The headline synthesis, stated before the evidence so the evidence can be checked against it:
**The cure for "a few notes slowly played back to back" is not more notes in the structural
line. It is a two-level line** — a slow structural skeleton carrying a fast divided surface —
plus a placed climax and a cadence. Schoenberg says exactly this (§2.3), the corpus
numbers are consistent with it (§3.1), and pass3_grammar.divide() already builds it while
instr_melodic_intelligence never checks whether it happened (§8, GAP-03).
---
Read-status is declared for every item, because a source consulted through a secondary summary is
weaker evidence than one read in full, and a research document that hides the difference is
lying by formatting.
| # | Source | Read status |
|---|---|---|
| S1 | **Arnold Schoenberg, *Fundamentals of Musical Composition*, ed. Gerald Strang with Leonard Stein (Faber & Faber, 1967).** Chapters III (The Motive), IV–VIII (Construction of Simple Themes), XI (Melody and Theme), XII (Advice for Self-Criticism). OCR scan used: <https://ia903208.us.archive.org/1/items/MusicTheoryGeorgeThaddeusJones1974/A.schoenberg-FundamentalsOfMusicalComposition_text.pdf>; alternate scan <https://archive.org/download/soi-book-collection-4/Schoenberg_Arnold_Fundamentals_of_Musical_Composition_no_OCR_text.pdf> | Read in full text (125pp OCR extracted and searched; the melody passages in §2 were read verbatim in context) |
| S2 | **William E. Caplin, *Classical Form: A Theory of Formal Functions for the Instrumental Music of Haydn, Mozart, and Beethoven* (OUP, 1998) and *Analyzing Classical Form* (OUP, 2013)**. <https://global.oup.com/academic/product/analyzing-classical-form-9780199747184> | Not read here. Already implemented in pass3_grammar.py; the sentence/period/continuation definitions used in §2.4 were checked against two teaching summaries: <http://shanahdt.github.io/MUSI4331/lessons/phrases1.html> and <https://smbutterfield.github.io/ibmt17-18/13-phrasing-texture/c2-tx-combperiodsent.html> |
| S3 | **Jack Perricone, *Melody in Songwriting: Tools and Techniques for Writing Hit Songs* (Berklee Press, 2000).** The Berklee songwriting-department melody text. <https://berkleepress.com/songwriting/melody-in-songwriting/> | Not read. Cited only for the fact that range, contour and melodic rhythm are the three axes the Berklee curriculum teaches; no numeric claim is attributed to it |
| S4 | **David Huron, *Sweet Anticipation: Music and the Psychology of Expectation* (MIT Press, 2006).** Source of the melodic "universals" — pitch proximity, step declination, step inertia, melodic arch, post-skip reversal. <https://books.google.com/books/about/Sweet_Anticipation.html?id=uyI_Cb8olkMC> | Not read in full. The five properties are taken as restated by Chiu & Temperley (S11, read in full) and the MUTOR unit <https://mutor-2.github.io/ScienceOfMusic/units/10/> |
| S5 | **Elizabeth Hellmuth Margulis, *On Repeat: How Music Plays the Mind* (Oxford, 2013; the review below gives 2013, not the 2014 date first recorded here).** <https://global.oup.com/academic/product/on-repeat-9780199990825> | Not read. Corrected 2026-08-08: this row previously claimed Margulis finds that repetition drives attention to the *sonic surface*. Albrecht's MTO review <https://mtosmt.org/issues/mto.14.20.4/mto.14.20.4.albrecht.html> reports the opposite direction — repeated exposure draws attention to *deeper structures*, and listeners got worse at identifying short surface repetitions with more hearings. No claim in this document rests on S5; the row is retained only so the reversal is on the record |
| S6 | **Winifred Phillips, *A Composer's Guide to Game Music* (MIT Press, 2014).** <https://mitpress.mit.edu/9780262534499/a-composers-guide-to-game-music/> | Not read. Her craft rules in §4.2 come from her own GDC write-ups (S18), which were read |
| S7 | **Tim Summers, *Understanding Video Game Music* (Cambridge University Press, 2016).** <https://books.google.com/books/about/Understanding_Video_Game_Music.html?id=gJXsDAAAQBAJ> | Not read. Cited only for the framing that game music carries proportionally more descriptive load than film music |
| # | Source | Read status |
|---|---|---|
| S8 | **Jakubowski, Finkel, Stewart & Müllensiefen (2017), "Dissecting an Earworm: Melodic Features and Song Popularity Predict Involuntary Musical Imagery", *Psychology of Aesthetics, Creativity, and the Arts* 11(2), 122–135.** Record: <https://research.gold.ac.uk/id/eprint/19405/>. Accepted manuscript: <https://research.gold.ac.uk/id/eprint/19405/1/Melodic_features_Paper_June2016%20%281%29.docx> | Read in full (accepted manuscript extracted and read; all §3.4 numbers are from its Results) |
| S9 | **Huron (1996), "The Melodic Arch in Western Folksongs", *Computing in Musicology* 10, 3–23.** <https://www.researchgate.net/publication/239063783_The_Melodic_Arch_in_Western_Folksongs> | Not read directly. Method and result taken as restated in Goldstein et al. (S10, read in full — this cell previously mis-pointed at S13) |
| S10 | Goldstein et al. (2023), "Exploring Melodic Contour: A Clustering Approach" (NYU / Hebrew University). <https://www.ripolleslab.com/uploads/1/2/6/7/126798162/goldstein_2023.pdf> | Read in full |
| S11 | **Chiu & Temperley (2024), "Melodic Differences Between Styles: Modeling Music With Step Inertia", *Music & Science* 7, 1–11.** DOI 10.1177/20592043231225731. Author PDF: <https://davidtemperley.com/wp-content/uploads/2024/06/chiu-temperley.pdf> | Read in full (Tables 1–2 transcribed into §3.2) |
| S12 | van Balen, Burgoyne, Bountouridis, Müllensiefen & Veltkamp (2015), "Corpus Analysis Tools for Computational Hook Discovery", ISMIR 2015, 227–233. <https://zenodo.org/records/1415038> | Read in full (Table 2 coefficients transcribed into §3.5) |
| S13 | Burgoyne, Bountouridis, van Balen & Honing (2013), "Hooked: A Game for Discovering What Makes Music Catchy", ISMIR 2013, 245–250. Definitions of catchiness and hook. <https://semanticscholar.org/paper/69f1faad66fb57bd09b2df7f0d90c735fd10d260> | Read via S14 ch.7 and S8's restatement |
| S14 | Jan van Balen, PhD thesis, chapter 7 ("Hooked"). <https://jvbalen.github.io/pdf/thesis-CH7.pdf> | Read in full (design chapter; results are in S12) |
| S15 | **"Trajectories and revolutions in popular melody based on U.S. charts from 1950 to 2023", *Scientific Reports* 14 (2024).** Madeline Hamilton & Marcus Pearce, *Scientific Reports* 14, art. 14749. The BiMMuDa corpus — 1131 MIDI files of the main melodies of the top-5 songs of every year 1950–2023, 366 manually transcribed melodies analysed. <https://www.nature.com/articles/s41598-024-64571-x> | Read (abstract, feature definitions and era means; figure-level values are approximate as noted in §3.1). Authors, article number and corpus counts added at the 2026-08-08 audit |
| S16 | Müllensiefen (2009), FANTASTIC: Feature ANalysis Technology Accessing STatistics (In a Corpus). The 82-feature melodic descriptor toolbox used by S8 and S12. Feature definitions as restated in S8 and in the Wagner leitmotive study <https://www.frontiersin.org/journals/psychology/articles/10.3389/fpsyg.2017.00662/full> | Read via S8 |
| S17 | **Müllensiefen & Halpern (2014), "The role of features and context in recognition of novel melodies", *Music Perception* 31(5), 418–435.** | Not read. Cited via S8 for the single finding that simple contours predict implicit melodic memory |
| # | Source | Read status |
|---|---|---|
| S18 | Winifred Phillips, three-part write-up of her GDC theme talks — "The Hook in Game Music" <https://winifredphillips.wpcomstaging.com/2020/06/16/game-composers-and-the-importance-of-themes-the-hook-in-game-music-pt-1/>, "Repetition in Game Music" <https://www.gamedeveloper.com/audio/game-composers-and-the-importance-of-themes-repetition-in-game-music-pt-2->, "Variation and Fragmentation in Game Music" <https://www.gamedeveloper.com/audio/variation-and-fragmentation-in-game-music-game-composers-and-the-importance-of-themes-pt-3-> | All three fetched and read in full 2026-08-08. Previously "pt 2 read via listing" — and a specific number had been attributed to pt 2 anyway; see the fabrication note in §4.2 |
| S19 | Koji Kondo, "Painting an Interactive Musical Landscape", GDC 2007. Recording: <https://archive.org/details/GDC2007Kondo>. Contemporaneous reports: <https://www.gamespot.com/articles/gdc-07-mario-composer-on-game-music/1100-6167047/>, <https://www.nintendoworldreport.com/feature/13118/koji-kondos-gdc-2007-presentation> | Talk not watched. The three-element framing (rhythm, balance, interactivity) is taken from the two reports and is confirmed verbatim in the Nintendo World Report piece (Aaron Kaluszka, 13 March 2007); the archive.org item confirms the talk title. The Mario tempo example is not in either report — see §4.1 |
| S20 | **"Analyzing Compositional Style in the Music of *The Legend of Zelda: Ocarina of Time*", Inquiry@Queen's Undergraduate Research Conference Proceedings. <https://ojs.library.queensu.ca/index.php/inquiryatqueens/article/view/16743> (DOI 10.24908/iqurcp16743). Author: Dominic Everitt**, 2023. | Abstract read only. The abstract confirms the "modes, chromatic harmony, parallel motion, proximal voice leading" phrase; it does not support the NES-timbre claim this doc formerly drew from it — struck in §4.1 |
| S21 | RPGFan, "The 'Eight Melodies' Template: How Koichi Sugiyama Shaped RPG Soundtracks". <https://www.rpgfan.com/feature/the-eight-melodies-template-how-koichi-sugiyama-shaped-rpg-soundtracks/> | Read. Note: contains no direct Sugiyama quotation on loop length or repeated listening — see §4.3 |
| S22 | Splice, "How the overworld theme of The Legend of Zelda takes us on a harmonic adventure". <https://splice.com/blog/legend-of-zelda-overworld-harmony/> | Fetched and read in full 2026-08-08 (was: search summary only — the summary produced a wrong modal reading, corrected in §4.1). Harrison Shimazu, 20 September 2019 |
| S23 | Hooktheory TheoryTab analyses — Zelda Overworld <https://www.hooktheory.com/theorytab/view/nintendo/the-legend-of-zelda---overworld-theme>, To Zanarkand <https://www.hooktheory.com/theorytab/view/nobuo-uematsu/to-zanarkand---final-fantasy-x>, Terra <https://www.hooktheory.com/theorytab/view/nobuo-uematsu/terra> | NOT read — the site returns HTTP 403 to automated fetch. Listed as a pointer for a human pass, not as evidence |
| S24 | **Jack F. Boss, "'Away with Motivic Working?' Not So Fast: Motivic Processes in Schoenberg's op. 11, no. 3", *Music Theory Online* 21.3 (September 2015).** Corroborates liquidation; the precise definition quoted in §2.5 is Schoenberg's own (S1 ch. VIII), read verbatim. <https://mtosmt.org/issues/mto.15.21.3/mto.15.21.3.boss.html> | Fetched and checked 2026-08-08 (title was previously recorded truncated) |
No published, numerically-grounded melodic analysis of a canonical game theme was obtainable.
Hooktheory (S23) has the note data and blocks automated access; the academic ludomusicology
literature (S7, S20) analyses harmony, mode and function rather than melodic statistics. Everything
this document says about Zelda, Final Fantasy and Dragon Quest is therefore **qualitative and
sourced to prose analysis, and §5 substitutes a public-domain reference set** — including one
melody that *is* a beloved game theme — so that the numeric side of the argument rests on melodies
this document could actually measure. A future pass with legitimate access to symbolic
transcriptions should redo §5 on the real targets.
---
Schoenberg states the requirements of a well-balanced melody in a single compact passage. Rendered
as conditions rather than quoted:
1. It progresses in waves — every rise is answered by a fall.
2. It approaches its high point through lesser high points, interrupted by recessions. The
climax is *arrived at*, never simply present.
3. Upward motion is balanced by downward motion.
4. Large intervals are compensated for by conjunct motion in the opposite direction. This is
the manual's version of what psychology later called post-skip reversal (§3.2).
5. It stays within a reasonable compass, not straying far from a central range.
Conditions 1–3 are contour conditions; 4 is an interval-order condition; 5 is an ambitus condition.
Our instrument measures none of the five directly. contour_depth
(instr_melodic_intelligence.py:480) measures hierarchy of a contour, not its shape; nothing
measures ambitus or wave balance at all. This is GAP-01 and GAP-02.
Schoenberg (S1, ch. III) defines the motive as intervals and rhythms combined into a memorable
shape or contour, which usually implies an inherent harmony. He then states the mechanism that
matters to us: a motive appears constantly and is therefore repeated; **repetition alone gives rise
to monotony, and monotony is overcome only by variation.** Variation means changing the
less-important features while preserving the more-important ones — change everything and you get
something foreign and incoherent; change nothing and you get monotony. He names the preservation of
rhythmic features as the cheapest way to buy coherence.
The corollary he draws in ch. VII is the operational one, and it is stated in capitals in the
source: preserving the rhythm permits extensive change to the melodic contour. In slow tempo
this allows far-reaching contour change; in fast tempo it protects comprehensibility.
CRAFT CONSENSUS. This is the same doctrine as pass3_grammar.OPERATIONS' spine_changing
flag, from the opposite direction: our flag asks whether the interval spine changed; Schoenberg
asks whether the *rhythm* survived. A variation that changes the pitches and keeps the rhythm is
the canonical developing variation. A variation that keeps the pitches and changes only the
register or instrument — six of round 2's ten transforms — is what he means by repetition.
In his analysis of Beethoven's Op. 2/1-II (Adagio) (S1, ch. VI, "Analysis of Periods from
Beethoven's Piano Sonatas"), Schoenberg makes an observation that is directly measurable and
directly on the point of Josh's complaint: the melodic approach to the climax in bar 6 is supported
by more frequent changes of harmony and an increasing use of small notes, and the downward
movement after the climax is equally significant. *(Correction 2026-08-08: this passage was
previously attributed here to Op. 2/2-IV. It is not — in the source the Op. 2/2-IV paragraph begins
immediately after this one and says something else entirely, about a caesura on V and a cadence in
the dominant. The observation itself is verbatim as stated.)*
And in ch. VII, on the cadence contour: if there is a climax, the melody is likely to recede from
it, balancing the compass by returning to the middle register — this decline, combined with
harmonic concentration and the liquidation of motivic obligations, is what delimits the structure.
This is the two-level line. The structural pitches may move slowly; the *surface* accelerates
into the climax and thins after it. "A few notes slowly played back to back" is precisely a line
that has the structural level and never grew the surface level. Schoenberg's remedy is not "write
more structural notes" — it is "divide the surface, and place the division where the climax is."
pass3_grammar.divide() (pass3_grammar.py:542) and _nct_figure() (:591) already implement
exactly this surface. **Nothing schedules the division against the climax, and nothing scores
whether it happened.** GAP-03.
Two further Schoenberg positions matter for a game score, both from ch. V (verified by chapter
offset in the scan; this section previously said ch. IV):
controls are delimitation, subdivision and simple repetition. A fast battle cue should carry
*fewer* distinct ideas than a slow one, stated more often — the opposite of the instinct that
says a busy cue needs busy material.
length and tempo require should be admitted. He explicitly notes that discretion of this kind is
most necessary where immediate intelligibility is the goal, and that the classical masters
themselves sought a "popular touch" in their themes — citing Mattheson's 1739 formulation that a
theme should contain something the whole world already knows.
That last position is the theoretical bridge to §3.5. Schoenberg in 1948 and van Balen's
regression in 2015 say the same thing: a memorable theme is partly conventional on purpose.
Our melodic_complexity axis has a variety half that rewards the opposite. GAP-05.
pass3_grammar.py implements Caplin (S2) faithfully — FUNCTION_NAMES at :131, THEME_TYPES at
:134, the cadence-strength ladder at :125-130, and builders sentence() (:992), period()
(:1093), hybrid_ant_cont() (:1193). The doctrine, restated only where §8 needs it:
continuation is defined by fragmentation (units get shorter) and harmonic acceleration
(chords change faster).
one. The antecedent's basic idea is juxtaposed against a *contrasting* idea; in a sentence it is
*repeated*.
uncharacteristic ones remain, which no longer demand continuation. It is generally supported by a
shortening of the phrase.
The grammar has all of this. The scorer has none of it — it re-derives "phrases" from rest gaps.
GAP-06 and GAP-07.
---
**EVIDENCE (S15, BiMMuDa, 366 manually transcribed lead-vocal melodies from the US top-5 of each
year 1950–2022).** Onset density, defined as notes per second, by era:
| era | onset density (notes/s) | mean pitch interval |
|---|---|---|
| 1950–1974 | ≈ 1.8 | ≈ 2.3 semitones |
| 1975–1999 | ≈ 2.0 | — |
| 2000–2022 | ≈ 2.8 | ≈ 2.0 semitones |
The paper also reports that in the most recent era **two-thirds of pitches lie within a
5.5-semitone range** — i.e. a pitch standard deviation of roughly 2.75 semitones — and that both
pitch-related and rhythmic information-theoretic complexity fell across the period, with revolutions
identified at 1975–76 and 2000–01 (tier-1 evidence) and 1996 (tier-2).
Two things follow, and the second is the one that stops this becoming a bad gate.
VOICE_MIN_RATE = 0.30 notes/s (instr_melodic_intelligence.py:153), sits **six to nine times
below** that. A line at 0.4 notes/s — unmistakably "a few notes slowly played back to back" —
passes our voice gate and is then never scored on density anywhere, because density is not in
METRIC_KEYS. GAP-01.
0.74 notes/s and Scarborough Fair at 0.66 — both far below the pop norm, both unimpeachable
melodies. A lyrical theme legitimately runs slow. The rule must therefore be **conditioned on the
cue's own idiom and stated at two levels** (structural vs surface), which is exactly rule R-03.
EVIDENCE (S11, Chiu & Temperley 2024, Table 1). Probability that a step is followed by another
step *in the same direction*, with a step defined as ≤ 2 semitones, pitch repetitions removed:
| corpus | pieces | total intervals | P(step inertia) |
|---|---|---|---|
| Essen Folksong Collection | 3,786 | 145,381 | 71% |
| Barlow & Morgenstern (classical instrumental themes) | 9,776 | 142,149 | 67% |
| Hymn Tune Index (1535–1820) | 17,683 | 901,372 | 63% |
| Rolling Stone "greatest songs" | 194 | 41,066 | 48% |
| Billboard Hot 100 | 214 | 68,493 | 43% |
With the step size widened to 3 semitones to accommodate pentatonic pop, the folk and classical
figures fall to 60% and 59% while pop is unchanged at 47% and 42%.
This is a style discriminator, not a quality bar, and that is the important reading. Folk,
classical-theme and hymn melodies *continue* through steps; rock and pop melodies *reverse* through
them. A scorer that hard-codes one target imposes a style. Our
melodic_complexity does exactly that with its step-to-leap term (§3.3).
The related Huron universals (S4, via S11 and S10) are CRAFT CONSENSUS supported by corpus work:
pitch proximity (melodies favour small intervals), step declination (descending steps outnumber
ascending), step inertia, the melodic arch, and post-skip reversal / gap fill (a large interval
tends to be followed by a change of direction). Chiu & Temperley note Huron's own finding that step
inertia may hold mainly for *descending* steps. Schoenberg's condition 4 (§2.1) is post-skip
reversal stated as a compositional instruction two decades before it was measured.
MEASURED HERE. instr_melodic_intelligence.melodic_complexity (:496) counts steps as
absolute intervals of 1–2 semitones (n_step) and leaps as ≥ 5 semitones (n_leap), forms the
ratio, and scores it as 1 − |log10(S:L) − log10(3.0)| / 1.3 — a target of S:L = 3.0, i.e. 75%
steps among counted motion. Evaluating that term:
| S:L ratio | structure-term score |
|---|---|
| 1 | 0.633 |
| 3 | 1.000 |
| 6 | 0.768 |
| 8.5 | 0.652 |
| 11 | 0.566 |
| 19 | 0.383 |
| 20 | 0.366 |
Measured S:L on real melodies (§5): Londonderry Air 8.5, Korobeiniki 11.0, Greensleeves 19.0,
Ode to Joy 20.0. Meanwhile the deliberately-bad control — eight slow notes, the Josh failure
mode — scores S:L 4.0 and therefore 0.904 on this term.
**Our instrument gives "a few notes slowly played back to back" a step-to-leap score of 0.904 and
gives Ode to Joy 0.366.** That is not a calibration wobble; the term is inverted relative to
practice over the range real melodies occupy. GAP-04.
Two causes, both fixable:
1. The dead zone. Steps are 1–2 semitones and leaps are ≥ 5, so **thirds (3–4 semitones) are
counted in neither.** Thirds are 17–30% of intervals in the §5 reference melodies. Excluding the
most common non-step interval from both sides of a ratio makes the ratio meaningless.
2. The target. 3.0 was chosen as "the one classical rule about order rather than variety"
(the function's own docstring). No corpus supports 3.0. The Essen/classical figure that *is*
supported is a step-inertia probability, not a step-to-leap ratio, and it is style-dependent
(§3.2).
EVIDENCE (S9 via S10; S10's own experiments). Huron reduced each Essen phrase to three points —
first note, mean of the interior, last note — and found the **convex (arch) contour most common,
followed by descending, ascending, and concave.** Goldstein et al. (S10) replicated without
assuming shapes in advance, clustering 35,793 European Essen phrases, and recovered the same four
families: convex, concave, descending, ascending. Arch and descending contours are common
cross-culturally and even in birdsong (Tierney et al. 2011, via S10). Complete short melodies tend
to describe an arch regardless of the shapes of their constituent phrases — the arch is a
multi-level property.
EVIDENCE (S8, Jakubowski et al. 2017). 100 tunes named as involuntary musical imagery matched
against 100 never so named, on 82 FANTASTIC melodic features plus tempo — 83 predictors in all
(this section previously said "83 FANTASTIC features"), second-order features computed against a
reference corpus of 14,063 commercial pop MIDI transcriptions (Geerdes). Random forest, then a
reduced three-predictor model. Results:
feature dens.step.cont.glob.dir exceeded 0.326 — i.e. the tune's overall melodic direction
profile was *common* relative to the pop corpus — approximately 80% of those tunes were
earworms.
those whose average gradient between melodic turning points was *less* like the corpus
(dens.int.cont.grad.mean ≤ 0.421) were more likely to be earworms.
the only tempo finding in the literature at the time and the authors flag it as novel.
predictive.
The shape of that result is the design instruction: **a conventional large-scale contour carrying
an unusual local gradient.** Familiar at the level you hum; distinctive at the level you notice.
Müllensiefen & Halpern (S17, via S8) independently found simple contours predictive of implicit
melodic memory.
EVIDENCE (S12, van Balen et al. 2015). Recognition data from the *Hooked!* game — 321 songs,
1,715 fifteen-second segments, catchiness operationalised as drift rate (reciprocal of median
recognition time under a linear-ballistic-accumulation model). Mixed-effects regression, stepwise
selection by Satterthwaite-adjusted F-tests at α = .005, coefficients on standardised component
scores.
Read the provenance column before the numbers. These coefficients are not one model. The
audio components come from the audio-only fit on the full 321 songs / 1,715 segments; the symbolic
components come from fits on the reduced set of 99 transcribed songs / 536 segments, because
transcriptions were unavailable for the rest. An earlier version of this table presented all rows
as a single regression, which overstated the symbolic side's sample.
| predictor | β̂ | 99.5% CI | which model |
|---|---|---|---|
| Vocal Prominence (audio) | 0.14 | [0.10, 0.18] | audio-only, 321 songs |
| Melodic Repetitivity (symbolic) | 0.12 | [0.06, 0.19] | symbolic-only, 99 songs |
| Timbral Conventionality | 0.09 | [0.05, 0.13] | audio-only, 321 songs |
| Melodic Conventionality | 0.06 | [0.02, 0.11] | audio-only, 321 songs |
| M/H Entropy Conventionality | 0.06 | [0.02, 0.10] | audio-only, 321 songs |
| Melodic/Bass Conventionality | 0.07 | [0.01, 0.13] | symbolic-only, 99 songs |
| Melodic Range Conventionality | 0.05 | [0.01, 0.08] | audio-only, 321 songs |
| Harmonic Conventionality | 0.05 | [0.01, 0.10] | audio-only, 321 songs |
| Timbral Recurrence | 0.05 | [0.02, 0.08] | audio-only, 321 songs |
| Sharpness Conventionality | 0.05 | [0.02, 0.09] | audio-only, 321 songs |
Marginal R² = .10 for the audio-only model on 321 songs and .10 for the combined model;
the audio-only and symbolic-only fits on the reduced 99-song set reach .06 and .07. Earlier
this table gave Melodic Range Conventionality as "0.05–0.07 [0.01, 0.13]", silently unioning three
different columns; the single-model value is now given. The authors' own reading of the symbolic side:
**melodic entropy and productivity both load negatively — recognisable melodies are more
repetitive** — and higher document frequency with lower second-order productivity means
recognisable melodies contain more typical motives.
Every melodic term that predicts catchiness points the same way: **less entropy, more repetition,
more conventional motives, more conventional range.** Our melodic_complexity axis is a
geometric mean of a five-term *variety* score and a three-term *structure* score
(instr_melodic_intelligence.py:496-560). There is no conventionality term anywhere in the
instrument, and the variety half rewards the exact quantity the catchiness literature finds
negative. GAP-05.
This is not an argument for blandness and must not be read as one. Schoenberg's own position (§2.4)
is that classical themes deliberately sought a popular touch; the earworm result (§3.4) pairs a
*conventional* global contour with an *unconventional* local gradient. The correct design is
conventional at the level of shape, distinctive at the level of detail — which is a two-band
requirement, not a single slider, and cannot be expressed by a scorer with only a variety axis.
CRAFT CONSENSUS, converging from three directions:
small-note activity, and the line recedes to the middle register afterwards.
two-thirds to three-quarters of the way through — as taught in songwriting curricula and
restated in accessible summaries (<https://hackmusictheory.com/blogs/theory/posts/6959928/melody-interval-rule>,
<https://champaignschoolofmusic.com/how-is-the-golden-ratio-present-in-music/>). Labelled
consensus, not evidence: no corpus study measuring peak position was obtainable in this pass.
the same statement in aggregate form.
MEASURED HERE (§5): Londonderry Air places its peak at 0.75 of the way through the strain —
squarely on the consensus figure. Ode to Joy places it at 0.10, and is a descending-family melody
with a repeated-period structure, which is a legitimate different design. So peak placement is a
per-theme design target to be declared, not a universal gate — see R-06.
EVIDENCE (S10) for the first half; CRAFT CONSENSUS restated by S10 for the second. In the Essen
database a phrase is *defined* as a unit marked by metrical quality, musical rests, and musical
syntax — that is the annotation scheme itself, read verbatim in S10. Separately, S10 lists as a
convention of Western melody, citing prior literature rather than measuring it here, that **notes at
the ends of phrases tend to be relatively long**. The two were previously presented as one
measurement; only the first is an Essen fact. That agogic cue is one of the two our
segment_phrases uses. The other class — a phrase that closes
harmonically inside an unbroken rhythmic surface — is invisible to a rest-and-agogic segmenter, and
segment_phrases' own docstring says so honestly (instr_melodic_intelligence.py:762-778). §5
shows what that costs in practice.
---
Kondo's GDC 2007 talk — its actual title is "Painting an Interactive Musical Landscape",
confirmed on the archive.org item — names three essentials for gameplay music: **rhythm, balance,
and interactivity.** That triad is directly confirmed in the Nintendo World Report write-up, as is
the contrast between Mario's music, which announces the action, and Zelda's, which is written for
ambience.
*Sourcing correction 2026-08-08:* this section previously offered *Super Mario Bros.*' tempo
increase as time runs out as Kondo's own illustration of interactivity. The tempo speed-up is a real
and well-known property of that game, but neither cited report attributes it to this talk, and
the talk itself was not watched. It is retained here as a general illustration of state-driven
music, not as something Kondo said.
CRAFT CONSENSUS with direct pipeline consequence: an "internal rhythm" the melody must ride is
a statement about onset density relative to the player's motion, not about tempo. A traversal
cue and a battle cue at the same BPM should not have the same note rate. That is R-03's second
level.
Ludomusicological analysis (S20 — Dominic Everitt, 2023) names Kondo's recurring devices as
modes, chromatic harmony, parallel motion and proximal voice leading; that phrase is verbatim in
the abstract. Two claims previously in this paragraph have been cut as unsupported:
— removed. S20 was read at abstract level only, and its abstract says merely "limitations of
the 1990s era game console hardware". Worse, S20 analyses *Ocarina of Time*, which is an N64
title, so an NES voice-count was never the right hardware claim for that source at all.
(S22)"* — corrected. S22 (Harrison Shimazu, Splice, 20 September 2019), now fetched in full
rather than read through a search summary, places the theme in B♭ major and explicitly
*rejects* the Mixolydian reading in favour of chords borrowed from the parallel minor. "B♭-centred"
survives; "Mixolydian inflection" and "consistently analysed" do not.
The most directly actionable studio practice found in this pass, and it is about **restatement
count**, which nothing in our stack currently plans:
Her worked example is the Star Wars main title: a four-bar melody with internal repetition,
alternating full statements with a contrast or bridge. Verbatim in Pt 1: this pattern happens
five times over the course of the theme.
pattern is heard 14 times in the main theme alone.
*Spore Hero* theme through the game's opening menu and its most frequently-used stinger.
different arrangement rather than being replaced.
FABRICATION REMOVED 2026-08-08. This section previously asserted that for *The Dark Eye: Book
of Heroes* Phillips "restated the theme in six variations, and used a roughly four-minute
dedicated main-theme track to assert the identity before propagating it." **Both numbers are
unsupported.** Pts 2 and 3 were fetched in full for this audit; the game is named in each as one of
her projects, but neither states a variation count and neither describes a main-theme track of any
duration. The doc's own §1 read-status for Pt 2 was "read via listing" — i.e. not read — and a
specific number was nonetheless attributed to it. Nothing downstream may cite "six variations".
**This is Schoenberg's monotony-and-variation doctrine (§2.2) restated by a working game composer,
plus two verified numbers: five full statements across a main theme (Star Wars), 14 restatements of
a four-chord signature (Liberation), and a five-note fragment as the smallest transportable unit.**
The Dragon Quest template — Overture, Castle, Town, Field, Dungeon, Battle, Final Battle, Ending —
is real and is correctly described as the template most RPG soundtracks inherited, with per-slot
stylistic commitments (Baroque contrapuntal castle music, romantic field music, deliberately
frantic and dissonant battle music), reused and expanded across the series.
Correction to the domain brief: this pass found no primary source for a Sugiyama doctrine
about eight-bar loops or composing specifically for repeated listening. The RPGFan feature contains
no such quotation. **That claim should not be repeated in our docs until a primary interview is
found.** What *is* supportable is the structural claim: a small fixed set of strongly-characterised
slot themes, each reused and varied across a long series.
Everything in §4 is craft consensus and prose analysis. **No number in this document about a
commercial game theme has been verified against its notes**, for the access reason declared in
§1.4. The numeric backbone of this document is §3 (published corpora) and §5 (our own scorers on
public-domain melodies). Any future doc that wants to assert "beloved game themes have property X
with value Y" must first obtain symbolic transcriptions and measure them.
---
MEASURED HERE. Five public-domain melodies were encoded as (MIDI pitch, duration in beats) with
a stated tempo, realised into feature_rig.Note objects, and passed through the *actual* functions
in harness/music_gen/instr_melodic_intelligence.py — melodic_complexity, contour_depth,
segment_phrases, phrase_sophistication. A sixth entry is a **deliberate control encoding Josh's
grade**: eight notes, one per bar at 70 bpm.
One of the five is a beloved game theme in its own right: *Korobeiniki* is the Russian folk song
(1861) used as the Tetris Type-A theme, and is public domain.
Reproduce with the script preserved at
C:/Users/joshu/AppData/Local/Temp/claude/C--dev-humanity-forgotten/c3c5f839-9982-4e1c-a28d-8fb8163ec031/scratchpad/probe.py
(transient; the encodings are restated in §5.3 so the measurement can be rebuilt anywhere).
| melody | n | notes/s | range (st) | mean abs iv | step% | 3rd% | leap% | S:L | inertia% | peak@ | contour | depth | H1 | complexity | phrases | phrase soph. |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Korobeiniki (Tetris A) | 20 | 1.67 | 7 | 1.95 | 58 | 21 | 5 | 11.0 | 78 | 0.00 | descending | 4.00 | 3.46 | 0.461 | 1 | *unavailable* |
| Ode to Joy | 30 | 1.88 | 7 | 1.24 | 69 | 0 | 0 | 20.0 | 92 | 0.10 | descending | 4.00 | 2.17 | 0.569 | 1 | *unavailable* |
| Londonderry Air | 25 | 0.74 | 14 | 2.21 | 71 | 17 | 8 | 8.5 | 64 | 0.75 | convex/arch | 5.00 | 2.82 | 0.512 | 1 | *unavailable* |
| Greensleeves | 28 | 1.30 | 11 | 2.07 | 70 | 30 | 0 | 19.0 | 71 | 0.15 | convex/arch | 4.50 | 3.00 | 0.699 | 1 | *unavailable* |
| Scarborough Fair | 13 | 0.66 | 9 | 2.33 | 50 | 25 | 8 | 6.0 | 50 | 0.17 | convex/arch | 5.00 | 3.05 | 0.657 | 1 | *unavailable* |
| CONTROL — "a few notes slowly" | 8 | 0.29 | 5 | 2.29 | 57 | 43 | 0 | 4.0 | 50 | 0.71 | convex/arch | 4.00 | 0.00 | 0.384 | 1 | *unavailable* |
Columns: n note count; notes/s onset density; range ambitus in semitones; step%/3rd%/leap%
share of intervals at 1–2 / 3–4 / ≥5 semitones (our own definitions); S:L our step-to-leap ratio;
inertia% computed by the Chiu & Temperley method (§3.2); peak@ position of the highest note as a
fraction of the line; contour Huron's three-point reduction; depth, H1 and complexity are
contour_depth, interval_entropy and melodic_complexity as our instrument returns them.
Six findings, each of which becomes a rule in §6 or a gap in §8.
1. melodic_complexity barely separates the failure mode from real melodies. The control scores
0.384; the Tetris theme scores 0.461; Londonderry Air scores 0.512. A 0.08 margin
between the thing Josh rejected and one of the most-recognised melodies of the last century is
not a working discriminator. Greensleeves at 0.699 does separate — but Greensleeves is the most
*ornamented* of the set, which is the two-level-line point (§2.3) arriving through the back door.
2. The step-to-leap term is inverted over the real range (§3.3). The control gets 0.904; Ode to
Joy gets 0.366.
3. contour_depth does not discriminate at all. The control scores 4.00 — identical to
Korobeiniki and Ode to Joy, and above Greensleeves' 4.50 only by rounding. Windowed
Douglas-Peucker depth saturates on any line that wanders.
4. phrase_sophistication was unavailable for every single melody in the set, including Ode to
Joy, which is the textbook example of a four-phrase double period. Cause in §5.4.
5. Density is genre-conditional. 0.66–1.88 notes/s across five unimpeachable melodies, against
a pop-corpus norm of 1.8–2.8 (§3.1). A flat floor would fail Scarborough Fair. The control's 0.29
is nonetheless below every real melody by a factor of 2.3.
6. Ambitus separates cleanly where nothing else does. Real melodies span 7–14 semitones; the
control spans 5. Range is currently measured nowhere in the instrument.
Stated as (MIDI pitch, beats), tempo in bpm. These are standard published forms of public-domain
melodies; they are given so the measurement can be rebuilt and challenged.
74·1, 72·1, 71·3, 72·1, 74·2, 76·2, 72·2, 69·2, 69·2.
(all others 1 beat), then the same with the close 74·1.5, 72·0.5, 72·2.
69·3, 67·1, 72·4, 72·1, 74·1, 76·2, 79·2, 81·2, 79·1, 76·1, 77·3, 76·1, 74·2, 72·2.
72·1.5, 71·0.5, 69·1, 68·2, then the strain repeats with the close 67·1, 69·1, 71·1, 68·1.5,
67·0.5, 65·1, 69·3.
72·3, 69·6.
Caveat, stated rather than buried: these are working encodings of the principal strain, not
critical editions. A wrong note would move a percentage by a point or two; it would not move any of
the six findings, all of which turn on gaps of 2× or more or on a metric returning *unavailable*.
Because finding 4 is severe, it was isolated. Ode to Joy was re-realised three ways and passed
through segment_phrases + phrase_sophistication:
| realisation | phrases found | phrase_sophistication |
|---|---|---|
| legato (notes 95% of their beat value) | 1 | unavailable — "fewer than two phrases" |
| detached (notes 60% of value) | 2 | available |
| with a real breath rest at each cadence | 4 | available |
The segmenter is doing what its docstring says it does — finding rests and agogic accents — and its
docstring is honest that a phrase carrying neither is invisible to it. But the consequence in
practice is that an entire scoring axis silently vanishes on any legato line, and legato lines
are most of a lyrical score. The agogic branch also fails on a knife edge here: it requires a note
strictly greater than 2.0× the local median duration, and Ode to Joy's cadential half-note
against a quarter-note median is exactly 2.0×.
The strongest single finding of this pass. A period was generated with the real builder —
pass3_grammar.period(cell, cell, mode, 0, "P1", weak_cadence="HC", strong_cadence="PAC") — and the
resulting note list was scored by the instrument.
pass3_grammar says: theme_type = period, 18 bars, 60 notes, cadence evaded_then_PAC at bar 17 with the weak HC recorded. By construction that is two phrases,
antecedent and consequent, with an expansion.
instr_melodic_intelligence.segment_phrases says: 11 phrases, median length 5 notes.phrase_sophistication then computed phrase_length_variety 0.5934, antecedent_consequent_rate 0.2, cadential_differentiation 0.6522, phrase_hierarchy 1.0, total 0.713 — every one
of those numbers computed over eleven fragments that the grammar does not consider phrases.
Worse, the tonic estimate was wrong. phrase_sophistication estimates the tonic as the most
frequent pitch class (instr_melodic_intelligence.py:842). On this grammar-generated period the
true tonic pitch class was 2; the estimate returned 4. Since "closed" is defined as ending
on the tonic pitch class, the entire open/closed classification — and therefore the
antecedent-consequent detection — was anchored to the wrong note.
And separately from the estimator being wrong, the *definition* is too narrow. Running each cadence
gesture the grammar can produce through the instrument's closed test with the true tonic:
pass3_grammar cadence | melodic goal | instrument verdict |
|---|---|---|
| PAC | scale degree 1 | CLOSED ✓ |
| IAC | scale degree 3 | OPEN ✗ |
| HC | scale degree 4 | OPEN (correct as a weak ending, but see GAP-08) |
| deceptive | scale degree 2 | OPEN ✓ |
An imperfect authentic cadence is a closed cadence in Caplin and in Schoenberg; our instrument
calls it open because it accepts only the tonic pitch class. A period built as
antecedent-HC → consequent-IAC — a completely standard construction our own grammar can emit —
scores an antecedent-consequent rate of zero.
---
Each rule names the file and the function or constant it lands in. Rules are ordered by how much
they close the gap Josh named. ENGINE/CANON declaration: every rule here is ENGINE — it changes
how we generate and score audio, and touches no world canon.
Where: instr_melodic_intelligence.py — add to METRIC_KEYS (:1767) and AXES (:1773).
note_rate_per_s is already computed at :350 and :390 and is currently used only as a *gate*
at :946. Promote it to a scored axis.
What: report structural note rate (rate of the line's own long/accented notes) and
surface note rate (rate of all notes) separately. Evidence targets (§3.1): pop-idiom surface
lives at 1.8–2.8 notes/s; a lyrical theme legitimately runs 0.6–1.5. **Do not gate on an absolute
floor.** Gate on the *ratio*: a line whose surface rate equals its structural rate has no divided
surface at all, which is precisely the control's signature.
Where: instr_melodic_intelligence.py — VOICE_MIN_SPREAD = 1.5 (:156) is a
voice-detection floor of one and a half semitones of pitch standard deviation. Add a scored
pitch_range_semitones and pitch_sd_semitones.
What: §5 measures real melodies at 7–14 semitones of ambitus against the control's 5.
BiMMuDa (§3.1) reports two-thirds of pitches inside 5.5 semitones for modern pop, i.e. SD ≈ 2.75.
Range is the single cleanest discriminator this pass found and we do not measure it.
Where: instr_melodic_intelligence.melodic_complexity (:496), the structure block.
What: three changes.
1. Close the dead zone. Thirds (3–4 semitones) are 17–30% of intervals in real melodies and are
counted as neither step nor leap. Either fold them into leaps (the conventional definition: a
step is a second, everything else is a leap) or report a three-way histogram.
2. Drop the 3.0 target. No corpus supports it; it scores Ode to Joy at 0.366 (§3.3).
3. Prefer step inertia (§3.2) as the order statistic, because it is measured, style-diagnostic,
and has published per-corpus values: 71% Essen, 67% classical themes, 63% hymns, 48% Rolling
Stone, 43% Billboard. Score against the declared idiom of the cue, not a universal target.
Where: instr_melodic_intelligence.melodic_complexity — the five-term variety half.
What: the catchiness regression (§3.5) finds melodic entropy negative and melodic
repetitivity, motive typicality, and range conventionality positive. The earworm study (§3.4)
finds conventional global contour the strongest single predictor. Our variety half rewards the
opposite of all four. Add a second-order term — how typical is this line's contour and interval
profile against our own corpus of authored cues — and make the axis a band, not a maximiser:
conventional at the level of shape, distinctive at the level of local gradient. This is the single
most important conceptual change in this document and it is the one most likely to be got wrong by
implementing it as another slider.
Where: new function in instr_melodic_intelligence.py; scheduling hook in
pass3_grammar.py's phrase builders.
What: report (a) the position of the highest note as a fraction of the line, (b) whether
density and interval activity increase into it, and (c) whether the line **recedes to the middle
register** after it. All three are Schoenberg's explicit conditions (§2.1, §2.3) and (a) is measured
here at 0.75 for Londonderry Air. Report per theme; do not gate universally, because a descending
melody like Ode to Joy legitimately peaks early (§3.6).
Where: instr_melodic_intelligence.contour_depth (:480) stays; add a shape classifier.
What: Huron's three-point reduction — first note, mean of interior, last note — classifying into
convex / concave / ascending / descending (§3.4). It is four lines of code, it is the reduction the
published corpus results use, and §5 shows contour_depth alone cannot tell the control from
Ode to Joy (both 4.00).
Where: the interface between pass3_grammar.Phrase (:921, with to_dict() carrying
functions, cadence, start_bar, bars, derived_from, operation) and
instr_melodic_intelligence.measure() (:1087), which today takes only a card and an audio path.
What: when we composed the cue, we know the phrase boundaries, the cadence types, and the
formal functions — they are in the plan. Re-deriving them from a rest-gap heuristic on rendered
audio throws away ground truth and then scores the noise (§5.5: 11 phrases where the grammar built
2). Pass the phrase plan through the card and let segment_phrases validate against it instead
of replacing it, keeping the audio path as the fallback for material we did not author.
Where: instr_melodic_intelligence.phrase_sophistication (:821).
What: accept scale degrees 1 and 3 as closed (an IAC is a closed cadence), and either take
the tonic from the plan — which we know exactly — or replace the most-frequent-pitch-class
heuristic with a Krumhansl-style key-profile correlation. §5.5 shows the current estimator returning
pitch class 4 for a period whose tonic is pitch class 2.
Where: segment_phrases (:762), constants PHRASE_GAP_MULT = 1.55 (:184) and the agogic
branch.
What: the agogic test requires a note strictly greater than 2.0× the local median, which a
textbook cadential half-note against a quarter-note median fails by exactly zero (§5.4). Relax to
≥ 1.8× and add a harmonic-arrival cue. Until then, report the axis as unavailable-with-reason
rather than letting a legato cue quietly carry no phrase score at all.
Where: the cue plan schema (build/audio/pass2/plans/*.json) and whatever authors it; the
hook.returns list already exists.
What: Phillips (§4.2) gives working numbers: five full statements of the melody across a
main-theme cue with contrast between them, 14 restatements of a short signature pattern within a
single main theme as the high end of the range, and a five-note fragment as the smallest
transportable unit for stingers and menus. (The "six variations" figure this rule previously carried
was a fabrication and has been struck — see §4.2.) Our round-2 cues carried 21 statement layers
across three cues of which literal alone was eight — restatement without variation, which is
Schoenberg's named monotony (§2.2).
Where: pass3_grammar.OPERATIONS (:63-123) and the spine_changing discriminator.
What: Schoenberg's stated corollary is that **preserving the rhythm licenses far-reaching
contour change** (§2.2). Our spine_changing flag asks the pitch question; add the rhythm question
as its partner, so an operation can be classified as *developing variation* (spine changed, rhythm
preserved) versus *re-dressing* (spine preserved) versus *foreign* (both changed). The middle
category is the one Schoenberg says produces coherence and it is not currently nameable.
Where: pass3_grammar.divide() (:542), _nct_figure() (:591), division_level on Phrase.
What: division_level is currently a phrase-wide constant. Schoenberg's climax passage (§2.3)
requires the small notes to increase into the high point and thin after it. Make division a
schedule across the phrase, keyed to the climax bar, rather than a flat setting. This is the
concrete mechanism that turns "a few notes slowly" into a real line without changing a single
structural pitch.
Where: cue planning; build/audio/music_cue_table.json.
What: earworm tunes averaged 124.10 bpm against 115.79 for matched non-earworms (§3.4),
and Kondo's first essential is that the music ride the game's internal rhythm (§4.1). Neither our
plans nor our scorers relate a cue's tempo to its melodic density or to the player action it
accompanies. At minimum, record the intended player motion per cue and check note rate against it.
Where: melodic_complexity (:496) and predictive_info (:420).
What: predictive_info returns all-zero for sequences shorter than 8 symbols. A melody of
exactly 8 notes yields 7 intervals, so interval_entropy is reported as 0.0 while
melodic_complexity still reports available: true — a measured-looking zero for something that
was never measured. Our control hit exactly this. Propagate an availability flag.
---
Ranked. Everything above the line is a defect fix and should land before any new melodic authoring
run; everything below is new capability.
Adopt now — these are corrections, not enhancements:
1. R-03 — fix the step-to-leap term. It currently scores the failure mode at 0.904 and Ode to
Joy at 0.366. Any melodic ranking produced while this stands is untrustworthy in a way that
points *toward* the material Josh rejected.
2. R-08 + R-09 — repair phrase_sophistication. It was unavailable for all five reference
melodies and produced numbers over 11 phantom phrases on our own grammar's period, off a wrong
tonic. It is currently reporting confidently and wrongly, which is worse than not reporting.
3. R-07 — hand the grammar's phrase plan to the scorer. This is the structural fix that makes
R-08 and R-09 mostly unnecessary for authored material, and it costs nothing but wiring: the
ground truth already exists in pass3_grammar.Phrase.to_dict().
4. R-14 — availability flags. Cheap, and it is the difference between a gate that is armed and
a gate that reads zero.
Adopt now — these are the three properties Josh's grade named that we do not measure:
5. R-01 — note density as a scored axis, at two levels. Directly names "slowly".
6. R-02 — ambitus. The cleanest discriminator found in §5, and directly names "a few notes".
7. R-12 — division scheduled against the climax. The mechanism that fixes "back to back"
without touching the structural line. Pairs with R-05.
Adopt next:
8. R-05 — climax measurement, reported per theme, gating nothing until we have a per-cue design
target to gate against.
9. R-06 — contour shape classification. Four lines, and it is the reduction every published
corpus result in §3.4 uses, which makes our numbers comparable to theirs.
10. R-04 — the conventionality axis. The highest-value item conceptually and the highest-risk
to implement. It requires a reference corpus of our own authored cues before it means anything,
and it must land as a band rather than a maximiser. Do not ship it as a slider.
11. R-10 + R-11 — restatement planning and the rhythm-preservation partner flag. These belong
with the motivic lane (MOTIVIC_DEVELOPMENT_AND_PHRASE_GRAMMAR.md §6) rather than here.
12. R-13 — tempo and player motion. Blocked on the cue table carrying an intended-motion field.
Explicitly NOT adopted:
the pop corpus would fail it. Density is conditional (R-01).
recognisability. The instinct that "more complex" means "higher entropy" is the instinct that
produced a variety half with no counterweight, and adding more of it would make the melodies
worse in the exact direction the literature predicts.
found (§4.3).
---
For the audit phase. Each gap names the file, the location, the practice it deviates from, and the
evidence. Severity is this document's own judgement and the audit is free to overturn it.
| ID | Severity | File and location | Deviation | Evidence |
|---|---|---|---|---|
| GAP-01 | HIGH | instr_melodic_intelligence.py — METRIC_KEYS :1767, AXES :1773; VOICE_MIN_RATE = 0.30 :153; note_rate_per_s computed :350, :390, used only as a gate :946 | Note density is computed and thrown away. It is a voice-detection gate, never a scored axis. The gate floor of 0.30 notes/s sits 6–9× below the pop-corpus norm, so a line at 0.4 notes/s — literally "a few notes slowly" — passes and is never assessed. | §3.1 (BiMMuDa 1.8–2.8 notes/s); §5.1 (control at 0.29 vs 0.66–1.88 for real melodies) |
| GAP-02 | HIGH | instr_melodic_intelligence.py — no ambitus metric; VOICE_MIN_SPREAD = 1.5 :156 is the only pitch-spread constant and is a gate | Pitch range is never measured. Schoenberg's fifth condition is a compass condition; the corpus literature reports range and range-conventionality as predictive. Our only range-adjacent number is a 1.5-semitone voice-detection floor. | §2.1 condition 5; §3.5 (Melodic Range Conventionality β̂ 0.05–0.07); §5.1 (real 7–14 st, control 5 st) |
| GAP-03 | HIGH | pass3_grammar.divide() :542, _nct_figure() :591, Phrase.division_level :921; nothing in instr_melodic_intelligence.py | The two-level line is buildable and neither scheduled nor scored. division_level is a flat per-phrase constant; Schoenberg requires small-note activity to increase into the climax and thin after. No metric checks whether a surface exists over the structural pitches. This is the direct mechanism for Josh's grade. | §2.3 (Schoenberg on Beethoven Op. 2/1-II); §5.2 finding 1 (Greensleeves, the most ornamented, scores highest) |
| GAP-04 | HIGH | instr_melodic_intelligence.melodic_complexity :496, the structure block: 1 − abs(log10(step_leap) − log10(3.0))/1.3 | The step-to-leap term is inverted over the range real melodies occupy, and has a dead zone. Target 3.0 is unsupported by any corpus. Thirds (3–4 st) are counted as neither step nor leap despite being 17–30% of intervals. Result: the failure-mode control scores 0.904 and Ode to Joy scores 0.366. | §3.3 (term table); §5.1 (S:L column: real 6.0–20.0, control 4.0) |
| GAP-05 | HIGH | instr_melodic_intelligence.melodic_complexity :496 — the five-term variety half; no conventionality term anywhere in the file | The axis rewards the quantity the catchiness literature finds negative. Melodic entropy loads negatively on recognisability; repetitivity, motive typicality and range conventionality load positively; conventional global contour is the strongest earworm predictor. Nothing in our instrument measures typicality against any reference corpus. | §3.4 (dens.step.cont.glob.dir > 0.326 → ~80% INMI); §3.5 (van Balen Table 2); §2.4 (Schoenberg's popular touch) |
| GAP-06 | HIGH | Interface gap: pass3_grammar.Phrase.to_dict() :921-960 carries functions, cadence, start_bar, bars, derived_from, operation; instr_melodic_intelligence.measure() :1087 accepts only a card and an audio path | The scorer never sees the grammar's ground truth and re-derives phrases from audio heuristics. On a real generated period the grammar built 2 phrases (18 bars, evaded_then_PAC) and the scorer found 11, median 5 notes. Every phrase-derived number was then computed over those 11 fragments. | §5.5 (measured end-to-end); grep confirms no pass3_grammar import in instr_melodic_intelligence.py |
| GAP-07 | HIGH | instr_melodic_intelligence.segment_phrases :762; PHRASE_GAP_MULT = 1.55 :184; agogic branch requires durs[i] > 2.0 * med_dur | An entire scoring axis silently disappears on legato lines. All five reference melodies returned 1 phrase and phrase_sophistication: unavailable. The agogic branch fails on a knife edge: a cadential half-note against a quarter-note median is exactly 2.0×, and the test is strictly greater. | §5.4 (legato 1 phrase / detached 2 / with rests 4) |
| GAP-08 | MEDIUM | instr_melodic_intelligence.phrase_sophistication :821, tonic estimate :842, closed test :844 | Two defects in one line. (a) Only the tonic pitch class counts as closed, so an IAC — a closed cadence in Caplin and Schoenberg — reads as OPEN, and a standard HC→IAC period scores an antecedent-consequent rate of zero. (b) The most-frequent-pitch-class tonic estimator returned pitch class 4 on a grammar period whose true tonic was pitch class 2, so even PACs can read open. | §5.5 (cadence verdict table; measured tonic mismatch) |
| GAP-09 | MEDIUM | instr_melodic_intelligence.contour_depth :480; no shape classifier anywhere | Contour hierarchy is measured; contour SHAPE is not. Depth saturates: the control, Korobeiniki and Ode to Joy all score 4.00. Every published corpus result on contour uses the three-point convex/concave/ascending/descending reduction, which we cannot produce, so our contour numbers are not comparable to any of them. | §3.4 (Huron's reduction; Goldstein's four clusters); §5.1 (depth column) |
| GAP-10 | MEDIUM | instr_melodic_intelligence.py — no climax metric | Climax position and approach are unmeasured. Schoenberg's conditions 1–3 are all about the approach to and recession from a high point; the practitioner consensus places it at two-thirds to three-quarters. Londonderry Air measures 0.75; we have no way to notice that or its absence. | §2.1, §2.3, §3.6; §5.1 (peak@ column, computed ad hoc for this document only) |
| GAP-11 | MEDIUM | instr_melodic_intelligence.melodic_complexity :496 + predictive_info :420 | A sub-threshold line reports a measured-looking zero. predictive_info returns all-zero below 8 symbols; melodic_complexity requires 8 *pitches*, so an 8-note melody yields 7 intervals and publishes interval_entropy: 0.0 with available: true. Our control hit exactly this boundary. | §5.1 (control H1 = 0.00 with complexity 0.384 reported as available) |
| GAP-12 | MEDIUM | pass3_grammar.OPERATIONS :63-123, SPINE_CHANGING :123 | The rhythm half of developing variation is not expressible. spine_changing asks whether the interval spine moved. Schoenberg's stated rule is that *preserving the rhythm* is what licenses large contour change and buys coherence. We cannot currently distinguish developing variation (spine changed, rhythm preserved) from foreign material (both changed). | §2.2 (S1 ch. III and ch. VII) |
| GAP-13 | LOW–MEDIUM | pass3_grammar._cadence_gesture :960, goal map {"PAC": 0, "IAC": 2, "HC": 3, ...} with Mode.pitch 0-indexed (:159) | The half cadence's melodic goal is scale degree 4. PAC→1 and IAC→3 are correct. Degree 4 as the plain, longest goal tone over a dominant is the seventh of V — an unusual choice where 2, 5 or 7 is conventional. Flagged for verification rather than asserted as a defect; it may be a deliberate colour. | §2.5; Mode.pitch indexing confirmed by reading :159-163 |
| GAP-14 | LOW | Cue planning — build/audio/music_cue_table.json; hook.returns in the plan schema | Restatement count and player-motion are not planned quantities. Working practice gives five full statements per main theme (Star Wars) up to 14 restatements of a short signature pattern (Liberation), and a five-note fragment as the transportable unit; Kondo's first essential is riding the game's internal rhythm. Neither is a field we carry, so neither can be authored to or checked. *(The "five to six" figure previously here rested on a fabricated variation count — see §4.2.)* | §4.1, §4.2; §3.4 (tempo 124.10 vs 115.79 bpm) |
GAP-04 and GAP-05 point the same direction and are the reason this document exists. Our melodic
scorer, as it stands, **rewards the failure mode Josh named on two of its three structure terms and
on its whole variety half.** That is not a claim that the instrument is worthless — its development
census and its shuffle-corrected predictive information are genuinely good work, and the file's own
mutation controls are exemplary. It is a claim that the axis most likely to be used as a ranking
signal is, over the range where real melodies live, pointing the wrong way. Nothing else in this
document matters as much as checking that claim, and §3.3 and §5.1 are written to be re-run rather
than believed.
---
1. No commercial game theme was measured. §1.4 and §4.4 declare this. Everything numeric here
comes from published corpora of folk, classical, hymn and pop melodies, or from public-domain
melodies measured in §5. The Zelda / Final Fantasy / Dragon Quest side of the domain brief is
served qualitatively only.
2. The §5 encodings are working transcriptions, not critical editions (§5.3). The findings are
robust to a wrong note; the individual percentages are not.
3. Five books were not read in full (S3, S4, S5, S6, S7) and are cited only for claims sourced
through readable secondary material, as marked in §1.
4. The peak-placement consensus (§3.6) has no corpus behind it in this pass — only practitioner
teaching and the aggregate arch results. It is labelled consensus and R-05 deliberately reports
rather than gates on it.
5. The conventionality axis (R-04) has no reference corpus yet. It cannot be implemented as
specified until we have enough authored cues to compute second-order features against, and
implementing it against an external pop corpus would import an idiom this game does not want.
6. Web search budget was exhausted during this pass, which is why the singability/vocal-range
literature and a primary Sugiyama interview are both absent. Both are named gaps for a follow-up
pass, not silent omissions.
7. This document shipped with two fabricated claims and three misattributions, found by the
2026-08-08 audit and repaired in §11. They were all in §4 and in the source table, never in the
corpus numbers — but that is a fact about where this pass was weakest, not a reason to trust §4
now. §4 remains prose analysis and carries no verified number about a commercial game theme.
Every number attributed to a published source in §3 was read out of the source document itself —
the Jakubowski accepted manuscript, the Chiu & Temperley author PDF, the van Balen ISMIR PDF, the
Goldstein preprint, the Schoenberg OCR scan — not out of an abstract or a search summary, except
where §1 marks otherwise. Every number in §3.3, §5 and §8 was produced by importing and running the
actual functions in harness/music_gen/instr_melodic_intelligence.py and
harness/music_gen/pass3_grammar.py at the repository state of this pass, and every one of them
should be re-derived by the audit rather than trusted, because a same-commit edit can invalidate a
true proof.
---
An independent reality check was run over this document's source set. Every source below was fetched
and read at the level stated; PDFs were text-extracted locally rather than trusted to a summariser.
The audit's own bias was to hunt for invention, so the clean results are stated as plainly as the
defects.
All five melodic conditions in §2.1 are a single verbatim passage. The monotony/variation doctrine,
the motive-as-shape definition, the rapidity and intelligibility positions, the Mattheson 1739
quotation, the climax/recession passage, the liquidation definition, and the capitalised sentence
in §2.2 — all present verbatim, and the capitalisation claim is literally true in the source.
Chapter attributions checked by offset: §2.1 ch. IV ✓, §2.2 ch. III and ch. VII ✓, §2.3 ch. VI and
ch. VII ✓, §2.5 ch. VIII ✓. Two errors found and fixed — see §11.2.
0.326, ~80%, 0.421, 124.10/28.73 vs 115.79/25.39 bpm, 62.5%, 100 vs 100 tunes, the 14,063-melody
Geerdes reference corpus — verified in the accepted manuscript. The novelty claim about the tempo
finding and the Müllensiefen & Halpern implicit-memory restatement are both there verbatim.
ascending ✓; Huron's three-point reduction and his convex > descending > ascending > concave
ordering ✓; short melodies arching regardless of constituent phrase shape ✓; Tierney birdsong ✓.
cell-for-cell: 194/41,066/48%, 214/68,493/43%, 3,786/145,381/71%, 9,776/142,149/67%,
17,683/901,372/63%. Table 2 at step size 3 verified: 47/42/60/59. Step ≤ 2 semitones with pitch
repetitions removed ✓. The Huron descending-steps caveat ✓.
linear ballistic accumulation ✓; α = .005 ✓; every coefficient in §3.5 verified against Table 2;
the authors' reading of the symbolic side (entropy and productivity both negative; higher DF with
lower second-order productivity ⇒ more typical motives) is verbatim.
via the reference list of S12.
and hooks, design chapter with results deferred to its ch. 8.
2.3 / 2.1 / 2.0 semitones verified; the 5.5-semitone two-thirds figure is correctly scoped to the
most recent era (the paper gives ~7 and ~6 semitones for the two earlier eras); revolutions 1975
and 2000 tier-1, 1996 tier-2 ✓; declining pitch and rhythmic complexity ✓.
and to say what is attributed to them, with the exceptions listed below.
refusal to cite it as evidence stands.
| # | Where | Defect | Disposition |
|---|---|---|---|
| 1 | §4.2, R-10, GAP-14 | *The Dark Eye: Book of Heroes* "six variations" and a "roughly four-minute" main-theme track, attributed to Phillips | FABRICATION — struck. Pts 2 and 3 read in full; neither states either number. Replaced with the verified Liberation figure (14 restatements) |
| 2 | §4.1 | "the NES's three available timbres forced melodic and rhythmic ingenuity", attributed to S20 | FABRICATION — struck. Not in the abstract (the only part read), and S20 analyses an N64 title |
| 3 | §4.1 | Zelda overworld "consistently analysed as B♭-centred with a Mixolydian inflection" (S22) | REVERSED — corrected. S22 gives B♭ major and explicitly rejects the Mixolydian reading |
| 4 | §2.3, GAP-03 | Schoenberg's climax-in-bar-6 observation attributed to Beethoven Op. 2/2-IV | MISATTRIBUTED — corrected to Op. 2/1-II (Adagio). The observation itself is verbatim |
| 5 | §2.4 | "both from ch. IV" | MISATTRIBUTED — corrected to ch. V |
| 6 | S5 row | Margulis cited for "repetition drives attention to the sonic surface" | REVERSED — corrected. The cited review reports attention moving to *deeper* structures. No body claim depended on it |
| 7 | S9 row | "restated in Goldstein et al. (S13…)" | Internal cross-reference error — Goldstein is S10. Fixed |
| 8 | §3.4 | "83 FANTASTIC features" | Imprecise — 82 FANTASTIC features plus tempo = 83 predictors. Fixed |
| 9 | §3.5 | Coefficients from three different model fits presented as one regression; Melodic Range Conventionality given as a "0.05–0.07 [0.01, 0.13]" union of columns | Provenance column added; single-model value restored |
| 10 | §4.1 | *Super Mario Bros.* tempo increase offered as Kondo's own illustration | Not in either cited report; re-labelled as a general illustration |
| 11 | §3.7 | An Essen annotation definition and a convention S10 restates from prior literature both labelled EVIDENCE | Labels split |
No number in §3 sourced to a published corpus was wrong. Every defect above is in the citation
apparatus or in §4's game-practice prose — the qualitative layer the document already declared
weakest in §1.4 and §4.4. The quantitative backbone (§3.1, §3.2, §3.4, §3.5) survived intact, which
means the §6 rules that rest on it — R-01, R-02, R-03, R-04, R-05, R-06 — are unaffected. R-10 is
the only rule whose stated numbers changed, and it was already ranked eleventh of twelve.
The §5 and §8 numbers are out of scope for this audit: they are MEASURED HERE, reproducible from
§5.3, and §10 already instructs that they be re-derived rather than trusted. Nothing in this pass
either confirms or disturbs them.
Standing lesson for the sibling volumes. Both fabrications and one reversal share a single
signature: **a specific claim attributed to a source whose own read-status line says it was not
read.** The read-status column is the control, and it works — it named all three defects before the
fetches did. The rule it implies is that a source marked unread may support framing and may not
support a number, a key, or a hardware fact.