MELODIC_INTELLIGENCE_VALIDATION.md

music/exemplars/instruments/MELODIC_INTELLIGENCE_VALIDATION.md

MELODIC INTELLIGENCE — the validation of instrument (e)

Tier: MEASURED INSTRUMENT. Produced by harness/music_gen/instr_melodic_intelligence.py, which writes this file and its machine record together and writes nothing else.
The law it measures: A melody must be complex without being random, must return DEVELOPED rather than restated, must be built of shaped phrases that answer one another, and must share the melodic lead with more than one voice across the cue.
The composition floor it serves: Josh's ROUND 2 GRADED ruling of 2026-08-07 night, at the tail of docs/spine/DECISIONS_PENDING_JOSH.md — "the melodies are way too simple. There needs to be a lot more complexity added and to put some intelligence into the melodies and how everything orchestrates ... do another pass through your grading criteria, and you probably need to add a few more axes." CANON SUBORDINATION: that ruling is the authority; this file is a measurement of it and changes no canon.
Machine record: build/audio/exemplars/instruments/MELODIC_INTELLIGENCE_VALIDATION.json.
Legal posture: analysis for understanding only; audio is read from disk, measured, and never copied, redistributed, or used as model input.

1. The one-paragraph answer

THESE AXES DO NOT YET SEPARATE THE CORPUS FROM ROUND 2, AND THE INSTRUMENT IS THEREFORE NOT YET REAL. The director's requirement was that the corpus hits score high and round 2 low. The best any of the five headline axes manages is orchestrational_dialogue at AUC 0.690, where 0.5 is indistinguishable; melodic_complexity reaches 0.621. Worse, 1 axis(es) point the WRONG WAY — counter_melody_presence 0.195, meaning round 2 scores HIGHER than the music Josh loves. An axis in that state cannot be used to grade round 3 no matter how well-motivated it is, and section 4 diagnoses why rather than leaving the null bare.

The separation question and the within-album prediction question are DIFFERENT questions and are answered separately below. Within-album prediction — can an axis tell a beloved track from competent filler on the same album — is priced against a multiplicity bar of 0.1962 (p95 deviation over 400 permutations), and nothing clears it. Four sibling instruments have returned honest nulls on that question and a fifth is not a disgrace — but it is also not what the director asked, and the two results must not be traded for one another.

THE HONESTY NUMBER, carried on every table below. Monophonic melodic-line extraction from a dense polyphonic mix is unreliable. On this file's own dense-mix control the same three planted voices read 4 clean and 4 under two full-band beds at 0.55, and melodic complexity moved from 0.641 to 0.574. Our own renders are dense on purpose, so that is the error bar these numbers carry, not a footnote.

2. The five axes, and why each is defined the way it is

Melodic complexity, and the entropy trap

Random notes maximise entropy and are not intelligent, so an entropy report would hand the composer the wrong instruction. The measure is the geometric mean of a VARIETY score (interval entropy, pitch-class alphabet, rhythmic-value entropy, contour depth, movement rate) and a STRUCTURE score built on PREDICTIVE INFORMATION, E = 2·H1 − H2 — the mutual information between one interval and the next. E is zero for an i.i.d. random line and zero for a constant one, and maximal for a line that is varied but constrained. Both entropies carry the Miller-Madow correction, and E additionally has the E of shuffles of the same interval multiset subtracted from it, because the plug-in bigram bias at these sample sizes points exactly at the trap.

controlinterval entropy H1structure termmelodic complexity
a developed tune3.7340.6970.812
RANDOM notes5.3150.0580.241
a scale run1.3130.365

The random line's raw entropy is HIGHER than the developed tune's, which is the trap working as advertised; its structure term collapses and the geometric mean ranks it below. Its predictive information reads 3.502 bits before the shuffle correction and 0.000 bits after — the correction is not cosmetic. Two mutation arms hold the two halves of that claim: M1 turns the shuffle correction off and the random line's structure term inflates from 0.058 past 0.25, and M2 removes the structure gate entirely so that complexity is variety alone, at which point the random line OUT-SCORES the developed tune and the trap re-opens. The gate is what closes it.

Motivic development

The cell is located at the card's own hook onset — disagreeing with the card here would make two instruments answer to two different hooks — and every recurrence of its INTERVAL sequence is classified: exact repetition, octave transfer, real or tonal transposition, inversion, retrograde, augmentation and diminution by measured ratio, fragmentation of a proper sub-cell, and ornamented variation (the cell's duration-weighted skeleton inside a busier span). n_development_ops counts only the operations that CHANGE the material. That boundary is not this file's invention: it is the boundary round 2's own plans draw between literal / octave_up / carrier_swap and fragment / inversion / augmentation / diminution.

Phrase sophistication

Phrases are opened at a rest — an inter-onset interval more than 1.55× the local median with a real gap after the note ends — or at an agogic accent followed by a contour reset. The method's error is measured, not asserted: against four planted boundaries the segmenter returns F1 1.000 (precision 1.000, recall 1.000). It CANNOT see a boundary that carries neither a rest nor an agogic accent — a purely harmonic cadence inside an unbroken texture is invisible to it, and that is a real class of music this instrument declines to score rather than guessing at. The four terms are phrase-length variety (a tune of nothing but four-bar units is the simplicity Josh named), antecedent-consequent pairing, cadential differentiation as the entropy of phrase-ending pitch classes, and whether phrase lengths group at more than one level.

Counter-melody presence

Not a layer count. instr_layer_census already measures spectral streams and its own validation records that it saturates near 8 where a real cue runs 24 to 40 parts; a pad, a drone and a shaker are three layers and none is a voice. A registral band counts as a melodic VOICE only if it is pitched, moves, moves at a melodic rate, and is not a DOUBLING of a voice already counted — doubling being the cheapest possible way to fake counter-melody. On the controls one voice reads 1, three independent voices read 4, and the same line doubled at the octave reads 1. The three-voice control reading 4 is an OVER-count of one — a synthesised tone's upper partials are loud enough in the band above to pass the movement gate — and it is printed rather than tuned away, because tuning the gate until the control read exactly three would be fitting the instrument to its own control. It is otherwise a LOWER bound: two instruments in one register are one band.

Orchestrational dialogue

In each 4-second window the melodic lead is the band with the most energy-weighted pitch movement; a hand-off is a change of lead that HOLDS for 2 windows. Without the dwell a lead that alternates every window reads as constant conversation (mutation arm M6, caught by the flicker control). The stated resolution limit that follows: an antiphonal exchange faster than about eight seconds reads as texture to this instrument. The score multiplies the hand-off rate by the entropy of how the lead is shared, because forty alternations between two bands is an ostinato with two colours, not a conversation. On the controls a static lead reads 0 hand-offs and a line handed between three registers reads 2.

3. THE HEADLINE TABLE

Corpus hits are the tracks Josh named as loved for decades. Round 2 is the three cues he turned down last night. AUC is the probability that a randomly chosen loved track scores above a randomly chosen round-2 cue: 0.5 is indistinguishable, 1.0 is perfect separation, below 0.5 means round 2 scores HIGHER.

axishit medianCACIRESTWATERFALLround-1 medianAUC hits over round 2
MELODIC COMPLEXITY0.5050.4470.3340.6110.4670.621
interval entropy H1 (bits)3.6644.5493.3084.1803.9060.356
interval bigram entropy H2 (bits)5.2866.8445.4956.1556.2330.333
contour depth7.0007.0006.0007.0007.2500.529
MOTIVIC DEVELOPMENT (operations)1.0000212.5000.565
development share of returns0.9360.0001.0001.0001.0000.405
PHRASE SOPHISTICATION0.7350.7540.7350.7300.7230.425
phrase count29.0003334738.5000.534
phrase length variety (CV)0.7830.8090.8330.7640.8370.425
COUNTER-MELODY PRESENCE0.5970.5740.7570.7500.6650.195
melodically active voices4.0003454.0000.408
ORCHESTRATIONAL DIALOGUE0.3290.0000.2900.2800.1850.690
lead hand-offs4.0000643.0000.563

Read the bottom row of that table against the top. Where an AUC sits below 0.5 the round Josh rejected scores HIGHER on that axis, and an axis in that state cannot be used to grade round 3 no matter how well-motivated it is.

One row of that table is worth reading on its own: raw interval entropy points the WRONG WAY. H1 reaches AUC 0.356 and H2 0.333 for the loved tracks over round 2 — which is to say the cues Josh called too simple carry MORE interval entropy in their extracted lines than the music he has loved for thirty years, and round 1 carries more still. Had this instrument reported entropy as complexity it would have told the composer that round 2 was already more complex than Chrono Trigger and that the fix was to make it simpler. That is the entropy trap, caught on the real pools rather than on a synthetic control, and it is the strongest evidence in this file that the variety-times-structure design was necessary rather than ornamental.

4. Round 2 against its own score — the known-answer fixture

Round 2 is the only pool in this program where we hold the SCORE as well as the render, which makes it a better calibration than any synthetic control: a real dense mix whose melodic content is known exactly. The plans' own census, and what the audio arm recovered from the renders without being told:

cuecellplan: real opsmeasured opsplan: literalplan: counterlinesmeasured voicesmeasured phrases
CACI6 notes / 2 bars3036333
REST6 notes / 2 bars3222434
WATERFALL6 notes / 2 bars111057

The gap between those columns IS this instrument's measured error on real dense audio, and stating it that way is worth more than a synthetic error bar. The plan census is also the sharpest statement of what round 2 actually is: the entire melodic content of every cue is one six-note, two-bar cell. Every layer whose role is carrier, statement, carrier_return or carrier_close plays that same cell again with a transform label attached. Across four minutes and thirty-two layers, PASS2_FLORES_REST contains three genuine development operations. There is no antecedent, no consequent, no continuation, and no phrase longer than two bars anywhere in any of the three cues. A motif was composed; a melody was not.

The counter-melody axis has a free three-point calibration in that table: the plans carry CACI 6, REST 2, WATERFALL 0 hand-written counterlines, so the measured ordering should be CACI > REST > WATERFALL. It measured REST > WATERFALL > CACI — the ordering is NOT recovered, and that disagreement is the axis's error on real audio, reported rather than reconciled.

The diagnosis: how much of each axis is the MIX rather than the WRITING

Every axis here is computed off an extracted melodic line, and the line is cleaner on a sparse mix than on a dense one. If an axis correlates strongly with HOW MUCH LINE THE EXTRACTOR FOUND, then it is partly a measurement of the arrangement's density rather than of the writing. This is the same class of defect instr_noise_structure found in its own headline axis and withdrew rather than published — its peak-stability separation of AUC 0.813 turned out to be reading codec provenance — so the check is run here and published whatever it says.

axisSpearman vs voiced fractionSpearman vs note countreading
MELODIC COMPLEXITY-0.0420.126clean
interval entropy H1 (bits)0.4330.357materially confounded
interval bigram entropy H2 (bits)0.4910.493materially confounded
contour depth0.3080.140clean
MOTIVIC DEVELOPMENT (operations)0.1120.397materially confounded
development share of returns0.2070.331clean
PHRASE SOPHISTICATION0.3700.467materially confounded
phrase count0.4670.928dominated by extraction density
phrase length variety (CV)-0.1700.124clean
COUNTER-MELODY PRESENCE0.6020.354materially confounded
melodically active voices0.4740.271materially confounded
ORCHESTRATIONAL DIALOGUE-0.050-0.097clean
lead hand-offs0.0290.195clean

This table is the finding. melodic_complexity is CLEAN — it is essentially uncorrelated with how much line the extractor found, so its weak separation is a weak result and not an artefact. phrase_count is not an axis at all: at a Spearman near 0.93 against note count it is a re-expression of extraction density wearing a musical name, and it is reported below only as a diagnostic. n_melodic_voices rising with voiced fraction is the direct explanation of the backwards counter-melody ordering above: PASS2_FLORES_WATERFALL is the SPARSEST of the three cues, so the extractor hears its registers most cleanly, so it reads the most voices — while PASS2_FLORES_CACI, which actually carries six hand-written counterlines, is the densest and reads the fewest. The axis is measuring how easy the mix is to hear into, not how many voices are thinking.

And the phrase axis has a second, independent falsification in section 4. Round 2's score contains no phrase longer than two bars anywhere in any of the three cues, and the segmenter reports dozens of phrases in each and a sophistication equal to the corpus median. It is segmenting the gaps in a predominant-pitch track that hops between instruments, not the phrasing of a melody. That is a defect in the segmenter, reported as one, and it is why phrase_sophistication must not be used to grade round 3 in its current form.

5. The corpus, track by track

idtitleclasscomplexityH1opsphrasessoph.voicesc-melodyhand-offs
EX_006Chrono Trigger Main ThemeTITLE_CHARACTER0.5322.19326330.74140.7031
EX_007Frog's ThemeTITLE_CHARACTER0.3433.503n/a30.45840.5793
EX_014Tekken 2 Attract Movie: Sound TrackTITLE_CHARACTERn/an/an/an/an/an/an/an/a
EX_017SwedenEXPLORATION0.5265.1360420.73540.6757
EX_018Subwoofer LullabyEXPLORATION0.4864.4441310.73240.7473
EX_022Wind SceneEXPLORATION0.5233.99822640.77040.5725
EX_030Route 209EXPLORATION0.3644.4240170.75040.6314
EX_032Terra's ThemeEXPLORATION0.5153.95825670.74330.4496
EX_033Corridors of TimeEXPLORATION0.5842.8574700.68950.6611
EX_046BFG DivisionBATTLE0.4281.93511240.60600.0004
EX_053Cynthia Battle ThemeBATTLE0.6162.760040.68820.2282
EX_055The Cyber GrindBATTLE0.4753.623070.65300.0003
EX_064Ludwig, the Holy BladeBOSS0.5243.1720140.75030.6107
EX_075Dancing MadBOSS0.5143.3301200.65440.56611
EX_082The Best Is Yet to ComeMELANCHOLIC0.4933.41039600.75450.6510
EX_088An Unwavering HeartMELANCHOLIC0.5624.6374260.74050.6622
EX_092Snake Eater (Ladder Climb)TENSION0.3914.5411320.75350.6318
EX_094Dark Creature PursuitTENSION0.5323.713040.60120.1845
EX_102Magus CastleTENSION0.3733.8510100.67950.6911
EX_110Want You GoneCREDITS_TRIUMPH0.4564.8022400.77840.5972
EX_113To Good FriendsCREDITS_TRIUMPH0.4075.1430410.74950.6595
EX_124Deus Ex Main Title (UNATCO)TITLE_CHARACTER0.5053.06920270.73830.4821
EX_136Humming the BasslineEXPLORATION0.6562.80631520.62110.0001
EX_142Shinshu FieldEXPLORATION0.6062.354050.54430.46612
EX_145Go StraightBATTLE0.6492.7370110.58410.0001
EX_147Fighting of the SpiritBATTLE0.4813.664070.55030.5908
EX_161Aria di Mezzo CarattereMELANCHOLIC0.4813.8343290.75940.5514
EX_162Schala's ThemeMELANCHOLIC0.3712.9774650.70340.6098
EX_179Ending Theme (Balance is Restored)CREDITS_TRIUMPH0.4504.5020380.76350.6405
EX_180Dreams DreamsCREDITS_TRIUMPH0.5603.9961380.76330.7050

14 of 30 loved tracks read ZERO development operations. That is the lesson instr_hook_presence recorded about its own return threshold, inherited here and reported rather than hidden: a threshold calibrated on synthetic controls does not transfer to real recordings, and a count of zero on a track that plainly develops its material is the extractor's failure, not the composer's.

6. The pool, read at HEAD

measured here
formula cards1180
corpus HITS30 (measured on audio: 30)
album siblings available927
siblings measured on audio132
albums carrying a hit18
round-2 cues3
round-1 cues32

Sibling rule: per album: sort by track_id, shuffle(seed=20260807), take 8 — the same rule instr_hook_presence uses, so the two instruments' audio subsets are the same tracks and their AUCs are comparable. Analysis cap: the first 180 s of every track, stated because a development operation after that point is invisible to this instrument.

7. Within-album prediction, and the multiplicity price

axiswithin-album AUCdeviationhits usedclears the bar
MELODIC COMPLEXITY0.58230.082327no
interval entropy H1 (bits)0.52830.028327no
interval bigram entropy H2 (bits)0.50170.001727no
contour depth0.53280.032827no
MOTIVIC DEVELOPMENT (operations)0.50150.001526no
development share of returns0.43080.069226no
PHRASE SOPHISTICATION0.52280.022827no
phrase count0.62590.125927no
phrase length variety (CV)0.48610.013927no
COUNTER-MELODY PRESENCE0.57970.079727no
melodically active voices0.47430.025727no
ORCHESTRATIONAL DIALOGUE0.57140.071427no
lead hand-offs0.55460.054627no

The bar is the permutation multiplicity null: within each album, re-draw which tracks are 'hits' at random, recompute every axis's within-album AUC, keep the best deviation of the whole bank, repeat 400 times. p95 = 0.1962, p99 = 0.2191. This is the identical apparatus that priced PATTERN_FINDINGS_V3, so these axes are directly comparable to it.

8. The controls, and which mutation caught which

controlplantedrecovered
C0 the pYIN-free linethe rig's own melodic control25 notes against the rig's 25 (agreement 1.000)
C1 the entropy traprandom must rank below developed0.241 < 0.812
C2 the scale runmaximally predictable, must score low0.365
C3 three identical statementszero development0 operations
C4 a developed tunesix named operations5 recovered: augmentation, diminution, fragmentation, inversion, retrograde
C5 planted inversionat 6.0 s[(5.979, 0.0)]
C6 planted augmentationat 13.0 s, ratio 2.0[(12.98, 2.0)]
C7 planted fragmentationat 24.0 s[(23.975, 4)]
C8 planted diminutionat 37.0 s, ratio 0.5[(36.978, 0.5)]
C9 one voice vs three1 vs 31 vs 4
C10 octave doublingmust read as ONE voice1, folded: [{'band': 'tenor', 'doubles': 'bass', 'onset_coincidence': 1.0, 'contour_r': 1.0}, {'band': 'alto', 'doubles': 'bass', 'onset_coincidence': 1.0, 'contour_r': 1.0}]
C11 phrase boundaries[3.15, 6.3, 10.35][3.135, 6.281, 10.333], F1 1.000
C12 the dense mix3 voices under two beds4 recovered
C13 static lead vs handed0 vs >0 hand-offs0 vs 2
C14 the four-second flickera two-colour alternation is texture0 hand-offs
C15 a degenerate cardavailable=False, never raisescard has no 'card' block
mutation armmust be caught byfired
M1 no shuffle correctionC1b the structure collapse on a random lineyes
M2 structure gate removed (complexity = variety alone)C1a the ranking of random below developedyes
M3 inversion removedC5 the planted inversionyes
M4 time-scale classification removedC6 the planted augmentationyes
M5 doubling test removedC10 the octave doublingyes
M6 dwell hysteresis removedC14 the four-second flickeryes
M7 phrase splitting removedC11 the planted phrase boundariesyes

A control without a mutation proof was never armed. Each arm above breaks one axis on purpose and the named control has to notice.

9. What this instrument refuses to claim

Generated by harness/site/structure_site.py — the URL path is the repo path. review root