music/exemplars/instruments/MELODIC_INTELLIGENCE_VALIDATION.md
Tier: MEASURED INSTRUMENT. Produced by harness/music_gen/instr_melodic_intelligence.py, which writes this file and its machine record together and writes nothing else.
The law it measures: A melody must be complex without being random, must return DEVELOPED rather than restated, must be built of shaped phrases that answer one another, and must share the melodic lead with more than one voice across the cue.
The composition floor it serves: Josh's ROUND 2 GRADED ruling of 2026-08-07 night, at the tail of docs/spine/DECISIONS_PENDING_JOSH.md — "the melodies are way too simple. There needs to be a lot more complexity added and to put some intelligence into the melodies and how everything orchestrates ... do another pass through your grading criteria, and you probably need to add a few more axes." CANON SUBORDINATION: that ruling is the authority; this file is a measurement of it and changes no canon.
Machine record: build/audio/exemplars/instruments/MELODIC_INTELLIGENCE_VALIDATION.json.
Legal posture: analysis for understanding only; audio is read from disk, measured, and never copied, redistributed, or used as model input.
THESE AXES DO NOT YET SEPARATE THE CORPUS FROM ROUND 2, AND THE INSTRUMENT IS THEREFORE NOT YET REAL. The director's requirement was that the corpus hits score high and round 2 low. The best any of the five headline axes manages is orchestrational_dialogue at AUC 0.690, where 0.5 is indistinguishable; melodic_complexity reaches 0.621. Worse, 1 axis(es) point the WRONG WAY — counter_melody_presence 0.195, meaning round 2 scores HIGHER than the music Josh loves. An axis in that state cannot be used to grade round 3 no matter how well-motivated it is, and section 4 diagnoses why rather than leaving the null bare.
The separation question and the within-album prediction question are DIFFERENT questions and are answered separately below. Within-album prediction — can an axis tell a beloved track from competent filler on the same album — is priced against a multiplicity bar of 0.1962 (p95 deviation over 400 permutations), and nothing clears it. Four sibling instruments have returned honest nulls on that question and a fifth is not a disgrace — but it is also not what the director asked, and the two results must not be traded for one another.
THE HONESTY NUMBER, carried on every table below. Monophonic melodic-line extraction from a dense polyphonic mix is unreliable. On this file's own dense-mix control the same three planted voices read 4 clean and 4 under two full-band beds at 0.55, and melodic complexity moved from 0.641 to 0.574. Our own renders are dense on purpose, so that is the error bar these numbers carry, not a footnote.
Random notes maximise entropy and are not intelligent, so an entropy report would hand the composer the wrong instruction. The measure is the geometric mean of a VARIETY score (interval entropy, pitch-class alphabet, rhythmic-value entropy, contour depth, movement rate) and a STRUCTURE score built on PREDICTIVE INFORMATION, E = 2·H1 − H2 — the mutual information between one interval and the next. E is zero for an i.i.d. random line and zero for a constant one, and maximal for a line that is varied but constrained. Both entropies carry the Miller-Madow correction, and E additionally has the E of shuffles of the same interval multiset subtracted from it, because the plug-in bigram bias at these sample sizes points exactly at the trap.
| control | interval entropy H1 | structure term | melodic complexity |
|---|---|---|---|
| a developed tune | 3.734 | 0.697 | 0.812 |
| RANDOM notes | 5.315 | 0.058 | 0.241 |
| a scale run | 1.313 | — | 0.365 |
The random line's raw entropy is HIGHER than the developed tune's, which is the trap working as advertised; its structure term collapses and the geometric mean ranks it below. Its predictive information reads 3.502 bits before the shuffle correction and 0.000 bits after — the correction is not cosmetic. Two mutation arms hold the two halves of that claim: M1 turns the shuffle correction off and the random line's structure term inflates from 0.058 past 0.25, and M2 removes the structure gate entirely so that complexity is variety alone, at which point the random line OUT-SCORES the developed tune and the trap re-opens. The gate is what closes it.
The cell is located at the card's own hook onset — disagreeing with the card here would make two instruments answer to two different hooks — and every recurrence of its INTERVAL sequence is classified: exact repetition, octave transfer, real or tonal transposition, inversion, retrograde, augmentation and diminution by measured ratio, fragmentation of a proper sub-cell, and ornamented variation (the cell's duration-weighted skeleton inside a busier span). n_development_ops counts only the operations that CHANGE the material. That boundary is not this file's invention: it is the boundary round 2's own plans draw between literal / octave_up / carrier_swap and fragment / inversion / augmentation / diminution.
Phrases are opened at a rest — an inter-onset interval more than 1.55× the local median with a real gap after the note ends — or at an agogic accent followed by a contour reset. The method's error is measured, not asserted: against four planted boundaries the segmenter returns F1 1.000 (precision 1.000, recall 1.000). It CANNOT see a boundary that carries neither a rest nor an agogic accent — a purely harmonic cadence inside an unbroken texture is invisible to it, and that is a real class of music this instrument declines to score rather than guessing at. The four terms are phrase-length variety (a tune of nothing but four-bar units is the simplicity Josh named), antecedent-consequent pairing, cadential differentiation as the entropy of phrase-ending pitch classes, and whether phrase lengths group at more than one level.
Not a layer count. instr_layer_census already measures spectral streams and its own validation records that it saturates near 8 where a real cue runs 24 to 40 parts; a pad, a drone and a shaker are three layers and none is a voice. A registral band counts as a melodic VOICE only if it is pitched, moves, moves at a melodic rate, and is not a DOUBLING of a voice already counted — doubling being the cheapest possible way to fake counter-melody. On the controls one voice reads 1, three independent voices read 4, and the same line doubled at the octave reads 1. The three-voice control reading 4 is an OVER-count of one — a synthesised tone's upper partials are loud enough in the band above to pass the movement gate — and it is printed rather than tuned away, because tuning the gate until the control read exactly three would be fitting the instrument to its own control. It is otherwise a LOWER bound: two instruments in one register are one band.
In each 4-second window the melodic lead is the band with the most energy-weighted pitch movement; a hand-off is a change of lead that HOLDS for 2 windows. Without the dwell a lead that alternates every window reads as constant conversation (mutation arm M6, caught by the flicker control). The stated resolution limit that follows: an antiphonal exchange faster than about eight seconds reads as texture to this instrument. The score multiplies the hand-off rate by the entropy of how the lead is shared, because forty alternations between two bands is an ostinato with two colours, not a conversation. On the controls a static lead reads 0 hand-offs and a line handed between three registers reads 2.
Corpus hits are the tracks Josh named as loved for decades. Round 2 is the three cues he turned down last night. AUC is the probability that a randomly chosen loved track scores above a randomly chosen round-2 cue: 0.5 is indistinguishable, 1.0 is perfect separation, below 0.5 means round 2 scores HIGHER.
| axis | hit median | CACI | REST | WATERFALL | round-1 median | AUC hits over round 2 |
|---|---|---|---|---|---|---|
| MELODIC COMPLEXITY | 0.505 | 0.447 | 0.334 | 0.611 | 0.467 | 0.621 |
| interval entropy H1 (bits) | 3.664 | 4.549 | 3.308 | 4.180 | 3.906 | 0.356 |
| interval bigram entropy H2 (bits) | 5.286 | 6.844 | 5.495 | 6.155 | 6.233 | 0.333 |
| contour depth | 7.000 | 7.000 | 6.000 | 7.000 | 7.250 | 0.529 |
| MOTIVIC DEVELOPMENT (operations) | 1.000 | 0 | 2 | 1 | 2.500 | 0.565 |
| development share of returns | 0.936 | 0.000 | 1.000 | 1.000 | 1.000 | 0.405 |
| PHRASE SOPHISTICATION | 0.735 | 0.754 | 0.735 | 0.730 | 0.723 | 0.425 |
| phrase count | 29.000 | 33 | 34 | 7 | 38.500 | 0.534 |
| phrase length variety (CV) | 0.783 | 0.809 | 0.833 | 0.764 | 0.837 | 0.425 |
| COUNTER-MELODY PRESENCE | 0.597 | 0.574 | 0.757 | 0.750 | 0.665 | 0.195 |
| melodically active voices | 4.000 | 3 | 4 | 5 | 4.000 | 0.408 |
| ORCHESTRATIONAL DIALOGUE | 0.329 | 0.000 | 0.290 | 0.280 | 0.185 | 0.690 |
| lead hand-offs | 4.000 | 0 | 6 | 4 | 3.000 | 0.563 |
Read the bottom row of that table against the top. Where an AUC sits below 0.5 the round Josh rejected scores HIGHER on that axis, and an axis in that state cannot be used to grade round 3 no matter how well-motivated it is.
One row of that table is worth reading on its own: raw interval entropy points the WRONG WAY. H1 reaches AUC 0.356 and H2 0.333 for the loved tracks over round 2 — which is to say the cues Josh called too simple carry MORE interval entropy in their extracted lines than the music he has loved for thirty years, and round 1 carries more still. Had this instrument reported entropy as complexity it would have told the composer that round 2 was already more complex than Chrono Trigger and that the fix was to make it simpler. That is the entropy trap, caught on the real pools rather than on a synthetic control, and it is the strongest evidence in this file that the variety-times-structure design was necessary rather than ornamental.
Round 2 is the only pool in this program where we hold the SCORE as well as the render, which makes it a better calibration than any synthetic control: a real dense mix whose melodic content is known exactly. The plans' own census, and what the audio arm recovered from the renders without being told:
| cue | cell | plan: real ops | measured ops | plan: literal | plan: counterlines | measured voices | measured phrases |
|---|---|---|---|---|---|---|---|
| CACI | 6 notes / 2 bars | 3 | 0 | 3 | 6 | 3 | 33 |
| REST | 6 notes / 2 bars | 3 | 2 | 2 | 2 | 4 | 34 |
| WATERFALL | 6 notes / 2 bars | 1 | 1 | 1 | 0 | 5 | 7 |
The gap between those columns IS this instrument's measured error on real dense audio, and stating it that way is worth more than a synthetic error bar. The plan census is also the sharpest statement of what round 2 actually is: the entire melodic content of every cue is one six-note, two-bar cell. Every layer whose role is carrier, statement, carrier_return or carrier_close plays that same cell again with a transform label attached. Across four minutes and thirty-two layers, PASS2_FLORES_REST contains three genuine development operations. There is no antecedent, no consequent, no continuation, and no phrase longer than two bars anywhere in any of the three cues. A motif was composed; a melody was not.
The counter-melody axis has a free three-point calibration in that table: the plans carry CACI 6, REST 2, WATERFALL 0 hand-written counterlines, so the measured ordering should be CACI > REST > WATERFALL. It measured REST > WATERFALL > CACI — the ordering is NOT recovered, and that disagreement is the axis's error on real audio, reported rather than reconciled.
Every axis here is computed off an extracted melodic line, and the line is cleaner on a sparse mix than on a dense one. If an axis correlates strongly with HOW MUCH LINE THE EXTRACTOR FOUND, then it is partly a measurement of the arrangement's density rather than of the writing. This is the same class of defect instr_noise_structure found in its own headline axis and withdrew rather than published — its peak-stability separation of AUC 0.813 turned out to be reading codec provenance — so the check is run here and published whatever it says.
| axis | Spearman vs voiced fraction | Spearman vs note count | reading |
|---|---|---|---|
| MELODIC COMPLEXITY | -0.042 | 0.126 | clean |
| interval entropy H1 (bits) | 0.433 | 0.357 | materially confounded |
| interval bigram entropy H2 (bits) | 0.491 | 0.493 | materially confounded |
| contour depth | 0.308 | 0.140 | clean |
| MOTIVIC DEVELOPMENT (operations) | 0.112 | 0.397 | materially confounded |
| development share of returns | 0.207 | 0.331 | clean |
| PHRASE SOPHISTICATION | 0.370 | 0.467 | materially confounded |
| phrase count | 0.467 | 0.928 | dominated by extraction density |
| phrase length variety (CV) | -0.170 | 0.124 | clean |
| COUNTER-MELODY PRESENCE | 0.602 | 0.354 | materially confounded |
| melodically active voices | 0.474 | 0.271 | materially confounded |
| ORCHESTRATIONAL DIALOGUE | -0.050 | -0.097 | clean |
| lead hand-offs | 0.029 | 0.195 | clean |
This table is the finding. melodic_complexity is CLEAN — it is essentially uncorrelated with how much line the extractor found, so its weak separation is a weak result and not an artefact. phrase_count is not an axis at all: at a Spearman near 0.93 against note count it is a re-expression of extraction density wearing a musical name, and it is reported below only as a diagnostic. n_melodic_voices rising with voiced fraction is the direct explanation of the backwards counter-melody ordering above: PASS2_FLORES_WATERFALL is the SPARSEST of the three cues, so the extractor hears its registers most cleanly, so it reads the most voices — while PASS2_FLORES_CACI, which actually carries six hand-written counterlines, is the densest and reads the fewest. The axis is measuring how easy the mix is to hear into, not how many voices are thinking.
And the phrase axis has a second, independent falsification in section 4. Round 2's score contains no phrase longer than two bars anywhere in any of the three cues, and the segmenter reports dozens of phrases in each and a sophistication equal to the corpus median. It is segmenting the gaps in a predominant-pitch track that hops between instruments, not the phrasing of a melody. That is a defect in the segmenter, reported as one, and it is why phrase_sophistication must not be used to grade round 3 in its current form.
| id | title | class | complexity | H1 | ops | phrases | soph. | voices | c-melody | hand-offs |
|---|---|---|---|---|---|---|---|---|---|---|
| EX_006 | Chrono Trigger Main Theme | TITLE_CHARACTER | 0.532 | 2.193 | 26 | 33 | 0.741 | 4 | 0.703 | 1 |
| EX_007 | Frog's Theme | TITLE_CHARACTER | 0.343 | 3.503 | n/a | 3 | 0.458 | 4 | 0.579 | 3 |
| EX_014 | Tekken 2 Attract Movie: Sound Track | TITLE_CHARACTER | n/a | n/a | n/a | n/a | n/a | n/a | n/a | n/a |
| EX_017 | Sweden | EXPLORATION | 0.526 | 5.136 | 0 | 42 | 0.735 | 4 | 0.675 | 7 |
| EX_018 | Subwoofer Lullaby | EXPLORATION | 0.486 | 4.444 | 1 | 31 | 0.732 | 4 | 0.747 | 3 |
| EX_022 | Wind Scene | EXPLORATION | 0.523 | 3.998 | 22 | 64 | 0.770 | 4 | 0.572 | 5 |
| EX_030 | Route 209 | EXPLORATION | 0.364 | 4.424 | 0 | 17 | 0.750 | 4 | 0.631 | 4 |
| EX_032 | Terra's Theme | EXPLORATION | 0.515 | 3.958 | 25 | 67 | 0.743 | 3 | 0.449 | 6 |
| EX_033 | Corridors of Time | EXPLORATION | 0.584 | 2.857 | 4 | 70 | 0.689 | 5 | 0.661 | 1 |
| EX_046 | BFG Division | BATTLE | 0.428 | 1.935 | 11 | 24 | 0.606 | 0 | 0.000 | 4 |
| EX_053 | Cynthia Battle Theme | BATTLE | 0.616 | 2.760 | 0 | 4 | 0.688 | 2 | 0.228 | 2 |
| EX_055 | The Cyber Grind | BATTLE | 0.475 | 3.623 | 0 | 7 | 0.653 | 0 | 0.000 | 3 |
| EX_064 | Ludwig, the Holy Blade | BOSS | 0.524 | 3.172 | 0 | 14 | 0.750 | 3 | 0.610 | 7 |
| EX_075 | Dancing Mad | BOSS | 0.514 | 3.330 | 1 | 20 | 0.654 | 4 | 0.566 | 11 |
| EX_082 | The Best Is Yet to Come | MELANCHOLIC | 0.493 | 3.410 | 39 | 60 | 0.754 | 5 | 0.651 | 0 |
| EX_088 | An Unwavering Heart | MELANCHOLIC | 0.562 | 4.637 | 4 | 26 | 0.740 | 5 | 0.662 | 2 |
| EX_092 | Snake Eater (Ladder Climb) | TENSION | 0.391 | 4.541 | 1 | 32 | 0.753 | 5 | 0.631 | 8 |
| EX_094 | Dark Creature Pursuit | TENSION | 0.532 | 3.713 | 0 | 4 | 0.601 | 2 | 0.184 | 5 |
| EX_102 | Magus Castle | TENSION | 0.373 | 3.851 | 0 | 10 | 0.679 | 5 | 0.691 | 1 |
| EX_110 | Want You Gone | CREDITS_TRIUMPH | 0.456 | 4.802 | 2 | 40 | 0.778 | 4 | 0.597 | 2 |
| EX_113 | To Good Friends | CREDITS_TRIUMPH | 0.407 | 5.143 | 0 | 41 | 0.749 | 5 | 0.659 | 5 |
| EX_124 | Deus Ex Main Title (UNATCO) | TITLE_CHARACTER | 0.505 | 3.069 | 20 | 27 | 0.738 | 3 | 0.482 | 1 |
| EX_136 | Humming the Bassline | EXPLORATION | 0.656 | 2.806 | 31 | 52 | 0.621 | 1 | 0.000 | 1 |
| EX_142 | Shinshu Field | EXPLORATION | 0.606 | 2.354 | 0 | 5 | 0.544 | 3 | 0.466 | 12 |
| EX_145 | Go Straight | BATTLE | 0.649 | 2.737 | 0 | 11 | 0.584 | 1 | 0.000 | 1 |
| EX_147 | Fighting of the Spirit | BATTLE | 0.481 | 3.664 | 0 | 7 | 0.550 | 3 | 0.590 | 8 |
| EX_161 | Aria di Mezzo Carattere | MELANCHOLIC | 0.481 | 3.834 | 3 | 29 | 0.759 | 4 | 0.551 | 4 |
| EX_162 | Schala's Theme | MELANCHOLIC | 0.371 | 2.977 | 4 | 65 | 0.703 | 4 | 0.609 | 8 |
| EX_179 | Ending Theme (Balance is Restored) | CREDITS_TRIUMPH | 0.450 | 4.502 | 0 | 38 | 0.763 | 5 | 0.640 | 5 |
| EX_180 | Dreams Dreams | CREDITS_TRIUMPH | 0.560 | 3.996 | 1 | 38 | 0.763 | 3 | 0.705 | 0 |
14 of 30 loved tracks read ZERO development operations. That is the lesson instr_hook_presence recorded about its own return threshold, inherited here and reported rather than hidden: a threshold calibrated on synthetic controls does not transfer to real recordings, and a count of zero on a track that plainly develops its material is the extractor's failure, not the composer's.
| measured here | |
|---|---|
| formula cards | 1180 |
| corpus HITS | 30 (measured on audio: 30) |
| album siblings available | 927 |
| siblings measured on audio | 132 |
| albums carrying a hit | 18 |
| round-2 cues | 3 |
| round-1 cues | 32 |
Sibling rule: per album: sort by track_id, shuffle(seed=20260807), take 8 — the same rule instr_hook_presence uses, so the two instruments' audio subsets are the same tracks and their AUCs are comparable. Analysis cap: the first 180 s of every track, stated because a development operation after that point is invisible to this instrument.
| axis | within-album AUC | deviation | hits used | clears the bar |
|---|---|---|---|---|
| MELODIC COMPLEXITY | 0.5823 | 0.0823 | 27 | no |
| interval entropy H1 (bits) | 0.5283 | 0.0283 | 27 | no |
| interval bigram entropy H2 (bits) | 0.5017 | 0.0017 | 27 | no |
| contour depth | 0.5328 | 0.0328 | 27 | no |
| MOTIVIC DEVELOPMENT (operations) | 0.5015 | 0.0015 | 26 | no |
| development share of returns | 0.4308 | 0.0692 | 26 | no |
| PHRASE SOPHISTICATION | 0.5228 | 0.0228 | 27 | no |
| phrase count | 0.6259 | 0.1259 | 27 | no |
| phrase length variety (CV) | 0.4861 | 0.0139 | 27 | no |
| COUNTER-MELODY PRESENCE | 0.5797 | 0.0797 | 27 | no |
| melodically active voices | 0.4743 | 0.0257 | 27 | no |
| ORCHESTRATIONAL DIALOGUE | 0.5714 | 0.0714 | 27 | no |
| lead hand-offs | 0.5546 | 0.0546 | 27 | no |
The bar is the permutation multiplicity null: within each album, re-draw which tracks are 'hits' at random, recompute every axis's within-album AUC, keep the best deviation of the whole bank, repeat 400 times. p95 = 0.1962, p99 = 0.2191. This is the identical apparatus that priced PATTERN_FINDINGS_V3, so these axes are directly comparable to it.
| control | planted | recovered |
|---|---|---|
| C0 the pYIN-free line | the rig's own melodic control | 25 notes against the rig's 25 (agreement 1.000) |
| C1 the entropy trap | random must rank below developed | 0.241 < 0.812 |
| C2 the scale run | maximally predictable, must score low | 0.365 |
| C3 three identical statements | zero development | 0 operations |
| C4 a developed tune | six named operations | 5 recovered: augmentation, diminution, fragmentation, inversion, retrograde |
| C5 planted inversion | at 6.0 s | [(5.979, 0.0)] |
| C6 planted augmentation | at 13.0 s, ratio 2.0 | [(12.98, 2.0)] |
| C7 planted fragmentation | at 24.0 s | [(23.975, 4)] |
| C8 planted diminution | at 37.0 s, ratio 0.5 | [(36.978, 0.5)] |
| C9 one voice vs three | 1 vs 3 | 1 vs 4 |
| C10 octave doubling | must read as ONE voice | 1, folded: [{'band': 'tenor', 'doubles': 'bass', 'onset_coincidence': 1.0, 'contour_r': 1.0}, {'band': 'alto', 'doubles': 'bass', 'onset_coincidence': 1.0, 'contour_r': 1.0}] |
| C11 phrase boundaries | [3.15, 6.3, 10.35] | [3.135, 6.281, 10.333], F1 1.000 |
| C12 the dense mix | 3 voices under two beds | 4 recovered |
| C13 static lead vs handed | 0 vs >0 hand-offs | 0 vs 2 |
| C14 the four-second flicker | a two-colour alternation is texture | 0 hand-offs |
| C15 a degenerate card | available=False, never raises | card has no 'card' block |
| mutation arm | must be caught by | fired |
|---|---|---|
| M1 no shuffle correction | C1b the structure collapse on a random line | yes |
| M2 structure gate removed (complexity = variety alone) | C1a the ranking of random below developed | yes |
| M3 inversion removed | C5 the planted inversion | yes |
| M4 time-scale classification removed | C6 the planted augmentation | yes |
| M5 doubling test removed | C10 the octave doubling | yes |
| M6 dwell hysteresis removed | C14 the four-second flicker | yes |
| M7 phrase splitting removed | C11 the planted phrase boundaries | yes |
A control without a mutation proof was never armed. Each arm above breaks one axis on purpose and the named control has to notice.