music/exemplars/PATTERN_FINDINGS_V2.md
SUPERSEDED BYPATTERN_FINDINGS_V3.md(2026-08-07). TheEX_014andEX_094quarantines
were retired by inspection and both rows re-attached, moving the pool from 28 hits / 851 controls
to 30 hits / 927 controls; two byte-identical UNMATCHED__ twins were retired in the same
sitting. At that denominator four axes clear the bar rather than six, none clears the 99th, and
the joint envelope changes two of its three axes — so §8's box, which the round-1 generation
filter reads, is the number to stop quoting first. V2 is not withdrawn: its method is
unchanged and it remains the correct answer for the pool it had, exactly as it says of V1.
Tier: MEASURED-SAMPLE FINDINGS, with six axes now clearing a calibrated multiplicity bar.
Everything below is derived from the formula cards already on disk. Nothing was purchased,
nothing was generated, no audio was decoded by this pass. The cards are read as data; the
exemplar bytes never move.
Supersedes: PATTERN_FINDINGS_V1.md (N = 11, zero axes clearing). V1 is not withdrawn —
its method is unchanged and its null is still the correct answer for the pool it had.
Produced by: harness/music_gen/derive_patterns.py (22 controls, mutation-proved).
Machine record: build/audio/exemplars/PATTERN_FINDINGS_V2.json.
Consumer: docs/proposals/music/MASTERPIECE_PROGRAM.md rung 2 (the archetype synthesis
toward MASTERPIECE_STANDARD) and rung 3 (the nostalgia predictor, whose floor this pass sets).
Legal posture: analysis for understanding only, unchanged from the cards themselves.
One thousand one hundred and eighty-one tracks are measured. Twenty-eight of them are corpus
rows — the tracks Josh named, the ones that stayed loved for decades — against eleven in V1. Eight
hundred and fifty-one are their own album siblings: same composer, same session, same production,
same codec. The design has not changed and neither has the feature bank. What changed is that the
multiplicity bar, which is measured rather than assumed, fell from 0.298 to 0.195 as the
positives went from eleven to twenty-eight, and six axes now stand above it where zero stood
above it before. The leave-one-out composite — axes and directions chosen fold by fold without ever
seeing the held-out track — moved from 0.586 ± 0.085 to 0.744 ± 0.045, twenty-two of
twenty-six corpus rows above their own album's median, binomial p = 0.0005. Held-out single rules
went from firing on 27.3% of unseen corpus rows (against a 43.9% sibling rate, worse than nothing)
to 67.9% against 35.9%. On the frozen train/test split, zero of six rules reverse sign,
where three of six reversed at N = 11. That is a real, replicated, out-of-sample signal, and it is
made of one thing: **structural size — how many distinct sections a track has, how long it runs,
how many bars, how much it repeats itself, how many event drops it contains.** The three
hypotheses V1 carried forward all failed at the larger N, and the length-independent half of
the bank collapsed to exactly chance (0.497 ± 0.046). The honest headline is therefore narrower
than the AUC suggests: *the measured card can pick a corpus row out of its own album at 0.74, and
everything it uses to do so is a measure of structural scale.*
Each corpus row is scored only against the tracks that shipped on its own record — composer,
scoring session, mixing chain, sample library, release year and audio codec held constant by
construction. Per-hit win rates are averaged with equal weight per hit, so a 208-track Pokémon
compilation cannot drown a four-track Bloodborne EP. This is the only framing the pool can answer
honestly, and it is V1's framing verbatim.
| V1 | V2 | note | |
|---|---|---|---|
| Formula cards measured | 914 | 1181 | formula_cards/*.json; repudiated/ excluded by construction |
| HITS — corpus rows, ACQUIRED or CARDED | 11 | 28 | the labelled positives |
| CONTROLS — siblings in an album containing a hit | 601 | 851 | across 15 albums with siblings |
| Measured but in an album with no hit | 301 | 301 | carried, never used as controls |
| Quarantined | 1 | 1 | EX_014 |
| Feature axes | 103 | 103 | unchanged bank, so V1 and V2 are directly comparable |
| Albums carrying a hit | 10 | 16 | one of them (Metal Gear Solid Vocal Tracks) contributes no siblings |
EX_014 is dropped from both sides, not merely from the positives, on the corpus
defect_register's standing instruction: the row says Tekken 2 "Opening Theme", the measured bytes
are a Pokémon X & Y opening movie, and a known-false label is no better as a negative than as a
positive.
Two hits sit on an album with no siblings at all — EX_082 and EX_092 are both corpus rows on the
four-track Metal Gear Solid Vocal Tracks compilation — so the within-album statistics are computed
on 26 of the 28. Every AUC below carries that denominator; the envelope and descriptive blocks,
which do not need siblings, use all 28.
| id | class | title | album measured against | siblings |
|---|---|---|---|---|
| EX_006 | TITLE_CHARACTER | Chrono Trigger Main Theme | Chrono Trigger (Original Soundtrack) [DS] | 71 |
| EX_007 | TITLE_CHARACTER | Frog's Theme | Chrono Trigger | 71 |
| EX_017 | EXPLORATION | Sweden | C418 — Minecraft Volume Alpha | 22 |
| EX_018 | EXPLORATION | Subwoofer Lullaby | Minecraft Volume Alpha | 22 |
| EX_022 | EXPLORATION | Wind Scene | Chrono Trigger | 71 |
| EX_030 | EXPLORATION | Route 209 | Pokémon Diamond & Pearl Super Music Collection | 145 |
| EX_032 | EXPLORATION | Terra's Theme | FINAL FANTASY VI | 57 |
| EX_033 | EXPLORATION | Corridors of Time | Chrono Trigger | 71 |
| EX_046 | BATTLE | BFG Division | DOOM (Original Game Soundtrack) | 30 |
| EX_053 | BATTLE | Cynthia Battle Theme | Pokémon Diamond & Pearl | 145 |
| EX_055 | BATTLE | The Cyber Grind | ULTRAKILL | 22 |
| EX_064 | BOSS | Ludwig, the Holy Blade | Bloodborne: The Old Hunters | 4 |
| EX_075 | BOSS | Dancing Mad | FINAL FANTASY VI | 57 |
| EX_082 | MELANCHOLIC | The Best Is Yet to Come | Metal Gear Solid Vocal Tracks | 0 |
| EX_088 | MELANCHOLIC | An Unwavering Heart | Pokémon X & Y Super Music Collection | 208 |
| EX_092 | TENSION | Snake Eater (Ladder Climb) | Metal Gear Solid Vocal Tracks | 0 |
| EX_102 | TENSION | Magus Castle | Chrono Trigger | 71 |
| EX_110 | CREDITS_TRIUMPH | Want You Gone | Portal 2 | 63 |
| EX_113 | CREDITS_TRIUMPH | To Good Friends | Chrono Trigger | 71 |
| EX_124 | TITLE_CHARACTER | Deus Ex Main Title (UNATCO) | Deus Ex | 39 |
| EX_136 | EXPLORATION | Humming the Bassline | Jet Set Radio Future | 16 |
| EX_142 | EXPLORATION | Shinshu Field | Ōkami OST Vol. 1 | 54 |
| EX_145 | BATTLE | Go Straight | Streets of Rage 2 (Remastered) | 24 |
| EX_147 | BATTLE | Fighting of the Spirit | Tales of Phantasia | 76 |
| EX_161 | MELANCHOLIC | Aria di Mezzo Carattere | FINAL FANTASY VI | 57 |
| EX_162 | MELANCHOLIC | Schala's Theme | Chrono Trigger | 71 |
| EX_179 | CREDITS_TRIUMPH | Ending Theme (Balance is Restored) | FINAL FANTASY VI | 57 |
| EX_180 | CREDITS_TRIUMPH | Dreams Dreams | NiGHTS into Dreams | 20 |
All seven purpose classes are now represented — EXPLORATION 8, BATTLE 5, MELANCHOLIC 4,
CREDITS_TRIUMPH 4, TITLE_CHARACTER 3, BOSS 2, TENSION 2 — where V1 had six of seven and four of
them on one or two tracks. No per-class claim in this document is a test; the per-class block in
the JSON remains a descriptive read and is labelled as one. Five classes still sit below the
five-row floor the MASTERPIECE_STANDARD's per-class bands need.
The bar is not guessed and it is not inherited from V1. Within each album, re-draw at random which
tracks are "hits" (keeping each album's count), recompute every feature's within-album AUC, and
keep the best deviation from 0.5 across the whole bank. Four hundred draws, same seed
discipline, same procedure as V1 — re-run from scratch on the new pool.
| statistic of best-of-bank \ | AUC − 0.5\ | V1 (N = 11) | V2 (N = 28) | V2 as AUC | |
|---|---|---|---|---|---|
| median of chance | 0.226 | 0.151 | 0.651 | ||
| 90th percentile | 0.280 | 0.185 | 0.685 | ||
| 95th percentile — the bar | 0.298 | 0.195 | 0.695 | ||
| 99th percentile | 0.330 | 0.220 | 0.720 | ||
| max over 400 draws | — | 0.240 | 0.740 | ||
| observed best | 0.242 (n_drops) | 0.250 (section_count) | 0.750 |
Features clearing the 95th-percentile bar: six of 103. Clearing the 99th: two. At N = 11 it was
zero and zero. The observed best barely moved (0.242 → 0.250); what moved is the null underneath
it, which is exactly the effect V1 §8 predicted and the reason its power table held the bar
variable rather than fixed. This is the single most important structural fact about the re-run: the
finding was bought by the bar falling, not by the effect growing.
| axis | within-album AUC | length-matched | ρ vs duration | hits above own-album median | binomial p | clears bar |
|---|---|---|---|---|---|---|
section_count | 0.750 | 0.677 | +0.78 | 22/26 | 0.0005 | yes (p99) |
total_length_s | 0.732 | 0.671 | +1.00 | 22/26 | 0.0005 | yes (p99) |
bars | 0.718 | 0.639 | +0.94 | 19/26 | 0.029 | yes (p95) |
repeat_sim_mean | 0.714 | 0.644 | +0.71 | 22/26 | 0.0005 | yes (p95) |
n_drops | 0.710 | 0.637 | +0.57 | 23/26 | 0.0001 | yes (p95) |
novelty_mean | 0.304 | 0.361 | −0.59 | 7/26 | 0.029 | yes (p95) |
note_count | 0.673 | 0.567 | +0.79 | 19/26 | 0.029 | no |
tempo_adjust_n | 0.670 | 0.608 | +0.70 | 20/26 | 0.009 | no |
n_raises | 0.669 | 0.595 | +0.56 | 18/26 | 0.076 | no |
breaks_per_min | 0.343 | 0.383 | −0.89 | 6/26 | 0.009 | no |
time_to_hook_frac | 0.344 | 0.384 | −0.56 | 7/26 | 0.029 | no |
silence_frac | 0.357 | 0.418 | −0.76 | 7/26 | 0.029 | no |
melody_span_semitones | 0.634 | 0.569 | +0.33 | 18/26 | 0.076 | no |
repeat_frac | 0.632 | 0.581 | +0.54 | 17/26 | 0.169 | no |
dynamic_range_db | 0.369 | 0.419 | −0.56 | 8/26 | 0.076 | no |
novelty_mean clears the bar in the negative direction and is the one non-size axis in the
group: it is the mean self-similarity novelty of the track's own structure, and corpus rows sit
*below* their siblings on it. Read together with repeat_sim_mean sitting *above*, the pair says
the same thing twice — a corpus row is more self-similar section-to-section than the album around
it. That is not "repetitive"; at these lengths and section counts it is *thematic return*, a track
that keeps coming back to its own material instead of moving on.
Tempo is still not the confound. Rescoring against siblings within ±20% bpm changes essentially
nothing (section_count 0.750 → 0.749, total_length_s 0.732 → 0.733, n_drops 0.710 → 0.717).
The pattern is not "battle themes are fast and battle themes are over-represented", and that was
true at N = 11 as well.
V1's §4 verdict was blunt and correct for its pool: *the top of the table is duration wearing a
costume*, and the duration effect is curation — Josh named finished themes while game albums are
full of jingles, stingers and menu blips. That remains true. The control pool still carries 161
tracks under 30 seconds and 267 under a minute, against a hit pool with exactly one row under a
minute (EX_102, Magus Castle, 28.6 s).
So the V1 rescoring control was run again, and a harder version of it was added: instead of only
matching each hit to siblings within ±35% of its own duration, the control pool is truncated from
below, admitting only siblings that are themselves finished-length tracks.
| axis | all 851 controls | ≥ 30 s (690) | ≥ 60 s (584) | ≥ 90 s (451) |
|---|---|---|---|---|
section_count | 0.750 | 0.723 | 0.686 | 0.652 |
n_drops | 0.710 | 0.690 | 0.664 | 0.645 |
repeat_sim_mean | 0.714 | 0.686 | 0.649 | 0.621 |
novelty_mean | 0.304 | 0.342 | 0.360 | 0.377 |
bars | 0.718 | 0.683 | 0.652 | 0.609 |
total_length_s | 0.732 | 0.701 | 0.663 | 0.581 |
note_count | 0.673 | 0.641 | 0.619 | 0.587 |
tempo_adjust_n | 0.670 | 0.645 | 0.618 | 0.586 |
Read the last column. Raw duration is the axis that dies — 0.732 down to 0.581, which is what
"curation, not composition" looks like when you take the stingers away. section_count, n_drops,
repeat_sim_mean and novelty_mean all decay much more slowly and are still around 0.62–0.65
against siblings that are every bit as long as the hits. That residual is the part of the effect
that is not length, and it is the part worth carrying into rung 2.
This truncation is a robustness read, not a test. The permutation bar in §3 was calibrated on
the full control pool; a truncated pool has fewer and differently-distributed siblings, so its own
bar would sit somewhat higher and is not computed here. Nothing in this table is claimed to clear
anything.
V1 §12 named three things to re-test, all of which had survived its length control. All three are
re-tested here on the same axes with no re-definition, and none survives.
| carried hypothesis | axis | V1 AUC | V1 length-matched | V2 AUC | V2 rank | verdict |
|---|---|---|---|---|---|---|
| Delayed self-recurrence — the theme waits longer before returning | recurrence_lag_norm | 0.662 | 0.703 | 0.505 | 102 / 103 | failed — dead at chance |
| Unbroken contour runs — the melody climbs or falls further before turning | contour_run_max | 0.672 | 0.602 | 0.536 | 66 / 103 | failed |
| Composed rest — silence budget plus drop rate | X_rest_grammar | 0.531 | 0.659 | 0.424 | 32 / 103 | failed, and reversed |
recurrence_lag_norm is the most instructive: it was V1's best surviving candidate, the one that
*rose* under length matching, and at N = 28 it is the second-worst axis in the entire bank. The
descriptive medians show why — hit median 0.033 against a sibling median 0.111, which is the
opposite sign to V1's reading. Seventeen new corpus rows were enough to erase and invert it.
This is the clearest vindication of V1's own honesty apparatus. V1 reported these three as *not
significant*, priced against a bar none of them cleared, and recommended them only as things to
re-test. Had they been promoted to MASTERPIECE_STANDARD bands on the strength of "survives the
length control", three false rules would now be in the generator.
**FINDING 2026-08-07 — THE RECURRENCE AXIS WAS MEASURING THE WRONG THING, AND CLOSING THAT DID NOT
RESCUE IT.** recurrence_lag_norm above, and the motif_economy and repeat_map blocks it comes
from, run on raw chroma and MFCC self-similarity. That means **a theme restated a fourth higher, at
half speed, or ornamented scores as NEW MATERIAL rather than as a return** — which is exactly
backwards for a leitmotif score, where the transformed return is the whole craft. So the axis failing
at 0.505 was never a clean test of the hypothesis: it was a test of literal, in-key, in-tempo
repetition, which is the one kind of return good music uses least.
That gap is now closed. harness/music_gen/instr_hook_presence.py matches over all twelve
pitch-class rotations (recording the winning rotation, so the INTERVAL of the return is itself
reported), over a seven-point time-ratio grid, and symbolically under an edit budget. **The
hypothesis fails again anyway, and harder: zero of its fourteen metrics clear the bar re-derived on
the live pool** (best deviation 0.175 against an instrument-only bar of 0.245 and a joint bar of
0.241, where the joint bar asks the harsher and more honest question of whether it beats everything
already measured), with n_returns at 0.510 and returns_per_min at 0.500 — chance to three
decimals. Against the eighteen candidates Josh rejected it is worse than useless: those score HIGHER
than the loved tracks on n_returns (3.0 against 1), transformation_share (1.000 against 0.893)
and hook_strength (0.736 against 0.691). Round 1 states hooks and brings them back. So this
row's verdict stands, now for a better reason: it is a second, invariance-corrected null on the same
family of ideas, not an artefact of a blunt ruler.
Two cautions travel with that, both from the instrument's own face. Its return threshold was
calibrated on synthetic controls and does not transfer to real recordings — nine of the thirty loved
tracks read zero returns, and a sensitivity sweep shows one of them going from 0 to 86 returns as the
threshold falls — so every ABSOLUTE return count it publishes is too low and the floor derived from
them is degenerate. The separation null survives that sweep at every threshold tested, because
loosening lifts the rejected candidates as much as it lifts the hits. And it cannot choose the motif:
it follows the card's recorded hook onset, so a track whose memorable idea is not its opening idea is
searched for the wrong phrase with perfect rigour.
| hypothesis | axis | V1 AUC | V2 AUC | V2 length-matched | verdict |
|---|---|---|---|---|---|
| Hook arrives inside the first phrase and repeats its own interval shape | X_hook_immediacy_economy | 0.531 | 0.508 | 0.461 | not separated |
| Singable = mostly steps, leaps answered against, compass narrow | X_singability | 0.575 | 0.455 | 0.448 | not separated |
| Arrangement keeps lanes in reserve, spends them at seams | X_reserve_and_spend | 0.505 | 0.444 | 0.455 | not separated |
| Rest is composed — silence budget plus drop rate | X_rest_grammar | 0.531 | 0.424 | 0.509 | not separated |
Their own four-feature null puts the bar at deviation 0.146 (AUC 0.646), down from 0.223 at N = 11.
None of the four comes near it, and three of the four moved *away* from separation as the pool
grew. **The four theories written into the extractor before the first run are, at 28 positives,
measurably not what distinguishes these tracks.**
Fifteen hits train, thirteen test. The six best training axes were turned into thresholded rules by
Youden's J on the training data, then scored once on the untouched test side.
| rule fitted on train | train AUC | test AUC | test TPR | test FPR | test lift | precision lift |
|---|---|---|---|---|---|---|
section_count ≥ 15 | 0.726 | 0.783 | 0.62 | 0.160 | 3.85× | 3.63× |
total_length_s ≥ 140.6 | 0.749 | 0.710 | 0.69 | 0.235 | 2.95× | 2.83× |
n_drops ≥ 17 | 0.720 | 0.696 | 0.62 | 0.185 | 3.32× | 3.16× |
bars ≥ 110 | 0.751 | 0.672 | 0.46 | 0.122 | 3.80× | 3.58× |
note_count ≥ 268 | 0.725 | 0.602 | 0.54 | 0.158 | 3.40× | 3.24× |
tempo_adjust_n ≥ 11 | 0.726 | 0.594 | 0.62 | 0.228 | 2.70× | 2.60× |
Zero of six rules reverse sign between train and test. At N = 11, three of six reversed and two
of those had the strongest training AUCs — the signature of selection noise. Here every rule holds
its direction and every one carries a 2.7–3.9× lift on the held-out side. This is the single
largest qualitative change from V1 and it is not a change of method: the same code, the same split
procedure, the same Youden thresholds.
The strongest axis is re-chosen from scratch on the other twenty-seven hits, thresholded on them,
and the untouched twenty-eighth is asked whether it fires.
| V1 (N = 11) | V2 (N = 28) | |
|---|---|---|
| held-out corpus rows firing | 3 of 11 — 27.3% | 19 of 28 — 67.9% |
| mean sibling fire rate for the same rules | 43.9% | 35.9% |
| ratio | 0.62× — worse than nothing | 1.89× |
At N = 11 the selected rule fired on held-out corpus rows *less often* than on the album filler it
was built to exclude. At N = 28 it fires nearly twice as often on the unseen corpus row. The rule
selected in almost every fold is section_count ≥ 14.
In each fold the top five axes and their directions are chosen on the other twenty-seven hits
alone, z-scored within album, summed, and the held-out track's percentile among its own siblings is
recorded.
| composite | V1 held-out AUC | V2 held-out AUC | SE | hits above own-album median | binomial p |
|---|---|---|---|---|---|
| over all 103 axes | 0.561 ± 0.071 | 0.744 | 0.045 | 22 / 26 | 0.0005 |
| over the length-independent axes | 0.586 ± 0.085 | 0.497 | 0.046 | 15 / 26 | 0.557 |
The length-independent bank is defined with no reference to the hit labels — the correlation
with duration is measured on the control pool alone (|ρ| < 0.35), which admits 70 of 103 axes here.
V1's leak (filtering the hit-ranked top-25 instead, worth about +0.06 AUC) stayed closed; the same
control-pool filter is used.
0.744 ± 0.045 is the honest realized predictive power of the whole measured card on this pool,
and 0.497 ± 0.046 is what is left of it once every axis correlated with duration is removed.
Both numbers must be quoted together. The composite works, and the composite is made of size.
Separation ("which side of a threshold") and containment ("inside which box") are different
questions, and the second is the one a generator can use even when the first is weak. Each envelope
is built from twenty-seven hits and validated by asking whether it contains the twenty-eighth.
| axis | hit envelope | held-out hits contained | filler admitted | lift |
|---|---|---|---|---|
melody_span_semitones | 19.7 – 57.0 st | 27 / 28 | 69.5% | 1.39× |
section_mean_s | 7.1 – 63.3 s | 26 / 28 | 70.0% | 1.33× |
X_rest_grammar | 2.68 – 17.43 | 26 / 28 | 70.0% | 1.33× |
section_count | 4 – 35 | 27 / 28 | 83.4% | 1.16× |
bars | 18 – 820 | 26 / 28 | 80.4% | 1.16× |
total_length_s | 28.6 – 1295.6 s | 26 / 28 | 81.3% | 1.14× |
The joint box over melody_span_semitones × section_mean_s × X_rest_grammar contains
25 of 28 held-out hits (89.3%) while admitting 44.8% of filler — a 1.99× enrichment.
V1's joint box contained 7 of 11 (63.6%) at 30.3% filler for 2.1×.
The enrichment is essentially unchanged; the recall is much better, which is the direction that
matters for a rejection filter. A box that rejects a third of real corpus rows rejects good
generated candidates at the same rate. Note that the per-axis envelopes are wider than V1's because
twenty-eight tracks span more than eleven do, and every individual axis's lift fell accordingly —
the box is doing its work jointly, not axis by axis. It remains a **necessary-not-sufficient
constraint**: a generated track outside the box is unlike every measured corpus row; a generated
track inside it has cleared a bar that almost half of album filler also clears.
X_rest_grammar earning a place in the joint box while failing badly as a separator is not a
contradiction — it is precisely the separation-versus-containment distinction. Corpus rows occupy a
*narrow band* of composed rest without sitting *high or low* on it.
V1 §8's sharpest structural complaint was that **the nostalgia lineage the whole thirty-year bar
rests on contributed zero measured tracks**: no SNES, no N64, no Zelda, no Kondo, twelve of fifteen
albums released in 2000 or later. The STARTER TEN closed half of that. Four 16-bit-era scores now
carry corpus rows — Chrono Trigger (1995, SNES), FINAL FANTASY VI (1994, SNES), Tales of Phantasia
(1995, Super Famicom) and Streets of Rage 2 (1992, Mega Drive) — **thirteen of the twenty-eight
hits and 228 of the 851 siblings.** The pool is now roughly half 16-bit era and half modern, which
makes the era question testable for the first time.
| axis | all 28 | 16-bit era (13 hits) | modern (13 hits used) |
|---|---|---|---|
section_count | 0.750 | 0.762 | 0.739 |
total_length_s | 0.732 | 0.720 | 0.745 |
bars | 0.718 | 0.739 | 0.697 |
repeat_sim_mean | 0.714 | 0.692 | 0.735 |
n_drops | 0.710 | 0.663 | 0.757 |
novelty_mean | 0.304 | 0.358 | 0.250 |
note_count | 0.673 | 0.705 | 0.640 |
tempo_adjust_n | 0.670 | 0.639 | 0.701 |
| — | |||
recurrence_lag_norm | 0.505 | 0.376 | 0.634 |
contour_run_max | 0.536 | 0.417 | 0.656 |
X_rest_grammar | 0.424 | 0.328 | 0.519 |
repeat_sim_max | 0.593 | 0.500 | 0.687 |
Two readings, and they point opposite ways.
same sign and roughly the same magnitude in both halves — section_count at 0.762 against 0.739,
bars at 0.739 against 0.697. A structural-size effect that reproduces independently in 1994
Super Famicom sample-ROM scores and in 2016 live-recorded soundtracks is not an artifact of
production era, budget or codec. This is the strongest single result in the re-run, and it is
the one V1 could not have obtained at any N, because it had only one era.
above 0.5 in the modern half at close to their V1 values (0.634, 0.656, 0.519) and *below* 0.5 in
the 16-bit half (0.376, 0.417, 0.328). Pooled, they cancel to chance. V1's pool was entirely
modern, so what V1 measured was a modern-era tendency and reported, correctly, as not
significant. The lineage that just arrived contradicts it.
The same split applied to the leave-one-out composite folds — the same percentiles the §7c headline
is built from, no re-selection — gives a mean held-out percentile of 0.726 for the 16-bit rows
against 0.763 for the modern rows on the all-axes composite: **the composite works about equally
well in both eras.** On the length-independent composite it is 0.599 against 0.395, two halves
pulling against each other around a pooled 0.497, which is a warning about that composite rather
than a finding about either era.
The lineage is still only half covered. Zero N64, zero Zelda, zero Kondo, zero Nintendo
first-party of any kind. The Kondo/Zelda/N64 gap is the same gap V1 named and it is still open; what
closed is the Square/Enix 16-bit JRPG and the Sega 16-bit action lineage.
Not discriminative findings unless marked as clearing the bar in §4. These are the measured profile
of the twenty-eight, usable as sanity bands for authored candidates.
| axis | hit median | hit range | sibling median | sibling p10–p90 | AUC |
|---|---|---|---|---|---|
| section count | 17 | 4 – 35 | 9 | 3 – 18 | 0.750 |
| length | 196.6 s | 28.6 – 1295.6 | 96.2 s | 13.6 – 223.7 | 0.732 |
| bars | 114 | 18 – 820 | 52 | 8 – 139 | 0.718 |
| mean section self-similarity | 0.902 | 0.580 – 0.952 | 0.855 | 0.496 – 0.921 | 0.714 |
| drops | 19.5 | 3 – 110 | 6 | 3 – 25 | 0.710 |
| structural novelty | 0.556 | 0.397 – 0.857 | 0.678 | 0.507 – 0.976 | 0.304 |
| note count | 316 | 13 – 2128 | 121 | 8 – 371 | 0.673 |
| tempo adjustments | 13.5 | 1 – 89 | 6 | 1 – 21 | 0.670 |
| melodic span | 43.0 st | 19.7 – 57.0 | 31.0 st | 9.0 – 48.0 | 0.634 |
| hook range | 22.0 st | 0.67 – 52.0 | 17.0 st | 5.0 – 38.0 | 0.589 |
| mean section length | 5.36 bars | 2.9 – 31.0 | 4.29 bars | 2.3 – 10.4 | 0.546 |
| leap fraction (≥5 st) | 0.400 | 0.00 – 0.87 | 0.400 | 0.07 – 0.80 | 0.546 |
| interval 3-gram repetition | 0.432 | 0.00 – 0.83 | 0.409 | 0.00 – 0.73 | 0.521 |
| gap-fill after leap | 0.000 | 0.00 – 1.00 | 0.000 | 0.00 – 1.00 | 0.533 |
| time to hook | 0.575 s | 0.04 – 61.6 | 0.615 s | 0.26 – 5.94 | 0.483 |
| melodic salience | 0.797 | 0.00 – 0.924 | 0.756 | 0.352 – 0.909 | 0.487 |
| stepwise interval fraction | 0.108 | 0.00 – 0.33 | 0.133 | 0.00 – 0.40 | 0.476 |
| contour turns per note | 0.367 | 0.00 – 0.667 | 0.400 | 0.133 – 0.667 | 0.470 |
| windowed key agreement | 0.411 | 0.109 – 0.929 | 0.500 | 0.212 – 1.000 | 0.460 |
| tempo | 120.3 bpm | 83.4 – 172.3 | 123.1 bpm | 99.4 – 152.0 | 0.441 |
| dynamic range | 11.7 dB | 4.4 – 50.7 | 13.7 dB | 6.6 – 69.8 | 0.369 |
| silence budget | 2.1% | 0.6 – 11.1 | 2.8% | 1.0 – 19.8 | 0.357 |
Four readings worth carrying into rung 2, each with its own confidence:
of hummability is built from — stepwise fraction, gap-fill after a leap, hook compass, contour
turn rate, interval 3-gram repetition, time to hook, melodic salience — still sits between AUC
0.46 and 0.59 with 2.5× the positives, and most of them moved *toward* 0.5, not away from it.
The most economical explanation is unchanged: the album siblings already have them. These are
well-written game scores throughout, and the corpus row is not the only competent track on its
record. What the card can see that separates them is size and return, not melodic craft. That
remains the most important sentence in this document for the program.
the section count, double the length, double the bars, triple the drop count, above its album on
self-similarity and below it on novelty. Six axes, one story, all six clearing a calibrated bar.
Part of it is curation (§5), and about 0.62–0.65 of it survives against equally long siblings.
N = 11). The effect weakened as the pool grew and is now nearly nothing. Do not carry it.
max_lanes is 9 for hits and siblings alike, at 28 positives as at 11. That is a measurement ceiling in the spectral_band_proxy method, not a finding, and any arrangement-width claim from
these cards is void until a real instrument roster exists. Two full re-runs have now confirmed
the ceiling; the cards' own declaration that the lane method is not a roster is due.
The model is re-fitted on the new pool: K_eff, the number of effectively independent axes this
bank behaves like, is chosen so the simulated bar at N = 28 reproduces the observed permutation
95th percentile, and only then extrapolated. Fitted K_eff = 103 (calibration error 0.007),
against 80 at N = 11 — the bank behaves like more independent axes now that the null is estimated
from more positives, which raises the bar slightly relative to a naive extrapolation. Real sibling
counts are used.
| carded corpus rows | multiplicity bar (AUC) | detection of a true 0.65 | of 0.70 | of 0.75 | of 0.80 |
|---|---|---|---|---|---|
| 11 | 0.797 | 3.0% | 12.0% | 29.5% | 53.5% |
| 16 | 0.759 | 6.0% | 19.0% | 47.0% | 78.0% |
| 22 | 0.719 | 12.2% | 34.0% | 74.8% | 95.0% |
| 28 (today) | 0.695 (observed) | — | — | — | — |
| 30 | 0.690 | 22.3% | 57.3% | 91.2% | 99.8% |
| 40 | 0.665 | 38.3% | 80.7% | 97.3% | 100% |
| 55 | 0.639 | 62.0% | 96.3% | 100% | 100% |
| 75 | 0.619 | 86.8% | 99.3% | 100% | 100% |
| 100 | 0.608 | 94.8% | 100% | 100% | 100% |
| 140 | 0.591 | 99.3% | 100% | 100% | 100% |
Rows needed for 80% power: about 24 for an effect of 0.75, 40 for 0.70, about 70 for 0.65
(interpolating the table — 0.75 reads 74.8% at 22 and 91.2% at 30; 0.65 reads 62.0% at 55 and
86.8% at 75). The pool has just passed the 0.75 threshold, which is why six axes appeared. The model is
generous by construction — a constant true effect on one axis, albums like the ones already
measured, no extra heterogeneity from a broader corpus — so read these as lower bounds.
What the next acquisitions are worth, in priority order:
zero Zelda, zero Kondo, zero Nintendo first-party. §9 shows the era split is measurable and that
it already overturned three hypotheses; a third era would be the strongest available test of
whether the six surviving axes are really era-invariant or merely invariant across the two eras
now held. This is the highest-value buy in the corpus and it is the same one V1 named.
TITLE_CHARACTER 3, MELANCHOLIC 4, CREDITS_TRIUMPH 4). The MASTERPIECE_STANDARD is specified in
per-class bands and no per-class band can be stated until each class carries five-plus rows.
BOSS and TENSION are the binding constraint.
are silently excluded from every within-album statistic here. Buying the Metal Gear Solid and
Metal Gear Solid 3 score albums would recover two positives already paid for.
multiplicity bar; the 4-sibling Bloodborne EP quantises one hit's AUC to fifths.
This is a measured-sample derivation, not a validated theory, and nothing here is yet a
MASTERPIECE_STANDARD band.
real and out-of-sample; it is also narrow. Nothing here says a long, many-sectioned, self-similar
track is *good* — it says the tracks Josh named are longer and more sectioned than the tracks
beside them, that this survives holdout, and that it survives partially even against
equally-long siblings.
0.742 → 0.750; the 95th-percentile null moved 0.798 → 0.695. Any future pass that reports a raw
AUC without re-deriving its own bar for its own N is reporting noise.
contour runs and composed rest were V1's survivors and are now dead. §9 explains why — they were
modern-era tendencies — and that explanation is itself a two-cell descriptive read, not a test.
length-independent ones. The single most quotable number in this document is only honest when
quoted with the second one.
measured is whether the tracks that have it share anything the cards can see; at N = 11 the
answer was no, and at N = 28 the answer is *yes, on structural size, and on nothing else the
card can see.*
851 controls; Chrono Trigger and FINAL FANTASY VI contribute 11 of the 28 hits between them, so
two records carry 39% of the positives. A per-album leave-one-album-out pass is the obvious next
robustness check and is not run here.
robustness reads on a finding whose bar was set on the full pool.
python harness/music_gen/derive_patterns.py --self-test — 22 controls. Two are regression teethfor defects this pass shipped and its own controls caught: a planted signal the first fixture
failed to plant, and a power model that treated one track's 145 sibling comparisons as 145
independent coin flips. A third was repaired in this re-run: C16 asserted len(hits) == 11,
which is not a control but the pool size of the day, and it went red the moment the pool grew.
It now asserts the SET identity its own name claims — every attached corpus row that has a card
and is not the known-false identity is a hit, and nothing else is — derived independently of
label_rows, so it survives the pool growing, which it must.
python harness/music_gen/derive_patterns.py --perm 400 --json build/audio/exemplars/PATTERN_FINDINGS_V2.json— the full derivation, about a minute, deterministic under seed 20260807.
held-out passes, the complete power curve and the per-class descriptive block. The §5 truncation
table and the §9 era split are derived from the same public functions
(load_cards → flatten → label_rows → within_album_auc) and from the JSON's own
loo_composite_* fold percentiles.
Section count, length, bars, self-similarity, drop count and structural novelty are the only
measured axes with an out-of-sample claim. A MASTERPIECE_STANDARD band on any of them is now
defensible at the honest tier "clears a calibrated multiplicity bar at N = 28, era-invariant
across the two eras measured, partially confounded with curation."
N = 28 as at N = 11. If the difference lives there, this card cannot see it, and the correct
action is a better feature — a real instrument roster, a real hook extractor — not a band.
they have now been tested at the N their own power model asked for.
labelled necessary-not-sufficient. It is materially better than V1's as a filter because it
rejects far fewer real corpus rows.
was no signal to floor it on. There is one now: 0.744 ± 0.045 held-out, from a composite of size
axes. Any predictor built on it must carry the 0.497 length-independent number beside it, or it
will be read as measuring more than it does.
actually lives, and §9 has just demonstrated that an incoming era can overturn a finding.