PATTERN_FINDINGS_V2.md

music/exemplars/PATTERN_FINDINGS_V2.md

PATTERN FINDINGS V2 — the derivation re-run at N = 28

SUPERSEDED BY PATTERN_FINDINGS_V3.md (2026-08-07). The EX_014 and EX_094 quarantines
were retired by inspection and both rows re-attached, moving the pool from 28 hits / 851 controls
to 30 hits / 927 controls; two byte-identical UNMATCHED__ twins were retired in the same
sitting. At that denominator four axes clear the bar rather than six, none clears the 99th, and
the joint envelope changes two of its three axes — so §8's box, which the round-1 generation
filter reads, is the number to stop quoting first. V2 is not withdrawn: its method is
unchanged and it remains the correct answer for the pool it had, exactly as it says of V1.
Tier: MEASURED-SAMPLE FINDINGS, with six axes now clearing a calibrated multiplicity bar.
Everything below is derived from the formula cards already on disk. Nothing was purchased,
nothing was generated, no audio was decoded by this pass. The cards are read as data; the
exemplar bytes never move.
Supersedes: PATTERN_FINDINGS_V1.md (N = 11, zero axes clearing). V1 is not withdrawn —
its method is unchanged and its null is still the correct answer for the pool it had.
Produced by: harness/music_gen/derive_patterns.py (22 controls, mutation-proved).
Machine record: build/audio/exemplars/PATTERN_FINDINGS_V2.json.
Consumer: docs/proposals/music/MASTERPIECE_PROGRAM.md rung 2 (the archetype synthesis
toward MASTERPIECE_STANDARD) and rung 3 (the nostalgia predictor, whose floor this pass sets).
Legal posture: analysis for understanding only, unchanged from the cards themselves.

1. The one-paragraph answer

One thousand one hundred and eighty-one tracks are measured. Twenty-eight of them are corpus

rows — the tracks Josh named, the ones that stayed loved for decades — against eleven in V1. Eight

hundred and fifty-one are their own album siblings: same composer, same session, same production,

same codec. The design has not changed and neither has the feature bank. What changed is that the

multiplicity bar, which is measured rather than assumed, fell from 0.298 to 0.195 as the

positives went from eleven to twenty-eight, and six axes now stand above it where zero stood

above it before. The leave-one-out composite — axes and directions chosen fold by fold without ever

seeing the held-out track — moved from 0.586 ± 0.085 to 0.744 ± 0.045, twenty-two of

twenty-six corpus rows above their own album's median, binomial p = 0.0005. Held-out single rules

went from firing on 27.3% of unseen corpus rows (against a 43.9% sibling rate, worse than nothing)

to 67.9% against 35.9%. On the frozen train/test split, zero of six rules reverse sign,

where three of six reversed at N = 11. That is a real, replicated, out-of-sample signal, and it is

made of one thing: **structural size — how many distinct sections a track has, how long it runs,

how many bars, how much it repeats itself, how many event drops it contains.** The three

hypotheses V1 carried forward all failed at the larger N, and the length-independent half of

the bank collapsed to exactly chance (0.497 ± 0.046). The honest headline is therefore narrower

than the AUC suggests: *the measured card can pick a corpus row out of its own album at 0.74, and

everything it uses to do so is a measure of structural scale.*

2. The design, unchanged, and what N = 28 bought

Each corpus row is scored only against the tracks that shipped on its own record — composer,

scoring session, mixing chain, sample library, release year and audio codec held constant by

construction. Per-hit win rates are averaged with equal weight per hit, so a 208-track Pokémon

compilation cannot drown a four-track Bloodborne EP. This is the only framing the pool can answer

honestly, and it is V1's framing verbatim.

The pool, V1 against V2

V1V2note
Formula cards measured9141181formula_cards/*.json; repudiated/ excluded by construction
HITS — corpus rows, ACQUIRED or CARDED1128the labelled positives
CONTROLS — siblings in an album containing a hit601851across 15 albums with siblings
Measured but in an album with no hit301301carried, never used as controls
Quarantined11EX_014
Feature axes103103unchanged bank, so V1 and V2 are directly comparable
Albums carrying a hit1016one of them (Metal Gear Solid Vocal Tracks) contributes no siblings

EX_014 is dropped from both sides, not merely from the positives, on the corpus

defect_register's standing instruction: the row says Tekken 2 "Opening Theme", the measured bytes

are a Pokémon X & Y opening movie, and a known-false label is no better as a negative than as a

positive.

Two hits sit on an album with no siblings at all — EX_082 and EX_092 are both corpus rows on the

four-track Metal Gear Solid Vocal Tracks compilation — so the within-album statistics are computed

on 26 of the 28. Every AUC below carries that denominator; the envelope and descriptive blocks,

which do not need siblings, use all 28.

The twenty-eight

idclasstitlealbum measured againstsiblings
EX_006TITLE_CHARACTERChrono Trigger Main ThemeChrono Trigger (Original Soundtrack) [DS]71
EX_007TITLE_CHARACTERFrog's ThemeChrono Trigger71
EX_017EXPLORATIONSwedenC418 — Minecraft Volume Alpha22
EX_018EXPLORATIONSubwoofer LullabyMinecraft Volume Alpha22
EX_022EXPLORATIONWind SceneChrono Trigger71
EX_030EXPLORATIONRoute 209Pokémon Diamond & Pearl Super Music Collection145
EX_032EXPLORATIONTerra's ThemeFINAL FANTASY VI57
EX_033EXPLORATIONCorridors of TimeChrono Trigger71
EX_046BATTLEBFG DivisionDOOM (Original Game Soundtrack)30
EX_053BATTLECynthia Battle ThemePokémon Diamond & Pearl145
EX_055BATTLEThe Cyber GrindULTRAKILL22
EX_064BOSSLudwig, the Holy BladeBloodborne: The Old Hunters4
EX_075BOSSDancing MadFINAL FANTASY VI57
EX_082MELANCHOLICThe Best Is Yet to ComeMetal Gear Solid Vocal Tracks0
EX_088MELANCHOLICAn Unwavering HeartPokémon X & Y Super Music Collection208
EX_092TENSIONSnake Eater (Ladder Climb)Metal Gear Solid Vocal Tracks0
EX_102TENSIONMagus CastleChrono Trigger71
EX_110CREDITS_TRIUMPHWant You GonePortal 263
EX_113CREDITS_TRIUMPHTo Good FriendsChrono Trigger71
EX_124TITLE_CHARACTERDeus Ex Main Title (UNATCO)Deus Ex39
EX_136EXPLORATIONHumming the BasslineJet Set Radio Future16
EX_142EXPLORATIONShinshu FieldŌkami OST Vol. 154
EX_145BATTLEGo StraightStreets of Rage 2 (Remastered)24
EX_147BATTLEFighting of the SpiritTales of Phantasia76
EX_161MELANCHOLICAria di Mezzo CarattereFINAL FANTASY VI57
EX_162MELANCHOLICSchala's ThemeChrono Trigger71
EX_179CREDITS_TRIUMPHEnding Theme (Balance is Restored)FINAL FANTASY VI57
EX_180CREDITS_TRIUMPHDreams DreamsNiGHTS into Dreams20

All seven purpose classes are now represented — EXPLORATION 8, BATTLE 5, MELANCHOLIC 4,

CREDITS_TRIUMPH 4, TITLE_CHARACTER 3, BOSS 2, TENSION 2 — where V1 had six of seven and four of

them on one or two tracks. No per-class claim in this document is a test; the per-class block in

the JSON remains a descriptive read and is labelled as one. Five classes still sit below the

five-row floor the MASTERPIECE_STANDARD's per-class bands need.

3. The multiplicity price, re-derived for the new N

The bar is not guessed and it is not inherited from V1. Within each album, re-draw at random which

tracks are "hits" (keeping each album's count), recompute every feature's within-album AUC, and

keep the best deviation from 0.5 across the whole bank. Four hundred draws, same seed

discipline, same procedure as V1 — re-run from scratch on the new pool.

statistic of best-of-bank \AUC − 0.5\V1 (N = 11)V2 (N = 28)V2 as AUC
median of chance0.2260.1510.651
90th percentile0.2800.1850.685
95th percentile — the bar0.2980.1950.695
99th percentile0.3300.2200.720
max over 400 draws0.2400.740
observed best0.242 (n_drops)0.250 (section_count)0.750

Features clearing the 95th-percentile bar: six of 103. Clearing the 99th: two. At N = 11 it was

zero and zero. The observed best barely moved (0.242 → 0.250); what moved is the null underneath

it, which is exactly the effect V1 §8 predicted and the reason its power table held the bar

variable rather than fixed. This is the single most important structural fact about the re-run: the

finding was bought by the bar falling, not by the effect growing.

4. The ranking

axiswithin-album AUClength-matchedρ vs durationhits above own-album medianbinomial pclears bar
section_count0.7500.677+0.7822/260.0005yes (p99)
total_length_s0.7320.671+1.0022/260.0005yes (p99)
bars0.7180.639+0.9419/260.029yes (p95)
repeat_sim_mean0.7140.644+0.7122/260.0005yes (p95)
n_drops0.7100.637+0.5723/260.0001yes (p95)
novelty_mean0.3040.361−0.597/260.029yes (p95)
note_count0.6730.567+0.7919/260.029no
tempo_adjust_n0.6700.608+0.7020/260.009no
n_raises0.6690.595+0.5618/260.076no
breaks_per_min0.3430.383−0.896/260.009no
time_to_hook_frac0.3440.384−0.567/260.029no
silence_frac0.3570.418−0.767/260.029no
melody_span_semitones0.6340.569+0.3318/260.076no
repeat_frac0.6320.581+0.5417/260.169no
dynamic_range_db0.3690.419−0.568/260.076no

novelty_mean clears the bar in the negative direction and is the one non-size axis in the

group: it is the mean self-similarity novelty of the track's own structure, and corpus rows sit

*below* their siblings on it. Read together with repeat_sim_mean sitting *above*, the pair says

the same thing twice — a corpus row is more self-similar section-to-section than the album around

it. That is not "repetitive"; at these lengths and section counts it is *thematic return*, a track

that keeps coming back to its own material instead of moving on.

Tempo is still not the confound. Rescoring against siblings within ±20% bpm changes essentially

nothing (section_count 0.750 → 0.749, total_length_s 0.732 → 0.733, n_drops 0.710 → 0.717).

The pattern is not "battle themes are fast and battle themes are over-represented", and that was

true at N = 11 as well.

5. Duration: still the confound, but no longer the whole story

V1's §4 verdict was blunt and correct for its pool: *the top of the table is duration wearing a

costume*, and the duration effect is curation — Josh named finished themes while game albums are

full of jingles, stingers and menu blips. That remains true. The control pool still carries 161

tracks under 30 seconds and 267 under a minute, against a hit pool with exactly one row under a

minute (EX_102, Magus Castle, 28.6 s).

So the V1 rescoring control was run again, and a harder version of it was added: instead of only

matching each hit to siblings within ±35% of its own duration, the control pool is truncated from

below, admitting only siblings that are themselves finished-length tracks.

axisall 851 controls≥ 30 s (690)≥ 60 s (584)≥ 90 s (451)
section_count0.7500.7230.6860.652
n_drops0.7100.6900.6640.645
repeat_sim_mean0.7140.6860.6490.621
novelty_mean0.3040.3420.3600.377
bars0.7180.6830.6520.609
total_length_s0.7320.7010.6630.581
note_count0.6730.6410.6190.587
tempo_adjust_n0.6700.6450.6180.586

Read the last column. Raw duration is the axis that dies — 0.732 down to 0.581, which is what

"curation, not composition" looks like when you take the stingers away. section_count, n_drops,

repeat_sim_mean and novelty_mean all decay much more slowly and are still around 0.62–0.65

against siblings that are every bit as long as the hits. That residual is the part of the effect

that is not length, and it is the part worth carrying into rung 2.

This truncation is a robustness read, not a test. The permutation bar in §3 was calibrated on

the full control pool; a truncated pool has fewer and differently-distributed siblings, so its own

bar would sit somewhat higher and is not computed here. Nothing in this table is claimed to clear

anything.

6. The three carried hypotheses — all three fail at N = 28

V1 §12 named three things to re-test, all of which had survived its length control. All three are

re-tested here on the same axes with no re-definition, and none survives.

carried hypothesisaxisV1 AUCV1 length-matchedV2 AUCV2 rankverdict
Delayed self-recurrence — the theme waits longer before returningrecurrence_lag_norm0.6620.7030.505102 / 103failed — dead at chance
Unbroken contour runs — the melody climbs or falls further before turningcontour_run_max0.6720.6020.53666 / 103failed
Composed rest — silence budget plus drop rateX_rest_grammar0.5310.6590.42432 / 103failed, and reversed

recurrence_lag_norm is the most instructive: it was V1's best surviving candidate, the one that

*rose* under length matching, and at N = 28 it is the second-worst axis in the entire bank. The

descriptive medians show why — hit median 0.033 against a sibling median 0.111, which is the

opposite sign to V1's reading. Seventeen new corpus rows were enough to erase and invert it.

This is the clearest vindication of V1's own honesty apparatus. V1 reported these three as *not

significant*, priced against a bar none of them cleared, and recommended them only as things to

re-test. Had they been promoted to MASTERPIECE_STANDARD bands on the strength of "survives the

length control", three false rules would now be in the generator.

**FINDING 2026-08-07 — THE RECURRENCE AXIS WAS MEASURING THE WRONG THING, AND CLOSING THAT DID NOT

RESCUE IT.** recurrence_lag_norm above, and the motif_economy and repeat_map blocks it comes

from, run on raw chroma and MFCC self-similarity. That means **a theme restated a fourth higher, at

half speed, or ornamented scores as NEW MATERIAL rather than as a return** — which is exactly

backwards for a leitmotif score, where the transformed return is the whole craft. So the axis failing

at 0.505 was never a clean test of the hypothesis: it was a test of literal, in-key, in-tempo

repetition, which is the one kind of return good music uses least.

That gap is now closed. harness/music_gen/instr_hook_presence.py matches over all twelve

pitch-class rotations (recording the winning rotation, so the INTERVAL of the return is itself

reported), over a seven-point time-ratio grid, and symbolically under an edit budget. **The

hypothesis fails again anyway, and harder: zero of its fourteen metrics clear the bar re-derived on

the live pool** (best deviation 0.175 against an instrument-only bar of 0.245 and a joint bar of

0.241, where the joint bar asks the harsher and more honest question of whether it beats everything

already measured), with n_returns at 0.510 and returns_per_min at 0.500 — chance to three

decimals. Against the eighteen candidates Josh rejected it is worse than useless: those score HIGHER

than the loved tracks on n_returns (3.0 against 1), transformation_share (1.000 against 0.893)

and hook_strength (0.736 against 0.691). Round 1 states hooks and brings them back. So this

row's verdict stands, now for a better reason: it is a second, invariance-corrected null on the same

family of ideas, not an artefact of a blunt ruler.

Two cautions travel with that, both from the instrument's own face. Its return threshold was

calibrated on synthetic controls and does not transfer to real recordings — nine of the thirty loved

tracks read zero returns, and a sensitivity sweep shows one of them going from 0 to 86 returns as the

threshold falls — so every ABSOLUTE return count it publishes is too low and the floor derived from

them is degenerate. The separation null survives that sweep at every threshold tested, because

loosening lifts the rejected candidates as much as it lifts the hits. And it cannot choose the motif:

it follows the card's recorded hook onset, so a track whose memorable idea is not its opening idea is

searched for the wrong phrase with perfect rigour.

The four pre-registered hypotheses, still all failed

hypothesisaxisV1 AUCV2 AUCV2 length-matchedverdict
Hook arrives inside the first phrase and repeats its own interval shapeX_hook_immediacy_economy0.5310.5080.461not separated
Singable = mostly steps, leaps answered against, compass narrowX_singability0.5750.4550.448not separated
Arrangement keeps lanes in reserve, spends them at seamsX_reserve_and_spend0.5050.4440.455not separated
Rest is composed — silence budget plus drop rateX_rest_grammar0.5310.4240.509not separated

Their own four-feature null puts the bar at deviation 0.146 (AUC 0.646), down from 0.223 at N = 11.

None of the four comes near it, and three of the four moved *away* from separation as the pool

grew. **The four theories written into the extractor before the first run are, at 28 positives,

measurably not what distinguishes these tracks.**

7. Held-out validation — three ways, all reported, all improved

7a. Frozen split, selection on train only

Fifteen hits train, thirteen test. The six best training axes were turned into thresholded rules by

Youden's J on the training data, then scored once on the untouched test side.

rule fitted on traintrain AUCtest AUCtest TPRtest FPRtest liftprecision lift
section_count ≥ 150.7260.7830.620.1603.85×3.63×
total_length_s ≥ 140.60.7490.7100.690.2352.95×2.83×
n_drops ≥ 170.7200.6960.620.1853.32×3.16×
bars ≥ 1100.7510.6720.460.1223.80×3.58×
note_count ≥ 2680.7250.6020.540.1583.40×3.24×
tempo_adjust_n ≥ 110.7260.5940.620.2282.70×2.60×

Zero of six rules reverse sign between train and test. At N = 11, three of six reversed and two

of those had the strongest training AUCs — the signature of selection noise. Here every rule holds

its direction and every one carries a 2.7–3.9× lift on the held-out side. This is the single

largest qualitative change from V1 and it is not a change of method: the same code, the same split

procedure, the same Youden thresholds.

7b. Leave-one-out, rule re-selected inside every fold

The strongest axis is re-chosen from scratch on the other twenty-seven hits, thresholded on them,

and the untouched twenty-eighth is asked whether it fires.

V1 (N = 11)V2 (N = 28)
held-out corpus rows firing3 of 11 — 27.3%19 of 28 — 67.9%
mean sibling fire rate for the same rules43.9%35.9%
ratio0.62× — worse than nothing1.89×

At N = 11 the selected rule fired on held-out corpus rows *less often* than on the album filler it

was built to exclude. At N = 28 it fires nearly twice as often on the unseen corpus row. The rule

selected in almost every fold is section_count ≥ 14.

7c. Leave-one-out composite — the honest headline

In each fold the top five axes and their directions are chosen on the other twenty-seven hits

alone, z-scored within album, summed, and the held-out track's percentile among its own siblings is

recorded.

compositeV1 held-out AUCV2 held-out AUCSEhits above own-album medianbinomial p
over all 103 axes0.561 ± 0.0710.7440.04522 / 260.0005
over the length-independent axes0.586 ± 0.0850.4970.04615 / 260.557

The length-independent bank is defined with no reference to the hit labels — the correlation

with duration is measured on the control pool alone (|ρ| < 0.35), which admits 70 of 103 axes here.

V1's leak (filtering the hit-ranked top-25 instead, worth about +0.06 AUC) stayed closed; the same

control-pool filter is used.

0.744 ± 0.045 is the honest realized predictive power of the whole measured card on this pool,

and 0.497 ± 0.046 is what is left of it once every axis correlated with duration is removed.

Both numbers must be quoted together. The composite works, and the composite is made of size.

8. The rejection envelope, re-derived

Separation ("which side of a threshold") and containment ("inside which box") are different

questions, and the second is the one a generator can use even when the first is weak. Each envelope

is built from twenty-seven hits and validated by asking whether it contains the twenty-eighth.

axishit envelopeheld-out hits containedfiller admittedlift
melody_span_semitones19.7 – 57.0 st27 / 2869.5%1.39×
section_mean_s7.1 – 63.3 s26 / 2870.0%1.33×
X_rest_grammar2.68 – 17.4326 / 2870.0%1.33×
section_count4 – 3527 / 2883.4%1.16×
bars18 – 82026 / 2880.4%1.16×
total_length_s28.6 – 1295.6 s26 / 2881.3%1.14×

The joint box over melody_span_semitones × section_mean_s × X_rest_grammar contains

25 of 28 held-out hits (89.3%) while admitting 44.8% of filler — a 1.99× enrichment.

V1's joint box contained 7 of 11 (63.6%) at 30.3% filler for 2.1×.

The enrichment is essentially unchanged; the recall is much better, which is the direction that

matters for a rejection filter. A box that rejects a third of real corpus rows rejects good

generated candidates at the same rate. Note that the per-axis envelopes are wider than V1's because

twenty-eight tracks span more than eleven do, and every individual axis's lift fell accordingly —

the box is doing its work jointly, not axis by axis. It remains a **necessary-not-sufficient

constraint**: a generated track outside the box is unlike every measured corpus row; a generated

track inside it has cleared a bar that almost half of album filler also clears.

X_rest_grammar earning a place in the joint box while failing badly as a separator is not a

contradiction — it is precisely the separation-versus-containment distinction. Corpus rows occupy a

*narrow band* of composed rest without sitting *high or low* on it.

9. The era mix — eleven-plus SNES-era hits entered a lineage that was empty

V1 §8's sharpest structural complaint was that **the nostalgia lineage the whole thirty-year bar

rests on contributed zero measured tracks**: no SNES, no N64, no Zelda, no Kondo, twelve of fifteen

albums released in 2000 or later. The STARTER TEN closed half of that. Four 16-bit-era scores now

carry corpus rows — Chrono Trigger (1995, SNES), FINAL FANTASY VI (1994, SNES), Tales of Phantasia

(1995, Super Famicom) and Streets of Rage 2 (1992, Mega Drive) — **thirteen of the twenty-eight

hits and 228 of the 851 siblings.** The pool is now roughly half 16-bit era and half modern, which

makes the era question testable for the first time.

axisall 2816-bit era (13 hits)modern (13 hits used)
section_count0.7500.7620.739
total_length_s0.7320.7200.745
bars0.7180.7390.697
repeat_sim_mean0.7140.6920.735
n_drops0.7100.6630.757
novelty_mean0.3040.3580.250
note_count0.6730.7050.640
tempo_adjust_n0.6700.6390.701
recurrence_lag_norm0.5050.3760.634
contour_run_max0.5360.4170.656
X_rest_grammar0.4240.3280.519
repeat_sim_max0.5930.5000.687

Two readings, and they point opposite ways.

same sign and roughly the same magnitude in both halves — section_count at 0.762 against 0.739,

bars at 0.739 against 0.697. A structural-size effect that reproduces independently in 1994

Super Famicom sample-ROM scores and in 2016 live-recorded soundtracks is not an artifact of

production era, budget or codec. This is the strongest single result in the re-run, and it is

the one V1 could not have obtained at any N, because it had only one era.

above 0.5 in the modern half at close to their V1 values (0.634, 0.656, 0.519) and *below* 0.5 in

the 16-bit half (0.376, 0.417, 0.328). Pooled, they cancel to chance. V1's pool was entirely

modern, so what V1 measured was a modern-era tendency and reported, correctly, as not

significant. The lineage that just arrived contradicts it.

The same split applied to the leave-one-out composite folds — the same percentiles the §7c headline

is built from, no re-selection — gives a mean held-out percentile of 0.726 for the 16-bit rows

against 0.763 for the modern rows on the all-axes composite: **the composite works about equally

well in both eras.** On the length-independent composite it is 0.599 against 0.395, two halves

pulling against each other around a pooled 0.497, which is a warning about that composite rather

than a finding about either era.

The lineage is still only half covered. Zero N64, zero Zelda, zero Kondo, zero Nintendo

first-party of any kind. The Kondo/Zelda/N64 gap is the same gap V1 named and it is still open; what

closed is the Square/Enix 16-bit JRPG and the Sega 16-bit action lineage.

10. What the numbers say at descriptive tier

Not discriminative findings unless marked as clearing the bar in §4. These are the measured profile

of the twenty-eight, usable as sanity bands for authored candidates.

axishit medianhit rangesibling mediansibling p10–p90AUC
section count174 – 3593 – 180.750
length196.6 s28.6 – 1295.696.2 s13.6 – 223.70.732
bars11418 – 820528 – 1390.718
mean section self-similarity0.9020.580 – 0.9520.8550.496 – 0.9210.714
drops19.53 – 11063 – 250.710
structural novelty0.5560.397 – 0.8570.6780.507 – 0.9760.304
note count31613 – 21281218 – 3710.673
tempo adjustments13.51 – 8961 – 210.670
melodic span43.0 st19.7 – 57.031.0 st9.0 – 48.00.634
hook range22.0 st0.67 – 52.017.0 st5.0 – 38.00.589
mean section length5.36 bars2.9 – 31.04.29 bars2.3 – 10.40.546
leap fraction (≥5 st)0.4000.00 – 0.870.4000.07 – 0.800.546
interval 3-gram repetition0.4320.00 – 0.830.4090.00 – 0.730.521
gap-fill after leap0.0000.00 – 1.000.0000.00 – 1.000.533
time to hook0.575 s0.04 – 61.60.615 s0.26 – 5.940.483
melodic salience0.7970.00 – 0.9240.7560.352 – 0.9090.487
stepwise interval fraction0.1080.00 – 0.330.1330.00 – 0.400.476
contour turns per note0.3670.00 – 0.6670.4000.133 – 0.6670.470
windowed key agreement0.4110.109 – 0.9290.5000.212 – 1.0000.460
tempo120.3 bpm83.4 – 172.3123.1 bpm99.4 – 152.00.441
dynamic range11.7 dB4.4 – 50.713.7 dB6.6 – 69.80.369
silence budget2.1%0.6 – 11.12.8%1.0 – 19.80.357

Four readings worth carrying into rung 2, each with its own confidence:

of hummability is built from — stepwise fraction, gap-fill after a leap, hook compass, contour

turn rate, interval 3-gram repetition, time to hook, melodic salience — still sits between AUC

0.46 and 0.59 with 2.5× the positives, and most of them moved *toward* 0.5, not away from it.

The most economical explanation is unchanged: the album siblings already have them. These are

well-written game scores throughout, and the corpus row is not the only competent track on its

record. What the card can see that separates them is size and return, not melodic craft. That

remains the most important sentence in this document for the program.

the section count, double the length, double the bars, triple the drop count, above its album on

self-similarity and below it on novelty. Six axes, one story, all six clearing a calibrated bar.

Part of it is curation (§5), and about 0.62–0.65 of it survives against equally long siblings.

N = 11). The effect weakened as the pool grew and is now nearly nothing. Do not carry it.

ceiling in the spectral_band_proxy method, not a finding, and any arrangement-width claim from

these cards is void until a real instrument roster exists. Two full re-runs have now confirmed

the ceiling; the cards' own declaration that the lane method is not a roster is due.

11. What more corpus would still sharpen — the power model, re-fitted

The model is re-fitted on the new pool: K_eff, the number of effectively independent axes this

bank behaves like, is chosen so the simulated bar at N = 28 reproduces the observed permutation

95th percentile, and only then extrapolated. Fitted K_eff = 103 (calibration error 0.007),

against 80 at N = 11 — the bank behaves like more independent axes now that the null is estimated

from more positives, which raises the bar slightly relative to a naive extrapolation. Real sibling

counts are used.

carded corpus rowsmultiplicity bar (AUC)detection of a true 0.65of 0.70of 0.75of 0.80
110.7973.0%12.0%29.5%53.5%
160.7596.0%19.0%47.0%78.0%
220.71912.2%34.0%74.8%95.0%
28 (today)0.695 (observed)
300.69022.3%57.3%91.2%99.8%
400.66538.3%80.7%97.3%100%
550.63962.0%96.3%100%100%
750.61986.8%99.3%100%100%
1000.60894.8%100%100%100%
1400.59199.3%100%100%100%

Rows needed for 80% power: about 24 for an effect of 0.75, 40 for 0.70, about 70 for 0.65

(interpolating the table — 0.75 reads 74.8% at 22 and 91.2% at 30; 0.65 reads 62.0% at 55 and

86.8% at 75). The pool has just passed the 0.75 threshold, which is why six axes appeared. The model is

generous by construction — a constant true effect on one axis, albums like the ones already

measured, no extra heterogeneity from a broader corpus — so read these as lower bounds.

What the next acquisitions are worth, in priority order:

zero Zelda, zero Kondo, zero Nintendo first-party. §9 shows the era split is measurable and that

it already overturned three hypotheses; a third era would be the strongest available test of

whether the six surviving axes are really era-invariant or merely invariant across the two eras

now held. This is the highest-value buy in the corpus and it is the same one V1 named.

TITLE_CHARACTER 3, MELANCHOLIC 4, CREDITS_TRIUMPH 4). The MASTERPIECE_STANDARD is specified in

per-class bands and no per-class band can be stated until each class carries five-plus rows.

BOSS and TENSION are the binding constraint.

are silently excluded from every within-album statistic here. Buying the Metal Gear Solid and

Metal Gear Solid 3 score albums would recover two positives already paid for.

multiplicity bar; the 4-sibling Bloodborne EP quantises one hit's AUC to fifths.

12. Honest register — what this is and is not

This is a measured-sample derivation, not a validated theory, and nothing here is yet a

MASTERPIECE_STANDARD band.

real and out-of-sample; it is also narrow. Nothing here says a long, many-sectioned, self-similar

track is *good* — it says the tracks Josh named are longer and more sectioned than the tracks

beside them, that this survives holdout, and that it survives partially even against

equally-long siblings.

0.742 → 0.750; the 95th-percentile null moved 0.798 → 0.695. Any future pass that reports a raw

AUC without re-deriving its own bar for its own N is reporting noise.

contour runs and composed rest were V1's survivors and are now dead. §9 explains why — they were

modern-era tendencies — and that explanation is itself a two-cell descriptive read, not a test.

length-independent ones. The single most quotable number in this document is only honest when

quoted with the second one.

measured is whether the tracks that have it share anything the cards can see; at N = 11 the

answer was no, and at N = 28 the answer is *yes, on structural size, and on nothing else the

card can see.*

851 controls; Chrono Trigger and FINAL FANTASY VI contribute 11 of the 28 hits between them, so

two records carry 39% of the positives. A per-album leave-one-album-out pass is the obvious next

robustness check and is not run here.

robustness reads on a finding whose bar was set on the full pool.

13. Reproduction

for defects this pass shipped and its own controls caught: a planted signal the first fixture

failed to plant, and a power model that treated one track's 145 sibling comparisons as 145

independent coin flips. A third was repaired in this re-run: C16 asserted len(hits) == 11,

which is not a control but the pool size of the day, and it went red the moment the pool grew.

It now asserts the SET identity its own name claims — every attached corpus row that has a card

and is not the known-false identity is a hit, and nothing else is — derived independently of

label_rows, so it survives the pool growing, which it must.

— the full derivation, about a minute, deterministic under seed 20260807.

held-out passes, the complete power curve and the per-class descriptive block. The §5 truncation

table and the §9 era split are derived from the same public functions

(load_cardsflattenlabel_rowswithin_album_auc) and from the JSON's own

loo_composite_* fold percentiles.

14. What rung 2 should do with this

Section count, length, bars, self-similarity, drop count and structural novelty are the only

measured axes with an out-of-sample claim. A MASTERPIECE_STANDARD band on any of them is now

defensible at the honest tier "clears a calibrated multiplicity bar at N = 28, era-invariant

across the two eras measured, partially confounded with curation."

N = 28 as at N = 11. If the difference lives there, this card cannot see it, and the correct

action is a better feature — a real instrument roster, a real hook extractor — not a band.

they have now been tested at the N their own power model asked for.

labelled necessary-not-sufficient. It is materially better than V1's as a filter because it

rejects far fewer real corpus rows.

was no signal to floor it on. There is one now: 0.744 ± 0.045 held-out, from a composite of size

axes. Any predictor built on it must carry the 0.497 length-independent number beside it, or it

will be read as measuring more than it does.

actually lives, and §9 has just demonstrated that an incoming era can overturn a finding.

Generated by harness/site/structure_site.py — the URL path is the repo path. review root