PATTERN_FINDINGS_V1.md

music/exemplars/PATTERN_FINDINGS_V1.md

PATTERN FINDINGS V1 — the derivation stage over the measured pool

Tier: MEASURED-SAMPLE FINDINGS. Everything below is derived from the formula cards already
on disk. Nothing was purchased, nothing was generated, no audio was decoded by this pass. The
cards are read as data; the exemplar bytes never move.
Produced by: harness/music_gen/derive_patterns.py (22 controls, mutation-proved).
Machine record: build/audio/exemplars/PATTERN_FINDINGS_V1.json.
Consumer: docs/proposals/music/MASTERPIECE_PROGRAM.md rung 2 (the archetype synthesis
toward MASTERPIECE_STANDARD) and rung 3 (the nostalgia predictor, whose floor this pass was
supposed to help set).
Legal posture: analysis for understanding only, unchanged from the cards themselves.

1. The one-paragraph answer

Nine hundred and fourteen tracks are measured. Eleven of them are corpus rows — the tracks Josh

named, the ones that stayed loved for decades. Six hundred and one are their own album siblings:

same composer, same session, same production, same codec. On one hundred and three measured axes

— contour classes, interval distributions, phrase lengths, hook repetition structure, tonal

stability, arrangement, event grammar, dynamics, timbre — **not one axis separates the eleven

from their siblings by more than label-shuffling produces on the same data.** The best single

axis reaches a within-album AUC of 0.742; the permutation null over the same feature bank puts

the 95th percentile of chance at 0.798. A composite chosen fold-by-fold without ever seeing the

held-out track lands at AUC 0.586 ± 0.085 — about one standard error from chance. That is

the finding, stated at the tier the sample supports: **at N = 11, on these axes, there is no

demonstrated formula.** The study is also not powerful enough to rule one out; the arithmetic in

§8 says a real effect of 0.75 would be detected here only about a third of the time. The

actionable output is therefore §8's number, not a rule: thirty carded corpus rows would make

an effect of that size visible, against eleven today.

2. The design, and why it is the only honest one available

A naive study compares corpus rows against arbitrary music and "discovers" that beloved game

themes are longer, better produced and more melodic than whatever else is lying around. Every

one of those differences is era, budget, genre and mastering, not composition.

So the control group here is the album sibling. Each of the eleven corpus rows is scored only

against the other tracks that shipped on its own record. Composer, scoring session, mixing chain,

sample library, release year and audio codec are held constant by construction. The question the

pool can actually answer is the sharp one: *what separates the track people still hum from the

forty competent tracks that shipped beside it?*

Per-hit win rates are then averaged with equal weight per hit, so a 208-track Pokémon compilation

cannot drown a four-track Bloodborne EP.

The pool

countnote
Formula cards measured914build/audio/exemplars/formula_cards/*.json; repudiated/ excluded by construction
HITS — corpus rows, ACQUIRED or CARDED11the labelled positives
CONTROLS — siblings in an album containing a hit601across 10 albums
Measured but in an album with no hit301carried, never used as controls
Quarantined1EX_014
Feature axes103scalars with ≥80% coverage and >2 distinct values

EX_014 is dropped from both sides, not merely from the positives. The corpus defect_register

names it a known-false identity — the row says Tekken 2 "Opening Theme", the measured bytes are a

Pokémon X & Y opening movie — and the ledger's instruction is that its ACQUIRED is not evidence

of anything until the album is bought and inspected. A known-false label is no better as a

negative than as a positive.

The eleven

idclasstitlealbum measured againstsiblings
EX_030EXPLORATIONRoute 209Pokémon Diamond & Pearl Super Music Collection145
EX_046BATTLEBFG DivisionDOOM (Original Game Soundtrack)30
EX_053BATTLECynthia Battle ThemePokémon Diamond & Pearl Super Music Collection145
EX_055BATTLEThe Cyber GrindULTRAKILL22
EX_064BOSSLudwig, the Holy BladeBloodborne: The Old Hunters4
EX_088MELANCHOLICAn Unwavering HeartPokémon X & Y Super Music Collection208
EX_110CREDITS_TRIUMPHWant You GonePortal 263
EX_124TITLE_CHARACTERDeus Ex Main Title (UNATCO)Deus Ex39
EX_136EXPLORATIONHumming the BasslineJet Set Radio Future16
EX_142EXPLORATIONShinshu FieldŌkami OST Vol. 154
EX_180CREDITS_TRIUMPHDreams DreamsNiGHTS into Dreams20

Seven purpose classes exist in the corpus; six are represented here, four of them by one or two

tracks. No per-class claim in this document is a test — the per-class block in the JSON is a

descriptive read and is labelled as one.

3. The multiplicity price, paid up front

With a hundred-odd axes and eleven positives, the best-looking feature will look good by luck.

The size of that luck is not guessed here, it is measured: within each album, re-draw at random

which tracks are "hits" (keeping each album's count), recompute every feature's within-album AUC,

and keep the best deviation from 0.5 across the whole bank. Four hundred draws give the

distribution a real finding has to clear.

statistic of best-of-bankAUC − 0.5valueequivalent AUC
median of chance0.2260.726
90th percentile0.2800.780
95th percentile — the bar0.2980.798
99th percentile0.3300.830
observed best (n_drops)0.2420.742

Features clearing the bar: zero of 103. The best real feature does not even reach the *median*

of what chance produces on this data by much. The coarseness is not an artifact of the method — it

is what eleven positives buy, most of it contributed by the small albums, where a single hit's

rank among four siblings is quantised to fifths.

4. The ranking, and the confound that explains most of it

axiswithin-album AUClength-matchedρ vs durationhits above own-album medianuncorrected pclears bar
n_drops0.7420.665+0.6010/110.012no
novelty_mean0.2670.319−0.641/110.012no
total_length_s0.7230.665+1.009/110.065no
repeat_sim_mean0.7200.632+0.739/110.065no
section_count0.7180.607+0.819/110.065no
recurrence_lag_s0.7000.689+0.338/110.227no
repeat_frac0.6970.627+0.588/110.227no
bars0.6950.626+0.958/110.227no
n_raises0.6890.616+0.628/110.227no
tempo_adjust_n0.6760.614+0.778/110.227no
contour_run_max0.6720.602+0.189/110.065no
recurrence_lag_norm0.6620.703−0.178/110.227no
bpm0.3440.388−0.133/110.227no
hook_vocab_size0.6220.635−0.127/110.549no
centroid_hz0.6190.601−0.068/110.227no

The top of that table is duration wearing a costume. The ρ column is the rank correlation of

each axis with track length, measured on the control pool alone with no hit label involved.

Everything above 0.69 AUC also sits above ρ = 0.58: more drops, more sections, more bars, more

tempo adjustments and higher self-similarity are all things that happen to a track simply by

lasting longer. Rescoring each hit only against siblings within ±35% of its own duration collapses

the whole leading group toward 0.61–0.67.

And the duration effect is itself curation, not composition. Josh named finished themes; game

albums are full of jingles, stingers, menu blips and fanfares. The control pool's 10th percentile

length is 10.0 s against a hit minimum of 82.3 s. "Corpus rows are longer than album filler" is a

true sentence about how the corpus was assembled, and it teaches the composer nothing.

Two things survive the length control rather than dying to it, and they are the only candidates

worth carrying forward:

how far into the track the principal material waits before returning, as a fraction of the

track. The hits let the listener wait longer before the theme comes back around. Stated as a

testable rule: *a corpus row delays its first strong self-recurrence past 6.4% of its length,

where its siblings return at 5.0%.* Not significant here. Worth re-testing at N = 30.

motion in the extracted hook: median 4 notes for the hits against 3 for the siblings, ranging to

12. The hits climb or fall further before turning. Directed melodic motion rather than zigzag.

Also not significant.

Tempo is not the confound. Rescoring against siblings within ±20% bpm changes essentially

nothing (n_drops 0.742 → 0.756, total_length_s 0.723 → 0.725). The pattern is not "battle

themes are fast and battle themes are over-represented."

5. The four pre-registered hypotheses, all failed

Four compound axes were written into the extractor as stated theories before the first run, so

they carry a multiplicity price of four rather than of the whole bank and get their own null

(95th percentile deviation 0.223, i.e. AUC 0.723).

hypothesisaxisAUClength-matchedverdict
The hook arrives inside the first phrase and repeats its own interval shapeX_hook_immediacy_economy0.5310.430not separated
Singable = mostly steps, leaps answered by a step against them, compass held narrowX_singability0.5750.554not separated
Arrangement keeps lanes in reserve and spends them at structural seamsX_reserve_and_spend0.5050.523not separated
Rest is composed — silence budget plus drop rateX_rest_grammar0.5310.659not separated

None clears its own four-feature bar. X_rest_grammar is the only one that strengthens under

length matching, and it is the one to re-state and re-test with more rows.

6. Held-out validation — three ways, all reported

6a. Frozen split, selection on train only

Six hits train (EX_046, EX_053, EX_055, EX_064, EX_110, EX_136), five test (EX_030, EX_088,

EX_124, EX_142, EX_180). The six best training axes were turned into thresholded rules by Youden's

J on the training data, then scored once on the untouched test side.

rule fitted on traintrain AUCtest AUCtest TPRtest FPRtest lift
time_to_full_texture_s ≥ 3.810.2410.6040.600.2302.61×
time_to_full_texture_bars ≥ 1.690.2630.5620.600.2662.26×
voiced_frac ≤ 0.1370.2430.5990.200.1931.04×
centroid_hz ≥ 23520.7450.4680.200.2150.93×
ss_baseline ≥ 0.8300.7550.3940.200.3240.62×
iv_mean_abs ≤ 3.070.2650.7070.000.2460.00×

Three of six rules reverse sign between train and test, and two of those had the *strongest*

training AUCs. That is the signature of selection noise, not of a weak-but-real effect. The two

rules that held (both forms of "takes longer to reach full texture") carry a 2.3–2.6× lift on the

held-out side, which is the single most encouraging number in this document and rests on three

firing tracks.

6b. Leave-one-out, rule re-selected inside every fold

The strongest axis is re-chosen from scratch on the other ten hits, thresholded on them, and the

untouched eleventh is asked whether it fires.

The rules fire on held-out corpus rows *less often* than on the album filler they were built to

exclude. A single-feature hit rule selected this way is worse than nothing.

6c. Leave-one-out composite — the honest headline

A predictor is used as several axes combined, so it is scored that way. In each fold the top five

axes and their directions are chosen on the other ten hits alone, z-scored within album, summed,

and the held-out track's percentile among its own siblings is recorded.

compositeheld-out AUCSEhits above own-album medianbinomial p
over all 103 axes0.5610.0715/111.000
over the 64 length-independent axes0.5860.0856/111.000

The length-independent bank is defined with no reference to the hit labels — the correlation

with duration is measured on the control pool alone. An earlier draft of this pass filtered the

hit-ranked top-25 instead, which leaks the held-out track into the candidate pool; that leak was

worth about +0.06 AUC and reported 0.648. It is named here because a leak found and closed is the

difference between a number and a claim.

0.586 ± 0.085 is the honest realized predictive power of the whole measured card on this pool.

7. The one thing with usable lift: the envelope, not the line

Separation ("which side of a threshold") and containment ("inside which box") are different

questions, and the second is the one a generator can use even when the first fails. Each envelope

below is built from ten hits and validated by asking whether it contains the eleventh.

axishit envelopeheld-out hits containedfiller admittedlift
repeat_sim_max0.989 – 1.00010/1153.1%1.71×
total_length_s82.3 – 506.8 s9/1147.1%1.74×
section_mean_s7.9 – 63.3 s9/1151.3%1.60×
bars39 – 3209/1153.6%1.53×
tempo_adjust_n5 – 379/1154.1%1.51×
time_to_hook_frac0.0002 – 0.01899/1157.4%1.43×

The joint box over duration, repeat_sim_max and section_mean_s contains 7 of 11 held-out

hits while admitting 30.3% of filler — a 2.1× enrichment. That is a genuine, validated,

modest result, and it is a necessary-not-sufficient constraint: useful as a rejection filter

on generated candidates, useless as a target to optimise. A generated track outside the box is

unlike every measured corpus row; a generated track inside it has cleared a bar that a third of

album filler also clears.

8. What more corpus would sharpen — with a number

Two things change as the corpus grows: the effect estimate gets more precise, **and the

multiplicity bar itself falls**, because the null of the best-of-bank statistic narrows. A power

model that holds the N = 11 bar fixed makes more data look useless, which is the opposite of the

truth.

The bar here is not assumed. K_eff — the number of effectively independent axes this bank

behaves like — is fitted so the simulated bar at N = 11 reproduces the observed permutation 95th

percentile (fitted K_eff = 80, calibration error 0.003), and only then extrapolated. Real

sibling counts are used, because a four-track EP quantises one hit's AUC to fifths and that

coarseness is most of the variance at N = 11.

carded corpus rowsmultiplicity bar (AUC)detection of a true 0.65of 0.70of 0.75of 0.80
11 (today)0.7913.5%12.5%33.5%58.8%
160.7586.2%17.7%47.5%76.2%
220.7209.8%36.0%74.3%95.0%
300.68429.7%64.2%93.2%99.3%
400.66438.5%79.2%99.0%99.8%
550.63962.3%95.0%99.5%100%
750.62377.5%99.3%100%100%
1000.60593.5%100%100%100%
1400.58599.8%100%100%100%

Rows needed for 80% power: 30 for an effect of 0.75, 55 for 0.70, 100 for 0.65. The model is

generous by construction — a constant true effect on one axis, albums like the ones already

measured, no extra heterogeneity from a broader corpus — so read these as lower bounds.

Three specific gaps make the next acquisitions worth more than their count suggests:

SNES-era, zero N64, zero Zelda, zero Kondo. Twelve of the fifteen measured albums are 2000-or-later

releases and three are Pokémon compilations; only two albums containing a hit predate 2001. The

corpus's own EXPANSION wave exists precisely to cover that lineage and none of it is on disk.

If the thirty-year property lives anywhere in these numbers, it lives in the tracks not yet

measured.

can be stated until each class carries five-plus rows, and the MASTERPIECE_STANDARD is specified

in per-class bands.

that sets the multiplicity bar. Full albums beat single-track purchases for this pass.

The nearest concrete step is the BUY_SHEET's STARTER TEN at $64.59, which is ten tracks spanning

all seven classes with four from the 90s/2000s console core — the exact shape this analysis is

starved of. It takes the pool from 11 to ~21 and roughly halves the distance to the 30-row bar.

9. What the numbers do say, at descriptive tier

These are not discriminative findings and must not be quoted as such. They are the measured

profile of the eleven, usable as sanity bands for authored candidates.

axishit medianhit rangesibling mediansibling p10–p90AUC
time to hook0.57 s0.04 – 9.580.59 s0.22 – 5.530.49
hook range14.0 st0.67 – 49.017.0 st5.0 – 37.70.49
stepwise interval fraction0.2000.00 – 0.330.1330.00 – 0.400.58
leap fraction (≥5 st)0.4670.00 – 0.670.4000.00 – 0.800.49
gap-fill after leap0.0000.00 – 1.000.0000.00 – 1.000.55
interval 3-gram repetition0.3290.14 – 0.800.3590.00 – 0.680.54
contour turns per note0.4000.00 – 0.670.4670.13 – 0.670.45
mean section length5.86 bars3.4 – 31.04.07 bars2.0 – 11.40.54
section count117 – 2682 – 160.72
length149.7 s82.3 – 506.878.7 s10.0 – 225.70.72
tempo112.4 bpm89.1 – 172.3123.1 bpm99.4 – 152.00.34
melodic salience0.7710.00 – 0.900.7540.38 – 0.910.42
windowed key agreement0.4740.14 – 0.930.5000.23 – 1.000.52
silence budget1.9%0.9 – 3.62.8%0.8 – 23.10.45
dynamic range12.5 dB5.2 – 25.814.4 dB7.1 – 76.40.38

Four readings worth carrying into rung 2, each stated with its own confidence:

after a leap, hook compass, contour turn rate, interval 3-gram and 4-gram repetition all sit

between AUC 0.45 and 0.58. The most economical explanation is not that these do not matter — it

is that the album siblings already have them. These are well-written game scores throughout;

the corpus row is not the only competent track on its record. Whatever the difference is, this

card does not measure it. That is the most important sentence in this document for the program.

median). Modest, uncorrected, and one of the few effects with no length dependence (ρ = −0.13).

0.453) — steadier, less gappy, more continuous than their siblings. Consistent with music built

to loop under play rather than to punctuate a scene.

spectral_band_proxy method, not a finding, and any arrangement-width claim from these cards is

void until a real instrument roster exists. The cards already declare the lane method is not a

roster; this is that declaration coming due.

10. Honest register — what this is and is not

a validated theory, and nothing here is a MASTERPIECE_STANDARD band.

about a third of the time. "No formula was found" and "no formula exists" are different

statements and only the first is supported.

bar. Any future pass that reports a single feature's raw AUC without that bar is reporting noise.

not 0.648.

measured is whether the tracks that have it share anything the cards can see, and at this N

the answer is no. MASTERPIECE_PROGRAM §8 already said the outcome cannot be measured; this pass

is the first evidence about how far the proxies fall short.

914 measured tracks and 3 of the 11 hits. Twelve of the fifteen are 2000-or-later releases.

11. Reproduction

regression teeth for defects this pass actually shipped and its own controls caught: a planted

signal that the first fixture failed to plant, and a power model that treated one track's

comparisons against 145 siblings as 145 independent coin flips, collapsing per-hit spread from

0.25 to 0.034 and making every power answer a fantasy. Restoring that model drops the control's

statistic to 0.034 against a required window of 0.21–0.29, so the teeth are armed by mutation

rather than by assertion.

— the full derivation, about a minute, deterministic under seed 20260807.

held-out passes, the complete power curve and the per-class descriptive block.

12. What rung 2 should do with this

synthesis; the synthesis does not yet have a signal to floor it on. Building the predictor now

would fit 103 axes to 11 tracks and produce a confident number with no out-of-sample meaning.

enrichment, labelled necessary-not-sufficient.

control: delayed normalized self-recurrence, longer unbroken contour runs, and composed rest

(X_rest_grammar, the only pre-registered hypothesis that strengthened under length matching).

lineage the whole bar rests on — SNES, N64, Kondo, Zelda — contributes zero measured tracks.

Generated by harness/site/structure_site.py — the URL path is the repo path. review root