music/exemplars/PATTERN_FINDINGS_V1.md
Tier: MEASURED-SAMPLE FINDINGS. Everything below is derived from the formula cards already
on disk. Nothing was purchased, nothing was generated, no audio was decoded by this pass. The
cards are read as data; the exemplar bytes never move.
Produced by: harness/music_gen/derive_patterns.py (22 controls, mutation-proved).
Machine record: build/audio/exemplars/PATTERN_FINDINGS_V1.json.
Consumer: docs/proposals/music/MASTERPIECE_PROGRAM.md rung 2 (the archetype synthesis
toward MASTERPIECE_STANDARD) and rung 3 (the nostalgia predictor, whose floor this pass was
supposed to help set).
Legal posture: analysis for understanding only, unchanged from the cards themselves.
Nine hundred and fourteen tracks are measured. Eleven of them are corpus rows — the tracks Josh
named, the ones that stayed loved for decades. Six hundred and one are their own album siblings:
same composer, same session, same production, same codec. On one hundred and three measured axes
— contour classes, interval distributions, phrase lengths, hook repetition structure, tonal
stability, arrangement, event grammar, dynamics, timbre — **not one axis separates the eleven
from their siblings by more than label-shuffling produces on the same data.** The best single
axis reaches a within-album AUC of 0.742; the permutation null over the same feature bank puts
the 95th percentile of chance at 0.798. A composite chosen fold-by-fold without ever seeing the
held-out track lands at AUC 0.586 ± 0.085 — about one standard error from chance. That is
the finding, stated at the tier the sample supports: **at N = 11, on these axes, there is no
demonstrated formula.** The study is also not powerful enough to rule one out; the arithmetic in
§8 says a real effect of 0.75 would be detected here only about a third of the time. The
actionable output is therefore §8's number, not a rule: thirty carded corpus rows would make
an effect of that size visible, against eleven today.
A naive study compares corpus rows against arbitrary music and "discovers" that beloved game
themes are longer, better produced and more melodic than whatever else is lying around. Every
one of those differences is era, budget, genre and mastering, not composition.
So the control group here is the album sibling. Each of the eleven corpus rows is scored only
against the other tracks that shipped on its own record. Composer, scoring session, mixing chain,
sample library, release year and audio codec are held constant by construction. The question the
pool can actually answer is the sharp one: *what separates the track people still hum from the
forty competent tracks that shipped beside it?*
Per-hit win rates are then averaged with equal weight per hit, so a 208-track Pokémon compilation
cannot drown a four-track Bloodborne EP.
| count | note | |
|---|---|---|
| Formula cards measured | 914 | build/audio/exemplars/formula_cards/*.json; repudiated/ excluded by construction |
| HITS — corpus rows, ACQUIRED or CARDED | 11 | the labelled positives |
| CONTROLS — siblings in an album containing a hit | 601 | across 10 albums |
| Measured but in an album with no hit | 301 | carried, never used as controls |
| Quarantined | 1 | EX_014 |
| Feature axes | 103 | scalars with ≥80% coverage and >2 distinct values |
EX_014 is dropped from both sides, not merely from the positives. The corpus defect_register
names it a known-false identity — the row says Tekken 2 "Opening Theme", the measured bytes are a
Pokémon X & Y opening movie — and the ledger's instruction is that its ACQUIRED is not evidence
of anything until the album is bought and inspected. A known-false label is no better as a
negative than as a positive.
| id | class | title | album measured against | siblings |
|---|---|---|---|---|
| EX_030 | EXPLORATION | Route 209 | Pokémon Diamond & Pearl Super Music Collection | 145 |
| EX_046 | BATTLE | BFG Division | DOOM (Original Game Soundtrack) | 30 |
| EX_053 | BATTLE | Cynthia Battle Theme | Pokémon Diamond & Pearl Super Music Collection | 145 |
| EX_055 | BATTLE | The Cyber Grind | ULTRAKILL | 22 |
| EX_064 | BOSS | Ludwig, the Holy Blade | Bloodborne: The Old Hunters | 4 |
| EX_088 | MELANCHOLIC | An Unwavering Heart | Pokémon X & Y Super Music Collection | 208 |
| EX_110 | CREDITS_TRIUMPH | Want You Gone | Portal 2 | 63 |
| EX_124 | TITLE_CHARACTER | Deus Ex Main Title (UNATCO) | Deus Ex | 39 |
| EX_136 | EXPLORATION | Humming the Bassline | Jet Set Radio Future | 16 |
| EX_142 | EXPLORATION | Shinshu Field | Ōkami OST Vol. 1 | 54 |
| EX_180 | CREDITS_TRIUMPH | Dreams Dreams | NiGHTS into Dreams | 20 |
Seven purpose classes exist in the corpus; six are represented here, four of them by one or two
tracks. No per-class claim in this document is a test — the per-class block in the JSON is a
descriptive read and is labelled as one.
With a hundred-odd axes and eleven positives, the best-looking feature will look good by luck.
The size of that luck is not guessed here, it is measured: within each album, re-draw at random
which tracks are "hits" (keeping each album's count), recompute every feature's within-album AUC,
and keep the best deviation from 0.5 across the whole bank. Four hundred draws give the
distribution a real finding has to clear.
| statistic of best-of-bank | AUC − 0.5 | value | equivalent AUC | |
|---|---|---|---|---|
| median of chance | 0.226 | 0.726 | ||
| 90th percentile | 0.280 | 0.780 | ||
| 95th percentile — the bar | 0.298 | 0.798 | ||
| 99th percentile | 0.330 | 0.830 | ||
observed best (n_drops) | 0.242 | 0.742 |
Features clearing the bar: zero of 103. The best real feature does not even reach the *median*
of what chance produces on this data by much. The coarseness is not an artifact of the method — it
is what eleven positives buy, most of it contributed by the small albums, where a single hit's
rank among four siblings is quantised to fifths.
| axis | within-album AUC | length-matched | ρ vs duration | hits above own-album median | uncorrected p | clears bar |
|---|---|---|---|---|---|---|
n_drops | 0.742 | 0.665 | +0.60 | 10/11 | 0.012 | no |
novelty_mean | 0.267 | 0.319 | −0.64 | 1/11 | 0.012 | no |
total_length_s | 0.723 | 0.665 | +1.00 | 9/11 | 0.065 | no |
repeat_sim_mean | 0.720 | 0.632 | +0.73 | 9/11 | 0.065 | no |
section_count | 0.718 | 0.607 | +0.81 | 9/11 | 0.065 | no |
recurrence_lag_s | 0.700 | 0.689 | +0.33 | 8/11 | 0.227 | no |
repeat_frac | 0.697 | 0.627 | +0.58 | 8/11 | 0.227 | no |
bars | 0.695 | 0.626 | +0.95 | 8/11 | 0.227 | no |
n_raises | 0.689 | 0.616 | +0.62 | 8/11 | 0.227 | no |
tempo_adjust_n | 0.676 | 0.614 | +0.77 | 8/11 | 0.227 | no |
contour_run_max | 0.672 | 0.602 | +0.18 | 9/11 | 0.065 | no |
recurrence_lag_norm | 0.662 | 0.703 | −0.17 | 8/11 | 0.227 | no |
bpm | 0.344 | 0.388 | −0.13 | 3/11 | 0.227 | no |
hook_vocab_size | 0.622 | 0.635 | −0.12 | 7/11 | 0.549 | no |
centroid_hz | 0.619 | 0.601 | −0.06 | 8/11 | 0.227 | no |
The top of that table is duration wearing a costume. The ρ column is the rank correlation of
each axis with track length, measured on the control pool alone with no hit label involved.
Everything above 0.69 AUC also sits above ρ = 0.58: more drops, more sections, more bars, more
tempo adjustments and higher self-similarity are all things that happen to a track simply by
lasting longer. Rescoring each hit only against siblings within ±35% of its own duration collapses
the whole leading group toward 0.61–0.67.
And the duration effect is itself curation, not composition. Josh named finished themes; game
albums are full of jingles, stingers, menu blips and fanfares. The control pool's 10th percentile
length is 10.0 s against a hit minimum of 82.3 s. "Corpus rows are longer than album filler" is a
true sentence about how the corpus was assembled, and it teaches the composer nothing.
Two things survive the length control rather than dying to it, and they are the only candidates
worth carrying forward:
recurrence_lag_norm rises under length matching, 0.662 → 0.703, with ρ = −0.17. This ishow far into the track the principal material waits before returning, as a fraction of the
track. The hits let the listener wait longer before the theme comes back around. Stated as a
testable rule: *a corpus row delays its first strong self-recurrence past 6.4% of its length,
where its siblings return at 5.0%.* Not significant here. Worth re-testing at N = 30.
contour_run_max, 0.672 → 0.602, ρ = +0.18. The longest unbroken run of same-directionmotion in the extracted hook: median 4 notes for the hits against 3 for the siblings, ranging to
12. The hits climb or fall further before turning. Directed melodic motion rather than zigzag.
Also not significant.
Tempo is not the confound. Rescoring against siblings within ±20% bpm changes essentially
nothing (n_drops 0.742 → 0.756, total_length_s 0.723 → 0.725). The pattern is not "battle
themes are fast and battle themes are over-represented."
Four compound axes were written into the extractor as stated theories before the first run, so
they carry a multiplicity price of four rather than of the whole bank and get their own null
(95th percentile deviation 0.223, i.e. AUC 0.723).
| hypothesis | axis | AUC | length-matched | verdict |
|---|---|---|---|---|
| The hook arrives inside the first phrase and repeats its own interval shape | X_hook_immediacy_economy | 0.531 | 0.430 | not separated |
| Singable = mostly steps, leaps answered by a step against them, compass held narrow | X_singability | 0.575 | 0.554 | not separated |
| Arrangement keeps lanes in reserve and spends them at structural seams | X_reserve_and_spend | 0.505 | 0.523 | not separated |
| Rest is composed — silence budget plus drop rate | X_rest_grammar | 0.531 | 0.659 | not separated |
None clears its own four-feature bar. X_rest_grammar is the only one that strengthens under
length matching, and it is the one to re-state and re-test with more rows.
Six hits train (EX_046, EX_053, EX_055, EX_064, EX_110, EX_136), five test (EX_030, EX_088,
EX_124, EX_142, EX_180). The six best training axes were turned into thresholded rules by Youden's
J on the training data, then scored once on the untouched test side.
| rule fitted on train | train AUC | test AUC | test TPR | test FPR | test lift |
|---|---|---|---|---|---|
time_to_full_texture_s ≥ 3.81 | 0.241 | 0.604 | 0.60 | 0.230 | 2.61× |
time_to_full_texture_bars ≥ 1.69 | 0.263 | 0.562 | 0.60 | 0.266 | 2.26× |
voiced_frac ≤ 0.137 | 0.243 | 0.599 | 0.20 | 0.193 | 1.04× |
centroid_hz ≥ 2352 | 0.745 | 0.468 | 0.20 | 0.215 | 0.93× |
ss_baseline ≥ 0.830 | 0.755 | 0.394 | 0.20 | 0.324 | 0.62× |
iv_mean_abs ≤ 3.07 | 0.265 | 0.707 | 0.00 | 0.246 | 0.00× |
Three of six rules reverse sign between train and test, and two of those had the *strongest*
training AUCs. That is the signature of selection noise, not of a weak-but-real effect. The two
rules that held (both forms of "takes longer to reach full texture") carry a 2.3–2.6× lift on the
held-out side, which is the single most encouraging number in this document and rests on three
firing tracks.
The strongest axis is re-chosen from scratch on the other ten hits, thresholded on them, and the
untouched eleventh is asked whether it fires.
The rules fire on held-out corpus rows *less often* than on the album filler they were built to
exclude. A single-feature hit rule selected this way is worse than nothing.
A predictor is used as several axes combined, so it is scored that way. In each fold the top five
axes and their directions are chosen on the other ten hits alone, z-scored within album, summed,
and the held-out track's percentile among its own siblings is recorded.
| composite | held-out AUC | SE | hits above own-album median | binomial p |
|---|---|---|---|---|
| over all 103 axes | 0.561 | 0.071 | 5/11 | 1.000 |
| over the 64 length-independent axes | 0.586 | 0.085 | 6/11 | 1.000 |
The length-independent bank is defined with no reference to the hit labels — the correlation
with duration is measured on the control pool alone. An earlier draft of this pass filtered the
hit-ranked top-25 instead, which leaks the held-out track into the candidate pool; that leak was
worth about +0.06 AUC and reported 0.648. It is named here because a leak found and closed is the
difference between a number and a claim.
0.586 ± 0.085 is the honest realized predictive power of the whole measured card on this pool.
Separation ("which side of a threshold") and containment ("inside which box") are different
questions, and the second is the one a generator can use even when the first fails. Each envelope
below is built from ten hits and validated by asking whether it contains the eleventh.
| axis | hit envelope | held-out hits contained | filler admitted | lift |
|---|---|---|---|---|
repeat_sim_max | 0.989 – 1.000 | 10/11 | 53.1% | 1.71× |
total_length_s | 82.3 – 506.8 s | 9/11 | 47.1% | 1.74× |
section_mean_s | 7.9 – 63.3 s | 9/11 | 51.3% | 1.60× |
bars | 39 – 320 | 9/11 | 53.6% | 1.53× |
tempo_adjust_n | 5 – 37 | 9/11 | 54.1% | 1.51× |
time_to_hook_frac | 0.0002 – 0.0189 | 9/11 | 57.4% | 1.43× |
The joint box over duration, repeat_sim_max and section_mean_s contains 7 of 11 held-out
hits while admitting 30.3% of filler — a 2.1× enrichment. That is a genuine, validated,
modest result, and it is a necessary-not-sufficient constraint: useful as a rejection filter
on generated candidates, useless as a target to optimise. A generated track outside the box is
unlike every measured corpus row; a generated track inside it has cleared a bar that a third of
album filler also clears.
Two things change as the corpus grows: the effect estimate gets more precise, **and the
multiplicity bar itself falls**, because the null of the best-of-bank statistic narrows. A power
model that holds the N = 11 bar fixed makes more data look useless, which is the opposite of the
truth.
The bar here is not assumed. K_eff — the number of effectively independent axes this bank
behaves like — is fitted so the simulated bar at N = 11 reproduces the observed permutation 95th
percentile (fitted K_eff = 80, calibration error 0.003), and only then extrapolated. Real
sibling counts are used, because a four-track EP quantises one hit's AUC to fifths and that
coarseness is most of the variance at N = 11.
| carded corpus rows | multiplicity bar (AUC) | detection of a true 0.65 | of 0.70 | of 0.75 | of 0.80 |
|---|---|---|---|---|---|
| 11 (today) | 0.791 | 3.5% | 12.5% | 33.5% | 58.8% |
| 16 | 0.758 | 6.2% | 17.7% | 47.5% | 76.2% |
| 22 | 0.720 | 9.8% | 36.0% | 74.3% | 95.0% |
| 30 | 0.684 | 29.7% | 64.2% | 93.2% | 99.3% |
| 40 | 0.664 | 38.5% | 79.2% | 99.0% | 99.8% |
| 55 | 0.639 | 62.3% | 95.0% | 99.5% | 100% |
| 75 | 0.623 | 77.5% | 99.3% | 100% | 100% |
| 100 | 0.605 | 93.5% | 100% | 100% | 100% |
| 140 | 0.585 | 99.8% | 100% | 100% | 100% |
Rows needed for 80% power: 30 for an effect of 0.75, 55 for 0.70, 100 for 0.65. The model is
generous by construction — a constant true effect on one axis, albums like the ones already
measured, no extra heterogeneity from a broader corpus — so read these as lower bounds.
Three specific gaps make the next acquisitions worth more than their count suggests:
SNES-era, zero N64, zero Zelda, zero Kondo. Twelve of the fifteen measured albums are 2000-or-later
releases and three are Pokémon compilations; only two albums containing a hit predate 2001. The
corpus's own EXPANSION wave exists precisely to cover that lineage and none of it is on disk.
If the thirty-year property lives anywhere in these numbers, it lives in the tracks not yet
measured.
can be stated until each class carries five-plus rows, and the MASTERPIECE_STANDARD is specified
in per-class bands.
that sets the multiplicity bar. Full albums beat single-track purchases for this pass.
The nearest concrete step is the BUY_SHEET's STARTER TEN at $64.59, which is ten tracks spanning
all seven classes with four from the 90s/2000s console core — the exact shape this analysis is
starved of. It takes the pool from 11 to ~21 and roughly halves the distance to the 30-row bar.
These are not discriminative findings and must not be quoted as such. They are the measured
profile of the eleven, usable as sanity bands for authored candidates.
| axis | hit median | hit range | sibling median | sibling p10–p90 | AUC |
|---|---|---|---|---|---|
| time to hook | 0.57 s | 0.04 – 9.58 | 0.59 s | 0.22 – 5.53 | 0.49 |
| hook range | 14.0 st | 0.67 – 49.0 | 17.0 st | 5.0 – 37.7 | 0.49 |
| stepwise interval fraction | 0.200 | 0.00 – 0.33 | 0.133 | 0.00 – 0.40 | 0.58 |
| leap fraction (≥5 st) | 0.467 | 0.00 – 0.67 | 0.400 | 0.00 – 0.80 | 0.49 |
| gap-fill after leap | 0.000 | 0.00 – 1.00 | 0.000 | 0.00 – 1.00 | 0.55 |
| interval 3-gram repetition | 0.329 | 0.14 – 0.80 | 0.359 | 0.00 – 0.68 | 0.54 |
| contour turns per note | 0.400 | 0.00 – 0.67 | 0.467 | 0.13 – 0.67 | 0.45 |
| mean section length | 5.86 bars | 3.4 – 31.0 | 4.07 bars | 2.0 – 11.4 | 0.54 |
| section count | 11 | 7 – 26 | 8 | 2 – 16 | 0.72 |
| length | 149.7 s | 82.3 – 506.8 | 78.7 s | 10.0 – 225.7 | 0.72 |
| tempo | 112.4 bpm | 89.1 – 172.3 | 123.1 bpm | 99.4 – 152.0 | 0.34 |
| melodic salience | 0.771 | 0.00 – 0.90 | 0.754 | 0.38 – 0.91 | 0.42 |
| windowed key agreement | 0.474 | 0.14 – 0.93 | 0.500 | 0.23 – 1.00 | 0.52 |
| silence budget | 1.9% | 0.9 – 3.6 | 2.8% | 0.8 – 23.1 | 0.45 |
| dynamic range | 12.5 dB | 5.2 – 25.8 | 14.4 dB | 7.1 – 76.4 | 0.38 |
Four readings worth carrying into rung 2, each stated with its own confidence:
after a leap, hook compass, contour turn rate, interval 3-gram and 4-gram repetition all sit
between AUC 0.45 and 0.58. The most economical explanation is not that these do not matter — it
is that the album siblings already have them. These are well-written game scores throughout;
the corpus row is not the only competent track on its record. Whatever the difference is, this
card does not measure it. That is the most important sentence in this document for the program.
median). Modest, uncorrected, and one of the few effects with no length dependence (ρ = −0.13).
0.453) — steadier, less gappy, more continuous than their siblings. Consistent with music built
to loop under play rather than to punctuate a scene.
max_lanes is 9 for hits and siblings alike. That is a measurement ceiling in the spectral_band_proxy method, not a finding, and any arrangement-width claim from these cards is
void until a real instrument roster exists. The cards already declare the lane method is not a
roster; this is that declaration coming due.
a validated theory, and nothing here is a MASTERPIECE_STANDARD band.
about a third of the time. "No formula was found" and "no formula exists" are different
statements and only the first is supported.
bar. Any future pass that reports a single feature's raw AUC without that bar is reporting noise.
not 0.648.
measured is whether the tracks that have it share anything the cards can see, and at this N
the answer is no. MASTERPIECE_PROGRAM §8 already said the outcome cannot be measured; this pass
is the first evidence about how far the proxies fall short.
914 measured tracks and 3 of the 11 hits. Twelve of the fifteen are 2000-or-later releases.
python harness/music_gen/derive_patterns.py --self-test — 22 controls. Two of them areregression teeth for defects this pass actually shipped and its own controls caught: a planted
signal that the first fixture failed to plant, and a power model that treated one track's
comparisons against 145 siblings as 145 independent coin flips, collapsing per-hit spread from
0.25 to 0.034 and making every power answer a fantasy. Restoring that model drops the control's
statistic to 0.034 against a required window of 0.21–0.29, so the teeth are armed by mutation
rather than by assertion.
python harness/music_gen/derive_patterns.py --perm 400 --json build/audio/exemplars/PATTERN_FINDINGS_V1.json— the full derivation, about a minute, deterministic under seed 20260807.
held-out passes, the complete power curve and the per-class descriptive block.
synthesis; the synthesis does not yet have a signal to floor it on. Building the predictor now
would fit 103 axes to 11 tracks and produce a confident number with no out-of-sample meaning.
enrichment, labelled necessary-not-sufficient.
control: delayed normalized self-recurrence, longer unbroken contour runs, and composed rest
(X_rest_grammar, the only pre-registered hypothesis that strengthened under length matching).
lineage the whole bar rests on — SNES, N64, Kondo, Zelda — contributes zero measured tracks.