music/exemplars/PATTERN_FINDINGS_V3.md
**Tier: MEASURED-SAMPLE FINDINGS. Four axes clear a re-calibrated multiplicity bar, down from
six, and the rejection envelope the generator is using has changed two of its three axes.**
Everything below is derived from the formula cards already on disk. Nothing was purchased,
nothing was generated, no audio was decoded by this pass. The cards are read as data; the
exemplar bytes never move.
Supersedes: PATTERN_FINDINGS_V2.md (N = 28, six axes clearing). V2 is not withdrawn — its
method is unchanged and it remains the correct answer for the pool it had, exactly as V2 says of
V1. What changed is the pool, not the apparatus.
Produced by: harness/music_gen/derive_patterns.py (26 controls, all green, mutation-proved).
Machine record: build/audio/exemplars/PATTERN_FINDINGS_V3.json.
Consumer: docs/proposals/music/MASTERPIECE_PROGRAM.md rung 2 (the archetype synthesis) and
rung 3 (the nostalgia predictor); and directly, harness/music_gen/slice_exemplars.py, whose
KEEP/REJECT filter reads the joint envelope this document re-derives.
Legal posture: analysis for understanding only, unchanged from the cards themselves.
Start with what got weaker, because that is the reason this re-run was owed. Two positives
were added — EX_014 and EX_094, whose quarantines were retired by inspection on 2026-08-07 —
and on the larger pool two of V2's six clearing axes fall back below the bar, bars and
n_drops; no axis clears the 99th percentile at all, where two did at N = 28; the leave-one-out
composite falls from 0.744 to 0.707; the length-independent composite falls from an already
sobering 0.497 to 0.426, which is now below chance rather than at it; and the joint rejection
envelope replaces two of its three axes, which means every KEEP/REJECT verdict in the round-1
generated slice was taken under a filter this document supersedes. One thousand one hundred and
eighty tracks are measured; thirty are corpus rows against twenty-eight in V2; nine hundred and
twenty-seven are their own album siblings. The bar itself fell from 0.195 to 0.184, so the
narrowing is not the bar rising — it is the observed effects shrinking faster than the bar fell.
The core V2 finding survives in shape: **structural size — how many sections a track has, how long
it runs, how much it repeats itself, how little it renews — still separates a corpus row from its
own album at 0.71 held-out, and still does so in both measured eras.** But it survives with less
room, and one 21.5-second attract sting is responsible for essentially all of the movement.
| headline figure | V1 (N = 11) | V2 (N = 28) | V3 (N = 30) | conclusion changed |
|---|---|---|---|---|
| formula cards measured | 914 | 1181 | 1180 | no |
| corpus rows (HITS) | 11 | 28 | 30 | no |
| album siblings (CONTROLS) | 601 | 851 | 927 | no |
| 95th-percentile bar, as deviation | 0.298 | 0.195 | 0.184 | no — it keeps falling |
| axes clearing the bar | 0 | 6 | 4 | yes |
| axes clearing the 99th | 0 | 2 | 0 | yes |
| best observed axis | 0.742 n_drops | 0.750 section_count | 0.715 section_count | no |
| leave-one-out composite, all axes | 0.561 ± 0.071 | 0.744 ± 0.045 | 0.707 ± 0.053 | no |
| leave-one-out composite, length-independent | 0.586 ± 0.085 | 0.497 ± 0.046 | 0.426 ± 0.046 | yes |
| held-out single-rule recall vs sibling rate | 27.3% / 43.9% | 67.9% / 35.9% | 63.3% / 34.3% | no |
| frozen-split rules reversing, by AUC sign | 6 of 6 | 0 of 6 | 0 of 6 | no |
| frozen-split rules reversing, by held-out lift | 3 of 6 | 0 of 6 | 1 of 6 | yes |
| joint envelope recall / filler / enrichment | 63.6% / 30.3% / 2.10× | 89.3% / 44.8% / 1.99× | 86.7% / 54.9% / 1.58× | yes |
| joint envelope axes | length, repeat-peak, section mean | span, section mean, rest | rest, section count, drop depth | yes |
Five rows changed a conclusion. The two that matter most are the last two, and they are the
subject of section 8.
Each corpus row is scored only against the tracks that shipped on its own record — composer,
scoring session, mixing chain, sample library, release year and audio codec held constant by
construction. Per-hit win rates are averaged with equal weight per hit, so a 208-track Pokémon
compilation cannot drown a four-track Bloodborne EP. This is the only framing the pool can answer
honestly, and it is V1's framing and V2's framing verbatim. The feature bank is the same 103 axes.
Nothing in the apparatus was touched for this pass.
| V1 | V2 | V3 | note | |
|---|---|---|---|---|
| Formula cards measured | 914 | 1181 | 1180 | formula_cards/*.json; repudiated/ and retired/ excluded by construction |
| HITS — corpus rows, ACQUIRED or CARDED | 11 | 28 | 30 | the labelled positives |
| CONTROLS — siblings in an album containing a hit | 601 | 851 | 927 | across 17 albums with siblings |
| Measured but in an album with no hit | 301 | 301 | 223 | carried, never used as controls |
| Quarantined | 1 | 1 | 0 | the KNOWN_FALSE set is now empty |
| Feature axes | 103 | 103 | 103 | unchanged bank, so all three passes are directly comparable |
| Albums carrying a hit | 10 | 16 | 18 | two of them contribute no siblings |
| Hits usable for within-album statistics | 11 | 26 | 28 | EX_082 and EX_092 have no siblings |
The card count falling by one while the pool grew is not an error and it reconciles exactly. Two
byte-identical UNMATCHED__ twins were retired by acquire_exemplar.py — one of the Tekken 2
attract cue, one of Rayman Legends' "Dark Creature Pursuit" — because each was a second card over
the same audio as a newly-attached hit, and scoring a hit against a copy of itself pulls every
within-album AUC toward the tie value. That is minus two. EX_094 gained a card it did not
previously have, which is plus one. EX_014's false card was repudiated and rewritten at net zero.
The control pool grew by exactly 76 — the Tekken 2 album's 31 siblings and Rayman Legends' 45 —
which moved out of the carried-but-unused pool the moment those albums carried a hit.
Both were quarantined in V2 and both quarantines were retired on evidence rather than assumption,
which this pass verified against CORPUS.json rather than taking on trust.
EX_014 — Josh named a Tekken 2 "Opening Theme". The album carries no such track. It was alreadyon disk in full from the starter-ten sweep, and its only attract cue is track 01, "Tekken 2
Attract Movie: Sound Track". The prior attach, which measured a Pokémon X & Y opening movie, was
demoted with its reason logged before the new one landed, so the false claim is repudiated rather
than silently overwritten. The row now reads CARDED on a human's inspection.
EX_094 — Josh named "Escape from Vector", which is on no Rayman Legends release. It was re-keyedonto "Dark Creature Pursuit", the album's longest timed-chase cue and the TENSION reading of the
row, with four runners-up named on the row itself and override_open set.
Both are legitimate positives and this pass treats them as such. It is nevertheless a measured fact,
not an opinion, that EX_014 is a **21.5-second attract sting with 4 sections, 8 bars, 8 notes and
a 7-semitone melodic span** — and that the entire V2 finding is that corpus rows are structurally
large. Section 8 quantifies what that one row does.
EXPLORATION 8, BATTLE 5, TITLE_CHARACTER 4, CREDITS_TRIUMPH 4, MELANCHOLIC 4, TENSION 3, BOSS 2.
All seven purpose classes remain represented; TITLE_CHARACTER and TENSION each gained a row. Five
of seven are still below the five-row floor the MASTERPIECE_STANDARD's per-class bands need, and
BOSS at two is still the binding constraint. No per-class claim in this document is a test; the
per-class block in the JSON remains a descriptive read and is labelled as one.
The bar is not guessed and it is not inherited. Within each album, re-draw at random which tracks
are "hits" (keeping each album's count), recompute every feature's within-album AUC, and keep the
best deviation from 0.5 across the whole bank. Four hundred draws under seed 20260807 — the
same draw count and the same seed discipline V2 used, stated here so the comparison is exact — and
re-run from scratch on the new pool rather than carried.
| statistic of best-of-bank \ | AUC − 0.5\ | V1 (N = 11) | V2 (N = 28) | V3 (N = 30) | V3 as AUC | |
|---|---|---|---|---|---|---|
| median of chance | 0.226 | 0.151 | 0.141 | 0.641 | ||
| 90th percentile | 0.280 | 0.185 | 0.176 | 0.676 | ||
| 95th percentile — the bar | 0.298 | 0.195 | 0.184 | 0.684 | ||
| 99th percentile | 0.330 | 0.220 | 0.216 | 0.716 | ||
| max over 400 draws | 0.358 | 0.240 | 0.238 | 0.738 | ||
| observed best | 0.242 (n_drops) | 0.250 (section_count) | 0.215 (section_count) | 0.715 |
Features clearing the 95th-percentile bar: four of 103. Clearing the 99th: none. At N = 28 it
was six and two; at N = 11 it was zero and zero.
The bar moved down, from 0.195 to 0.184, and the reason is the one V1 §8 predicted and V2 §3
restated: the null distribution of a best-of-bank statistic narrows as the number of positives
grows, because each permuted draw averages over more hits. Two more positives is a small change and
it bought a small fall. What did not happen is the effects growing to meet it. The observed best
fell from 0.250 to 0.215 — more than the bar fell — and that is the whole arithmetic of losing two
axes.
| axis | within-album AUC | length-matched | tempo-matched | ρ vs duration | hits above own-album median | binomial p | clears bar |
|---|---|---|---|---|---|---|---|
section_count | 0.715 | 0.663 | 0.717 | +0.77 | 22/28 | 0.0037 | yes |
total_length_s | 0.700 | 0.670 | 0.710 | +1.00 | 23/28 | 0.0009 | yes |
novelty_mean | 0.313 | 0.360 | 0.309 | −0.57 | 8/28 | 0.036 | yes |
repeat_sim_mean | 0.685 | 0.636 | 0.681 | +0.69 | 23/28 | 0.0009 | yes (by 0.0009) |
n_drops | 0.676 | 0.634 | 0.693 | +0.57 | 23/28 | 0.0009 | no (by 0.0078) |
bars | 0.675 | 0.625 | 0.691 | +0.93 | 19/28 | 0.087 | no (by 0.0090) |
n_raises | 0.651 | 0.597 | — | +0.55 | 19/28 | 0.087 | no |
tempo_adjust_n | 0.648 | 0.612 | — | +0.70 | 21/28 | 0.013 | no |
breaks_per_min | 0.362 | 0.376 | — | −0.87 | 7/28 | 0.013 | no |
time_to_hook_frac | 0.365 | 0.403 | — | −0.52 | 8/28 | 0.036 | no |
note_count | 0.628 | 0.546 | — | +0.78 | 19/28 | 0.087 | no |
drop_depth_max_db | 0.373 | 0.390 | — | −0.05 | 9/28 | 0.087 | no |
silence_frac | 0.382 | 0.419 | — | −0.74 | 8/28 | 0.036 | no |
repeat_frac | 0.613 | 0.584 | — | +0.53 | 18/28 | 0.185 | no |
melody_span_semitones | 0.600 | 0.552 | — | +0.34 | 18/28 | 0.185 | no |
dynamic_range_db | 0.401 | 0.427 | — | −0.53 | 10/28 | 0.185 | no |
Do not read this table as "the set shrank from six to four." Read the margins. The six axes span
deviations from 0.2150 down to 0.1751 — a band 0.04 wide — and the bar at 0.1841 falls inside
it. repeat_sim_mean clears by 0.0009. n_drops misses by 0.0078. bars misses by 0.0090. These
are not four findings and two refutations; they are six axes smeared across a threshold, and at
this N the threshold is a coin toss for the middle four of them. The honest statement is that
section_count and total_length_s clear with room, and that the other four sit close enough to
the bar that their status will change again with the next two acquisitions in either direction.
novelty_mean still clears in the negative direction and is still the one non-size axis in the
group. Corpus rows sit below their siblings on mean structural novelty while sitting above them on
repeat_sim_mean; the pair says the same thing twice, and it is thematic return rather than
repetition — a track that keeps coming back to its own material instead of moving on.
Tempo is still not the confound, and this is the one control that got slightly stronger.
Rescoring against siblings within ±20% bpm (517 controls) moves section_count 0.715 → 0.717,
total_length_s 0.700 → 0.710, n_drops 0.676 → 0.693 and bars 0.675 → 0.691. Every size axis
holds or improves. "Battle themes are fast and battle themes are over-represented" has now failed
to explain the pattern at three separate values of N.
V1's §4 verdict was blunt and correct for its pool: the top of the table is duration wearing a
costume, and the duration effect is curation — Josh named finished themes while game albums are
full of jingles, stingers and menu blips. That remains true. The control pool still carries 162
tracks under 30 seconds and 270 under a minute, against 927 in total.
What changed is on the other side of the comparison. V2 could say the hit pool held "exactly one row
under a minute". It now holds two: EX_102 (Magus Castle, 28.6 s) and EX_014 (21.5 s). The
corpus is no longer a set of uniformly finished-length themes, and the axis that most directly
encodes finished length is the axis that suffers most.
| axis | all 927 controls | ≥ 30 s (765) | ≥ 60 s (657) | ≥ 90 s (498) |
|---|---|---|---|---|
section_count | 0.715 | 0.690 | 0.655 | 0.618 |
n_drops | 0.676 | 0.658 | 0.633 | 0.611 |
repeat_sim_mean | 0.685 | 0.659 | 0.625 | 0.592 |
novelty_mean | 0.313 | 0.349 | 0.368 | 0.390 |
bars | 0.675 | 0.643 | 0.614 | 0.571 |
total_length_s | 0.700 | 0.671 | 0.636 | 0.553 |
note_count | 0.628 | 0.597 | 0.577 | 0.545 |
tempo_adjust_n | 0.648 | 0.624 | 0.599 | 0.561 |
Read the last column. Raw duration is still the axis that dies — 0.700 down to 0.553, closer to
chance than V2's 0.581 — which is what "curation, not composition" looks like when you take the
stingers away. section_count, n_drops, repeat_sim_mean and novelty_mean again decay more
slowly and sit around 0.59–0.62 against siblings every bit as long as the hits, against V2's
0.62–0.65. The residual that is not length is real and it is smaller than it was.
This truncation is a robustness read, not a test. The permutation bar in §3 was calibrated on
the full control pool; a truncated pool has fewer and differently-distributed siblings, so its own
bar would sit somewhat higher and is not computed here. Nothing in this table is claimed to clear
anything.
| carried hypothesis | axis | V1 AUC | V2 AUC | V3 AUC | V3 rank | verdict |
|---|---|---|---|---|---|---|
| Delayed self-recurrence | recurrence_lag_norm | 0.662 | 0.505 | 0.500 | 103 / 103 | dead at chance, now last in the bank |
| Unbroken contour runs | contour_run_max | 0.672 | 0.536 | 0.541 | 55 / 103 | failed |
| Composed rest | X_rest_grammar | 0.531 | 0.424 | 0.445 | 40 / 103 | failed, still reversed |
recurrence_lag_norm is now the single worst-separating axis of 103, at 0.4997. Its descriptive
medians are hit 0.019 against sibling 0.084 — the same inversion V2 found, deeper. V1's best
surviving candidate is now the bank's floor. Nothing here argues for a fourth test.
A parallel lane reached the same verdict from the opposite direction on the same day, and the two
results should be read together. Its finding, recorded in the 2026-08-07 block appended to V2 §6 and
implemented in harness/music_gen/instr_hook_presence.py, is that this axis was never a clean test:
it runs on raw chroma and MFCC self-similarity, so a theme restated a fourth higher or at half speed
scored as new material rather than as a return, which is backwards for a leitmotif score. Rebuilt
with transposition and tempo invariance it fails again and harder — zero of fourteen metrics clear
that lane's own re-derived bar. **The axis is dead both as measured and as it should have been
measured**, which is a stronger retirement than either pass could issue alone.
| hypothesis | axis | V1 AUC | V2 AUC | V3 AUC | V3 length-matched |
|---|---|---|---|---|---|
| Hook arrives early and repeats its interval shape | X_hook_immediacy_economy | 0.531 | 0.508 | 0.477 | 0.448 |
| Singable = mostly steps, leaps answered, narrow compass | X_singability | 0.575 | 0.455 | 0.442 | 0.433 |
| Arrangement keeps lanes in reserve, spends them at seams | X_reserve_and_spend | 0.505 | 0.444 | 0.443 | 0.461 |
| Rest is composed — silence budget plus drop rate | X_rest_grammar | 0.531 | 0.424 | 0.445 | 0.505 |
Their own four-feature null puts the bar at deviation 0.150 (AUC 0.650). None comes near it. The
four theories written into the extractor before the first run are, at thirty positives, measurably
not what distinguishes these tracks.
A director derivation this sitting proposed structural yield — novelty-peak boundaries divided
by texture drops, over cards of 45 seconds or more — reporting AUC 0.911 separating the tracks Josh
named from eighteen rejected generated candidates, and 0.437 against those candidates' own siblings.
It was offered with the prediction that it would fail the within-album bar. It was added to the
feature bank and priced through permutation_null_best exactly like every other axis, alongside
boundaries_per_min. The prediction is correct and the axes fail, plainly.
| new axis | within-album AUC | deviation | bar | length-matched | tempo-matched | ρ vs duration | above own-album median | clears |
|---|---|---|---|---|---|---|---|---|
structural_yield | 0.462 | 0.038 | 0.184 | 0.476 | 0.447 | −0.04 | 11/26 | no |
boundaries_per_min | 0.467 | 0.033 | 0.184 | 0.536 | 0.467 | −0.71 | 14/28 | no |
structural_yield with no length gate | 0.495 | 0.005 | 0.184 | 0.480 | 0.482 | +0.16 | 11/28 | no |
Neither reaches a quarter of the bar. structural_yield is 0.038 against a required 0.184, and its
ungated form is 0.005 — the closest any axis in this document comes to exact chance. Both hold near
chance in both eras (structural_yield 0.488 sixteen-bit against 0.440 modern). Adding both to the
bank and re-running the four-hundred-draw null changes the bar not at all: median 0.141, 95th
percentile 0.184, maximum 0.238, identical to four decimal places, because neither axis ever wins a
permuted draw.
Two further facts belong on the record rather than in a footnote.
structural_yield would not have been admitted to the bank at all under the tool's own rule. Its 45-second gate leaves it undefined on 23% of the pool, and feature_bank requires 80%
coverage. It was priced here by explicit construction, not by admission.
45 seconds or more is scored against a length-truncated sibling pool, which is a different and
easier control group than every other axis faces. It fails anyway.
The director's reading of why is the correct one and this pass confirms it: a ratio of boundaries to
drops separates our generator from real music, not loved music from filler. Against album siblings —
real music by the same composer in the same session — there is nothing there. A director's finding
gets the same bar as anyone's, and this one does not clear it.
Sixteen hits train, fourteen test. The six best training axes were turned into thresholded rules by
Youden's J on the training data, then scored once on the untouched test side.
| rule fitted on train | train AUC | test AUC | test TPR | test FPR | test lift | precision lift |
|---|---|---|---|---|---|---|
section_count ≥ 14 | 0.766 | 0.656 | 0.50 | 0.191 | 2.62× | 2.54× |
repeat_sim_mean ≥ 0.8964 | 0.739 | 0.623 | 0.43 | 0.192 | 2.23× | 2.18× |
drop_depth_max_db ≤ −59.67 | 0.291 | 0.467 | 0.14 | 0.259 | 0.55× | 0.56× |
novelty_mean ≤ 0.6265 | 0.295 | 0.335 | 0.79 | 0.387 | 2.03× | 1.99× |
n_drops ≥ 17 | 0.702 | 0.647 | 0.57 | 0.215 | 2.65× | 2.57× |
bars ≥ 110 | 0.693 | 0.654 | 0.43 | 0.138 | 3.10× | 2.98× |
**Zero of six rules reverse sign, and one of six now fires on unseen corpus rows less often than on
filler.** Those two sentences are not in conflict; they are two different definitions of reversal,
and V2 used the second without naming it, so both are reported here and in the JSON.
0 of 6, and V3 reverses 0 of 6. Unchanged and good.
V2 failed 0 of 6, and V3 fails 1 of 6. drop_depth_max_db ≤ −59.67 carries a held-out lift of
0.55×. This is V2's own "three of six reversed at N = 11" statistic, and it has gone from zero
back to one.
The selection also got noisier in a way worth naming: two of the six axes chosen on the training
half are now negative-direction axes, where V2's six were all positive-direction size axes. Four
rules still carry held-out lifts of 2.0–3.1×, against V2's 2.7–3.9×.
section_count ≥ 15 survive at N = 30?It is the strongest held-out rule in the program and everything downstream leans on it, so it was
tested as written rather than as refitted.
| V2 (N = 28) | V3 (N = 30), rule refitted | V3, the shipped ≥ 15 scored as written | |
|---|---|---|---|
| fitted threshold | 15 | 14 | 15 |
| test TPR | 0.62 | 0.50 | 0.50 |
| test FPR | 0.160 | 0.191 | 0.159 |
| test lift | 3.85× | 2.62× | 3.15× |
| precision lift | 3.63× | 2.54× | 3.02× |
| axis test AUC | 0.783 | 0.656 | 0.656 |
The rule survives, and it survives better than the threshold Youden's J now prefers. At N = 30
the training half has sixteen hits and ≥ 14 wins J there on recall (train TPR 0.81 against 0.75),
but on the untouched test side ≥ 15 is the better rule on every metric that matters — same recall,
lower false-positive rate, 3.15× lift against 2.62×. Across the whole pool it fires on 19 of 30
corpus rows and 21.3% of the 927 siblings, a 2.98× lift.
What did weaken is its recall and its axis. Test TPR fell from 0.62 to 0.50, and section_count's
held-out AUC fell from 0.783 to 0.656. The rule is still the best single statement this program has;
it now catches half the unseen corpus rows rather than five-eighths. **Keep the threshold at 15. Do
not refit it to 14 on this pool.**
| V1 (N = 11) | V2 (N = 28) | V3 (N = 30) | |
|---|---|---|---|
| held-out corpus rows firing | 3 of 11 — 27.3% | 19 of 28 — 67.9% | 19 of 30 — 63.3% |
| mean sibling fire rate for the same rules | 43.9% | 35.9% | 34.3% |
| ratio | 0.62× — worse than nothing | 1.89× | 1.85× |
Essentially unchanged, and this is the most stable result in the document. The rule selected is
section_count ≥ 14 in 28 of the 30 folds.
In each fold the top five axes and their directions are chosen on the other twenty-nine hits alone,
z-scored within album, summed, and the held-out track's percentile among its own siblings recorded.
| composite | V1 held-out AUC | V2 held-out AUC | V3 held-out AUC | SE | above own-album median | binomial p |
|---|---|---|---|---|---|---|
| over all 103 axes | 0.561 ± 0.071 | 0.744 ± 0.045 | 0.707 | 0.053 | 22 / 28 | 0.0037 |
| over the length-independent axes | 0.586 ± 0.085 | 0.497 ± 0.046 | 0.426 | 0.046 | 9 / 28 | 0.087 |
The length-independent bank is defined with no reference to the hit labels — the correlation with
duration is measured on the control pool alone (|ρ| < 0.35), admitting 70 of 103 axes, the same
count and the same filter as V2. V1's leak stayed closed.
**0.707 ± 0.053 is the honest realized predictive power of the whole measured card on this pool, and
0.426 ± 0.046 is what is left of it once every axis correlated with duration is removed.** Both
numbers must be quoted together, and the second one is now worse than V2's honest headline rather
than merely as bad. At N = 28 the length-independent composite was exactly chance; at N = 30 it sits
below chance with nine of twenty-eight rows above their album median. It is not significantly
inverted — p = 0.087 — but the direction of travel across two re-runs is unmistakable, and the
axes those folds select (drop_depth_max_db, melody_span_semitones, dyn_arc_range_db,
hook_range_semitones) are exactly the ones that will not hold. **The composite works, the
composite is made of size, and stripping size does not leave a weaker predictor — it leaves
something that is beginning to point the wrong way.**
Separation ("which side of a threshold") and containment ("inside which box") are different
questions, and the second is the one a generator can use even when the first is weak. Each envelope
is built from twenty-nine hits and validated by asking whether it contains the thirtieth.
| axis | V3 hit envelope | held-out hits contained | filler admitted | lift | V2 lift |
|---|---|---|---|---|---|
X_rest_grammar | 2.68 – 17.43 | 28 / 30 | 71.6% | 1.30× | 1.33× |
section_count | 4 – 35 | 30 / 30 | 84.6% | 1.18× | 1.16× |
drop_depth_max_db | −101.86 – −10.73 | 28 / 30 | 79.0% | 1.18× | — |
breaks_per_min | 0.22 – 4.20 | 28 / 30 | 79.1% | 1.18× | — |
section_mean_s | 5.38 – 63.35 | 28 / 30 | 85.2% | 1.10× | 1.33× |
total_length_s | 21.5 – 1295.6 | 28 / 30 | 84.8% | 1.10× | 1.14× |
melody_span_semitones | 7.0 – 57.0 | 29 / 30 | 92.5% | 1.05× | 1.39× |
The joint box changed two of its three axes.
| V1 | V2 | V3 | |
|---|---|---|---|
| axis 1 | total_length_s 82.3 – 506.8 | melody_span_semitones 19.7 – 57.0 | X_rest_grammar 2.68 – 17.43 |
| axis 2 | repeat_sim_max 0.9889 – 0.9999 | section_mean_s 7.1 – 63.3 | section_count 4 – 35 |
| axis 3 | section_mean_s 7.9 – 63.3 | X_rest_grammar 2.68 – 17.43 | drop_depth_max_db −101.86 – −10.73 |
| held-out hits contained | 7 / 11 — 63.6% | 25 / 28 — 89.3% | 26 / 30 — 86.7% |
| filler admitted | 30.3% | 44.8% | 54.9% |
| enrichment | 2.10× | 1.99× | 1.58× |
Only X_rest_grammar survives from V2's box. melody_span_semitones fell from the top of the
envelope ranking to 17th, and section_mean_s from 2nd to 9th. As a filter the box got
worse in both directions at once: it now admits 54.9% of album filler instead of 44.8%, for
slightly *lower* recall, and its enrichment fell from 1.99× to 1.58×.
A leave-one-hit-out sensitivity attributes the change with no ambiguity. This is a robustness read
with no null of its own; the N = 30 bar of 0.184 is quoted only as a reference line.
| pool | axes above the 0.184 reference line | joint box top three |
|---|---|---|
| all 30 | section_count, total_length_s, novelty_mean, repeat_sim_mean | rest, section count, drop depth |
minus EX_014 (N = 29) | all six of V2's, bars and n_drops restored | span, rest, section mean — V2's box, to four decimals |
minus EX_094 (N = 29) | five — n_drops still falls out | section count, breaks per min, drop depth |
| minus both (N = 28) | all six of V2's | V2's box exactly |
The last row is a positive control and it lands: dropping both new positives reproduces V2's numbers
to four decimal places — section_count 0.7501, total_length_s 0.7323, repeat_sim_mean 0.7138,
n_drops 0.7099, bars 0.7177 — which confirms the derivation is stable and that the pool is the
only thing that changed. The second row is the finding: **remove EX_014 alone and V2's six axes
and V2's exact box both come back.**
EX_014 measures melody_span_semitones 7.0 against a previous hit minimum of 19.67, and
section_mean_s 5.375 against a previous minimum of 7.145. Widening those two lower bounds pushed
melody_span_semitones's filler admission from 69.5% to 92.5% and section_mean_s's from 70.0% to
85.2%, which collapsed both lifts and demoted both out of the top three. A single 21.5-second attract
sting, correctly attached, is the entire cause.
build/audio/generated/slice_v1/verdicts.json names V2's box as its _envelope_source, so every
verdict in the round-1 slice was taken under a filter this document supersedes. All 32 candidate
cards were re-judged under both boxes with the same rule the filter uses — reject if any box axis is
outside or unmeasurable.
That is the positive control for this count.
KEEP → REJECT, both rejected on drop_depth_max_db, whose deepest drop must reach −10.73 dB and
which measures −9.45 and −7.14 on the two candidates.
SLICE_T1_EXPLORATION_FLORES_a_41022 and SLICE_T1_EXPLORATION_FLORES_a_41033,and this is the part that matters: **they are ranks 1 and 2 of that theme and both carry
kept_for_review = true.** The EXPLORATION theme's three-candidate review shortlist loses its top
two picks and drops to two candidates. No other theme is affected.
Two flips of thirty-two is a small number. Two flips out of the twelve candidates actually shortlisted
for review, both at the top of their theme, is not. **The round-1 EXPLORATION shortlist should be
re-cut against the V3 box before anyone listens to it.**
X_rest_grammar earning a place in the joint box while failing badly as a separator remains the
separation-versus-containment distinction, not a contradiction: corpus rows occupy a narrow band of
composed rest without sitting high or low on it. The box remains a **necessary-not-sufficient
constraint**, and at 54.9% filler admission it is now a weak one.
Four sixteen-bit-era scores carry corpus rows — Chrono Trigger, FINAL FANTASY VI, Tales of Phantasia
and Streets of Rage 2 — thirteen of the thirty hits and 228 of the 927 siblings. Both new positives
are modern-era, so the split is now 13 sixteen-bit against 17 modern.
| axis | all 30 | 16-bit era (13 hits) | modern (17 hits) |
|---|---|---|---|
section_count | 0.715 | 0.762 | 0.675 |
total_length_s | 0.700 | 0.720 | 0.683 |
bars | 0.675 | 0.739 | 0.620 |
repeat_sim_mean | 0.685 | 0.692 | 0.679 |
n_drops | 0.676 | 0.663 | 0.688 |
novelty_mean | 0.313 | 0.358 | 0.275 |
note_count | 0.628 | 0.706 | 0.560 |
tempo_adjust_n | 0.648 | 0.639 | 0.657 |
| — | |||
recurrence_lag_norm | 0.500 | 0.376 | 0.607 |
contour_run_max | 0.541 | 0.417 | 0.648 |
X_rest_grammar | 0.445 | 0.328 | 0.547 |
structural_yield | 0.462 | 0.488 | 0.440 |
boundaries_per_min | 0.467 | 0.458 | 0.474 |
Both V2 readings reproduce.
magnitude in both halves. A structural-size effect that reproduces independently in 1994 Super
Famicom sample-ROM scores and in 2016 live-recorded soundtracks is not an artifact of production
era, budget or codec. This remains the strongest single result in the program, and V2's numbers
for the sixteen-bit half are essentially untouched (section_count 0.762 both times) because
neither new positive is sixteen-bit.
half and below 0.5 in the sixteen-bit half, and cancel to chance pooled. Unchanged from V2.
The leave-one-out composite folds split by era give a mean held-out percentile of 0.724 for the
sixteen-bit rows against 0.692 for the modern rows on the all-axes composite, against V2's 0.726
and 0.763. **The sixteen-bit half is unchanged to three decimals and the entire drop in the headline
composite is in the modern half**, which is where both new positives landed. On the
length-independent composite it is 0.479 sixteen-bit against 0.380 modern, both down from V2's 0.599
and 0.395, and both now below chance.
The lineage is still only half covered. Zero N64, zero Zelda, zero Kondo, zero Nintendo
first-party of any kind. It is the same gap V1 named, V2 named, and this pass names again.
Not discriminative findings unless marked as clearing the bar in §4. These are the measured profile
of the thirty, usable as sanity bands for authored candidates.
| axis | hit median (V2 → V3) | hit range | sibling median | sibling p10–p90 | AUC |
|---|---|---|---|---|---|
| section count | 17 → 16 | 4 – 35 | 9 | 3 – 18 | 0.715 |
| length | 196.6 → 186.1 s | 21.5 – 1295.6 | 96.5 s | 14.3 – 222.0 | 0.700 |
| structural novelty | 0.556 → 0.556 | 0.397 – 0.857 | 0.667 | 0.504 – 0.951 | 0.313 |
| mean section self-similarity | 0.902 → 0.901 | 0.580 – 0.952 | 0.854 | 0.534 – 0.921 | 0.685 |
| drops | 19.5 → 19 | 2 – 110 | 7 | 3 – 26 | 0.676 |
| bars | 114 → 110.5 | 8 – 820 | 53 | 8 – 138 | 0.675 |
| tempo adjustments | 13.5 → 13.5 | 1 – 89 | 6 | 1 – 21 | 0.648 |
| note count | 316 → 298 | 8 – 2128 | 123 | 8 – 369 | 0.628 |
| melodic span | 43.0 → 42.0 st | 7.0 – 57.0 | 31.0 st | 9.0 – 48.0 | 0.600 |
| hook range | 22.0 → 22.0 st | 0.67 – 52.0 | 17.0 st | 5.0 – 38.0 | 0.589 |
| interval leap fraction | — | 0.00 – 0.87 | — | — | 0.548 |
| mean section length in bars | — | — | — | — | 0.524 |
| gap-fill after leap | — | 0.00 – 1.00 | — | — | 0.516 |
| interval 3-gram repetition | — | 0.00 – 0.83 | — | — | 0.488 |
| time to hook | — | 0.04 – 61.6 s | — | — | 0.487 |
| stepwise interval fraction | — | 0.00 – 0.33 | — | — | 0.474 |
| melodic salience | — | 0.00 – 0.924 | — | — | 0.462 |
| windowed key agreement | — | 0.109 – 0.929 | — | — | 0.462 |
| contour turns per note | — | 0.00 – 0.667 | — | — | 0.459 |
| tempo | 120.3 → 117.5 bpm | 83.4 – 172.3 | 123.1 bpm | 95.7 – 152.0 | 0.416 |
| dynamic range | 11.7 → 12.23 dB | 4.4 – 68.2 | 13.67 dB | 6.6 – 67.3 | 0.401 |
| silence budget | 2.1% → 2.1% | 0.6 – 11.1 | 2.8% | 1.0 – 19.0 | 0.382 |
Four readings worth carrying into rung 2.
program.** Every axis the folk theory of hummability is built from — stepwise fraction, gap-fill
after a leap, hook compass, contour turn rate, interval 3-gram repetition, time to hook, melodic
salience — sits between AUC 0.46 and 0.59 at eleven positives, at twenty-eight and now at thirty.
Not one has moved meaningfully in three passes. The most economical explanation is unchanged:
the album siblings already have them. These are well-written game scores throughout, and the
corpus row is not the only competent track on its record. What the card can see that separates
them is size and return, not melodic craft.
roughly the same margins — nearly double the section count, double the length, double the bars,
nearly triple the drop count. Part of it is curation, and about 0.59–0.62 of it survives against
equally long siblings.
cleared anything and it should not be carried.
max_lanes is 9 for every one of the thirty hits, at eleven positives, at twenty-eight and atthirty**, and takes only the values 8 or 9 across all 927 controls — which is why it fails the
bank's own distinctness admission and is not one of the 103. **Three consecutive derivations have
now confirmed this is a measurement ceiling in the spectral_band_proxy method, not a finding.**
Any arrangement-width claim from these cards is void until a real instrument roster exists. The
cards' own declaration that the lane method is not a roster is overdue and should be written.
K_eff, the number of effectively independent axes this bank behaves like, is chosen so the
simulated bar at N = 30 reproduces the observed permutation 95th percentile, and only then
extrapolated. Real sibling counts are used.
| carded corpus rows | multiplicity bar (AUC) | detection of a true 0.65 | of 0.70 | of 0.75 | of 0.80 |
|---|---|---|---|---|---|
| 11 | 0.780 | 5.2% | 20.0% | 34.0% | 66.5% |
| 16 | 0.742 | 10.0% | 24.0% | 52.0% | 87.3% |
| 22 | 0.705 | 15.8% | 48.7% | 78.2% | 97.0% |
| 30 (today) | 0.684 (observed) | 30.2% | 62.3% | 90.0% | 100% |
| 40 | 0.652 | 47.7% | 87.5% | 98.5% | 100% |
| 55 | 0.634 | 61.0% | 98.0% | 99.8% | 100% |
| 75 | 0.612 | 88.2% | 99.8% | 100% | 100% |
| 100 | 0.602 | 97.5% | 100% | 100% | 100% |
| 140 | 0.588 | 99.3% | 100% | 100% | 100% |
Rows needed for 80% power: 16 for an effect of 0.80, 30 for 0.75, 40 for 0.70, 75 for 0.65. The
pool sits exactly at the 0.75 threshold, which is why the clearing set is unstable — four axes above
and two just below is what a study sitting on its own detection boundary looks like.
One caution the model earns on itself: **the fitted K_eff moved from 103 at N = 28 to 50 at
N = 30**, with a better calibration error (0.0013 against 0.007). That is a large swing from two
positives, and it means the extrapolated columns are softer than their decimals suggest. Read the
whole table as lower bounds, as V2 did, and read the fitted constant as unstable.
What the next acquisitions are worth, in priority order.
zero Zelda, zero Kondo, zero Nintendo first-party. §9 shows the era split is measurable and that
it already overturned three hypotheses. A third era is the strongest available test of whether the
surviving axes are era-invariant or merely invariant across the two eras now held.
the bar. The next two positives will move bars and n_drops back across it or push
repeat_sim_mean and novelty_mean out. Nothing built on the four-versus-six distinction should
be treated as settled; build on the ordering, which is stable across all three passes.
MELANCHOLIC 4, TITLE_CHARACTER 4. No per-class band can be stated until each carries five-plus
rows. BOSS is the binding constraint.
EX_082 and EX_092 sit on a four-track vocalcompilation and are excluded from every within-album statistic here. The Metal Gear Solid score
albums would recover two positives already paid for.
of the 30 hits between them; two Pokémon compilations carry 353 of the 927 controls. §8's
leave-one-hit-out sensitivity shows how much a single row can move this pool, and albums are
heavier than rows.
is a measured-sample derivation, not a validated theory, and nothing here is yet a
MASTERPIECE_STANDARD band.
correct reading is not that two findings were refuted but that six axes are smeared across a
threshold in a 0.04-wide band, with the bar inside it. repeat_sim_mean clears by 0.0009 and
n_drops misses by 0.0078. Treat the ordering as the finding and the membership as provisional.
with the old one.** Two of thirty-two verdicts flip, both KEEP → REJECT, and both were shortlisted
top-two picks in the EXPLORATION theme. This is the most consequential change in the document.
EX_014 — a correctly attached 21.5-second attractsting — restores V2's six clearing axes and V2's exact joint box. Removing both new positives
reproduces V2 to four decimals. That is a clean attribution and also a warning: at N = 30 a single
atypical positive can rewrite the generator's filter.
length-independent ones, against V2's 0.744 and 0.497. The single most quotable number in this
document is only honest when quoted with the second one, and the second one got worse.
drop_depth_max_db ≤ −59.67 carries a held-outlift of 0.55×. Zero of six reverse by AUC sign; one of six reverses by the lift definition V2's
own "three of six at N = 11" used.
section_count ≥ 15 survives and should not be refitted to 14. Its held-out recall fell from0.62 to 0.50 and its axis AUC from 0.783 to 0.656, but as written it still carries a 3.15×
held-out lift, better than the refitted threshold on every test-side metric.
structural_yield scores 0.462 within-album against a required 0.184 deviation, boundaries_per_min 0.467, and the ungated ratio 0.495. Neither would move the
bar if adopted. The 0.911 that motivated the proposal was measured against generated candidates,
which is a different question from the one this pool answers.
measured is whether the tracks that have it share anything the cards can see; the answer is yes,
on structural size, and on nothing else the card can see.
carrying 353 of 927 controls, two records carrying 11 of 30 positives.
their own.** They are robustness reads on a finding whose bar was set on the full pool, and the
0.184 line is quoted in them as a reference only.
python harness/music_gen/derive_patterns.py --self-test — 26 controls, all green. The count has grown from V2's 22: C12b asserts no quarantined id is an attached corpus row, C18b proves
no hit leaked into the controls by BYTES rather than by name, C18c proves that dedup armed by
construction with a synthetic twin and a one-byte-mutated sibling as its mutation arm, and C18d
proves the two historical twins left the live pool and were retired rather than deleted. C18c
is worth noting as a pattern: it previously censused the disk for real twins and went red when
acquire_exemplar.py retired them, which is a control expiring because the defect it proved was
fixed. It now makes its claim by construction and cannot expire.
python harness/music_gen/derive_patterns.py --perm 400 --json build/audio/exemplars/PATTERN_FINDINGS_V3.json— the full derivation, deterministic under seed 20260807, the same draw count and seed discipline
V2 used.
passes, the complete power curve and the per-class descriptive block. Its auxiliary key carries
the §5 truncation table, the §6c new-axis pricing, the §8 sensitivity and slice re-judgement, the
§9 era split, the reversal census under both definitions, and the lane-ceiling re-verification.
Everything under auxiliary is derived from the same public functions — load_cards → flatten
→ label_rows → within_album_auc / permutation_null_best / matched_ctrl / rule_scores —
and is labelled there as a robustness read where it has no null of its own. Everything outside
auxiliary is derive()'s own unmodified output.
len(duration_and_structure.novelty_peak_strength) divided by len(event_grammar.drops) for structural_yield, gated to cards of 45 s or more, and by minutes
of runtime for boundaries_per_min. Card drops lists and counts.drops agree on all 1180 cards,
which was checked rather than assumed. Folding both into derive_patterns.flatten is a one-line
change and is left to whoever next edits that file; they are priced here and they fail, so nothing
downstream depends on it.
verdicts flip and both were shortlisted top-two EXPLORATION picks. Re-point
slice_exemplars.py's FINDINGS constant at PATTERN_FINDINGS_V3.json, re-run filter and
sheet, and let _envelope_source name V3. Until that is done the review sheet is quoting a box
that no longer exists.
section_count, total_length_s, repeat_sim_mean and novelty_mean clear at N = 30; n_drops and bars sit
within 0.009 of the bar and cleared at N = 28. All six are the same story and their rank order is
stable across three passes. A band on any of them is defensible at the honest tier "clears or sits
at a calibrated multiplicity bar at N = 30, era-invariant across the two eras measured, partially
confounded with curation." A band that depends on the four-versus-six distinction is not.
section_count ≥ 15 as written. It is still the program's best single statement and itbeats the threshold Youden's J now prefers on every held-out metric.
at 28 as at 11. If the difference lives there, this card cannot see it, and the correct action is
a better feature — a real instrument roster, a real hook extractor — not a band.
recurrence_lag_norm is now the worst axis of103. They have been tested at the N their own power model asked for.
structural_yield or boundaries_per_min as corpus axes. They separate thegenerator from real music, which is a useful thing and a different thing. If a generator-detection
filter is wanted, build it as one and name it as one; do not let it enter a corpus derivation
where it measures 0.462.
composite of size axes, with 0.426 length-independent beside it. Any predictor built on this must
carry the second number, which is now below chance rather than at it.
just demonstrated that one row can rewrite the filter; an era can do more.