NOISE_STRUCTURE_VALIDATION.md

music/exemplars/instruments/NOISE_STRUCTURE_VALIDATION.md

Noise versus structure — the stacked-noise detector, validated

Tier: MEASURED INSTRUMENT. Every number below is computed from formula cards and audio already on disk. Nothing was purchased and nothing was generated by this pass.
Instrument: harness/music_gen/instr_noise_structure.py (v1).
Machine record: build/audio/exemplars/instruments/NOISE_STRUCTURE_VALIDATION.json.
The law it measures, from Josh's round-1 grade (docs/spine/DECISIONS_PENDING_JOSH.md, commit da627060): the tracks carry dozens of layers, and there cannot just be stacked noise.
Legal posture: Analysis for understanding only. Audio lawfully obtained by Josh, read from disk, measured, and never copied, redistributed, or used as model input of any kind.
Audio tier provenance: the raw Tier-2 measurements were taken in a full decoding pass and REUSED here; every verdict, percentile and bar below was recomputed at the constants in force. Measurements are facts and do not change; verdicts are derived and are re-derived whenever a bar moves.

1. The one-paragraph answer

The instrument measures 1180 of 1180 formula cards on the card tier and 98 corpus tracks plus 32 generated candidates on the audio tier. On the card tier it does what it was built to do — it orders real released music the way an ear does, it never calls a track Josh named stacked noise, and its worst-window pointer lands inside a wall spliced into a control at a timestamp it was never told. What it does not do is see the defect Josh heard. The eighteen turned-down round-one candidates sit at a density-without-structure conjunction of 0.0415 against a corpus median of 0.2277 and a loved-set median of 0.2332: on every card-only axis the rejected tracks are LESS wall-like than the average shipped soundtrack, and 0 of 18 fire the verdict. That is a negative result and it is reported as one. Whatever "stacked noise" names in Josh's ear, it is not broadband band co-activity.

The audio tier answers the same way. Sensory roughness — the measurement two research lanes independently named as the cheapest high-value addition available to this program, and the only thing that can grade harmonies and clashes on a render at all — lands at AUC 0.4963 between the rejected round and the loved set. That is chance. The rejected tracks are exactly as rough as the music Josh has loved for decades, and they carry LESS literal noise content, not more — residual share 0.0664 against 0.1123. One measure did appear to separate them, peak_stability at AUC 0.813, and it did not survive its own control: held to the corpus's lossless stratum it falls to 0.5208, so it was reading file provenance rather than music. Section 9 says what the defect actually is and which instrument should own it.

2. What the two indices are

Crowding is the contrast between near-band and far-band co-activity: the geometric mean of how much energy is simultaneously alive across four octave bands or more, and how much of the adjacent-band cohesion survives that distance. A wall keeps its cohesion at distance; an arrangement of separate voices does not.

Separation is the mean of four corpus percentiles — melodic salience, voiced fraction, spectral peakiness and lane reserve — each of which asks whether there is anything in the texture a listener can take hold of. The cards' own register spacing was measured for this and is carried at weight zero: it takes nine distinct values across the whole corpus and reads 1.006 at both the fifth and the ninety-fifth percentile, so it cannot separate anything.

3. The corpus rows, by name

30 rows carry a usable card, measured at HEAD on disk. PATTERN_FINDINGS_V2 reported twenty-eight against 1,181 cards; the pool has moved since, and the numbers here are what the files say today rather than what the earlier document said. Crowding and separation are corpus-referenced; the conjunction is the minimum of its four legs' percentiles and its bar is 0.7536, the corpus's own ninety-eighth percentile of that statistic.

idtitleclasscrowdingseparationconjunctionwall scoreverdict
EX_053Cynthia Battle ThemeBATTLE0.9670.12840.67440.8954clean
EX_142Shinshu FieldEXPLORATION0.96670.22130.42810.8487clean
EX_145Go StraightBATTLE0.93590.22820.48280.7995clean
EX_006Chrono Trigger Main ThemeTITLE_CHARACTER0.92720.39150.5430.7058clean
EX_102Magus CastleTENSION0.92530.6280.07590.5849clean
EX_014Tekken 2 Attract Movie: Sound TrackTITLE_CHARACTER0.9070.40480.1650.6697clean
EX_007Frog's ThemeTITLE_CHARACTER0.88940.28960.6320.7005clean
EX_161Aria di Mezzo CarattereMELANCHOLIC0.88120.61130.23550.5278clean
EX_055The Cyber GrindBATTLE0.87320.12940.64430.7574clean
EX_147Fighting of the SpiritBATTLE0.86230.18230.61380.7157clean
EX_180Dreams DreamsCREDITS_TRIUMPH0.81890.40980.44580.5508clean
EX_110Want You GoneCREDITS_TRIUMPH0.79020.54840.30050.4503clean
EX_136Humming the BasslineEXPLORATION0.78310.41580.43590.51clean
EX_032Terra's ThemeEXPLORATION0.77780.58050.27950.4228clean
EX_064Ludwig, the Holy BladeBOSS0.7670.38210.31570.512clean
EX_162Schala's ThemeMELANCHOLIC0.73610.58620.35550.3865clean
EX_046BFG DivisionBATTLE0.73450.32980.35680.5135clean
EX_094Dark Creature PursuitTENSION0.72440.41860.15560.46clean
EX_113To Good FriendsCREDITS_TRIUMPH0.71770.82430.15350.2504clean
EX_033Corridors of TimeEXPLORATION0.6610.63610.23090.2974clean
EX_075Dancing MadBOSS0.64660.67180.09820.2686clean
EX_088An Unwavering HeartMELANCHOLIC0.61290.86190.09090.1527clean
EX_030Route 209EXPLORATION0.59180.64870.14320.2472clean
EX_082The Best Is Yet to ComeMELANCHOLIC0.57650.8030.12270.1622clean
EX_092Snake Eater (Ladder Climb)TENSION0.56440.57410.1150.2704clean
EX_017SwedenEXPLORATION0.56080.90970.00160.1008clean
EX_179Ending Theme (Balance is Restored)CREDITS_TRIUMPH0.52310.76280.0820.1596clean
EX_124Deus Ex Main Title (UNATCO)TITLE_CHARACTER0.35630.79960.03360.117clean
EX_018Subwoofer LullabyEXPLORATION0.35130.96430.00490.0343clean
EX_022Wind SceneEXPLORATION0.25370.63420.01950.1926clean

The most crowded track Josh named reaches 0.967 (EX_053, Cynthia Battle Theme). The highest conjunction any of them reaches is 0.6744, below the bar of 0.7536 — and the bar was set from all 1180 cards without reference to which of them were loved, so the loved set clearing it is a measurement rather than a construction.

4. The calibration case: BFG Division against Sweden

BFG Division is supposed to be a wall of sound and it measures like one: crowding 0.7345 against Sweden's 0.5608, with The Cyber Grind higher still at 0.8732 and Subwoofer Lullaby lowest at 0.3513. The index is doing its job. The verdict is what must not fire on them, and it does not: EX_046 conjunction 0.3568, EX_055 conjunction 0.6443, EX_017 conjunction 0.0016, EX_018 conjunction 0.0049. Density is not the defect. Density without structure is, and DOOM has structure — its peak-versus-valley contrast and its melodic salience hold up underneath the loudness, which is exactly what separates a wall of sound from a wall.

5. Does it see what Josh heard

The eighteen turned-down candidates and the fourteen further generations beside them are measured on the same instrument. On the card tier, 0 of 18 fire the verdict, and the AUC of a rejected track outranking a loved one on the conjunction is 0.1407 — the wrong side of chance. The card tier does not see it.

On the audio tier, every measure with both medians. An AUC above 0.5 means the rejected candidates read HIGHER on that measure than the tracks Josh loved; below 0.5 means lower. Neither direction is automatically the defect — the reading is in the column, not in the ranking.

measureAUC reject over lovedloved-set medianreject mediann lovedn reject
peak_stability0.8130.49870.62233018
residual_share0.28430.11230.06643018
wall_score0.30.38290.27873018
peak_structure_score0.3524.580523.68353018
masked_fraction0.37220.08970.08713018
onset_coherence0.52220.62150.61543018
voice_independence0.47780.37850.38463018
roughness_mean0.49630.13310.13253018

On the full audio verdict, 0 of 18 rejected candidates fire STACKED NOISE.

The confound that has to be removed before any of that is read

Every generated candidate is a lossless wav and most of the corpus is a lossy rip, so any measure that differs between those populations separates the rejects from the loved set for a reason that is not musical. The comparison is therefore re-run against the corpus's lossless stratum alone — 16 tracks, same file format on both sides.

measureAUC against all 30 lovedAUC against the lossless stratumseparation lost
peak_stability0.8130.52080.2921
residual_share0.28430.3090.0248
wall_score0.30.31250.0125
peak_structure_score0.350.45830.1083
masked_fraction0.37220.184-0.1882
onset_coherence0.52220.6111-0.0889
voice_independence0.47780.3889-0.0889
roughness_mean0.49630.4097-0.0866

A bitrate sweep on one generated track moved peak_stability by 0.0072 from 320 kbps to 64 kbps (0.6237 to 0.6165) and roughness by 0.0002, so the ENCODER is not the mechanism. The lossless stratum of this corpus is modern FLAC rips -- Minecraft, ULTRAKILL, Streets of Rage 2 Remastered -- and the lossy stratum is older rips of older music, so the split is a PROVENANCE proxy rather than a compression artefact. Either way it confounds the comparison and either way stratifying is the fix.

The worst window per rejected candidate — the stretch a composer would be pointed at, which is the instrument's most actionable output whatever the verdict says:

candidateworst windowroughnessresidualpeak structure dBmasked
SLICE_T1_EXPLORATION_FLORES_a_4101135.76–39.75 s0.15270.071222.4870.0687
SLICE_T1_EXPLORATION_FLORES_a_410221.02–5.02 s0.15510.07823.3360.0761
SLICE_T1_EXPLORATION_FLORES_a_41033181.86–185.85 s0.13450.068323.1670.09
SLICE_T1_EXPLORATION_FLORES_a_41044208.42–212.42 s0.18540.044322.1080.0821
SLICE_T2_BATTLE_CACI_a_4201188.89–92.88 s0.16710.150522.0880.0846
SLICE_T2_BATTLE_CACI_b_42033189.01–193.0 s0.1890.120922.0870.0721
SLICE_T2_BATTLE_CACI_b_42044167.56–171.55 s0.16460.136822.3090.0838
SLICE_T3_TITLE_PROTAGONIST_a_4301128.61–32.6 s0.11530.077224.1360.0879
SLICE_T3_TITLE_PROTAGONIST_a_4302282.76–86.75 s0.14080.088723.7030.0873
SLICE_T3_TITLE_PROTAGONIST_a_43033168.58–172.57 s0.12090.062223.8150.0892
SLICE_T3_TITLE_PROTAGONIST_b_43033169.6–173.59 s0.13340.064524.1130.0941
SLICE_T4_MELANCHOLIC_HOME_AND_LOSS_a_44011111.36–115.36 s0.10120.026724.5750.0974
SLICE_T4_MELANCHOLIC_HOME_AND_LOSS_a_440224.09–8.08 s0.11040.023924.1150.0881
SLICE_T4_MELANCHOLIC_HOME_AND_LOSS_a_440330.0–3.99 s0.10360.020524.410.0628
SLICE_T4_MELANCHOLIC_HOME_AND_LOSS_a_44044198.21–202.2 s0.09030.012524.4920.0717
SLICE_T4_MELANCHOLIC_HOME_AND_LOSS_b_44011153.25–157.25 s0.11680.039924.2360.0909
SLICE_T4_MELANCHOLIC_HOME_AND_LOSS_b_4402235.76–39.75 s0.13160.059523.4820.0868
SLICE_T4_MELANCHOLIC_HOME_AND_LOSS_b_4403344.95–48.95 s0.12530.071823.6640.094

6. Does anything clear the multiplicity bar

Each corpus row is scored only against the tracks that shipped on its own record, and the bar is the best-of-bank deviation a within-album label permutation reaches by chance over 107 axes. That bar is 0.1868899450475604.

axiswithin-album AUCdeviationclears the bar
ns_crowding0.47640.0236no
ns_separation0.52910.0291no
ns_conjunction0.49760.0024no
ns_wall0.47550.0245no

No axis clears it. Stated plainly: none of the noise-versus-structure axes separates a loved track from its own album siblings. That is the honest result and it is the expected one — this instrument was built to detect a DEFECT, not to predict affection, and a corpus of released soundtracks does not contain the defect. A second null reported honestly is worth more than an inflated claim.

7. The two walls, and why the verdict has two conjunctions

The first version of the audio verdict ANDed roughness, residual content, unpeakedness and mutual masking into one statistic. Control C3 — four layers of broadband noise, the unambiguous wall — did not fire it, and the reason was physics rather than a threshold. Layered noise has low roughness, because its loudest partials land far apart and the Plomp-Levelt curve is near zero there, and it has low masking, because every critical band is equally loud so nothing sits under anyone's threshold. Demanding all four demanded a signature no real wall has.

So the verdict names two walls, because they are two different sounds. A NOISE WALL is high residual content and high flatness with no peaks standing above the floor; roughness and masking are not required. A CRAMMED WALL is high roughness and high mutual masking with no peaks; residual content is not required, because crammed voices are perfectly harmonic — they are simply all in the same place. Each carries its own bar at its own ninety-eighth percentile of the audio subset, and a track fires if either one does.

8. What it cannot see

The card-only crowding index cannot see register cramming. Four voices squeezed into one octave light fewer octave bands, not more, so the index falls rather than rises. Control C2c pins that behaviour with its number rather than deleting the control; cramming is caught in the audio tier, by roughness and by mutual masking, which is where the physics of it lives. A card-only run has not ruled cramming out.

Onset coherence separates an ensemble from a single block, not a wall from music. A solo piano reads near one and is not a wall; steady broadband noise reads near zero because it has no onsets to correlate. It is reported beside the verdict and is deliberately not one of its legs.

Melodic salience is exactly zero on ninety-one cards, two of them corpus rows — Frog's Theme and Shinshu Field. Frog's Theme has a melody. A zero there is the extractor declining, so this instrument treats it as missing rather than as maximal evidence of a wall, and any downstream reader should do the same.

Most of the corpus is lossy and its high-band figures carry codec artefacts, which the cards declare themselves. This instrument analyses at 22.05 kHz, so its whole working range sits below 11 kHz — under the lowpass of every mp3 in the pool — and it uses no brilliance or air term. Generations are lossless and exemplars are not, and this is how that comparison is kept honest.

Finally, and most importantly: this instrument answers whether a texture resolves into streams. It cannot hear whether those streams are worth hearing. A clean verdict here is necessary and nowhere near sufficient, and the hooks, the melody and the three-to-five cycle variation law are other instruments' business.

9. What "stacked noise" actually names, and which instrument should own it

This section replaces an open question with an answer, because three instruments have now looked for the defect and the pattern of where they failed is itself the evidence.

What was ruled out, and by what. The card tier says the rejected round is LESS spectrally crowded than the average shipped soundtrack and better separated than most of it. The audio tier says it is no rougher than the tracks Josh has loved for decades — sensory roughness lands at almost exactly chance — and carries LESS literal noise content, not more. The one measure that appeared to separate them turned out to be reading file provenance and collapsed to chance once the comparison was held to a single format. The sibling cycle-law instrument returns its own null: 17 of 18 rejected candidates pass Josh's five-cycle floor. Three independent instruments, three honest negatives.

So the defect is not spectral. It is not that too many things are sounding at once — by measurement, fewer things are sounding at once than in the music he loves. The strongest positive evidence anywhere in the program is temporal: the round-1 tracks carry roughly half the structural boundaries per minute of the corpus while chopping the texture two and a half times as often and firing lane-adding raises nearly twice as often. Motion without change. Constant activity, and almost nothing new ever arrives.

That reconciles every negative in this document. A listener asked to describe a texture that never stops moving and never goes anywhere reaches for "stacked noise" because it is what the experience feels like from the inside — layers piling up to no purpose. It is a description of futility, not of spectral density, and an instrument built to measure spectral density will keep returning clean verdicts on it forever. The phrase was a true report of a listening experience and a false lead about its cause, and taking it literally is what produced three nulls.

The owner is the cycle-law instrument, on the structural-change-per-texture-event axis. That is where the effect actually is, it is where a threshold would bind on something a generator can be steered by, and it is the same axis Josh named himself when he wrote that he dislikes repetitious loops lasting more than three to five cycles without a new instrument or catch or melody or beat drop. He described the disease and the cure in one sentence; this instrument was measuring a different organ.

This instrument's honest scope is the spectral half of the question, and the corpus says the generator is not currently failing it. That is worth keeping rather than retiring, for two reasons. The first is that it is a live regression guard: the moment the fix for the temporal defect is "add more layers," the spectral failure mode becomes reachable, and this is the instrument that would catch it — with a bar already calibrated so that DOOM passes. The second is that a null is only informative if it was capable of firing, and the controls prove this one is: it fires on layered broadband noise, it separates a crammed chorale from a spread one on the same notes, and it points at a spliced-in wall to the second. It looked, it could have found something, and it did not.

Generated by harness/site/structure_site.py — the URL path is the repo path. review root