DYNAMIC_ARC_VALIDATION.md

music/exemplars/instruments/DYNAMIC_ARC_VALIDATION.md

THE DYNAMIC ARC — the validation of instrument (f)

Tier: MEASURED INSTRUMENT. Produced by harness/music_gen/instr_dynamic_arc.py, which writes this file and its machine record together and writes nothing else.
The law it measures: A cue must BUILD to a peak that stands above its surroundings, HOLD there long enough to be heard as arrival rather than as a spike, and keep something worth following in its valleys, so the listener is lost neither at the top nor at the bottom.
CANON SUBORDINATION. The composition floor this instrument serves is Josh's ROUND-2 FLOOR ADDENDUM of 2026-08-07 night, at the tail of docs/spine/DECISIONS_PENDING_JOSH.md — "The songs dont build and hold with proper peaks and valleys that still never lose the listeners." That ruling is the authority; this file is a measurement of it, it changes no canon, and where it disagrees with canon, canon wins.
Machine record: build/audio/exemplars/instruments/DYNAMIC_ARC_VALIDATION.json.
Legal posture: analysis for understanding only; audio is read from disk, measured, and never copied, redistributed, or used as model input.

1. The one-paragraph answer

This instrument sees what Josh heard. The music he has loved for decades spends 0.407 of its running time in sustained high-energy HOLDS, with the longest single hold running 41.6 s; the three cues he turned down spend 0.092 and hold for 13.4 s. The AUC of the loved tracks over round 2 is 0.964 on the energy curve and 0.982 on the lane-count curve, which contains no loudness term at all. Round 2's peak distinction is HIGHER than the corpus's (0.718 against 0.324) — its peak is a spike in an empty field rather than an arrival, which is precisely a cue that does not build and hold. And the second finding, which is printed with equal weight: nothing here clears the within-album multiplicity bar. The best axis reaches a deviation of 0.1527 against a p95 permutation bar of 0.1750. Like structural_yield before it, this is a GENERATION-DEFECT DETECTOR, not a taste oracle: it separates our generator's output from real released music and it cannot tell a beloved track from competent filler on the same album.

2. What was measured, and against what

2.1 The mastering panel — the honesty bar, run rather than asserted

Our renders are dry sampler mixes and the corpus is mastered commercial audio, so any loudness-domain comparison across those two populations is suspect. A sibling lane watched an apparently-strong axis collapse when it turned out to be reading codec provenance. Every headline axis here is therefore computed on four curves, two of which contain no loudness term whatsoever.

curveloudness contenthold share, hitshold share, round 2AUC hits over round 2
lane simultaneitynone0.4760.0000.982
flow curve (primary)45% RMS0.4070.0920.964
complexity indexnone0.4050.1030.845
arc_curve_dbpure0.3950.1120.714

Read that panel in the right direction. If the finding were a mastering artifact, the pure-loudness curve would separate best and the loudness-free curves would collapse. The observed ordering is the exact opposite: the two curves with no loudness term are the strongest and the pure-loudness curve is the weakest of the four. The effect is in the arrangement — how many voices sound together and for how long — not in the gain staging. Control C11 additionally limits the same synthetic arch to within 3 dB of flat and confirms the primary axis moves less than the between-pool gap.

3. The corpus arc typology

poolsingle_archdouble_archterraced_ascentplateaurampflatarch_with_false_peakmulti_arch
corpus HITS140410018
album siblings6710582061237308
round 200000003
round 1014210222

multi_arch is the one type added beyond the seven the brief named, and it is named rather than hidden: it is the majority of the corpus, and folding three-or-more-peak cues into "double arch" would have been a false reading of most of the pool.

3.1 Per purpose class — DESCRIPTIVE READS, never tests

A battle arc is not a lament arc, and per-class is the only honest unit. Five of the seven classes carry four rows or fewer. Nothing in this table is a test and no threshold is derived from it.

classnbuildsholdsreleasesvalleyslongest holdhold sharevalley engagementdominant shape
EXPLORATION82.01.01.50.049.4 s0.3351.496multi_arch
BATTLE52.02.01.00.057.2 s0.3662.380multi_arch
MELANCHOLIC41.52.01.00.059.4 s0.5010.774plateau
CREDITS_TRIUMPH42.02.51.50.069.8 s0.5271.054multi_arch
TITLE_CHARACTER32.01.02.00.033.4 s0.2631.661multi_arch
BOSS25.03.03.50.052.4 s0.3451.238multi_arch
TENSION22.51.52.50.541.2 s0.3530.896multi_arch

4. The three round-2 cues, measured

cueshapebuildsholdsreleasesvalleyslongest holdhold sharelane hold sharevalley engagementpeak distinction
PASS2_FLORES_CACImulti_arch212123.8 s0.1230.0000.7240.524
PASS2_FLORES_RESTmulti_arch10210.0 s0.0000.0000.6150.718
PASS2_FLORES_WATERFALLmulti_arch413013.4 s0.0920.0000.6060.813
corpus HITS median2.02.01.50.041.6 s0.4070.4761.1990.324

The diagnosis is coherent across every axis and it is Josh's sentence in numbers. The cues reach a peak that is MORE distinct than the corpus's and then leave it immediately: the corpus spends 0.274 of its length within reach of its own peak and round 2 spends 0.071. On the lane-count curve, round 2's hold share is 0.000 — the number of simultaneously sounding voices never settles at a high level for even one window. That is "does not build and hold" in the most literal available sense.

Round 2's dynamic range is 25.1 dB against the corpus's 11.7 dB. Read alone that would look like a virtue. Read beside the hold share it is the defect: the range is spent on isolated spikes rather than on an arrival that is reached and then inhabited.

4.1 Intended against measured — the known-answer fixture

Round 2 is the only pool where we hold the SCORE as well as the render: each plan declares an intensity per section and a cycle length, so the INTENDED energy curve is on disk. The gap is this instrument's error and a real finding about the realiser at once.

cueintended shapeintended hold sharemeasured shapemeasured hold sharegap
PASS2_FLORES_CACImulti_arch0.000multi_arch0.123-0.123
PASS2_FLORES_RESTmulti_arch0.000multi_arch0.0000.000
PASS2_FLORES_WATERFALLmulti_arch0.061multi_arch0.092-0.031

THIS IS THE MOST ACTIONABLE FINDING IN THE DOCUMENT, AND IT DOES NOT SAY WHAT WE EXPECTED. The fixture was built to measure the realiser's error — how far the render drifts from the score. The intended hold shares are 0.000, 0.000, 0.061, against a corpus median of 0.407. The plans themselves contain almost no holds. The renders are not failing to deliver the written arc; they are delivering it faithfully, and the measured hold share comes out slightly HIGHER than intended in two of three cues. The defect is upstream of the realiser, in the plan.

The mechanism is visible in the schedule. One cue declares 21 sections across 195 s — roughly nine seconds each — and assigns a DIFFERENT intensity to almost every one (0.55, 0.62, 0.66, 0.70, 0.74, 0.72, 0.70, 0.78, 0.80, 0.84, 0.86, 0.28, …). No two consecutive sections sit at the same level, so the energy staircase never has a flat step and no HOLD can exist anywhere in the cue by construction. The composer wrote continuous motion and got continuous motion; what Josh asked for is arrival.

So the pass-3 instruction this instrument supports is narrow and checkable: plan the holds explicitly. Consecutive sections must be allowed to share an intensity, and the plan can be checked against this instrument BEFORE a single sample is rendered, because the intended curve is computable from the schedule alone. That turns the floor into a pre-render gate rather than a post-mortem.

5. What the pricing says

Every axis was priced against the same permutation multiplicity null the published derivations use: within each album, re-draw which tracks are hits at random, recompute every axis's within-album AUC, keep the best deviation, repeat 400 times. The p95 bar is 0.1750.

No axis clears that bar. The strongest within-album separator falls short of the price of having looked at this many axes. That is the honest result and it is stated here rather than buried: this instrument cannot rank real music by quality.

These are two different questions and the artifact keeps them apart. Within-album prediction asks whether an axis can pick the beloved track out of its own album — it cannot. Hits-over-round-2 asks whether an axis separates real released music from our generator's output — it does, at 0.982 on the loudness-free curve. The director asked the second question. The first is reported anyway, because an instrument that answered only the flattering question would be worth nothing.

axiswithin-album AUCdeviationclears p95hitssiblingsround 2AUC hits over round 2
hold_share0.5340.0344no0.4070.3640.0920.964
longest_hold_s0.6000.1004no41.61033.43013.3800.952
n_holds0.6530.1527no2.0001.0001.0000.869
valley_engagement0.5280.0283no1.1991.4680.6150.857
peak_distinction0.5500.0502no0.3240.2620.7180.119
sim_hold_share0.5230.0227no0.4760.4420.0000.982
sim_longest_hold0.6290.1288no64.04039.6300.0000.982
flow_p100.5160.0157no0.3520.4140.0750.952
dynamic_range_db0.4140.0865no11.71511.93025.0700.083

6. The controls

15 of 15 control arms green. Every arm plants an answer the instrument must recover, and two mutation arms break the measurement to prove the arms are load-bearing.

7. A live defect in the formula-card feature bank

tempo_and_time_grammar.groove.onset_rate_per_s is pinned and unusable as a pulse or density proxy. Flagged by a peer lane on 2026-08-07 and verified here independently against all 1180 cards: it runs 10.7040 to 10.8900, a span of 1.73% of its own median, with twenty distinct values at two decimal places and a correlation of −0.038 against texture_and_density.note_rate_per_s. At the cards' analysis rate (22050 Hz, hop 512) the frame rate is 43.07 Hz and the pinned value is 43.07 / 4 = 10.77. It is measuring the onset detector's own output density, not the music. No axis in this file reads it; the PULSE term that would naturally have used it is computed from onset detection on the audio instead. Control C12 is the standing guard, and it is written to fire if the field is ever repaired so the pulse term can be revisited.

8. The energy floor, and a correction to the framing it arrived with

A peer lane relayed a further Josh ruling on 2026-08-07 — the serene-waterfall distinction, "a track that's peaceful and still uplifting" against one that puts you to sleep, and "we are building an ARPG ... needs to have energy across the board in every track. Scaled by theme as well." Provenance, stated plainly: that quotation was not in docs/spine/DECISIONS_PENDING_JOSH.md at the HEAD this file was written against. It is carried here as a peer relay, not as canon, and the canonical authority for this instrument remains the ROUND-2 FLOOR ADDENDUM. The capability was built regardless, because "never lose the listener in the valley" is already this instrument's brief.

The relayed framing proposed a single global floor of flow_p10 >= 0.34, derived from six hand-picked rows, and it does not survive the full pool. Measured across all 28 carded hits, the rows falling below 0.20 are not only the two Minecraft ambient tracks:

rowclassflow p10in genre?
EX_017 SwedenEXPLORATION0.050Minecraft — ambient, no combat loop
EX_018 Subwoofer LullabyEXPLORATION0.050Minecraft — ambient, no combat loop
EX_088 An Unwavering HeartMELANCHOLIC0.106yes
EX_124 Deus Ex Main Title (UNATCO)TITLE_CHARACTER0.137yes

2 of those 4 rows are squarely in genre — An Unwavering Heart (MELANCHOLIC), Deus Ex Main Title (UNATCO) (TITLE_CHARACTER) — so the low-energy tail is not an ambient-genre artifact that can be excluded. More decisively, 12 of the 28 measured hits sit below the proposed 0.34 floor, including classes whose own median is far under it: TENSION's median flow_p10 is 0.218 and MELANCHOLIC's is 0.265. A global 0.34 floor would reject a large share of the music Josh has loved for decades. This is exactly the failure MASTERPIECE_PROGRAM.md section 11.4 already ruled against — the class curve is the unit and a global scalar is banned. The floor is therefore published PER CLASS, as the hits' own 25th percentile, and it is a descriptive read at these row counts:

classnenergy floor (flow p10, hits p25)
EXPLORATION80.274
BATTLE50.502
MELANCHOLIC40.222
CREDITS_TRIUMPH40.293
TITLE_CHARACTER30.293
BOSS20.350
TENSION20.213

measure() emits timestamped SEDATION SPANS — the stretches where a cue's raw energy sits below its own class floor for 12 s or more — so a composer is pointed at a span rather than handed a scalar. Note also that flow_p10 is partly a measure of dynamic RANGE rather than of sedation: a cue with one huge climax and a quiet body scores low without being sedative. That is why the headline valley axis is ENGAGEMENT (is anything still moving down there) and not energy level, and why the two are reported separately.

9. What this instrument does not say

Generated 2026-08-08T01:37:29Z in 27.7 s.

Generated by harness/site/structure_site.py — the URL path is the repo path. review root