PROTOTYPE_VERDICT.md

music/pass2/PROTOTYPE_VERDICT.md

Pass 2 prototype — the verdict

Honest tier: STRUCTURE. Nobody has heard these tracks — not Josh, not the director, not the lane

that made them. Every claim below is either a property of the written plan or a number measured off

the rendered audio by the same rulers the corpus is measured with. There is no listening verdict

here and none should be read into it. Nothing from this lane goes on the review queue; publishing is

the director's act.

1. What was planned

One theme for REGION_FLORES_ISLAND, realised as three cues that are transformations of it rather

than three pieces. The theme is authored once and shared: one head cell, one mode, one pokok, one

harmony frame, one care block, one derivation. pass2_flores.py --check verifies that sharing

rather than asserting it.

VERBATIM, licensed +9 leap and all, which is what

T0_Theme_Registry :: REGION_FLORES_ISLAND.transformation_plan requires ("the cell must QUOTE the

parent head at least once or the payoff has no soil to grow in"). Notes five and six are the

region's own answer, stepping down off the leap. Checked in self_test against

author_head_cell.THEMES['JOURNEY_WORLD'] rather than transcribed.

reference-composition choice, not an ethnographic claim. It contains every note of the parent head.

rejected SLICE_T1_EXPLORATION_FLORES. PASS2_FLORES_CACI (BATTLE, 148 bpm, 4-bar cycle)

directly comparable to SLICE_T2_BATTLE_CACI. PASS2_FLORES_WATERFALL (SANCTUARY, 66 bpm, 2-bar

cycle), the sparse third variant, serving MC_0051.

Plans: build/audio/pass2/plans/*.json. All three pass all thirty-seven validation predicates.

2. What was rendered

Audio on D:/audio/pass2/<cue>/ — one mixdown plus one stem per part, kept on disk. The repo carries

the plans, the stem manifests, the formula cards and the measurements.

The chain ran end to end with no manual step: plan, validate, compile, one MIDI per layer, one sfizz

stem per layer, mix with declared gains and pans, formula card, planned-versus-measured.

THREE RENDER EPOCHS, and the middle one is a refutation rather than a step.

measurements/v2_sustain_and_harmony_variants/ with its own README, because the harmony

hypothesis was tested there and measured worse. The exploration and waterfall cues stayed at this

epoch, which is where they clear the bar.

3. The numbers

Emitted by python harness/music_gen/pass2_realise.py --md-table, read from the cards and

measurement docs on disk rather than transcribed.

FIRST RENDER

cuelen sparts plannedparts renderedsections plannedsections measuredboundaries/mindrops/minstructural yieldmelody span stsection mean scomposed restjoint boxtime to hook s
CACI194.637372182.169.560.22650.024.3211.85True1.683
REST240.0323219266.258.750.71452.09.238.75True0.07
WATERFALL145.5111115166.192.892.14357.339.092.89False0.058

CURRENT RENDER

cuelen sparts plannedparts renderedsections plannedsections measuredboundaries/mindrops/minstructural yieldmelody span stsection mean scomposed restjoint boxtime to hook s
CACI194.639392192.4710.480.23554.021.6210.48True1.695
REST240.0323219235.55.01.152.3310.435.0True0.081
WATERFALL145.5111115124.542.891.57148.012.122.89True0.07

The caci cue's middle epoch, kept as evidence at measurements/v2_sustain_and_harmony_variants/:

5 sections measured, 1.23 boundaries per minute, structural yield 0.118.

reference poolstructural yieldboundaries/mindrops/minsection count
corpus hits (n=28)0.765.095.5617 median
album siblings (n=686)1.005.774.91--
round 1, all rejected (n=32)0.222.7814.0811 median

Floor battery, python harness/music_gen/floor_instruments.py --cards build/audio/pass2/cards:

PASS2_FLORES_CACI

structural yield 0.235 -- reads like our generator

boundaries/min 2.47 drops/min 10.48 lane-adding raises/min 11.10

cycle_law pass max_cycles_unchanged=0.297, cycle_period_s=84.6, n_cycles=2.3

hook_presence no verdict time_to_hook_s=1.7, hook_strength=0.517

noise_structure pass crowding=0.434, separation=0.617, stacked_noise=0

PASS2_FLORES_REST

structural yield 1.100 -- reads like real released music

boundaries/min 5.50 drops/min 5.00 lane-adding raises/min 7.75

cycle_law pass max_cycles_unchanged=1.06, cycle_period_s=10.2, n_cycles=23.5

hook_presence no verdict time_to_hook_s=0.081, hook_strength=0.511

noise_structure pass crowding=0.185, separation=0.911, stacked_noise=0

PASS2_FLORES_WATERFALL

structural yield 1.571 -- reads like real released music

boundaries/min 4.54 drops/min 2.89 lane-adding raises/min 2.89

cycle_law pass max_cycles_unchanged=2.23, cycle_period_s=8.22, n_cycles=17.7

hook_presence no verdict time_to_hook_s=0.07, hook_strength=0.496

noise_structure pass crowding=0.416, separation=0.96, stacked_noise=0

READ AGAINST THE TARGETS IN THE BRIEF, one by one and without rounding up.

3-8 band for the sparse class, so the sparse cue is three parts over its own band and that is

stated rather than absorbed.

max_cycles_unchanged planned at 1, 1 and 2; the cycle-law instrument measures 1.06, 0.297 and

2.23 and passes all three.

TRANSFORMED. MET in the plan: 7, 8 and 3 returns at 0.714, 0.625 and 0.667 transformed. Measured

time_to_hook_s 0.081 and 0.07 on the two cues that open on a solo carrier, against a corpus

median of 0.575 — and 1.695 on the caci cue, which is late and is a real miss.

crowding index 0.0 on all three, no stratum over 0.60 of summed gain. The noise-structure

instrument passes all three and reports no stacked noise.

rejection fell outside this band.

the waterfall at 12 and badly missed on the caci cue at 9. The strongest held-out rule in the

derivation is section_count >= 15, and one of three cues fires it.

median and at the album-sibling level. MISSED on the caci cue at 0.235, which is where round 1

sits, and the floor battery says so in exactly those words: "reads like our generator."

4. The sighted-structure read

Plan open beside measured card. What follows names every place they disagree.

4.1 What the render does realise, exactly

architecture's central claim and it is the one that a single-shot generator cannot make at all:

the layer count is a field that was written down, not a number inferred from a mixdown.

is arithmetic on the section map rather than a target handed to a model.

cue, 0.8947 to 0.85 on the waterfall. Parts enter across the whole cue rather than at the top.

that refuses a statement whose interval shape does not come back intact. That guard fired on the

first caci build, where the octave-up return was written for a trumpet too small to play it.

compiled output: the of-the-place layers emit exactly the mode's five pitch classes

{2, 4, 6, 9, 11} and nothing else; the five broker layers emit exactly {0, 3, 5} and nothing

else. The two sets are disjoint. The menace cannot have been made by darkening the ensemble's

material because it is not made of the ensemble's material.

/ 0.05 on the waterfall. No stratum over 0.60.

that open on a solo carrier, against a corpus median of 0.575.

4.2 What the render does not realise, and why

19 planned against 23 measured on the exploration cue, 15 against 12 on the waterfall, and 21

against 5 on the second caci render. The plan's sections are where the COMPOSER intends a

boundary; the card's are where the SIGNAL changes. Only some of the first become the second, and

the property that decides which is density swing, not harmonic motion. That is the prototype's

single largest finding and section 4.3 is about it.

planned at 24 and measured at 9; the two are different units — planned simultaneous PARTS against

measured spectral BANDS — and the floor battery's own layer census says so in its verdict line:

it "saturates near 8 streams where a real cue runs 24 to 40 parts, so it cannot adjudicate the

layer clause at all." The manifest is the only honest answer to "how many layers", which is

precisely why the architecture makes it a field.

doc pairs the planned cycle against the card's dominant_onset_period_s, which is the

subdivision, not the cycle: 10.4348 planned against 1.3003 measured. The right ruler exists and

agrees — instr_cycle_law reports cycle_period_s=10.2 and n_cycles=23.5 for that cue against

a planned 10.4348 and 23. The pairing in planned_vs_measured should be re-pointed at the

instrument; it is left as it stands and named here rather than quietly corrected.

counts PARTS WITH AN ENTRY rather than activation transitions, so a part that rests and returns is

counted once against the plan's two. The planned number is right and the measured one answers a

different question.

stratum-lanes holding three or more non-doubling pitched parts, which is zero by design; measured

is the card's mean band interlap. They are reported side by side as corroboration and neither

should be read as validating the other.

on the first caci render; seven against 20 on the exploration cue. A planned removal is one a

composer intended; a measured drop is any moment the envelope or the band count falls. Roughly

eight of every ten measured drops on the first render were sampler artefacts, and section 4.4 has

the timestamp evidence.

4.3 The finding: measured sections track density swing, not the section map

The exploration cue declared 19 sections and measured 26. The caci cue declared 21 and measured 8.

Same schema, same compiler, same theme, same harmony frame. The difference is not in the plans'

quality; it is that the exploration cue's active part count swings from 2 to 24 while the caci cue's

floor never fell below about ten.

That was a hypothesis after the first render. It was then tested twice, and the test is the reason

to believe it:

cue, so no two consecutive sections share a harmonic frame or a structural line. Reharmonisation

is on the enumerated colour-change device list, and this was the obvious repair. It measured

WORSE: 5 sections against 8, structural yield 0.118 against 0.226. In a 39-part mix the layers

carrying the harmony sit twenty decibels under the ones carrying the timbre, and a chroma-and-MFCC

self-similarity matrix reads the ensemble as one object however much its inner harmony moves.

improved decisively, from 0.714 to 1.100 structural yield. So the variants are not harmful in

general; they simply do not address what a dense cue's segmenter is reading.

down to a named few, each rebuilding on the next cycle — taking its planned density floor from ten

to one, mean 15.7, peak 31. Subtraction is the strongest event available anyway; this is Bolero's

promotion law and Barber's climax into silence, and it is a real composed removal with a real

re-entry rather than an analytical trick. Result in the table above.

THE THIN POINTS CONFIRMED THE DIRECTION AND DID NOT CLOSE THE GAP. Measured sections went 8 (first

render) to 5 (harmony variants) to 9 (thin points); structural yield went 0.226 to 0.118 to 0.235;

boundaries per minute 2.16 to 1.23 to 2.47. So the density lever moves the number the harmony lever

moved backwards, and it moves it in the right direction — but six single-cycle emptyings in a

thirty-cycle cue are not enough. The planned density profiles say why: the caci cue's floor is 1 but

its MEAN is 15.7 of a peak 31, while the exploration cue's mean is 13.6 of a peak 24. Between its

thin points the battle cue is still a wall, and it is the time spent at low density rather than the

existence of a low point that the segmenter integrates.

The next rung for this cue is therefore not more thin points but a genuinely two-phase form: the

whole ensemble moving register and losing half its parts for a stretch of cycles rather than for one,

which is what the class research prescribes for a boss cue anyway (each phase gets its own build and

its own release, and the last phase withholds the full harmonised statement). That is a composition

change of real size and it is the honest next step rather than a tuning pass.

4.4 The drop count was mostly a realisation artefact, checked against timestamps

The first caci render planned four removals and the card counted 31 drops. Read against the score:

The detector fires on d_db <= -6.0 OR d_lanes <= -1.5, so these are firing on the LANE arm — the

spectral band count sagging once per cycle. The cause is that a cue built almost entirely from

struck transients empties its high and low bands between strokes: sfizz stops a sample at note-off,

so a vibraphone bar written as a 0.45-beat note is cut after 0.45 beats while a real bar rings for

seconds.

The repair is in the realisation, not the composition, and it is simply what a real score does: let

ringing instruments ring, and hold a quiet sustained bed under the percussion. THE REPAIR WORKED AND

THE MEASUREMENT SAYS SO. With no change to its composition at all, the exploration cue's drop count

fell from 35 to 20 — 8.75 per minute to 5.00, against a corpus median of 5.56 — and its structural

yield rose from 0.714 to 1.100, clearing the floor and landing at the album-sibling level.

5. Is the architecture proved?

THE COMPOSE-THEN-REALISE CHAIN IS PROVED AS A MECHANISM. Thirty-two, thirty-nine and eleven planned

parts became thirty-two, thirty-nine and eleven rendered stems, at the planned lengths, at the

planned tempi, with the planned entrance spread and the head cell intact note for note. The layer

count is a field that was written down rather than a number inferred from a mixdown, and the care

line is a property of the compiled pitch classes rather than a sentence in a document. Neither of

those is available to single-shot generation at all.

More than that: the plan was FALSIFIED BY ITS OWN RENDER, three times, and each falsification named

a cause that could be tested. The section map turned out not to predict the measured section count.

The drop count turned out to be eight-tenths sampler artefact. The harmony hypothesis turned out to

be wrong and the density hypothesis turned out to be right in direction and short in magnitude. A

prompt cannot be wrong in any of those ways, because a prompt asserts nothing checkable. That

capacity to be shown wrong is the architecture's actual product, and it is why this is a result

rather than a pile of audio.

TWO OF THE THREE CUES CLEAR THE MEASURED BAR AND ONE DOES NOT. The exploration cue — the one

directly comparable to round 1's rejected SLICE_T1_EXPLORATION_FLORES — reads at structural yield

1.100 against a corpus median of 0.76 and round 1's 0.22, at 5.50 boundaries per minute against a

corpus 5.09, at 5.00 drops per minute against a corpus 5.56, inside all three joint-box axes, with

23 measured sections against round 1's exploration median of 9 and a hook at 0.081 seconds. The

waterfall cue reads at 1.571 yield and clears the box. The caci cue — the one comparable to the

rejected SLICE_T2_BATTLE_CACI — reads at 0.235, which is round 1's own territory, with a named,

tested, directionally-confirmed cause and a next step that is a real composition change rather than

a tuning pass.

A PROTOTYPE THAT LANDS THE ARCHITECTURE AND SOUNDS LIKE A SAMPLER IS A SUCCESS AT THIS RUNG; A

PROTOTYPE THAT PRODUCES A BEAUTIFUL MIXDOWN BY ABANDONING THE PLAN IS A FAILURE. This is the first

kind. The plan was followed exactly — the compiler refuses to depart from it, and the two places

where it tried to (an out-of-range head note, a silent part) are now hard refusals rather than

silent repairs.

AND THE SOUND IS NOT CLAIMED. A CC0 community sampler with one dynamic layer per note, no round

robins, no legato transitions, no expression curves and no room is not a scoring session, and this

lane cannot hear what it made. The timbre-critic gap the round-1 review sheet records — the

authored-plus-sampler route at 1.932 against the generated route at 0.659 — is real, is not

addressed here, and closes only with a commercial library or with players.

THE NEXT RUNG, in the order I would take it:

1. The two-phase form for the battle cue, described in section 4.3. It is the one open failure and

it has a named repair.

2. Re-point the cycle_period_s and n_entrances comparisons at the instruments that measure them

properly. Both are my rulers being wrong, not the music.

3. Expression: CC curves per part in the emitted MIDI. Every phrase in these renders is dynamically

flat within a note, which no human performance is, and sfizz supports the fix.

4. A room. The mixdown is a sum of dry stems; a convolution stage is small code and a licensing

question about impulse responses.

5. Then, and only then, the twenty-cue sweep for this region. The architecture's economic claim —

136 lines of theme amortised across twenty cues of 120 to 160 lines each — is untested at

anything past three, and three is a small n for a claim about twenty.

THE SINGLE BIGGEST RISK to the architecture is that section_count and structural_yield, the two

axes it is now being steered by, are measured off a self-similarity matrix that reads DENSITY far

more strongly than it reads harmony, melody or form. Optimising against them could produce cues that

lurch between full and empty because that is what the ruler rewards, which is a different failure

from round 1's and no better. The corpus band is a containment check and was never meant to be a

target. The mitigation is that every one of these numbers stays necessary-and-never-sufficient, and

that Josh's ear remains the first ear.

6. The care line, checked on the compiled notes

The registry's deny_register_list requires that the region's own register is never the threat

signal, and that the darkening operations travel with the antagonist material rather than with the

place. pass2_flores.py --check verifies four things, and the fourth matters most because it reads

the NOTES the compiler emitted rather than the labels the author wrote.

Twenty of the cue's layers are marked of the place; the overlap with the threat set is empty, and

the plan validator refuses a plan where it is not.

D E F# A B. The broker's five layers use {0, 3, 5} — C, E-flat, F natural, none of which the

mode contains. Measured on the compiled output rather than on the plan: the of-the-place layers

emit exactly {2, 4, 6, 9, 11} and nothing else; the broker layers emit exactly {0, 3, 5} and

nothing else.

where BE_0003 turns from the ritual rounds to the unmasking.

least one is octave_up. The ensemble answers higher, never darker. That is CH03_B10's own

resolution — the ritual victory with the broker exposed and no kill — as a compositional act

rather than as a note in a document.

The palette extension carries its own care argument, in the block appended to sfz_palette.py: the

world-instrument half of VCSL stays out; two added instruments that carry a living-tradition name

(conga, bongos) are marked as such and are not used by these plans; and the honest limits — no

pitched gong exists in either CC0 set, the tuning is 12-TET and a gong waning is not, the bamboo

flute is a baroque recorder, what is taken is the technique and not a timbre claim — are stated on

every plan rather than left to be discovered.

7. Reproduction

python harness/music_gen/pass2_sfz_probe.py --check

python harness/music_gen/pass2_sfz_probe.py --all

python harness/music_gen/pass2_plan.py --check

python harness/music_gen/pass2_realise.py --check

python harness/music_gen/pass2_flores.py --check

python harness/music_gen/pass2_flores.py --emit

python harness/music_gen/pass2_realise.py --all

python harness/music_gen/pass2_realise.py --md-table

python harness/music_gen/floor_instruments.py --cards build/audio/pass2/cards

Self-test totals at the time of writing: probe 4/4, plan 14/14, realise 15/15, flores 13/13. Every

one of those includes mutation controls that break the rule under test and check that it fires; a

rule that has never been made to fail is not armed.

Determinism was NOT re-proved for this chain. orchestrate.py established that two renders of the

same plan through the same sfizz invocation are bit-identical on this palette, and pass 2 imports

that invocation unchanged, so the claim is inherited rather than re-measured.

pass2_realise.py --determinism exists and has not been run on these cues; that is a gap and it is

named as one.

Generated by harness/site/structure_site.py — the URL path is the repo path. review root