music/pass2/PROTOTYPE_VERDICT.md
Honest tier: STRUCTURE. Nobody has heard these tracks — not Josh, not the director, not the lane
that made them. Every claim below is either a property of the written plan or a number measured off
the rendered audio by the same rulers the corpus is measured with. There is no listening verdict
here and none should be read into it. Nothing from this lane goes on the review queue; publishing is
the director's act.
One theme for REGION_FLORES_ISLAND, realised as three cues that are transformations of it rather
than three pieces. The theme is authored once and shared: one head cell, one mode, one pokok, one
harmony frame, one care block, one derivation. pass2_flores.py --check verifies that sharing
rather than asserting it.
JOURNEY_WORLD's headVERBATIM, licensed +9 leap and all, which is what
T0_Theme_Registry :: REGION_FLORES_ISLAND.transformation_plan requires ("the cell must QUOTE the
parent head at least once or the payoff has no soil to grow in"). Notes five and six are the
region's own answer, stepping down off the leap. Checked in self_test against
author_head_cell.THEMES['JOURNEY_WORLD'] rather than transcribed.
reference-composition choice, not an ethnographic claim. It contains every note of the parent head.
PASS2_FLORES_REST (EXPLORATION, 92 bpm, 4-bar cycle) directly comparable to round 1's rejected SLICE_T1_EXPLORATION_FLORES. PASS2_FLORES_CACI (BATTLE, 148 bpm, 4-bar cycle)
directly comparable to SLICE_T2_BATTLE_CACI. PASS2_FLORES_WATERFALL (SANCTUARY, 66 bpm, 2-bar
cycle), the sparse third variant, serving MC_0051.
Plans: build/audio/pass2/plans/*.json. All three pass all thirty-seven validation predicates.
Audio on D:/audio/pass2/<cue>/ — one mixdown plus one stem per part, kept on disk. The repo carries
the plans, the stem manifests, the formula cards and the measurements.
The chain ran end to end with no manual step: plan, validate, compile, one MIDI per layer, one sfizz
stem per layer, mix with declared gains and pans, formula card, planned-versus-measured.
THREE RENDER EPOCHS, and the middle one is a refutation rather than a step.
measurements/v1_first_render/. measurements/v2_sustain_and_harmony_variants/ with its own README, because the harmony
hypothesis was tested there and measured worse. The exploration and waterfall cues stayed at this
epoch, which is where they clear the bar.
Emitted by python harness/music_gen/pass2_realise.py --md-table, read from the cards and
measurement docs on disk rather than transcribed.
FIRST RENDER
| cue | len s | parts planned | parts rendered | sections planned | sections measured | boundaries/min | drops/min | structural yield | melody span st | section mean s | composed rest | joint box | time to hook s |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| CACI | 194.6 | 37 | 37 | 21 | 8 | 2.16 | 9.56 | 0.226 | 50.0 | 24.32 | 11.85 | True | 1.683 |
| REST | 240.0 | 32 | 32 | 19 | 26 | 6.25 | 8.75 | 0.714 | 52.0 | 9.23 | 8.75 | True | 0.07 |
| WATERFALL | 145.5 | 11 | 11 | 15 | 16 | 6.19 | 2.89 | 2.143 | 57.33 | 9.09 | 2.89 | False | 0.058 |
CURRENT RENDER
| cue | len s | parts planned | parts rendered | sections planned | sections measured | boundaries/min | drops/min | structural yield | melody span st | section mean s | composed rest | joint box | time to hook s |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| CACI | 194.6 | 39 | 39 | 21 | 9 | 2.47 | 10.48 | 0.235 | 54.0 | 21.62 | 10.48 | True | 1.695 |
| REST | 240.0 | 32 | 32 | 19 | 23 | 5.5 | 5.0 | 1.1 | 52.33 | 10.43 | 5.0 | True | 0.081 |
| WATERFALL | 145.5 | 11 | 11 | 15 | 12 | 4.54 | 2.89 | 1.571 | 48.0 | 12.12 | 2.89 | True | 0.07 |
The caci cue's middle epoch, kept as evidence at measurements/v2_sustain_and_harmony_variants/:
5 sections measured, 1.23 boundaries per minute, structural yield 0.118.
| reference pool | structural yield | boundaries/min | drops/min | section count |
|---|---|---|---|---|
| corpus hits (n=28) | 0.76 | 5.09 | 5.56 | 17 median |
| album siblings (n=686) | 1.00 | 5.77 | 4.91 | -- |
| round 1, all rejected (n=32) | 0.22 | 2.78 | 14.08 | 11 median |
Floor battery, python harness/music_gen/floor_instruments.py --cards build/audio/pass2/cards:
PASS2_FLORES_CACI
structural yield 0.235 -- reads like our generator
boundaries/min 2.47 drops/min 10.48 lane-adding raises/min 11.10
cycle_law pass max_cycles_unchanged=0.297, cycle_period_s=84.6, n_cycles=2.3
hook_presence no verdict time_to_hook_s=1.7, hook_strength=0.517
noise_structure pass crowding=0.434, separation=0.617, stacked_noise=0
PASS2_FLORES_REST
structural yield 1.100 -- reads like real released music
boundaries/min 5.50 drops/min 5.00 lane-adding raises/min 7.75
cycle_law pass max_cycles_unchanged=1.06, cycle_period_s=10.2, n_cycles=23.5
hook_presence no verdict time_to_hook_s=0.081, hook_strength=0.511
noise_structure pass crowding=0.185, separation=0.911, stacked_noise=0
PASS2_FLORES_WATERFALL
structural yield 1.571 -- reads like real released music
boundaries/min 4.54 drops/min 2.89 lane-adding raises/min 2.89
cycle_law pass max_cycles_unchanged=2.23, cycle_period_s=8.22, n_cycles=17.7
hook_presence no verdict time_to_hook_s=0.07, hook_strength=0.496
noise_structure pass crowding=0.416, separation=0.96, stacked_noise=0
READ AGAINST THE TARGETS IN THE BRIEF, one by one and without rounding up.
3-8 band for the sparse class, so the sparse cue is three parts over its own band and that is
stated rather than absorbed.
max_cycles_unchanged planned at 1, 1 and 2; the cycle-law instrument measures 1.06, 0.297 and
2.23 and passes all three.
TRANSFORMED. MET in the plan: 7, 8 and 3 returns at 0.714, 0.625 and 0.667 transformed. Measured
time_to_hook_s 0.081 and 0.07 on the two cues that open on a solo carrier, against a corpus
median of 0.575 — and 1.695 on the caci cue, which is late and is a real miss.
crowding index 0.0 on all three, no stratum over 0.60 of summed gain. The noise-structure
instrument passes all three and reports no stacked noise.
rejection fell outside this band.
the waterfall at 12 and badly missed on the caci cue at 9. The strongest held-out rule in the
derivation is section_count >= 15, and one of three cues fires it.
median and at the album-sibling level. MISSED on the caci cue at 0.235, which is where round 1
sits, and the floor battery says so in exactly those words: "reads like our generator."
Plan open beside measured card. What follows names every place they disagree.
architecture's central claim and it is the one that a single-shot generator cannot make at all:
the layer count is a field that was written down, not a number inferred from a mixdown.
is arithmetic on the section map rather than a target handed to a model.
cue, 0.8947 to 0.85 on the waterfall. Parts enter across the whole cue rather than at the top.
that refuses a statement whose interval shape does not come back intact. That guard fired on the
first caci build, where the octave-up return was written for a trumpet too small to play it.
compiled output: the of-the-place layers emit exactly the mode's five pitch classes
{2, 4, 6, 9, 11} and nothing else; the five broker layers emit exactly {0, 3, 5} and nothing
else. The two sets are disjoint. The menace cannot have been made by darkening the ensemble's
material because it is not made of the ensemble's material.
/ 0.05 on the waterfall. No stratum over 0.60.
time_to_hook_s of 0.081 and 0.07 on the two cuesthat open on a solo carrier, against a corpus median of 0.575.
19 planned against 23 measured on the exploration cue, 15 against 12 on the waterfall, and 21
against 5 on the second caci render. The plan's sections are where the COMPOSER intends a
boundary; the card's are where the SIGNAL changes. Only some of the first become the second, and
the property that decides which is density swing, not harmonic motion. That is the prototype's
single largest finding and section 4.3 is about it.
est_layers_max wasplanned at 24 and measured at 9; the two are different units — planned simultaneous PARTS against
measured spectral BANDS — and the floor battery's own layer census says so in its verdict line:
it "saturates near 8 streams where a real cue runs 24 to 40 parts, so it cannot adjudicate the
layer clause at all." The manifest is the only honest answer to "how many layers", which is
precisely why the architecture makes it a field.
cycle_period_s COMPARISON USES THE WRONG RULER AND SHOULD BE READ AS VOID. The measurement doc pairs the planned cycle against the card's dominant_onset_period_s, which is the
subdivision, not the cycle: 10.4348 planned against 1.3003 measured. The right ruler exists and
agrees — instr_cycle_law reports cycle_period_s=10.2 and n_cycles=23.5 for that cue against
a planned 10.4348 and 23. The pairing in planned_vs_measured should be re-pointed at the
instrument; it is left as it stands and named here rather than quietly corrected.
n_entrances MEASURED PROXY IS WEAK. Planned 81 against measured 32, but the measured figurecounts PARTS WITH AN ENTRY rather than activation transitions, so a part that rests and returns is
counted once against the plan's two. The planned number is right and the measured one answers a
different question.
crowding_index PLANNED AND MEASURED ARE DIFFERENT DEFINITIONS. Planned is the fraction ofstratum-lanes holding three or more non-doubling pitched parts, which is zero by design; measured
is the card's mean band interlap. They are reported side by side as corroboration and neither
should be read as validating the other.
on the first caci render; seven against 20 on the exploration cue. A planned removal is one a
composer intended; a measured drop is any moment the envelope or the band count falls. Roughly
eight of every ten measured drops on the first render were sampler artefacts, and section 4.4 has
the timestamp evidence.
The exploration cue declared 19 sections and measured 26. The caci cue declared 21 and measured 8.
Same schema, same compiler, same theme, same harmony frame. The difference is not in the plans'
quality; it is that the exploration cue's active part count swings from 2 to 24 while the caci cue's
floor never fell below about ten.
That was a hypothesis after the first render. It was then tested twice, and the test is the reason
to believe it:
cue, so no two consecutive sections share a harmonic frame or a structural line. Reharmonisation
is on the enumerated colour-change device list, and this was the obvious repair. It measured
WORSE: 5 sections against 8, structural yield 0.118 against 0.226. In a 39-part mix the layers
carrying the harmony sit twenty decibels under the ones carrying the timbre, and a chroma-and-MFCC
self-similarity matrix reads the ensemble as one object however much its inner harmony moves.
improved decisively, from 0.714 to 1.100 structural yield. So the variants are not harmful in
general; they simply do not address what a dense cue's segmenter is reading.
down to a named few, each rebuilding on the next cycle — taking its planned density floor from ten
to one, mean 15.7, peak 31. Subtraction is the strongest event available anyway; this is Bolero's
promotion law and Barber's climax into silence, and it is a real composed removal with a real
re-entry rather than an analytical trick. Result in the table above.
THE THIN POINTS CONFIRMED THE DIRECTION AND DID NOT CLOSE THE GAP. Measured sections went 8 (first
render) to 5 (harmony variants) to 9 (thin points); structural yield went 0.226 to 0.118 to 0.235;
boundaries per minute 2.16 to 1.23 to 2.47. So the density lever moves the number the harmony lever
moved backwards, and it moves it in the right direction — but six single-cycle emptyings in a
thirty-cycle cue are not enough. The planned density profiles say why: the caci cue's floor is 1 but
its MEAN is 15.7 of a peak 31, while the exploration cue's mean is 13.6 of a peak 24. Between its
thin points the battle cue is still a wall, and it is the time spent at low density rather than the
existence of a low point that the segmenter integrates.
The next rung for this cue is therefore not more thin points but a genuinely two-phase form: the
whole ensemble moving register and losing half its parts for a stretch of cycles rather than for one,
which is what the class research prescribes for a boss cue anyway (each phase gets its own build and
its own release, and the last phase withholds the full harmonised statement). That is a composition
change of real size and it is the honest next step rather than a tuning pass.
The first caci render planned four removals and the card counted 31 drops. Read against the score:
The detector fires on d_db <= -6.0 OR d_lanes <= -1.5, so these are firing on the LANE arm — the
spectral band count sagging once per cycle. The cause is that a cue built almost entirely from
struck transients empties its high and low bands between strokes: sfizz stops a sample at note-off,
so a vibraphone bar written as a 0.45-beat note is cut after 0.45 beats while a real bar rings for
seconds.
The repair is in the realisation, not the composition, and it is simply what a real score does: let
ringing instruments ring, and hold a quiet sustained bed under the percussion. THE REPAIR WORKED AND
THE MEASUREMENT SAYS SO. With no change to its composition at all, the exploration cue's drop count
fell from 35 to 20 — 8.75 per minute to 5.00, against a corpus median of 5.56 — and its structural
yield rose from 0.714 to 1.100, clearing the floor and landing at the album-sibling level.
THE COMPOSE-THEN-REALISE CHAIN IS PROVED AS A MECHANISM. Thirty-two, thirty-nine and eleven planned
parts became thirty-two, thirty-nine and eleven rendered stems, at the planned lengths, at the
planned tempi, with the planned entrance spread and the head cell intact note for note. The layer
count is a field that was written down rather than a number inferred from a mixdown, and the care
line is a property of the compiled pitch classes rather than a sentence in a document. Neither of
those is available to single-shot generation at all.
More than that: the plan was FALSIFIED BY ITS OWN RENDER, three times, and each falsification named
a cause that could be tested. The section map turned out not to predict the measured section count.
The drop count turned out to be eight-tenths sampler artefact. The harmony hypothesis turned out to
be wrong and the density hypothesis turned out to be right in direction and short in magnitude. A
prompt cannot be wrong in any of those ways, because a prompt asserts nothing checkable. That
capacity to be shown wrong is the architecture's actual product, and it is why this is a result
rather than a pile of audio.
TWO OF THE THREE CUES CLEAR THE MEASURED BAR AND ONE DOES NOT. The exploration cue — the one
directly comparable to round 1's rejected SLICE_T1_EXPLORATION_FLORES — reads at structural yield
1.100 against a corpus median of 0.76 and round 1's 0.22, at 5.50 boundaries per minute against a
corpus 5.09, at 5.00 drops per minute against a corpus 5.56, inside all three joint-box axes, with
23 measured sections against round 1's exploration median of 9 and a hook at 0.081 seconds. The
waterfall cue reads at 1.571 yield and clears the box. The caci cue — the one comparable to the
rejected SLICE_T2_BATTLE_CACI — reads at 0.235, which is round 1's own territory, with a named,
tested, directionally-confirmed cause and a next step that is a real composition change rather than
a tuning pass.
A PROTOTYPE THAT LANDS THE ARCHITECTURE AND SOUNDS LIKE A SAMPLER IS A SUCCESS AT THIS RUNG; A
PROTOTYPE THAT PRODUCES A BEAUTIFUL MIXDOWN BY ABANDONING THE PLAN IS A FAILURE. This is the first
kind. The plan was followed exactly — the compiler refuses to depart from it, and the two places
where it tried to (an out-of-range head note, a silent part) are now hard refusals rather than
silent repairs.
AND THE SOUND IS NOT CLAIMED. A CC0 community sampler with one dynamic layer per note, no round
robins, no legato transitions, no expression curves and no room is not a scoring session, and this
lane cannot hear what it made. The timbre-critic gap the round-1 review sheet records — the
authored-plus-sampler route at 1.932 against the generated route at 0.659 — is real, is not
addressed here, and closes only with a commercial library or with players.
THE NEXT RUNG, in the order I would take it:
1. The two-phase form for the battle cue, described in section 4.3. It is the one open failure and
it has a named repair.
2. Re-point the cycle_period_s and n_entrances comparisons at the instruments that measure them
properly. Both are my rulers being wrong, not the music.
3. Expression: CC curves per part in the emitted MIDI. Every phrase in these renders is dynamically
flat within a note, which no human performance is, and sfizz supports the fix.
4. A room. The mixdown is a sum of dry stems; a convolution stage is small code and a licensing
question about impulse responses.
5. Then, and only then, the twenty-cue sweep for this region. The architecture's economic claim —
136 lines of theme amortised across twenty cues of 120 to 160 lines each — is untested at
anything past three, and three is a small n for a claim about twenty.
THE SINGLE BIGGEST RISK to the architecture is that section_count and structural_yield, the two
axes it is now being steered by, are measured off a self-similarity matrix that reads DENSITY far
more strongly than it reads harmony, melody or form. Optimising against them could produce cues that
lurch between full and empty because that is what the ruler rewards, which is a different failure
from round 1's and no better. The corpus band is a containment check and was never meant to be a
target. The mitigation is that every one of these numbers stays necessary-and-never-sufficient, and
that Josh's ear remains the first ear.
The registry's deny_register_list requires that the region's own register is never the threat
signal, and that the darkening operations travel with the antagonist material rather than with the
place. pass2_flores.py --check verifies four things, and the fourth matters most because it reads
the NOTES the compiler emitted rather than the labels the author wrote.
Twenty of the cue's layers are marked of the place; the overlap with the threat set is empty, and
the plan validator refuses a plan where it is not.
{2, 4, 6, 9, 11} — D E F# A B. The broker's five layers use {0, 3, 5} — C, E-flat, F natural, none of which the
mode contains. Measured on the compiled output rather than on the plan: the of-the-place layers
emit exactly {2, 4, 6, 9, 11} and nothing else; the broker layers emit exactly {0, 3, 5} and
nothing else.
where BE_0003 turns from the ritual rounds to the unmasking.
octave_up, augmentation or literal, and at least one is octave_up. The ensemble answers higher, never darker. That is CH03_B10's own
resolution — the ritual victory with the broker exposed and no kill — as a compositional act
rather than as a note in a document.
The palette extension carries its own care argument, in the block appended to sfz_palette.py: the
world-instrument half of VCSL stays out; two added instruments that carry a living-tradition name
(conga, bongos) are marked as such and are not used by these plans; and the honest limits — no
pitched gong exists in either CC0 set, the tuning is 12-TET and a gong waning is not, the bamboo
flute is a baroque recorder, what is taken is the technique and not a timbre claim — are stated on
every plan rather than left to be discovered.
python harness/music_gen/pass2_sfz_probe.py --check
python harness/music_gen/pass2_sfz_probe.py --all
python harness/music_gen/pass2_plan.py --check
python harness/music_gen/pass2_realise.py --check
python harness/music_gen/pass2_flores.py --check
python harness/music_gen/pass2_flores.py --emit
python harness/music_gen/pass2_realise.py --all
python harness/music_gen/pass2_realise.py --md-table
python harness/music_gen/floor_instruments.py --cards build/audio/pass2/cards
Self-test totals at the time of writing: probe 4/4, plan 14/14, realise 15/15, flores 13/13. Every
one of those includes mutation controls that break the rule under test and check that it fires; a
rule that has never been made to fail is not armed.
Determinism was NOT re-proved for this chain. orchestrate.py established that two renders of the
same plan through the same sfizz invocation are bit-identical on this palette, and pass 2 imports
that invocation unchanged, so the claim is inherited rather than re-measured.
pass2_realise.py --determinism exists and has not been run on these cues; that is a gap and it is
named as one.