music/RECOMPOSE_RUNG_TWO_FINDINGS.md
CANON SUBORDINATION — this document is PROPOSAL-TIER: it serves canon and never outranks it.
The canon: the CVD · the T1 foundation docs · the T0 registries (registries/) · the spine
(docs/spine/CH_*.md) · the region pages (_source/02_Tier_2_Region_Pages/). Authority order:docs/DOC_MAP.md§ 0.
Canon served: NONE FOUND — this document cites no canon anchor anywhere in its body.
Do not build from it alone: open the spine node (docs/spine/CH_NN.md), the region page
(_source/02_Tier_2_Region_Pages/), and the T0 registry your work targets, and derive from those.
READ THAT CANON FIRST — open it and derive from it before you build anything from this document.
If this document disagrees with canon, CANON WINS and this document is the defect — fix the
document, never the canon. Nothing here is applied until it is ratified into canon.
Status: OPEN BOARD (2026-08-05). Owner: the music lane. This is the standing input to the next
music rung, which is FROZEN behind the world pipeline by the depth law and does not start here.
---
The nostalgia-bar lane built an exemplar-anchored ruler, scored the fourteen landed themes against
it, recomposed the flagship, and put the result through a fresh-context cold verification the same
day. The verdict was NOT-YET, on eight findings. The apparatus was kept; the pick was demoted.
The findings were about specific, measured things, and re-deriving them next time would cost the
same verification spend twice — so they are written down here rather than left in a run journal.
A rung that opens without answering the ones marked BINDING below is repeating a known failure.
WHERE THE VERDICT ITSELF LIVES: on the artifact, not in this file.
build/audio/authored/journey_world/JOURNEY_WORLD_RECOMPOSED_RUNG_TWO.json carries a
verifier_verdict block, harness/music_gen/emit_authored_picks.py rule R7 withdraws any rung
whose verdict has not cleared, and harness/music_gen/realise_recompose.py re-states the block on
regeneration. So the demotion cannot be undone by re-running a script; it takes a deliberate edit at
the source, in the same commit as whatever answered the finding. C07 stays on disk in full — audio,
stems, MIDI, records — as the next rung's measured baseline and as the winning arm of the A/B
against the theme Josh rejected.
---
These three are not defects in the pick. They are defects in the ruler and its generator, which
means the next rung inherits them unless they are fixed first.
P4 and L1 cannot both be satisfied by an honest declarationP4 asks sentence_bar_count == 8 (PHRASE-02 / ECON-10, the classical sentence). L1, a HARD
gate, asks loop_length_bars >= 24 (FORM-05). Both cite the SAME exemplar: Zelda's Theme, read as a
24-bar binary made of a 16-bar A of two 8-bar sentences plus an 8-bar B.
The recomposition declared 24 bars of composed material as one unit, so P4 failed — and that
single failure is the only feature on which the pick lost to the theme Josh rejected. It is a
REPRESENTATIONAL artifact, not a musical one: declaring bars_per_sentence = 8 and tiling three
times would have passed both gates and produced an 8-bar loop heard three times, which is worse
music and a better score. A ruler that pays better for the worse answer is the defect.
WHAT THE FIX HAS TO DO: make the schema able to say what the exemplar actually is — a loop composed
of N sentences, each with its own bar count — so P4 scores the SENTENCE and L1 scores the LOOP,
and the exemplar's own form passes both without a declaration trick. Until then, neither feature's
verdict on any theme means what it appears to mean.
recompose_ladder.md states that the rubric saturates and that the pick was therefore made outside
it. That declaration is honest and it stops one sentence too early: it never says WHY.
The measured cause: **32 of 57 features return identical values across all 16 candidates, and only 2
distinct verdict vectors exist across the whole ladder.** Every HARD gate except H4 is constant.
The entire LOOP FORM axis except L3/L6 is constant — the axis carrying the census's single
biggest failure (L1, 0 of 14) contributes zero discriminating information. The reason is that
recompose_journey.py builds all sixteen arms from ONE fixed 24-bar assembly plus four plans
authored against the rubric's own form features. Sixteen names, one template.
WHY THAT MATTERS BEYOND THE LADDER: "7 of 7 axes across 16 candidates" reads as sixteen independent
results and is one result. The pick then had to be made on deciding axes D1-D9 that sit outside the
rubric — and D9 is a self-declared DERIVED proxy, which is the exact metric class Josh's ruling
named as the original defect. The lane reintroduced a proxy at the last step of removing proxies.
WHAT THE FIX HAS TO DO: vary the ASSEMBLY per candidate so the arms differ structurally, then
re-emit the ladder and report how many distinct verdict vectors it produced. A ladder that cannot
rank is not a ladder, and the honest report of one is a number, not a paragraph.
L2 has a target the corpus can ignore, and no feature catches the form gapL2 motif_count_per_loop reads >= 3 (target 4), carrying Kondo's verbatim four-motifs quote as
its grounding. The pick scores 3. The theme Josh rejected scores 4. On the rubric's own
headline anti-fatigue mechanism, the recomposition is worse than the thing it replaced — and it
passed, because the band's floor is 3 and only the floor is enforced.
Hand-verified structure, not a metric artifact: 9 distinct bars in 24; bars 9-16 are a verbatim
repeat of bars 1-8; one 1-bar cell occupies bars 17, 19, 21 and 23.
And the form gap underneath it: L1's own citation reads the exemplar's A section as TWO 8-bar
SENTENCES, not one 8-bar unit stated twice. The pick does the second thing. **No feature in the
rubric can tell those two apart**, which is why nothing caught it — and it is the same hole as 1.1,
seen from the other side.
WHAT THE FIX HAS TO DO: compose to the target rather than the floor (four mutually-unlike motifs),
retire the 8-bar verbatim restatement and the four-fold 1-bar cell, and add the feature that
distinguishes a restated sentence from a second sentence. If a target is never enforced, it is a
comment; either arm it or state that it is advisory.
---
resolved to exactly one file: the rubric itself. The three research returns now live at
docs/proposals/music/school_grammars/ (Kondo · SNES-JRPG · earworm/loop science), extracted
verbatim, and all 49 cited ids resolve. The rubric § 0 states the correction in place.
and Brame appear nowhere in this repo outside the rubric and the grammars, several as bare
author-year strings. LAW 1 ("every band names its exemplar") is what makes this ruler non-proxy;
a named exemplar a later reader cannot resolve is asserted rather than demonstrated. OWED: a
resolvable bibliography with DOI/URL and full titles.
the pick's own melody: exact and transposed quotes are REFUSED (six fire cases), but a one-interior-
note edit, a final-interval edit, or an ornamented passing note all read CLEAR. The doc declares
its SEED-SET gap (14 absent exemplars) and never declares its MATCH-TOLERANCE gap. OWED: declare
the tolerance beside the seed set, in the same sentence that says what CLEAR means.
S16 checks the feature ID SET only. § 3 says "Seven features are hard" and lists EIGHT (H1 H4 C5 C6 P1 Y1 L1 L2; the code agrees with
the list). DOC_MAP says 50 features; the code emits 57. LAW 3 says nine are NOT_COMPUTABLE,
which is theme-dependent (8 for the pick). OWED: fix the three numbers and widen S16 to cover
counts and hard flags, so the prose cannot drift from the instrument again.
24-bar score (51.4 s); the artifact on disk is 26 bars (55.7 s) with a 2-bar turnaround appended
by realise_recompose.py. The verifier scored the shipped loop: 7 of 7 axes, zero hard fails,
unchanged axis verdicts (Y1 modal_weakened → half_cadence, still not a PAC; L6 seam 0 → 2, an
improvement; L4 0.625 → 0.577, in band). OWED: score the SHIPPED artifact by default, so the
claim covers the object a listener hears.
rubric is symbolic and reads no audio; feature_rig.py is spectral and asks whether a signal
carries a melody at all. Neither asks whether the result SOUNDS like the exemplar class. Josh's
complaint was about what he heard. No symbolic gate can discharge an auditory complaint — the
next rung needs a listening protocol before any further orchestration spend.
---
span and a textbook PAC, and it scores 1 of 7 axes with 6 hard fails. Self-test 26/26, exit 0.
ARM_ZERO and loses exactly one (P4, the representational artifact of § 1.1). The gains land on
the features the census named as the cohort's failures. That result stands; it is simply not the
same claim as "this clears the bar".
DOC_MAP rows present.---
1. A ladder whose arms differ STRUCTURALLY, with the distinct-verdict-vector count reported (§ 1.2).
2. A schema that can express a multi-sentence loop, so P4 and L1 stop contradicting each other
and the restated-sentence hole closes (§ 1.1, § 1.3).
3. A loop composed to L2's target of four mutually-unlike motifs, with no 8-bar verbatim
restatement (§ 1.3).
4. A resolvable bibliography (F4), a declared match-tolerance (F5), the three numbers fixed and
S16 widened (F6), and the SHIPPED artifact scored (F7).
5. A listening pass in front of Josh before any further orchestration spend (F8).
---
Josh listened to the shipped cohort and ruled, verbatim: *"Why is all this music the fuxking same,
all just 30 seconds each the exact same track, whys it all just one type of sound on a 5 second
loop, like wtf? Were supposed to be creating at least 12 to 20 complete tracks with actual full
orchestra complexity like Star Wars and Harry Potter and super Mario 64 and sonic etc. even
undertale was a solo dev project that created good music for their game."*
That is F8 discharging itself. The board said no symbolic gate can answer an auditory complaint;
the auditory verdict came in anyway, and it is worse than NOT-YET — the whole cohort reads as one
track.
build/audio/complete_tracks/SAMENESS_BASELINE.mdEmitted by harness/music_gen/sameness_baseline.py, measured off the artifacts on disk. Fifteen
orchestral rung-one renders, on the eight axes a listener hears first:
distinct_composed_sections: 1 distinct value across 15 of 15 tracks. Every track is one8-bar sentence plus a one-bar turnaround. No track in the cohort has a B section.
15 of 15 are under a minute and 0 of 15 reach the 2.5-minute floor a complete track needs.
two string-ensemble frame parts and the same contrabass, every time.
percussion anywhere in the cohort**.
orchestral cue moves through. Machine-flat, measured.
The lane's defence had been that the themes are DISTINCT — different pitch sets, different keys,
different head cells, each measured and gated. That defence is true and irrelevant. **Pitch-set
distinctness is not an axis a listener hears in the first four seconds**, and the axes that are
were the ones nobody was scoring. That is the same defect class as § 1.2's ladder: a real
measurement of the wrong thing.
One cause is not compositional and has to be said plainly. The landed CPU stack is sfizz over
VSCO2-CE and VCSL — 23 instruments, and it contains **no percussion, no choir, no trumpet, no tuba
and no period instrument of any kind**. Part of "one type of sound" is what the renderer could
physically play. Extending the stack is a wave-two entry condition, not an afterthought.
build/audio/complete_tracks/cards/HF_MT_*.json, emitted by
harness/music_gen/emit_identity_cards.py, proof at build/audio/complete_tracks/VARIETY_MATRIX.md.
Twenty per-track cards — ten function cues (title, exploration/traversal, ordinary battle, two boss
tiers, sanctuary, trades and festival, sorrow, wonder, credits) and ten slice region cues, each
region's instrumentation copied from that page's own Section 10 rather than composed. Each card
carries cue role, derived duration, tempo band, key and modal colour, instrumentation family,
energy class, a real multi-section form with a bar map, the motif it states or develops, and its
stack_coverage.
Seven teeth, every one mutation-proved rather than asserted (--selftest, 18/18):
tempo_band + instrumentation_familyenergy_class. Measured: 20 distinct triples over 20 cards, 7 of 7 tempo bands, 12 of 12instrumentation families, 19 of 19 energy classes; of 190 pairs, **zero differ on fewer than two
of the three axes**.
stays unnamed and no card names the House parent theme.
in-game loop span is legal only with a written justification.
finding closed: L2's target of four is no longer a comment.
Measured 5-12 sections per card, mean 7.4, against the cohort's 2.
inside its own declared band.
that is tiling by another name.
§ 1.1's schema defect is closed by construction. A card says *a track composed of N named
sections, each with its own bar count*, plus a separately declared in_game_loop_span_s governed
by LOOP-01's per-cue-class band. P4 can score the SENTENCE and L1 the LOOP without the two
contradicting, and the declaration trick that scored better for worse music is no longer
expressible: duration is DERIVED from total_bars x beats_per_bar x 60 / bpm, so a card whose
form and tempo do not produce its stated length fails V5 rather than being believed.
rung; the listening pass in front of Josh still gates any orchestration spend, and now gates the
stack acquisition as well.
recompose_journey.py still builds sixteen arms from oneassembly. The cards give a later ladder structurally different arms to build from, which is a
precondition and not the fix.
Valley region cue, because that page's Section 10 opens by stating no instruments are attested in
the corpus and an identity card's central field is instrumentation; and the fairy realm, because
Josh's slice ruling designs it last. Both name the chapters that cover them meanwhile and both
name an entry condition.
SAMENESS_BASELINE.md, VARIETY_MATRIX.md, the card directory andthe two emitters. They are not written here because this lane's declared write scope is
build/audio/ and docs/proposals/music/.