MASTERPIECE_PROGRAM.md

music/MASTERPIECE_PROGRAM.md

# THE MASTERPIECE PROGRAM — the score, per chapter, to the 30-year bar

CANON SUBORDINATION — this document is PROPOSAL-TIER: it serves canon and never outranks it.
The canon: the CVD · the T1 foundation docs · the T0 registries (registries/) · the spine
(docs/spine/CH_*.md) · the region pages (_source/02_Tier_2_Region_Pages/). Authority order: docs/DOC_MAP.md § 0.
Canon served: _source/02_Tier_2_Region_Pages/flores_island.md · docs/spine/CH_02.md · registries/T0_Chapter_Index [ACTIVE v1.2] · registries/T0_Region_Index [ACTIVE v1.2] · _source/01_Tier_1_Foundation/ (the CVD AAA direction + the scoring law) — the per-region palette (§5 layer 2) derives from each region page's own attested music; the per-chapter cue set derives from the chapter's own spine rows.
READ THAT CANON FIRST — open it and derive from it before you build anything from this document.
If this document disagrees with canon, CANON WINS and this document is the defect.
Tier: PROPOSAL (planning artifact; ratifies nothing, generates nothing).
Canon parent: docs/proposals/music/NOSTALGIA_RUBRIC.md + the school grammars +
[[music-direction-melody-first-30-year-bar]]. Supersedes the complete-tracks charter's
SCOPE (20 tracks total) — not its laws.

0. THE RULING THIS EXECUTES (Josh, 2026-08-05, verbatim)

"I didn't really like the music. it was a much better improvement, but I need 12-20+

DIFFERENT & UNIQUE tracks, 2-7 minutes long, PER CHAPTER. They should be inspired by the

cultural region and authenticity but then made into a masterpiece. Here are inspirations in

the list below. We need to evaluate what each track has and their purpose, range, complexity,

etc. by actually listening to them and creating their formulas (this evaluation and standard

needs to be another science) so that we can do these things properly and actually build

masterpieces people will be nostalgiac and listen to 20, 30, or even 50+ years later and will

play in orchestras and do broadways. Needs to be amazing and it'll take plenty of iterations."

Plus the 100-track exemplar list across seven purpose classes (kept verbatim in §7).

1. THE SCOPE, STATED HONESTLY

79 nodes x 12-20 tracks = 950 to 1,600 tracks, each 2-7 minutes. Against 20 tracks landed

today. This is not 20x the work — done wrong it is 50x and produces 1,600 unrelated pieces

that sound like a library. Done right it is a LEITMOTIF ECONOMY (§5) where the corpus shares

DNA and every chapter is a treatment, not an invention. That is how Uematsu, Mitsuda and

Shimomura shipped 60-90 track scores that cohere: a few dozen ideas, ruthlessly varied.

THE ONE-LINE ARCHITECTURE: a track = MOTIF x REGIONAL PALETTE x PURPOSE CLASS, each of

the three authored once and combined per chapter.

2. PHASE A — THE MEASUREMENT RIG (the science, part one)

The rubric shipped today scores DIALECT CONFORMANCE. It cannot say what a MASTERPIECE has,

because it was derived from published analyses, not from the recordings. Phase A fixes that.

Legal posture (settled by Josh's own Sonniss line, 2026-08-05): analysis for

UNDERSTANDING is ordinary use — listening with a tool is listening. The bytes NEVER become

model input: no training, no fine-tuning, no conditioning, no derivation. We measure them the

way a musicologist measures them, and we compose our own music. Source audio is what Josh

already owns or buys; nothing is redistributed and no exemplar audio ever enters a generator.

Per-exemplar measurement (the FORMULA CARD):

DimensionMeasured
Duration & structuretotal length; section boundaries (novelty-curve segmentation); section count; repeat map
TempoBPM; tempo curve (rubato, accelerando, phase shifts)
Metresignature; metric changes; dominant onset period (measured, never assumed)
Tonalitykey; mode; modulation map with pivot points; harmonic rhythm (chords/bar)
The hookonset time (seconds AND bars); pitch contour; interval vocabulary; range in semitones; the ONE distinctive deviation
Motif economydistinct motif count; recurrence map; transformation types (inversion, augmentation, fragmentation, reharmonisation)
Orchestrationinstrument roster; entry timeline (who plays when); voice count over time; doubling map
Texture & densitynotes-per-second curve; polyphony curve; tutti vs solo ratio
Dynamicsdynamic range (p95-p10 dB); the arc shape; where the peak sits proportionally
Registerspan; the melody's tessitura; bass floor
Cadence & endingcadence census; final cadence type; loop or through-composed; how it ends
Timbrespectral centroid curve; the signature colour

Output per track: formula_cards/<track_id>.json + a human page. ~100 cards.

Instrument controls (the standard's own honesty): every extractor self-tests against

synthetic material with a KNOWN answer (a 32-bar AABA at 120 BPM in D minor with a planted

hook at bar 3), and against two of OUR tracks whose ground truth we authored. An extractor

that cannot recover a known answer does not get to describe Uematsu.

3. PHASE B — THE ARCHETYPE SYNTHESIS (the science, part two)

Cluster the 100 cards by PURPOSE CLASS (Josh's own seven headings) and derive, per class,

the TARGET BANDS and the RECURRING MECHANISMS:

theme carries identity in <=8 bars (Aerith, Ezio's Family, Dearly Beloved).

hours, the ambient-to-melodic ratio (Ard Skellig, Sweden, Kokiri, Stickerbush, Dire Docks).

(Rivers in the Desert, BFG Division, Battle on the Big Bridge, Megalovania).

(One-Winged Angel, Radagon, Soul of Cinder, Dancing Mad, Hopes and Dreams).

carries (To Zanarkand, Midna's Lament, Weight of the World, Lavender Town).

(Suicide Mission, Snake Eater, Life Will Change).

(Answers, Baba Yetu, Setting Sail Coming Home, Scars of Time).

Deliverable: MASTERPIECE_STANDARD.md — per class, the measured bands and the named

mechanisms, each citing the exemplars it came from. THIS REPLACES the proxy rubric as the

scoring bar.

4. PHASE C — THE NOSTALGIA PREDICTOR (the hard part, honestly framed)

The 30-year bar is a claim about MEMORY, and no measurement proves it. What CAN be built is a

scorer whose hypotheses are testable against the corpus:

VALIDATION IS THE WHOLE POINT: the scorer must rank the 100 exemplars ABOVE our 20 landed

tracks (a positive control it can fail), and must rank sensibly WITHIN the exemplars against

a small human-ranked subset. A scorer that cannot separate Uematsu from our own first pass is

not a bar; it is a mirror.

5. PHASE D — THE LEITMOTIF ECONOMY (how 1,600 tracks stay coherent)

Layer 1 — THE MOTIF BANK (~40-60 motifs, authored once, at the masterpiece bar):

the protagonist; the companion(s); the healer's child; the antagonist pattern (held unnamed

pre-Ch-38); the realm; each of the 22 threads that carries musical identity; the major

factions; the recurring creatures/bosses. Each motif: a head cell, its transformation family,

and its declared emotional register.

Layer 2 — THE REGIONAL PALETTE (one per region, ~13 for the slice, ~76 for the arc):

period-and-culture-true instrumentation, modal language, rhythmic cells, and ornament

practice — grounded in that region's own attested music (the scoring law), then RAISED: the

palette is the vocabulary, not a folk-recording pastiche. This is where "inspired by the

cultural region and authenticity but then made into a masterpiece" lives.

Layer 3 — THE PURPOSE CLASSES per chapter (the 12-20): exploration, traversal, town/

settlement, sanctuary/rest, discovery/wonder, ordinary battle, boss (per phase), stealth/

tension, sorrow/loss, festival/trade, night, the chapter's own set-piece, the chapter theme

itself. Which classes a chapter gets is READ from its own canon (a combat-light chapter

trades battle tracks for discovery and ambient; a boss chapter spends on phases).

A track is then specified as: motif(s) x palette x purpose x the formula card its class

demands. Unique because the combination is unique; coherent because the motifs recur; and

authentically regional because the palette is.

6. PHASE E — PRODUCTION & ITERATION

1. Chapter-set batching — a chapter's 12-20 tracks are composed as a SET so they cohere

and share the palette; never one-off.

2. Compose symbolically to the card (the proven authored path), then realise.

3. Measure against the MASTERPIECE_STANDARD, not the proxy rubric; iterate until the

class bands are met and the predictor clears its floor.

4. Josh's ear is the final gate, on the site, per chapter set.

5. The realisation ladder — today's sfizz + CC0 stack got us to "complete track". The

masterpiece bar likely needs a better sample library rung (BBC SO Discover is free and

already queued as rung 2) and, for hero cues, eventually real players. NAMED as a decision

with a cost, not assumed.

6. The concert deliverable — every hero cue emits real notation (MusicXML + parts) so an

orchestra can play it. The Broadway/concert ambition is a FORMAT requirement from day one,

not a post-hoc export.

7. SEQUENCING (against the weekly limit)

10-track PILOT spanning all seven classes; prove the rig recovers known answers.

whole program, because everything downstream inherits it.

8. WHAT THIS DOES NOT DO

It does not ratify any exemplar's material into our score; it does not put exemplar audio into

any generator; it does not claim the 30-year outcome can be measured — only that the features

the exemplars share can be, and that failing to share them is disqualifying.

9. JOSH CONFIRMS THE POSTURE (2026-08-05, verbatim)

"I'm pretty sure I can buy or download them. You can listen to them and build the equations to

understand what is good. this is math and theory - we aren't training the generator. it's so

you can know what to prompt so we can make our original video game music masterpieces"

RULED AND SETTLED — this is exactly the posture in section 2: MEASUREMENT AND THEORY, never

model input. The exemplars are STUDIED to derive the equations; the equations tell the

factory what to COMPOSE. No exemplar audio touches a generator, is redistributed, or is

conditioned on. Same shape as the Sonniss line (listening is use; bytes-as-model-input is

the only prohibition), and the same shape as any composer studying scores.

ONE HONESTY NOTE ON WHAT "LISTENING" MEANS HERE: the director does not hear audio

subjectively. Its listening is MEASUREMENT — spectral, structural, harmonic and dynamic

extraction — plus published theory. For deriving equations this is a STRENGTH (it recovers

exact numbers a human ear estimates), but it means the predictor of section 4 can never

certify "this is beautiful"; it certifies "this has what the beautiful ones measurably have."

JOSH'S EAR REMAINS THE ONLY SUBJECTIVE GATE, and the predictor's validation set is calibrated

against it.

10. WHAT JOSH ACQUIRES (the only human step in rung 1)

Audio files, purchased or downloaded, dropped at D:/audio/exemplars/ in any common format

(FLAC/WAV preferred for measurement fidelity; MP3/OGG works — the rig declares format-induced

limits per card). Naming does not matter; the rig fingerprints and matches against the corpus.

Partial sets are fine: the rig scores what it has and declares what is missing rather than

guessing, and the pilot needs only 10 tracks spanning the seven classes.

THE LIST IS NO LONGER IN THIS DOCUMENT. It lives in two wired artifacts, and this section is

subordinate to them:

116-track list (three culled by his ruling, recorded in a binding one-directional excluded

array that no wave may reverse) plus the 68-row 1990-2008 expansion wave, each expansion row

carrying the compositional GRAMMAR it contributes that the base list lacked. Seven purpose

classes. Each row's s field is the acquisition state and flips MISSING → ACQUIRED → CARDED.

https://claude.ai/code/artifact/ab157f80-9a79-4b5c-876f-08ae9c41e137) — the verified route

from every row to a file on disk, organised by storefront. Carries the STARTER TEN that

satisfies the pilot rule above for $64.59, and — the part that matters to the rig's honesty —

the 17 rows with NO legitimate route at any price, each with a named substitute and an

instruction to card it SUBSTITUTED or ARRANGEMENT-TIER rather than measure a cover and report

it as the source.

Rung 1 therefore does NOT wait on a complete corpus and must never claim one. It waits on ten

files, and every card it emits declares its own tier.

10.1 THE ACQUISITION PROOF — the one command that turns a purchase into a card

Section 10 above names exactly one human step. Everything after the drop is

harness/music_gen/acquire_exemplar.py, so the FIRST purchase is a one-command test rather than

an afternoon of finding out whether the file even decodes.

THE COMMAND, verbatim, after dropping files at D:/audio/exemplars/:

D:\audio\rig\venv\Scripts\python.exe harness\music_gen\acquire_exemplar.py --scan

Run it from the repo root. It proofs every new file in the drop folder and leaves already-proofed

files alone, so it is the standing command and not a one-shot — buy three more next month, run the

same line, and only the three new files cost anything.

WHAT IT DOES, IN ORDER, EVERY LINK ABLE TO FAIL LOUDLY:

is a synthetic signal with a planted 120 bpm, planted lane entries at 8 s and 16 s, a planted

drop and a 2-second break at 24 s, encoded to AAC and decoded back through the same ffmpeg path

a purchased .m4a takes. A negative control (one steady tone) must yield zero events, and a

zero-guard requires every reported curve to be populated and non-degenerate — together they

close both doors, so the rig can neither pass by reporting zeros nor by reporting events

everywhere. This is section 2's instrument-controls rule, armed rather than described.

the system PATH is never modified). AAC/M4A/ALAC need this — soundfile cannot read them.

the BUY_SHEET's DRM-free route, not with a silently-empty measurement.

section 11 into curves and events: the tempo function, the lane entry/exit map and interlap

matrix, the event grammar (drops, raises, breaks, tutti, silence budget), the onset type and

flow curve, and the complexity step map.

for beside it.

guesses. An unmatched file is still measured and still carded, under an UNMATCHED__ id — the

work is never thrown away, it is reported, and --id EX_NNN attaches it.

(sha256, source path, format, codec, sample rate, channels, duration, acquired_at, and the paths

to the card and the proof record).

build/audio/records/exemplars/<track_id>.acquisition.json.

WHAT IT PRINTS. On success, per file: the matched row, the measured length / tempo / key / section

count / lane count, the event counts and silence budget, the card path, and MISSING -> ACQUIRED.

It closes with the corpus tally. On failure it prints the verdict, the reason, and the top scoring

candidate rows so the next move is obvious. On an empty drop folder it says so and reports how many

rows are still MISSING rather than exiting silently.

EXIT CODES ARE HONEST: 0 proofed (or nothing to do), 2 the control failed and nothing was

measured, 3 the file was measured and carded but no corpus row could be claimed, 4 the file

could not be decoded or is protected.

OTHER INVOCATIONS: --self-test runs the control alone (this is gate music_acquire);

--audio <path> proofs one file; --id EX_NNN overrides the match; --dry-run measures and cards

without flipping; --no-corpus is the decode-path test for audio that is not an exemplar at all.

THE LEGAL POSTURE IS THE CORPUS'S OWN AND IS RESTATED ON EVERY CARD AND EVERY RECORD: analysis for

understanding only; the audio is read and measured, never copied, never redistributed, and never

model input of any kind.

11. JOSH EXPANDS THE SCIENCE (2026-08-05, verbatim) — AND IT IS THE FLOOR

"yeah you can analyze and mathematically create the patterns behind all different lanes and

interlaps on a track, how many drops or raises or when and by how much of tempo adjustments,

how they start and then pick up and how they flow throughout, which tonal harmonies and

instruments and how things play together, when to introduce more and provide complexity, what

types of scenes and areas require which, what types of instruments and ambience to combine,

how to do all of this - what i mentioned is the floor, it's our standing ruling you expand

within our vision to create our masterpiece"

THE STRUCTURAL UPGRADE THIS FORCES: section 2's card was SCALAR (a tempo, a count, a range).

Josh is describing CURVES AND EVENTS — what enters when, what drops where, by how much, and

why there. A masterpiece is not "120 BPM with eight instruments"; it is a TIMELINE. Every

dimension below is therefore measured as a FUNCTION OF POSITION (normalised 0-1 through the

track AND in bars), not as a number. The formula card becomes a SCORE-SHAPED OBJECT.

11.1 LANE ANALYSIS (Josh's "all different lanes and interlaps")

Measured per LANE, not per mix — source-separated stems where separation is clean, symbolic

transcription where it is not, and the method DECLARED per card (a lane read off a muddy

separation is marked low-confidence, never reported at equal authority).

arpeggio · percussion (pulse vs colour) · texture/atmosphere · vocal · solo feature.

load-bearing artefact — it is literally the arrangement.

Masterpieces have STRUCTURE here (pads never coexist with the ostinato; the countermelody

only enters after the lead has stated twice), and the matrix is where that shows.

arrangement's story.

doubling intervals when two lanes share a line (octaves, thirds, sixths — Wise's signature).

11.2 THE EVENT GRAMMAR (Josh's "drops or raises, when and by how much")

Every discontinuity detected, classified, timed and MAGNITUDE-MEASURED:

delta, bars held, what re-enters first}.

restraint is a masterpiece mechanism and the current factory has no vocabulary for it.

11.3 TEMPO & TIME GRAMMAR ("when and by how much")

Not a BPM — a tempo FUNCTION: baseline, every adjustment event {position, delta BPM, ramp vs

step, duration}, rubato depth per section, fermatas, metric modulation, and any metre change

with its pivot. Plus the GROOVE layer: swing ratio, syncopation density, and the rhythmic cell

that repeats.

11.4 THE OPENING & FLOW LAW ("how they start and then pick up and how they flow")

fade-in · pickup/anacrusis · silence-then-hit.

class we derive the ARCHETYPAL SHAPE — the exploration arc is not the boss arc is not the

lament arc — and a composed track is scored against its class's curve, not a scalar band.

**FINDING 2026-08-07 — DO NOT PIN THE CLIMAX, AND DO NOT LET THE GOLDEN SECTION IN THROUGH THIS

SECTION. Measured on the corpus cards at HEAD, the dynamic-peak position has a median of 0.575**

and the flow-peak a median of 0.606, which sits temptingly near 0.618 and would read as a

confirmation of the golden-section climax claim if only the median were printed. The spread is the

answer: p10 to p90 runs 0.16 to 0.87, and only seven of the corpus rows land inside 0.55 to 0.70.

Pinning every cue's peak near the golden section would therefore make our output MORE UNIFORM THAN

THE MUSIC WE ARE TRYING TO MATCH — the tracks Josh has loved for decades put their loudest moment

almost anywhere, and the ones that put it early are not doing it wrongly.

This is the exact failure mode this section's own last sentence already guards against ("scored

against its class's curve, not a scalar band"), and the finding is recorded here so that a later lane

reaching for a single number has the measurement in front of it. **The class curve stays the unit.

A single global climax position is banned, and the median above must never be quoted without its

spread.** A related caution from the same pass: the energy peak and the loudness peak are not the

same event and do not co-locate — on EX_124 the flow peak sits at 0.272 and the dynamic peak at

0.4595 — so "the peak" must always say which one it means. Evidence:

docs/proposals/music/MUSIC_CRAFT_DOCTORATE.md §4 and

docs/proposals/music/craft_research/HOOK_AND_TENSION_ARCHITECTURE.md §8.

11.5 HARMONY & COLOUR ("which tonal harmonies and instruments and how things play together")

Chord vocabulary with frequencies; borrowed/modal-mixture inventory; the SURPRISE CHORD and

its normalised position (masterpieces place it, they do not sprinkle it); harmonic rhythm as a

curve; modulation map with pivots; pedal/drone usage; cadence census per section boundary.

INSTRUMENT COMBINATORICS: which pairings actually occur and in what register relationship —

mined as an association table so the composer inherits real orchestration practice

("clarinet + harp at a sixth in the lower-middle" is a fact we can measure, not a guess).

11.6 THE COMPLEXITY LAW ("when to introduce more and provide complexity")

A single COMPLEXITY INDEX per bar (lane count + harmonic rate + rhythmic subdivision +

register span + polyphony), and its STEP MAP: where it increases, by how much, and what

mechanism does it (a new lane? a subdivision change? a modulation?). The finding we expect and

must verify: masterpieces step complexity at STRUCTURAL boundaries and hold it flat inside

sections — variation comes from ORCHESTRATION, not from constant churn. If the corpus says

otherwise, the corpus wins.

**MEASURED 2026-08-07, AND THE CORPUS SAYS OTHERWISE. THIS SECTION'S OWN LAST SENTENCE THEREFORE

SETTLES IT AGAINST THE EXPECTATION ABOVE.** Over the formula cards at HEAD, restricted to tracks of

45 seconds or more, corpus rows put 0.358 of their complexity steps at a section boundary

against their own album siblings' 0.500 — the loved tracks land a SMALLER share of their steps

on the seams, not a larger one. If change were carried by the form, the steps would sit on the

boundaries; they do not, they sit inside the sections. Six further readings agree with it: a corpus

row carries about half again as many structural boundaries as its album's filler (16 against 11)

while making only 17.9% of them strong against 36.4%, fires strong events at 0.88 per minute

against 1.89, and brings a new voice in more often and more quietly (4.55 lane-adding raises per

minute at 7.49 dB against 3.10 at 10.09 dB).

The prediction was half right and its halves were swapped. Variation does come from

orchestration rather than churn — that part holds, and holds strongly. What is wrong is the

LOCATION: re-treatment is CONTINUOUS and lives inside sections, under a sparse ridge of about three

genuine structural events, with a hierarchy ratio (largest event over typical) of 1.91 against the

siblings' 1.65. The corrected law is therefore two-tier — a TIDE of small treatment changes every

four to eight bars (the corpus does it every 5.79) under a handful of LANDMARKS — and it is the same

fact as Josh's three-to-five-cycle floor read at a second scale.

Derivation and full numbers: docs/proposals/music/craft_research/REPETITION_AND_VARIATION_FORM.md

§1 and docs/proposals/music/MUSIC_CRAFT_DOCTORATE.md §2 and §4. Every figure is re-derivable from

complexity_law.step_map, duration_and_structure.novelty_peak_strength and

event_grammar.raises on the existing cards; nothing new had to be measured. **Honest register: these

are post-hoc descriptive reads with no calibrated multiplicity null of their own — the boundary-count

row is a restatement of section_count, which does clear the bar at the 99th percentile, and the

rest agree with it directionally.** The expectation above is superseded as an EXPECTATION; it is left

standing in the text because a program that quietly deletes its wrong predictions cannot be audited.

11.7 THE SCENE→FORMULA MAP ("what types of scenes and areas require which")

Each exemplar carries its GAME CONTEXT (what it plays under). Cluster those contexts into

scene archetypes — first-arrival, safe town, hostile wilds, dungeon interior, sacred site,

travel/traverse, ambush, duel, multi-phase boss, revelation, loss, farewell, credits — and

derive, per archetype, which formula the exemplars actually use. THIS IS THE TABLE THE GAME

QUERIES: a Humanity scene declares its archetype (from its own canon rows) and the formula

follows. It is also the bridge to the cue system: our EG rows, encounter phases and region

pages already declare scene type, so the map has a real consumer on day one.

11.8 SCORE + AMBIENCE ("what types of instruments and ambience to combine")

The game-audio half, which no rung has touched: how score and world sound share the frequency

and attention budget. Per scene archetype — the ambience bed's spectral occupancy, where the

score sits above/below it, ducking behaviour, the diegetic/non-diegetic seam (a musician IN

the scene versus the score), and the SILENCE-FOR-AMBIENCE budget (the moments the score

deliberately yields). Composes with the 680-row SFX registry and the 650-cue system already

in the engine.

11.9 ADAPTIVE TRANSITIONS (beyond-floor addition, NAMED for veto)

Not in Josh's list, added under the expand-within-vision ruling: how a cue BECOMES another

cue — the transition grammar (stinger, bridge, layer-swap, hard cut on a downbeat), and the

multitrack/stem architecture that makes it possible. Reason it belongs: 12-20 tracks per

chapter only feel like one world if the seams between them are composed. If Josh vetoes,

tracks stay independent and the seams are handled by crossfade.

11.10 WHAT THE FORMULAS ACTUALLY ARE

Not prose. Per purpose class, a PARAMETRIC SPECIFICATION a composer (human or ours) can

target: the energy curve shape with tolerance bands · the lane entry schedule as a function of

form position · the event schedule (how many drops/raises/tutti and where) · the complexity

step map · the harmonic vocabulary with the surprise slot · the instrument-combination table ·

the tempo function · the silence budget · the onset type · the ending type. A composition is

scored by DISTANCE FROM ITS CLASS SPECIFICATION, per dimension, with the misses named — the

same shape as every other bar in this factory, but pointed at what actually makes music good.

Generated by harness/site/structure_site.py — the URL path is the repo path. review root