music/MASTERPIECE_PROGRAM.md
# THE MASTERPIECE PROGRAM — the score, per chapter, to the 30-year bar
CANON SUBORDINATION — this document is PROPOSAL-TIER: it serves canon and never outranks it.
The canon: the CVD · the T1 foundation docs · the T0 registries (registries/) · the spine
(docs/spine/CH_*.md) · the region pages (_source/02_Tier_2_Region_Pages/). Authority order:docs/DOC_MAP.md§ 0.
Canon served:_source/02_Tier_2_Region_Pages/flores_island.md·docs/spine/CH_02.md·registries/T0_Chapter_Index [ACTIVE v1.2]·registries/T0_Region_Index [ACTIVE v1.2]·_source/01_Tier_1_Foundation/(the CVD AAA direction + the scoring law) — the per-region palette (§5 layer 2) derives from each region page's own attested music; the per-chapter cue set derives from the chapter's own spine rows.
READ THAT CANON FIRST — open it and derive from it before you build anything from this document.
If this document disagrees with canon, CANON WINS and this document is the defect.
Tier: PROPOSAL (planning artifact; ratifies nothing, generates nothing).
Canon parent: docs/proposals/music/NOSTALGIA_RUBRIC.md + the school grammars +
[[music-direction-melody-first-30-year-bar]]. Supersedes the complete-tracks charter's
SCOPE (20 tracks total) — not its laws.
"I didn't really like the music. it was a much better improvement, but I need 12-20+
DIFFERENT & UNIQUE tracks, 2-7 minutes long, PER CHAPTER. They should be inspired by the
cultural region and authenticity but then made into a masterpiece. Here are inspirations in
the list below. We need to evaluate what each track has and their purpose, range, complexity,
etc. by actually listening to them and creating their formulas (this evaluation and standard
needs to be another science) so that we can do these things properly and actually build
masterpieces people will be nostalgiac and listen to 20, 30, or even 50+ years later and will
play in orchestras and do broadways. Needs to be amazing and it'll take plenty of iterations."
Plus the 100-track exemplar list across seven purpose classes (kept verbatim in §7).
79 nodes x 12-20 tracks = 950 to 1,600 tracks, each 2-7 minutes. Against 20 tracks landed
today. This is not 20x the work — done wrong it is 50x and produces 1,600 unrelated pieces
that sound like a library. Done right it is a LEITMOTIF ECONOMY (§5) where the corpus shares
DNA and every chapter is a treatment, not an invention. That is how Uematsu, Mitsuda and
Shimomura shipped 60-90 track scores that cohere: a few dozen ideas, ruthlessly varied.
THE ONE-LINE ARCHITECTURE: a track = MOTIF x REGIONAL PALETTE x PURPOSE CLASS, each of
the three authored once and combined per chapter.
The rubric shipped today scores DIALECT CONFORMANCE. It cannot say what a MASTERPIECE has,
because it was derived from published analyses, not from the recordings. Phase A fixes that.
Legal posture (settled by Josh's own Sonniss line, 2026-08-05): analysis for
UNDERSTANDING is ordinary use — listening with a tool is listening. The bytes NEVER become
model input: no training, no fine-tuning, no conditioning, no derivation. We measure them the
way a musicologist measures them, and we compose our own music. Source audio is what Josh
already owns or buys; nothing is redistributed and no exemplar audio ever enters a generator.
Per-exemplar measurement (the FORMULA CARD):
| Dimension | Measured |
|---|---|
| Duration & structure | total length; section boundaries (novelty-curve segmentation); section count; repeat map |
| Tempo | BPM; tempo curve (rubato, accelerando, phase shifts) |
| Metre | signature; metric changes; dominant onset period (measured, never assumed) |
| Tonality | key; mode; modulation map with pivot points; harmonic rhythm (chords/bar) |
| The hook | onset time (seconds AND bars); pitch contour; interval vocabulary; range in semitones; the ONE distinctive deviation |
| Motif economy | distinct motif count; recurrence map; transformation types (inversion, augmentation, fragmentation, reharmonisation) |
| Orchestration | instrument roster; entry timeline (who plays when); voice count over time; doubling map |
| Texture & density | notes-per-second curve; polyphony curve; tutti vs solo ratio |
| Dynamics | dynamic range (p95-p10 dB); the arc shape; where the peak sits proportionally |
| Register | span; the melody's tessitura; bass floor |
| Cadence & ending | cadence census; final cadence type; loop or through-composed; how it ends |
| Timbre | spectral centroid curve; the signature colour |
Output per track: formula_cards/<track_id>.json + a human page. ~100 cards.
Instrument controls (the standard's own honesty): every extractor self-tests against
synthetic material with a KNOWN answer (a 32-bar AABA at 120 BPM in D minor with a planted
hook at bar 3), and against two of OUR tracks whose ground truth we authored. An extractor
that cannot recover a known answer does not get to describe Uematsu.
Cluster the 100 cards by PURPOSE CLASS (Josh's own seven headings) and derive, per class,
the TARGET BANDS and the RECURRING MECHANISMS:
theme carries identity in <=8 bars (Aerith, Ezio's Family, Dearly Beloved).
hours, the ambient-to-melodic ratio (Ard Skellig, Sweden, Kokiri, Stickerbush, Dire Docks).
(Rivers in the Desert, BFG Division, Battle on the Big Bridge, Megalovania).
(One-Winged Angel, Radagon, Soul of Cinder, Dancing Mad, Hopes and Dreams).
carries (To Zanarkand, Midna's Lament, Weight of the World, Lavender Town).
(Suicide Mission, Snake Eater, Life Will Change).
(Answers, Baba Yetu, Setting Sail Coming Home, Scars of Time).
Deliverable: MASTERPIECE_STANDARD.md — per class, the measured bands and the named
mechanisms, each citing the exemplars it came from. THIS REPLACES the proxy rubric as the
scoring bar.
The 30-year bar is a claim about MEMORY, and no measurement proves it. What CAN be built is a
scorer whose hypotheses are testable against the corpus:
VALIDATION IS THE WHOLE POINT: the scorer must rank the 100 exemplars ABOVE our 20 landed
tracks (a positive control it can fail), and must rank sensibly WITHIN the exemplars against
a small human-ranked subset. A scorer that cannot separate Uematsu from our own first pass is
not a bar; it is a mirror.
Layer 1 — THE MOTIF BANK (~40-60 motifs, authored once, at the masterpiece bar):
the protagonist; the companion(s); the healer's child; the antagonist pattern (held unnamed
pre-Ch-38); the realm; each of the 22 threads that carries musical identity; the major
factions; the recurring creatures/bosses. Each motif: a head cell, its transformation family,
and its declared emotional register.
Layer 2 — THE REGIONAL PALETTE (one per region, ~13 for the slice, ~76 for the arc):
period-and-culture-true instrumentation, modal language, rhythmic cells, and ornament
practice — grounded in that region's own attested music (the scoring law), then RAISED: the
palette is the vocabulary, not a folk-recording pastiche. This is where "inspired by the
cultural region and authenticity but then made into a masterpiece" lives.
Layer 3 — THE PURPOSE CLASSES per chapter (the 12-20): exploration, traversal, town/
settlement, sanctuary/rest, discovery/wonder, ordinary battle, boss (per phase), stealth/
tension, sorrow/loss, festival/trade, night, the chapter's own set-piece, the chapter theme
itself. Which classes a chapter gets is READ from its own canon (a combat-light chapter
trades battle tracks for discovery and ambient; a boss chapter spends on phases).
A track is then specified as: motif(s) x palette x purpose x the formula card its class
demands. Unique because the combination is unique; coherent because the motifs recur; and
authentically regional because the palette is.
1. Chapter-set batching — a chapter's 12-20 tracks are composed as a SET so they cohere
and share the palette; never one-off.
2. Compose symbolically to the card (the proven authored path), then realise.
3. Measure against the MASTERPIECE_STANDARD, not the proxy rubric; iterate until the
class bands are met and the predictor clears its floor.
4. Josh's ear is the final gate, on the site, per chapter set.
5. The realisation ladder — today's sfizz + CC0 stack got us to "complete track". The
masterpiece bar likely needs a better sample library rung (BBC SO Discover is free and
already queued as rung 2) and, for hero cues, eventually real players. NAMED as a decision
with a cost, not assumed.
6. The concert deliverable — every hero cue emits real notation (MusicXML + parts) so an
orchestra can play it. The Broadway/concert ambition is a FORMAT requirement from day one,
not a post-hoc export.
10-track PILOT spanning all seven classes; prove the rig recovers known answers.
whole program, because everything downstream inherits it.
It does not ratify any exemplar's material into our score; it does not put exemplar audio into
any generator; it does not claim the 30-year outcome can be measured — only that the features
the exemplars share can be, and that failing to share them is disqualifying.
"I'm pretty sure I can buy or download them. You can listen to them and build the equations to
understand what is good. this is math and theory - we aren't training the generator. it's so
you can know what to prompt so we can make our original video game music masterpieces"
RULED AND SETTLED — this is exactly the posture in section 2: MEASUREMENT AND THEORY, never
model input. The exemplars are STUDIED to derive the equations; the equations tell the
factory what to COMPOSE. No exemplar audio touches a generator, is redistributed, or is
conditioned on. Same shape as the Sonniss line (listening is use; bytes-as-model-input is
the only prohibition), and the same shape as any composer studying scores.
ONE HONESTY NOTE ON WHAT "LISTENING" MEANS HERE: the director does not hear audio
subjectively. Its listening is MEASUREMENT — spectral, structural, harmonic and dynamic
extraction — plus published theory. For deriving equations this is a STRENGTH (it recovers
exact numbers a human ear estimates), but it means the predictor of section 4 can never
certify "this is beautiful"; it certifies "this has what the beautiful ones measurably have."
JOSH'S EAR REMAINS THE ONLY SUBJECTIVE GATE, and the predictor's validation set is calibrated
against it.
Audio files, purchased or downloaded, dropped at D:/audio/exemplars/ in any common format
(FLAC/WAV preferred for measurement fidelity; MP3/OGG works — the rig declares format-induced
limits per card). Naming does not matter; the rig fingerprints and matches against the corpus.
Partial sets are fine: the rig scores what it has and declares what is missing rather than
guessing, and the pilot needs only 10 tracks spanning the seven classes.
THE LIST IS NO LONGER IN THIS DOCUMENT. It lives in two wired artifacts, and this section is
subordinate to them:
build/audio/exemplars/CORPUS.json — the machine-readable corpus. 182 rows: Josh's own 116-track list (three culled by his ruling, recorded in a binding one-directional excluded
array that no wave may reverse) plus the 68-row 1990-2008 expansion wave, each expansion row
carrying the compositional GRAMMAR it contributes that the base list lacked. Seven purpose
classes. Each row's s field is the acquisition state and flips MISSING → ACQUIRED → CARDED.
build/audio/exemplars/BUY_SHEET.md and its phone-first twin buy_sheet.html (published at https://claude.ai/code/artifact/ab157f80-9a79-4b5c-876f-08ae9c41e137) — the verified route
from every row to a file on disk, organised by storefront. Carries the STARTER TEN that
satisfies the pilot rule above for $64.59, and — the part that matters to the rig's honesty —
the 17 rows with NO legitimate route at any price, each with a named substitute and an
instruction to card it SUBSTITUTED or ARRANGEMENT-TIER rather than measure a cover and report
it as the source.
Rung 1 therefore does NOT wait on a complete corpus and must never claim one. It waits on ten
files, and every card it emits declares its own tier.
Section 10 above names exactly one human step. Everything after the drop is
harness/music_gen/acquire_exemplar.py, so the FIRST purchase is a one-command test rather than
an afternoon of finding out whether the file even decodes.
THE COMMAND, verbatim, after dropping files at D:/audio/exemplars/:
D:\audio\rig\venv\Scripts\python.exe harness\music_gen\acquire_exemplar.py --scan
Run it from the repo root. It proofs every new file in the drop folder and leaves already-proofed
files alone, so it is the standing command and not a one-shot — buy three more next month, run the
same line, and only the three new files cost anything.
WHAT IT DOES, IN ORDER, EVERY LINK ABLE TO FAIL LOUDLY:
is a synthetic signal with a planted 120 bpm, planted lane entries at 8 s and 16 s, a planted
drop and a 2-second break at 24 s, encoded to AAC and decoded back through the same ffmpeg path
a purchased .m4a takes. A negative control (one steady tone) must yield zero events, and a
zero-guard requires every reported curve to be populated and non-degenerate — together they
close both doors, so the rig can neither pass by reporting zeros nor by reporting events
everywhere. This is section 2's instrument-controls rule, armed rather than described.
harness/music_gen/rig_config.json; the system PATH is never modified). AAC/M4A/ALAC need this — soundfile cannot read them.
REJECT_DRM and a pointer tothe BUY_SHEET's DRM-free route, not with a silently-empty measurement.
section 11 into curves and events: the tempo function, the lane entry/exit map and interlap
matrix, the event grammar (drops, raises, breaks, tutti, silence budget), the onset type and
flow curve, and the complexity step map.
build/audio/exemplars/formula_cards/<track_id>.json plus the human page section 2 asksfor beside it.
CORPUS.json row on its tags and filename, and REJECTS rather than guesses. An unmatched file is still measured and still carded, under an UNMATCHED__ id — the
work is never thrown away, it is reported, and --id EX_NNN attaches it.
MISSING → ACQUIRED, keeping every existing field and adding an acq block(sha256, source path, format, codec, sample rate, channels, duration, acquired_at, and the paths
to the card and the proof record).
build/audio/records/exemplars/<track_id>.acquisition.json.
WHAT IT PRINTS. On success, per file: the matched row, the measured length / tempo / key / section
count / lane count, the event counts and silence budget, the card path, and MISSING -> ACQUIRED.
It closes with the corpus tally. On failure it prints the verdict, the reason, and the top scoring
candidate rows so the next move is obvious. On an empty drop folder it says so and reports how many
rows are still MISSING rather than exiting silently.
EXIT CODES ARE HONEST: 0 proofed (or nothing to do), 2 the control failed and nothing was
measured, 3 the file was measured and carded but no corpus row could be claimed, 4 the file
could not be decoded or is protected.
OTHER INVOCATIONS: --self-test runs the control alone (this is gate music_acquire);
--audio <path> proofs one file; --id EX_NNN overrides the match; --dry-run measures and cards
without flipping; --no-corpus is the decode-path test for audio that is not an exemplar at all.
THE LEGAL POSTURE IS THE CORPUS'S OWN AND IS RESTATED ON EVERY CARD AND EVERY RECORD: analysis for
understanding only; the audio is read and measured, never copied, never redistributed, and never
model input of any kind.
"yeah you can analyze and mathematically create the patterns behind all different lanes and
interlaps on a track, how many drops or raises or when and by how much of tempo adjustments,
how they start and then pick up and how they flow throughout, which tonal harmonies and
instruments and how things play together, when to introduce more and provide complexity, what
types of scenes and areas require which, what types of instruments and ambience to combine,
how to do all of this - what i mentioned is the floor, it's our standing ruling you expand
within our vision to create our masterpiece"
THE STRUCTURAL UPGRADE THIS FORCES: section 2's card was SCALAR (a tempo, a count, a range).
Josh is describing CURVES AND EVENTS — what enters when, what drops where, by how much, and
why there. A masterpiece is not "120 BPM with eight instruments"; it is a TIMELINE. Every
dimension below is therefore measured as a FUNCTION OF POSITION (normalised 0-1 through the
track AND in bars), not as a number. The formula card becomes a SCORE-SHAPED OBJECT.
Measured per LANE, not per mix — source-separated stems where separation is clean, symbolic
transcription where it is not, and the method DECLARED per card (a lane read off a muddy
separation is marked low-confidence, never reported at equal authority).
arpeggio · percussion (pulse vs colour) · texture/atmosphere · vocal · solo feature.
load-bearing artefact — it is literally the arrangement.
Masterpieces have STRUCTURE here (pads never coexist with the ostinato; the countermelody
only enters after the lead has stated twice), and the matrix is where that shows.
arrangement's story.
doubling intervals when two lanes share a line (octaves, thirds, sixths — Wise's signature).
Every discontinuity detected, classified, timed and MAGNITUDE-MEASURED:
delta, bars held, what re-enters first}.
restraint is a masterpiece mechanism and the current factory has no vocabulary for it.
Not a BPM — a tempo FUNCTION: baseline, every adjustment event {position, delta BPM, ramp vs
step, duration}, rubato depth per section, fermatas, metric modulation, and any metre change
with its pivot. Plus the GROOVE layer: swing ratio, syncopation density, and the rhythmic cell
that repeats.
fade-in · pickup/anacrusis · silence-then-hit.
class we derive the ARCHETYPAL SHAPE — the exploration arc is not the boss arc is not the
lament arc — and a composed track is scored against its class's curve, not a scalar band.
**FINDING 2026-08-07 — DO NOT PIN THE CLIMAX, AND DO NOT LET THE GOLDEN SECTION IN THROUGH THIS
SECTION. Measured on the corpus cards at HEAD, the dynamic-peak position has a median of 0.575**
and the flow-peak a median of 0.606, which sits temptingly near 0.618 and would read as a
confirmation of the golden-section climax claim if only the median were printed. The spread is the
answer: p10 to p90 runs 0.16 to 0.87, and only seven of the corpus rows land inside 0.55 to 0.70.
Pinning every cue's peak near the golden section would therefore make our output MORE UNIFORM THAN
THE MUSIC WE ARE TRYING TO MATCH — the tracks Josh has loved for decades put their loudest moment
almost anywhere, and the ones that put it early are not doing it wrongly.
This is the exact failure mode this section's own last sentence already guards against ("scored
against its class's curve, not a scalar band"), and the finding is recorded here so that a later lane
reaching for a single number has the measurement in front of it. **The class curve stays the unit.
A single global climax position is banned, and the median above must never be quoted without its
spread.** A related caution from the same pass: the energy peak and the loudness peak are not the
same event and do not co-locate — on EX_124 the flow peak sits at 0.272 and the dynamic peak at
0.4595 — so "the peak" must always say which one it means. Evidence:
docs/proposals/music/MUSIC_CRAFT_DOCTORATE.md §4 and
docs/proposals/music/craft_research/HOOK_AND_TENSION_ARCHITECTURE.md §8.
Chord vocabulary with frequencies; borrowed/modal-mixture inventory; the SURPRISE CHORD and
its normalised position (masterpieces place it, they do not sprinkle it); harmonic rhythm as a
curve; modulation map with pivots; pedal/drone usage; cadence census per section boundary.
INSTRUMENT COMBINATORICS: which pairings actually occur and in what register relationship —
mined as an association table so the composer inherits real orchestration practice
("clarinet + harp at a sixth in the lower-middle" is a fact we can measure, not a guess).
A single COMPLEXITY INDEX per bar (lane count + harmonic rate + rhythmic subdivision +
register span + polyphony), and its STEP MAP: where it increases, by how much, and what
mechanism does it (a new lane? a subdivision change? a modulation?). The finding we expect and
must verify: masterpieces step complexity at STRUCTURAL boundaries and hold it flat inside
sections — variation comes from ORCHESTRATION, not from constant churn. If the corpus says
otherwise, the corpus wins.
**MEASURED 2026-08-07, AND THE CORPUS SAYS OTHERWISE. THIS SECTION'S OWN LAST SENTENCE THEREFORE
SETTLES IT AGAINST THE EXPECTATION ABOVE.** Over the formula cards at HEAD, restricted to tracks of
45 seconds or more, corpus rows put 0.358 of their complexity steps at a section boundary
against their own album siblings' 0.500 — the loved tracks land a SMALLER share of their steps
on the seams, not a larger one. If change were carried by the form, the steps would sit on the
boundaries; they do not, they sit inside the sections. Six further readings agree with it: a corpus
row carries about half again as many structural boundaries as its album's filler (16 against 11)
while making only 17.9% of them strong against 36.4%, fires strong events at 0.88 per minute
against 1.89, and brings a new voice in more often and more quietly (4.55 lane-adding raises per
minute at 7.49 dB against 3.10 at 10.09 dB).
The prediction was half right and its halves were swapped. Variation does come from
orchestration rather than churn — that part holds, and holds strongly. What is wrong is the
LOCATION: re-treatment is CONTINUOUS and lives inside sections, under a sparse ridge of about three
genuine structural events, with a hierarchy ratio (largest event over typical) of 1.91 against the
siblings' 1.65. The corrected law is therefore two-tier — a TIDE of small treatment changes every
four to eight bars (the corpus does it every 5.79) under a handful of LANDMARKS — and it is the same
fact as Josh's three-to-five-cycle floor read at a second scale.
Derivation and full numbers: docs/proposals/music/craft_research/REPETITION_AND_VARIATION_FORM.md
§1 and docs/proposals/music/MUSIC_CRAFT_DOCTORATE.md §2 and §4. Every figure is re-derivable from
complexity_law.step_map, duration_and_structure.novelty_peak_strength and
event_grammar.raises on the existing cards; nothing new had to be measured. **Honest register: these
are post-hoc descriptive reads with no calibrated multiplicity null of their own — the boundary-count
row is a restatement of section_count, which does clear the bar at the 99th percentile, and the
rest agree with it directionally.** The expectation above is superseded as an EXPECTATION; it is left
standing in the text because a program that quietly deletes its wrong predictions cannot be audited.
Each exemplar carries its GAME CONTEXT (what it plays under). Cluster those contexts into
scene archetypes — first-arrival, safe town, hostile wilds, dungeon interior, sacred site,
travel/traverse, ambush, duel, multi-phase boss, revelation, loss, farewell, credits — and
derive, per archetype, which formula the exemplars actually use. THIS IS THE TABLE THE GAME
QUERIES: a Humanity scene declares its archetype (from its own canon rows) and the formula
follows. It is also the bridge to the cue system: our EG rows, encounter phases and region
pages already declare scene type, so the map has a real consumer on day one.
The game-audio half, which no rung has touched: how score and world sound share the frequency
and attention budget. Per scene archetype — the ambience bed's spectral occupancy, where the
score sits above/below it, ducking behaviour, the diegetic/non-diegetic seam (a musician IN
the scene versus the score), and the SILENCE-FOR-AMBIENCE budget (the moments the score
deliberately yields). Composes with the 680-row SFX registry and the 650-cue system already
in the engine.
Not in Josh's list, added under the expand-within-vision ruling: how a cue BECOMES another
cue — the transition grammar (stinger, bridge, layer-swap, hard cut on a downbeat), and the
multitrack/stem architecture that makes it possible. Reason it belongs: 12-20 tracks per
chapter only feel like one world if the seams between them are composed. If Josh vetoes,
tracks stay independent and the seams are handled by crossfade.
Not prose. Per purpose class, a PARAMETRIC SPECIFICATION a composer (human or ours) can
target: the energy curve shape with tolerance bands · the lane entry schedule as a function of
form position · the event schedule (how many drops/raises/tutti and where) · the complexity
step map · the harmonic vocabulary with the surprise slot · the instrument-combination table ·
the tempo function · the silence budget · the onset type · the ending type. A composition is
scored by DISTANCE FROM ITS CLASS SPECIFICATION, per dimension, with the misses named — the
same shape as every other bar in this factory, but pointed at what actually makes music good.