MUSIC_CRAFT_DOCTORATE.md

music/MUSIC_CRAFT_DOCTORATE.md

MUSIC CRAFT — THE DOCTORATE PASS

CANON SUBORDINATION — this document is PROPOSAL-TIER: it serves canon and never outranks it.
Canon served: the composition floor ruled by Josh on 2026-08-07 -- docs/spine/DECISIONS_PENDING_JOSH.md,
MUSIC ROUND 1 GRADED + THE COMPOSITION FLOOR (commit da627060); the theme rows of
registries/T0_Theme_Registry [DRAFT v0.1].
If this document disagrees with canon, CANON WINS and this document is the defect.
Tier: RESEARCH SYNTHESIS, proposal-tier. This document contains no canon and rules on nothing.
It is the craft layer underneath docs/proposals/MUSIC_COMPOSITION_DOCTRINE.md — the reasons its
rules are the rules, the places the reasons do not support them, and the single prioritised list of
what to build next.
Chartered by: Josh, 2026-08-07, at the tail of docs/spine/DECISIONS_PENDING_JOSH.md
(commit da627060), grading round 1 of the generated music and turning all eighteen candidates
down: *"Learn real music theory and what someone who studied music and got their phd would focus
on and improve."* Under the standing law of 2026-08-01 that ruling is a FLOOR TO ENHANCE, not a
specification to meet.
What it is not. It is not a summary of the six lane documents and it does not restate them.
Sixty-two thousand cited words sit under docs/proposals/music/craft_research/ and each lane is
authoritative for its own domain. This document carries only what no single lane could produce:
where six independent lanes CONVERGED, where they CONTRADICT what this program already believes,
the thirteen principles that actually change what we build, and the measurement debt in the order
it should be paid.
Every number about our own corpus in this document was re-derived at HEAD through
harness/music_gen/derive_patterns.py, not quoted from a prior artifact. Where a lane's figure and
a re-derivation disagree, both are printed with their denominators.

1. The six lanes, and what each is authoritative for

voice-leading derived from perception rather than authority, an ERB-computable low-interval-limit

floor, dissonance treatment and the deliberate clash, and non-Western polyphony read as

counterpoint with its own rules. Authoritative for: why many layers do or do not become a wall.

in four different units, texture types, the staged entrance schedule, density management, and the

non-orchestral layers Josh named. Authoritative for: how many layers, and when each one arrives.

architecture question, deployment across a long game, and the measurable signature of a motif.

Authoritative for: what a theme is and how it is transformed.

minimalism's written-down rules for literal repetition, cyclic traditions as engineering, and game

loop form. Authoritative for: the 3-to-5-cycle law and the novelty schedule.

memorability literature and its limits, the tension parameters one by one, arch form and climax

placement, and the drop. Authoritative for: what makes a cue land and what makes it exhausting.

practice, voice layer, environmental layer and care posture for all thirteen slice nodes.

Authoritative for: what may sound in which chapter, and what must never.

2. The one thing worth reading if you read nothing else

Josh's floor contains what looks like a contradiction with our own measured derivation, and it is

not one. build/audio/exemplars/PATTERN_FINDINGS_V2.md found that the tracks Josh named as lifelong

favourites sit BELOW their own albums on structural novelty and ABOVE them on self-similarity — they

repeat more and change less than the filler beside them. Josh's floor says change every three to five

cycles. The repetition lane resolved it by measurement rather than argument, and the resolution is

the governing idea of the whole pass 2.

The loved tracks repeat their MATERIAL and change their TREATMENT — continuously, in small

increments, under a sparse ridge of about three genuine structural events. Measured at HEAD on tracks

of forty-five seconds or more: a corpus row carries about half again as many structural boundaries as

its own album's filler (16 against 11) and makes each one SMALLER — its share of boundaries at or

above 0.7 of its own maximum is 17.9% against the siblings' 36.4%, and its strong boundaries arrive

at 0.88 per minute against 1.89. It brings a new voice in more often and raises the level less when

it does (4.55 lane-adding raises per minute at 7.49 dB, against 3.10 at 10.09 dB). And the ratio of

its largest event to its typical one is steeper, 1.91 against 1.65, which is what a hierarchy looks

like when it is real rather than flat.

So the law and the measurement are the same fact read at two scales. Josh's floor governs the dense

tier: something must change every few cycles, and in the corpus something changes every 5.79 bars.

The measurement governs the sparse tier: most of those changes must be SMALL, because a track that

makes every change a big one has spent its hierarchy and has nothing left to arrive at.

That single sentence reframes what round 1 got wrong. Round 1 was not short of events. It had far too

many of the wrong size — a median of 14.08 drops per minute against the corpus median of 5.57

(2.53x, re-derived at HEAD), which is a track shouting a structural event two and a half times more

often than any music Josh loves, while changing its treatment almost never in between.

Two corroborations are worth recording beside it. First, Josh's number was arrived at independently

by the game-audio literature: a 2012 Game Developer treatment of loop fatigue gives one to two

repeats as tolerable, THREE TO FIVE as problematic and more than five as significant fatigue — the

same band, neither source citing the other. His floor is not a preference we are accommodating; it is

a threshold that has been found twice. Second, a corpus study across 21,391 pieces puts bar-level

repetition at 62.79% and finds it flat from the NES era through the PlayStation era, across roughly a

tenfold jump in storage. Looping is an aesthetic choice, not a memory constraint, and it has never

stopped being one.

3. Where six independent lanes converged

Convergence across lanes that never read each other is the strongest evidence this pass produced.

Six items, each reached by at least two lanes independently.

counterpoint lane names accidental surface-texture integration as the mechanism and declared

stratification as the fix. The orchestration lane reaches it from Huron's finding that listener

counting accuracy collapses past three or four concurrent voices, and concludes that dozens of

parts must group into three or four perceptual STRATA. The organology lane finds the same shape in

the traditions themselves. The reconciliation of Josh's two sentences is therefore structural: many

PARTS, few STRATA. Enforce the strata count, not the part count.

audio-side methods and found none defensible for a part count; the counterpoint lane reached the

same conclusion from the card's own method note; the repetition lane found the proof, which is that

max_lanes is pinned at 9 for hits and siblings alike — a ceiling in the ten-band proxy method

rather than a fact about any music. The layer count must be a FIELD we wrote, with every audio

figure as corroboration only.

colour_change field drawn from an enumerated eleven-device list; the repetition lane derives the

same rule from the variation canon and from irama; the leitmotif lane's transformation budget is the

same discipline applied to a theme. This is the 3-to-5-cycle law made mechanical, and it is

checkable before a sample is rendered.

separable density levels — is named by the organology lane as the cross-cutting device and by the

repetition lane as the mechanism this project needs most. Two of our slice chapters are literally

gamelan cultures, and our own layer_mode: swap already means this without knowing it did.

ambience in his floor. The orchestration lane treats them as parts with register slots and entry

schedules; the organology lane independently attests body percussion in six of twelve slice nodes.

This is not decoration and it is not a texture layer; it is scored material with a frequency slot.

our own rejection filter all land on the same axis, and it is the one axis on which round 1 failed

unanimously: every one of the fourteen round-1 rejections fell out of the envelope on composed rest

and nothing else.

4. Where the research CONTRADICTS what this program already believes

These are findings for the program to rule on. Nothing here was applied.

complexity AT structural boundaries and hold it flat inside sections. Measured at HEAD, corpus rows

put a SMALLER share of their complexity steps at boundaries than their siblings do — 0.358 against

0.500. Change is not carried by the form; it happens inside the sections, which is exactly what

continuous re-treatment over a returning cycle looks like from outside. §11.6's own instruction is

"if the corpus says otherwise, the corpus wins." The corpus says otherwise.

0.575 and a flow-peak median of 0.606, which looks like a confirmation until the spread is printed:

p10 to p90 runs 0.16 to 0.87 and only seven of twenty-eight rows land inside 0.55 to 0.70. Pinning

every cue's peak at 0.618 would make our output MORE uniform than the music we are trying to match.

to 17 semitones and the card's the_hook.range_semitones, whose corpus median is 22, are not the

same object, and a cue can pass one while failing the other for no musical reason.

found this independently. motif_economy and repeat_map run on raw chroma and MFCC

self-similarity, so a theme returning a fourth higher scores LOWER than one returning in the same

key. For a leitmotif score that is exactly backwards: the transposed return is the craft, and our

instrument reads it as new material.

controls and 1181 cards with EX_014 quarantined. The live pool at HEAD is 30 / 927 / 1180 —

EX_014 and EX_094 had their quarantines retired by inspection on 2026-08-07, both legitimately.

Nothing is wrong; the derivation is simply two positives stale, and the drift is material (corpus

median dynamic range moves from 11.7 to 12.23 dB). Any lane quoting V2's tables must say which

denominator it ran at, and a V3 re-run is owed.

5. The thirteen principles that change what we build

Ordered by how much they change the next track, not by how interesting they are.

1. COMPOSE THE STRUCTURE BEFORE THE SOUND. A track is a section map, a cycle grid, an entrance

schedule, a novelty schedule and a hook plan before it is audio. Everything else here is

downstream of this one. PIPELINE: the pass-2 track plan; single-shot generation cannot satisfy any

clause of the floor because it places nothing.

2. MANY PARTS, FEW STRATA. Twenty-four to forty parts for a full cue, three to eight for the sparse

nature class, grouped into three or four perceptual strata. Our ten authored tracks currently

average 9.3 parts, range 4 to 19 — about three times short on the full-cue class, and the gap is

known exactly in the right unit before anything renders. PIPELINE: stratum and

part_count_floor on the manifest; the battery refuses out-of-band.

3. THE MANIFEST IS A SCHEDULE, NOT A ROSTER. Every part carries entry, exit and rest spans. Three

predicates then become checkable before rendering: no part runs the whole cue except a declared

ostinato, at least one real subtraction occurs after the midpoint, and the peak part count lands in

the last third. PIPELINE: entry_bar / exit_bar / rest_spans.

4. CHANGE THE TREATMENT CONTINUOUSLY, THE FORM RARELY — AND ON TWO TIERS. A TIDE of small treatment

changes every four to eight bars (the corpus does it every 5.79), under about three LANDMARKS per

track arriving at roughly 0.88 per minute where the filler beside them fires 1.89. Josh's law

governs the tide; the measurement governs the landmarks; the hierarchy between them is the whole

deliverable, and it has a number — largest event over typical event is 1.91 for corpus rows against

1.65 for filler. PIPELINE: the novelty schedule carries a MAGNITUDE and a TIER per event, not just

a time.

5. THE LAW IS NON-DECAY, NOT ESCALATION. This corrects an assumption the program was carrying.

Corpus strong events are close to uniformly spread — median position 0.477, with 47.2% in the

second half — so a cue that escalates monotonically toward its loop point collapses at the wrap.

What separates the loved tracks is that they do not DECAY: complexity in the last third minus the

first third is -0.020 for corpus rows against -0.094 for siblings. Gate the non-decay, never the

escalation.

6. A NOVELTY EVENT ONLY COUNTS IF IT ADDS OR REMOVES A FOLLOWABLE VOICE, or changes the rule the

music is running under. A filter sweep, a reverb change, a volume ramp and a pan move are CHEAP and

satisfy nothing. The warrant is measured: corpus raises average 7.49 dB against the siblings' 10.09

while firing lane-ADDING raises at 4.55 per minute against 3.10 — novelty comes from arrivals, not

from loudness. PIPELINE: classify on lanes_added and lanes_removed; anything rounding to zero

is cheap and cannot close a cycle.

7. TAKE THE DENSITY-LEVEL SHIFT AS THE DEFAULT LANDMARK. Javanese irama re-tiles a fixed cycle at

separable density levels and Carnatic laya reached the identical ladder independently; it is a real

structural event that needs NO new material, and it leaves a measurable signature — a step in note

rate and per-bar complexity while repeat similarity STAYS high. Two of our slice chapters are

gamelan cultures and our own layer_mode: swap already means this.

8. REPETITION IS LICENSED ONLY BY A NAMED COLOUR CHANGE. Carrier, family, register, solo or tutti,

mute, attack mode, articulation, added counter-line, reharmonisation, composite timbre, timbral

echo. A verbatim repeat with no named device is a battery failure. PIPELINE: colour_change

required on every licensed repeat.

9. SUBTRACTION IS THE STRONGEST EVENT AVAILABLE. Removal creates more energy than addition, and the

loudest moment is followed by the least. Round 1 dropped 2.53 times more often than the corpus and

was rejected unanimously on that axis, which is what happens when the strongest device is spent as

small change. PIPELINE: a required removal event of declared depth after the flow peak; composed

rest held inside the corpus band of 2.68 to 17.43.

10. STATE THE THEME PLAINLY ONCE, THEN NEVER NEUTRALLY AGAIN. The plant is unhurried and in the home

carrier; every later appearance is fragment, transposition, augmentation, reharmonisation, carrier

migration or mode change. If it needs its orchestration to be recognisable it is not a motif — test

it through three carriers before spending anything else on it. PIPELINE: plant_node, the

transformation budget, and motif_compare.py, which exists today and is not wired.

11. COMMON CONTOUR, ONE UNCOMMON LEAP. The single melodic prescription the empirical work actually

supports. Everything else about melodic memorability in the literature is weaker than it is usually

quoted as being, and this document says so rather than dressing craft consensus as evidence.

PIPELINE: the_hook.pitch_contour plus one licensed distinctive_deviation.

12. AUTHENTICITY IS A CONSTRAINT ON MECHANISM, NOT ON TIMBRE. The traps the organology lane found are

not wrong-sounding instruments; they are right-sounding instruments from the wrong century — gong

kebyar is 700 years late for Bali, the djembe is a twentieth-century global export, the likembe

arrived with colonial labour migration, and a bone flute at Olduvai is 1.76 million years early.

Where the record is genuinely absent, invention is allowed and required; it just has to be declared

as invention rather than dressed as evidence.

13. MENACE NEVER RIDES THE PLACE'S OWN REGISTER. Already canon in every theme row's

deny_register_list; the exoticism literature independently supports it as craft. The antagonist

material carries its own darkening operations and travels with him. PIPELINE: a per-cue check that

a region's instrumentation set never co-occurs with the antagonist's darkening operations — which

is checkable from the cue table as soon as cues carry both fields, and is not checked today.

6. The measurement debt, in the order it should be paid

The four instruments this sitting built cover the first four. Items five through seven are named and

NOT built, which is the honest state.

from an STFT with no transcription. Two lanes independently named it the cheapest high-value

addition available, and it is the only thing that would let us grade "harmonies and clashes" on a

render at all. Built this sitting inside the noise-versus-structure instrument.

twelve circular rotations, with the winning rotation recorded so the INTERVAL of the return is

itself reported. This turns "the theme comes back in a new key" from invisible into a measured event

with a number attached. Built this sitting inside the hook-presence instrument.

from audio is corroboration whose error must be published beside every number it produces. Built

this sitting across the layer-census instrument and the pass-2 manifest.

needs anything else, and the period must be cross-checked between methods rather than asserted from

one. Built this sitting inside the cycle-law instrument.

2's resolution: take every pair of sections whose repeat similarity is 0.95 or higher and ask

whether the set of sounding instruments differs. If corpus rows re-orchestrate across identical

music more than their siblings do, the reconciliation is confirmed rather than inferred. It needs a

real instrument roster — source separation or a trained recogniser — and it is the highest-value

measurement this pass could not build.

mode and raga and qenet identification, and expanding the anti-plagiarism signature set past its

fourteen declared gaps. Polyphonic transcription of a forty-layer orchestral render will not become

reliable soon; separate-then-transcribe-per-stem is the honest partial path, with the caveat that

the available separators are trained on popular-music stems and orchestral material is out of their

distribution.

that restores what it took. Our card carries every component — raises[].end_s, drops[].at_s,

bars_held_s, re_enters_at_s — and nothing computes the PAIRING. It is the highest-value derived

feature available from data we already have.

7. What this contracts pass 2 to

declares its cycle period, its novelty schedule, its layer manifest with entries and exits, and its

hook statement and return times; the instruments measure the render; the two are compared. A

prototype that produces a good mixdown by abandoning its plan is a FAILURE of this architecture even

if it sounds better, because the thing being proved is that the structure survives realisation.

allocated so no two parts in a stratum fight for the same band.

and about three of them large.

its care line quoted, and the menace, where a cue has any, carried by material that is not of the

place.

structure read. The honest tier at this rung is STRUCTURE. A CC0 sampler orchestra is not a live

session and must never be labelled as one.

7.5 THE UPBEAT REGISTER — Josh's taste, calibrated against the corpus before it was applied

Josh, 2026-08-07 night, in the round-2 floor addendum: "I like more upbeat." His taste runs

energetic and adventurous — the classic JRPG overworld optimism. Under the standing law that is a

floor to enhance rather than a preference to accommodate, so it lands here as a REGISTER DEFAULT

per purpose class.

It was measured before it was applied, and the measurement stopped a naive over-correction that

would have made the score worse. Across the thirty corpus rows — the tracks Josh named as loved for

decades — the register reads:

purpose classnmajor-mode sharemedian bpmmedian spectral centroid Hz
EXPLORATION84 of 8114.91845
BATTLE50 of 5117.52575
TITLE_CHARACTER40 of 4117.72277
MELANCHOLIC42 of 4136.41665
CREDITS_TRIUMPH44 of 499.41977
TENSION31 of 3117.51843
BOSS22 of 2126.11551
ALL3013 of 30 (43%)

THE FINDING, and it is the reason this section exists rather than a one-line instruction to write

in major: the music Josh loves is NOT predominantly major. It is 43% major overall, BATTLE is zero

of five, and TITLE_CHARACTER is zero of four. A pass that answered "more upbeat" by defaulting to

major keys would have moved AWAY from the corpus on two of its most energetic classes, and it would

have been done in Josh's name.

So UPBEAT IS NOT A MODE. In the corpus it is carried by three other things, and BATTLE is the proof:

entirely minor, and the BRIGHTEST class in the whole corpus at a median centroid of 2575 Hz against

an all-class spread of 1551 to 2575, with driving tempo and continuous forward motion. That is what

energetic sounds like in the music that stayed loved. The register is ENERGY, BRIGHTNESS and FORWARD

MOTION, and mode is nearly orthogonal to it.

The per-class defaults

actually names, it is the class most of our region cues belong to, and it is the one where the

corpus most supports him: half of it is already major and its tempo sits mid-band. Default to a

major or bright-modal centre, forward motion in the accompaniment, and a melody that climbs. This

is the strongest single application of his note.

it; the corpus is unanimously minor here and unmistakably upbeat anyway.

scale and ambition rather than by brightness of key.

Slowest median tempo of any class at 99.4 bpm, which says triumph is BROAD rather than fast.

upbeat is a lament destroyed. Josh's note is a default, not a floor to apply everywhere, and the

class whose entire purpose is sorrow does not receive it. Our own Home and Loss cue is in this

class.

The honest limits of this table

and non-Western material — several of our slice cultures are not in a major/minor system at all,

so "major share" for a gong-chime or pentatonic cue is a category error rather than a measurement.

Read the mode column as a weak signal and the centroid and tempo columns as the strong ones.

no per-class number here has a calibrated multiplicity null.

the ABSOLUTE centroid figures are not directly comparable to ours. The ORDERING across classes is

the usable part.

7.6 THE SCORE-FIRST RULING — where this program measures quality, and where it stopped being able to

DIRECTOR RULING, 2026-08-07, applied rather than flagged (methodology is creative-within-vision;

the path is shown here so it can be argued with). It supersedes the working assumption that the

formula card is where music gets graded.

The evidence that forced it

Five measurement instruments now exist against Josh's floor, every one of them built with real

controls and mutation arms, and NOT ONE separates the music he loves from the music he rejects.

AUC 0.501. Eight threshold settings swept; none separates.

it cannot adjudicate its own clause at all.

two independent research lanes as the highest-value measurement available to this program, came

back at AUC 0.4963. Its one apparently-strong separator collapsed to 0.5208 when held to the

lossless stratum and was withdrawn as codec provenance.

the WRONG side of chance — those tracks carry more restatement and better textbook hook geometry

than the corpus.

counter-melody presence lands at AUC 0.195, meaning the rejected round scores HIGHER.

Against that, the SCORE-side census separated round 2 instantly, trivially, and correctly: the

entire melodic content of every cue is one six-note two-bar cell, three genuine development

operations across four minutes, no phrase longer than two bars, no cadence anywhere. That census

took one pass over a JSON file. It is exactly what Josh heard, and it needed no DSP at all.

The ruling

authored symbolically before it is played, so for our own output the note list is available,

exact, and free. Melodic complexity, development operations and their chain depth, phrase

grammar and cadences, counter-melody independence, voice-leading, range and playability, the

ornament layer, and the layer count are ALL exactly computable from the score. Measuring them by

extracting a monophonic line from our own finished dense mix — and then reporting the extraction

error as the headline caveat — is measuring a thing we already know, worse.

and become the RENDER-FIDELITY CHECK: does the mixdown express the structure the score declares?

That is the question they are actually good at, it is the question the pass-2 plan-versus-render

contract already asks, and it is the one place their measured error is tolerable because both

sides of the comparison are ours.

audio is the only way to measure them and every corpus number stays audio-derived. What changes

is the inference direction: the corpus sets BANDS and shows what the masters' structures look

like; it does not certify our cues. A corpus band is containment, never a target — a rule this

program already learned once and now applies to a second class of number.

hook presence's negative axes are reported as measurements and are not permitted to grade round 3.

Tuning an inverted axis until it points the right way is fitting to three data points.

What this does not license

listener knows whether it works, and nobody in this program has heard any of this music. Josh

remains the first ear and the honest tier stays STRUCTURE.

and they bound where the defect is NOT, which is why round 3's defect had to be melodic: the

structural instruments had already cleared that ground.

corpus. Their calibration comes from the craft literature and from our own before-and-after, and

that weaker warrant must be stated wherever they are quoted.

8. Honest limits of this document

checked against a human ear, and the program's own record is that the ear is the thing it keeps

failing. Josh is the first ear and remains so.

found craft consensus rather than evidence it labelled it, and those labels were preserved here.

regions, and the reasons differ by region — Bali fails on tuning, West Africa on mechanism (the

dundun's defining property is a continuous tension glide no fixed-pitch sample produces), Flores on

simple absence. The single highest-leverage acquisition named is a CC0 voice set, whose absence

alone blocks four nodes.

arithmetic on fields we already have. None of them measures whether the music is beautiful. That

remains outside every instrument in this program, and every artifact it produces says so.

Generated by harness/site/structure_site.py — the URL path is the repo path. review root