music/COUNTER_MELODY_AND_DIALOGUE.md
CANON SUBORDINATION — this document is PROPOSAL-TIER: it serves canon and never outranks it.
Canon served: the composition floor and the ROUND 2 GRADED ruling at
docs/spine/DECISIONS_PENDING_JOSH.md (tail) — Josh's verbatim words being that the melodies are
way too simple and that intelligence is owed both to the melodies AND to how everything
orchestrates. This lane owns the second half of that sentence.
If this document disagrees with canon, CANON WINS and this document is the defect.
TIER: RESEARCH SYNTHESIS — PROPOSAL-TIER. This document is HOW, never WHAT. It sets no canon, names
no region content, and changes no spine or registry row. Every substantive claim carries a source
(URL, or book plus author plus chapter); where sources disagree the disagreement is stated; where a
statement is craft consensus rather than evidence it is labelled CRAFT CONSENSUS. Every principle
carries a PIPELINE HOOK line marked CODE-ABLE or MEASURE-ONLY, naming the field and the operation.
WHAT THIS LANE ANSWERS. Volume 1 of this craft series —
docs/proposals/music/craft_research/COUNTERPOINT_AND_VOICE_LEADING.md — taught how simultaneous
voices avoid DESTROYING each other: species discipline, Huron's six perceptual principles, the fusion
hazard in parallel perfects, the ERB spacing floor, the stream ceiling. That is the negative craft and
it is done. This document teaches the positive craft, which is a different subject: how voices TALK to
each other. A texture can satisfy every rule in volume 1 and still contain exactly one musical idea,
and a listener will call that simple, because it is. Where this document touches a topic volume 1
opened — invertible counterpoint, heterophony, the deliberate clash — it says so and supplies the
writing constraints and operations rather than the definition.
The finding that motivates the lane, computed from build/audio/pass2/plans/PASS2_FLORES_REST.json
rather than asserted.
The cue carries 32 layers across 24 cycles. Its pattern.kind census: 8 statement, 7 ostinato, 4
pokok, 3 pad, 3 kotekan, 2 colotomic, 2 arpeggio, 2 counterline, 1 drone. All eight statement layers
compile from the SAME six-note head cell — the compiler at harness/music_gen/pass2_realise.py reads
plan["hook"]["head_cell"] for every one of them and applies a transform. The kotekan layers are
deterministic splits of the pokok, the pads and arpeggios come from the harmony grid, the ostinati and
colotomic strokes are fixed-key figures. So the cue contains exactly ONE invented melodic idea plus two
authored counterlines; everything else is that idea, its skeleton, or the frame under both.
The two counterlines are the whole of the cue's answering material, and both fail the tests of section
2, measurably:
runs 7, 6, 7, 8, 6. The head cell spans sixteen semitones and contains a licensed plus-nine leap. One
is a tune; the others are a scale with a rhythm.
onsets inside that window fall at 0, 2, 4; F07's at 0, 4, 6. Every counterline onset lands on a head
onset. By McAdams's concurrent-grouping account onset synchrony is a primary fusion cue, so these
lines are written to fuse with the tune rather than to answer it.
eight beats, the fraction on which the counterline sounds and the head does not is zero. Neither ever
speaks while the tune is silent. The head cell sounds continuously for all eight of its beats, so no
complementarity was available to be written — a fact about the HOOK that section 2.4 turns into a
constraint.
cycles, and two cycles have none. The maximum is three, in five cycles, and every one of those
includes at least two carriers stating the same head cell.
That is the technical content of "the melodies are way too simple". It is not a layer-count problem —
32 layers is inside the class band and the strata are balanced. The cue has one thing to say and
thirty-one ways of accompanying it.
The umbrella term is counter-melody: a sequence of notes written to be played simultaneously with a
more prominent lead melody, a secondary melody in counterpoint with the primary one, whose presence
makes the texture polyphonic rather than homophonic
(https://en.wikipedia.org/wiki/Counter-melody). Beneath the umbrella sit four objects with different
contracts, and conflating them is how a pipeline ends up emitting a scale and calling it a
counter-melody.
| object | contract | independence | the clause that binds our pipeline |
|---|---|---|---|
| descant | above the tune, often sharing its harmonic rhythm; unison or tonic at both seams, departure in the middle | lowest | it must still have melodic interest in its own right |
| obbligato | a prominent counter-melody the performance may NOT omit, usually above the theme | medium | an obbligato belongs in the SAME adaptive layer as its tune, or the layer system will let a player mute it |
| countersubject | new material stated against the answer and recurring against every later subject entry, contrasting in phrase shape and rhythm | high | its counterpoint against the subject must be INVERTIBLE — either line workable as the lowest part |
| second tune | a line of equal melodic weight, introduced separately and combined later | highest | it costs a real invention; introduce it alone first |
Sources: https://forums.bobdunsire.com/forum/general-discussion/music/57899-counter-melody;
https://en.wikipedia.org/wiki/Obbligato and https://tristanartsonline.com/obbligato/ and
https://wikidiff.com/countermelody/obbligato; https://www.britannica.com/art/countersubject and
https://viva.pressbooks.pub/openmusictheory/chapter/high-baroque-fugal-exposition/;
https://www.jwfan.com/forums/index.php?%2Ftopic%2F37504-john-williams-best-horn-counter-melodies%2F=
and https://www.production-expert.com/production-expert-1/5-composer-tricks-to-sound-like-john-williams
(the last two are practitioner sources, CRAFT CONSENSUS).
Two honest notes on the roster. Many fugues have no countersubject at all and Britannica says so
plainly — the accompanying counterpoint is then free and does not recur, so a countersubject is a
choice with a cost rather than a default. And Golden-Age film practice puts the second tune in the
horns against a big string melody, with the counter-material sometimes carrying a darker minor colour
than the theme's own harmony to mark danger while the theme itself stays intact.
Assembled from the craft sources above and from Sevsay's framing of the secondary melody's degree of
independence — that it should function as a singable, playable tune with its own phrasing and its own
single climax, while remaining subordinate through dynamic and timbral restraint rather than through
melodic poverty (Ertugrul Sevsay, The Cambridge Guide to Orchestration, CUP 2013;
https://books.google.com/books/about/The_Cambridge_Guide_to_Orchestration.html?id=yBwu2t_5ZVIC, scanned
copy at https://archive.org/details/cambridgeguideto0000sevs). That last clause is what our pipeline
got backwards: subordination is a MIXING and SCORING decision, never a licence to write worse material.
source states this in the same words: the counter-melody is active when the main melody rests and
vice versa (https://medium.com/@NickEss/a-beginners-guide-to-counter-melodies-ebc5ae8b10cd,
https://composingthescore.wordpress.com/2016/01/08/creating-counter-melodies/). Volume 1 section 2.1
gives the mechanism: shared onsets are a fusion cue, so a line sharing the tune's onsets is written
to disappear into it.
the mechanism: parts moving together in the same direction and by similar amounts fuse.
rule, already carried in ORCHESTRATION_TEXTURE_AND_DENSITY.md section 5.4: content occupying the
melody's frequency range prevents the melody being perceived clearly in the full mix.
is the two-tunes-that-fight failure of section 8.
the four properties above plus the mix: the material is genuinely good, and the separation cues plus
gain and timbre decide which line is foreground. Volume 1 section 8's doctorate drill — singability
of every part, not just the tune — pointed at the specific case.
The two legitimate origins, and the choice has consequences.
intervals. It coheres automatically. It is also the source of the everything-sounds-the-same
complaint when it is the only method used, because it shares the tune's intervallic fingerprint and
the two can read as one idea heard twice.
second-tune object of section 2.1. It costs an invention and risks incoherence, which is why the film
practice introduces it separately and combines it only after the listener knows both.
CRAFT CONSENSUS, and this lane's ruling for our factory: a cue needs BOTH and the split should be
declared. Derived counter-material is the cheap default for middleground layers; at least one
independent second idea per cue is what makes the cue not-simple.
This is the deliverable of the section. Given a head cell H as a list of (midi, onset_beats,
dur_beats) — which is exactly the shape plan["hook"]["head_cell"] already has — produce a
counter-melody C in five deterministic steps, each with a licence to be overridden by a declared
reason.
span. Place C's onsets in the LONGEST silent or held runs of that mask and end C's notes before H's
next onset. Where H has no rests at all — the case for our current head cell, sounding continuously
across all eight beats — the rule is not that C must overlap; it is that H must be rewritten to
breathe. A hook with no gaps forbids dialogue by construction, and that is a finding about the hook.
climbs, C descends by a comparable but NOT identical amount, because exact mirroring at a fixed
interval is parallel motion in disguise and fuses.
retrograde of one of H's interval pairs; from the third onward C is free within the mode. This buys
audible relatedness at the entrance and independence thereafter — both origins of section 2.3 in one
line.
register_slot away from thecarrier's, on the side where the harmony frame leaves room. Above the tune it is a descant or
obbligato; below it is the horn-against-strings tenor counter-melody of section 2.1.
after H's apex, so the two peaks read as an answer rather than a unison shout.
A line produced by these five steps, then checked against section 5's tension table for its vertical
intervals against H at every coincidence, is a counter-melody. A line produced without steps 1 and 2 is
what our manifest currently contains.
pattern.kind of counter_melody, required fields derive_from (a layer id, normally the carrier), rhythm (complement or explicit onsets), contour (reflect
or explicit), seed (inversion, retrograde, augmentation, none) and slot_offset. The
compiler in pass2_realise.py gains one branch and derives notes from the referenced layer instead
of reading a hand-typed degree list. The existing counterline kind STAYS, because it carries the
care-line's absolute flag that lets antagonist material sit outside the region's mode.
pass2_plan.py, continuing the V-numbering: the counter-melody's onset-coincidence share with its derive_from layer stays under a declared ceiling;
its exclusive-sounding fraction clears a declared floor; its pitch range in scale steps is at least
half its carrier's. Each takes a mutation control, per the house rule that a rule never made to fail
is not armed.
Polychoral practice splits singers, often with instrumentalists, into two or more groups that engage in
lively dialogue and then join in tutti climaxes; antiphonal psalm chanting, dialogue and canon all fed
it, and it peaked with Gabrieli at St Mark's and with Schutz (Anthony F. Carver, Cori Spezzati vol. 1,
CUP 1988; https://www.amazon.com/Cori-Spezzati-Development-Sacred-Polychoral/dp/0521303982;
https://digital.library.unt.edu/ark:/67531/metadc33220/m2/1/high_res_d/dissertation.pdf). The
structural detail worth importing: Gabrieli's O Domine Jesu Christe juxtaposes a LOW choir and a HIGH
choir in a dialogue of extended passages against short phrases, and the two merge into full eight-voice
counterpoint especially at phrase ENDINGS (http://www.operatoday.com/content/2007/02/cori_spezzati_v.php).
Three transferable laws, none of which requires a cathedral.
Spatial separation is the amplifier, not the mechanism — which matters because our renderer has
pan per part and can supply the amplifier cheaply.
in length reads as repetition; one that compresses or extends it reads as a reply.
schedule, and a schedule is a field.
The smaller-scale devices, stated as operations.
a different instrument. It costs one cell and makes a phrase feel answered rather than merely
finished. Williams supplies the game-scoring form: at the second statement of the E.T. flying theme,
staccato downward figures in flutes and bells are overlaid as another point of interest while the
main theme repeats, and the horns imitate the theme one bar after the rest of the orchestra
(https://qualifications.pearson.com/content/dam/pdf/A%20Level/Music/2013/Teaching%20and%20learning%20materials/Unit-6-45-John-Williams-ET-1982-Flying-Theme.pdf,
via indexed search results — the PDF did not parse to text at time of writing, so this citation is
weaker than a direct read and is declared rather than hidden).
change, already enumerated in ORCHESTRATION_TEXTURE_AND_DENSITY.md section 7 as a colour-change
device. What this lane adds is that it is also a DIALOGUE device when the two statements are placed
as call and response rather than as a repeat: same material, different entry_cycle arithmetic,
completely different musical meaning.
the timbre stream changes. Volume 1 section 3.2 names this as dovetailing; the operation here is
that the hand-off must fall inside a phrase, because a hand-off at the boundary is a carrier change.
Our cues are adaptive and cyclic. A call-and-response pair is two layers that are individually sparse,
never simultaneously dominant, and jointly continuous — so the pair degrades gracefully when one is
muted by the layer system, which is exactly what Phillips's vertical-layering doctrine demands. A
tutti statement of the same material has none of those properties.
answers field on a layer, naming the layer it responds to plus a delay_beats. The compiler places the answering material after the referenced layer's phrase
ending. A validator predicate then checks that answering pairs never share more than a declared
fraction of their sounding time, which is the machine-checkable definition of a dialogue.
Volume 1 section 4 named invertible counterpoint as our highest-value classical device and gave the
octave and twelfth inversion arithmetic in six lines. This section supplies the WRITING CONSTRAINTS
that make an invertible pair actually work, because the arithmetic is the easy half.
A countersubject is not merely a counter-melody that recurs. It is a counter-melody that recurs
against every subject entry, contrasts with the subject in rhythm and phrase shape, and is invertible
against it (https://www.britannica.com/art/countersubject,
https://www.pianistmagazine.com/blogs/understanding-theory-part-14-fugue-2/). The rhythmic-contrast
clause is the one our pipeline needs most: it is section 2.2's complementarity rule stated as a
formal requirement rather than as advice, and it is why fugal writing does not produce the wall our
manifest produced.
At each inversion interval, some vertical intervals become dissonances under inversion and are
therefore restricted to passing or suspended treatment in the original. The rules, from the standard
references (https://www.teoria.com/en/reference/i/invertible-counterpoint.php,
https://rothfarb.faculty.music.ucsb.edu/courses/103/invertible-cpt.html,
https://www.teoria.com/en/articles/BWV885/02.php, and the Grove article at
https://docdrop.org/download_annotation_doc/invertible-counterpoint-0g1kc.pdf):
| inversion | the mapping that matters | the interval to police | the writing constraint |
|---|---|---|---|
| at the octave | 3rd becomes 6th, 6th becomes 3rd, 5th becomes 4th | the perfect 5th, which inverts to a 4th | treat the 5th as a dissonance in the original — pass through it or prepare it, never let it sit on a strong beat |
| at the tenth | 3rd becomes 8ve, 6th becomes 5th | parallel 3rds and parallel 6ths | reach 3rds and 6ths by CONTRARY motion only, or the inversion produces parallel octaves and fifths |
| at the twelfth | 6th becomes 7th, 3rd becomes 10th, 5th becomes 8ve | the 6th, which inverts to a dissonant 7th | use 6ths only as passing or suspended dissonances; the 3rd is the safe workhorse |
Bach's BWV 885 is the worked case: its countersubject appears against the subject in invertible
counterpoint at the octave, tenth, twelfth and thirteenth
(https://www.teoria.com/en/articles/BWV885/, https://www.teoria.com/en/articles/BWV885/01.php). The
horizontal-shifting generalisation — inverting in TIME as well as in register — is treated at
https://mtosmt.org/issues/mto.18.24.4/mto.18.24.4.collins.html and is a further multiplier we can take
later.
A pair written invertibly at the twelfth yields, from ONE invention: the pair as written, the pair
inverted, each line alone, and each line against a transformed statement of the other. Six
distinguishable textures. Our architecture's central economic claim — a theme authored once and
amortised across twenty cues, recorded in build/audio/pass2/PROTOTYPE_VERDICT.md section 5 as
untested past three — is exactly what invertible writing strengthens, because the multiplier applies
per cue rather than per theme.
invertible_pair block in the plan naming two layer ids and aninversion interval. A validator walks every simultaneity between the two and refuses the pair if a
policed interval sits on a strong beat unprepared. The compiler then gains a swap transform that
emits the inverted arrangement, and the section map may schedule it as a colour_change device
with no new material. This is the single highest ratio of textures gained to lines of code in this
document.
The single most important honesty point in this section. There are two separate rankings of vertical
interval tension and they disagree.
and beating. Plomp and Levelt (1965), computationally standardised by Sethares (Tuning Timbre
Spectrum Scale, https://sethares.engr.wisc.edu/consemi.html) and refined by Vassilakis, whose model
adds terms for the amplitudes of the interfering sines and for the difference between
amplitude-modulation depth and degree of amplitude fluctuation, summing pairwise roughness over all
unique two-tone pairs (http://acousticslab.org/papers/ASA142.htm,
http://www.acousticslab.org/learnmoresra/moremodel.html,
https://www.acousticslab.org/papers/VassilakisP2001Dissertation.pdf). It is a property of the SIGNAL
and the TIMBRE, and it changes with register.
the boundary precisely: in the Western tradition, where roughness is generally avoided as dissonant,
the consonance hierarchy of harmonic intervals corresponds to variations in roughness degree — but
dissonance judgements are also culturally and historically mediated and sometimes bypass roughness.
society with limited exposure to Western music, rated consonant and dissonant chords and vocal
harmonies as EQUALLY pleasant, while Bolivian town-dwellers preferred consonance less strongly than
US listeners — despite the Tsimane showing Western-like preferences on other aesthetic contrasts and
Western-like discrimination ability (Nature 535, 2016, https://www.nature.com/articles/nature18635).
A 2025 follow-up finds the preference tracks level of global integration rather than being absent
outright (https://www.sciencedirect.com/science/article/pii/S0010027725002744). The modelling
literature agrees the split is real: consonance models divide into periodicity/harmonicity,
interference/roughness and cultural-familiarity families, and roughness alone does not account for
consonance judgements (Harrison and Pearce, Simultaneous Consonance in Music Perception and
Composition, https://pure.mpg.de/rest/items/item_3257902_1/component/file_3257903/content).
For a corpus spanning ancient, mythological, modern and technological registers across dozens of
cultures this is not a footnote. The sensory ranking may be treated as physical and applied
everywhere. The stylistic ranking is a per-idiom setting, and applying the Western one to a gong-chime
or forest-polyphony cue would be a category error our own care doctrine forbids.
Rather than cite a picture of somebody's chart, compute ours, the way volume 1 computed its ERB
spacing floor. Method: the standard Sethares dissmeasure formulation with published constants (Dstar
0.24, S1 0.0207, S2 18.96, C1 5, C2 -5, A1 -3.51, A2 -5.75), a six-partial harmonic timbre rolling off
as 0.88 to the partial index, minimum-amplitude pairwise weighting, summed over all partial pairs.
Normalised within each register so the minor second reads 1.000.
| interval | roughness at C3 | at C4 | at C5 | tension class | the craft reading |
|---|---|---|---|---|---|
| unison | 0.222 | 0.045 | 0.012 | fusion | collapses the voice count; use only as a declared doubling |
| minor 2nd | 1.000 | 1.000 | 1.000 | maximum | the bite; register does not soften it |
| major 2nd | 0.968 | 0.707 | 0.477 | high, register-sensitive | brutal low, expressive high — the biggest register swing in the table |
| minor 3rd | 0.831 | 0.559 | 0.390 | moderate | unusable as a low doubling, safe from the mid up |
| major 3rd | 0.754 | 0.513 | 0.404 | moderate | the same, and the classic over-doubled interval |
| perfect 4th | 0.634 | 0.347 | 0.214 | low | sensorily mild; its stylistic tension is entirely learned |
| tritone | 0.709 | 0.501 | 0.410 | moderate | sensorily unremarkable; its tension is 100 percent stylistic |
| perfect 5th | 0.417 | 0.180 | 0.109 | fusion hazard | smooth, and therefore a fusion risk, not a tension tool |
| minor 6th | 0.632 | 0.471 | 0.405 | low-moderate | the warm counter-melody interval |
| major 6th | 0.513 | 0.307 | 0.219 | low | the safest independent-line interval in the table |
| minor 7th | 0.564 | 0.371 | 0.240 | low-moderate | far milder than its stylistic reputation |
| major 7th | 0.542 | 0.490 | 0.458 | moderate, rising with register | the one interval whose relative tension INCREASES upward |
| octave | 0.126 | 0.026 | 0.007 | fusion | reduces the voice count outright |
Four findings that change how a cue should be written, all four of which the stylistic ranking alone
would have got wrong.
above the major third. Everything a tritone does in a Western cue it does because the listener has
learned a syntax. In a pentatonic gong-chime cue that syntax is absent, so reaching for a tritone
there does not buy the tension it buys in a Western cue.
The same written interval is a mud event low and a shimmer high — the quantitative form of volume
1's minor-second ERB finding, extended to the interval a pentatonic mode produces constantly.
holds 0.542, 0.490, 0.458. It is the one clash that does not go away when you move it up, and
therefore the reliable choice when a clash must survive a bright orchestration.
is why volume 1 classes it as a fusion hazard. An arranger reaching for a fifth to add edge is
adding thickness instead.
Two honest limits. The table is computed for a single six-partial harmonic timbre, and real
instruments differ — a clarinet's odd-harmonic spectrum and a struck idiophone's inharmonic spectrum
both produce different curves, which is Sethares's central point rather than a caveat on it. And
roughness is one of three model families: the physical one, not the complete one.
CRAFT CONSENSUS, and it is per-idiom. In common-practice Western terms the ordering is perfect
consonances (unison, octave, fifth), imperfect consonances (thirds and sixths), the conditionally
consonant fourth, and the dissonances (seconds, sevenths, tritone), with Huron's aggregate dyadic
consonance the standard numeric aggregation of the perceptual literature into a per-interval-class
index (Huron 1994, Interval-Class Content in Equally Tempered Pitch-Class Sets; discussed and tabulated
in the modelling literature at
https://pure.mpg.de/rest/items/item_3257902_1/component/file_3257903/content). Our slice idioms carry
their own orderings, and volume 1 section 7.1 already records the relevant one for the Aka repertoire:
in an anhemitonic pentatonic system the melodic minor second never appears as such, so the whole top
of the sensory table is unreachable by construction and the mode is doing the tension governance.
tension_profile on the plan, valued per idiom, that maps eachvertical interval to a permitted-position class: free, weak-beat only, prepared-and-resolved only,
or forbidden. The default profile is derived from the mode's own interval content rather than from
the Western ordering, which is what keeps the care line intact on a non-Western cue.
The ranking is inert until it is placed. The shape, drawn from the species vocabulary in volume 1
section 6.1 and from phrase-arc practice in the treatises:
at their register-appropriate roughness — establishing the reference against which sharper is heard.
where the phrase's sharpest interval belongs: not after it, which reads as a stumble, and not before
it, which spends the charge early.
prepared as a consonance in the voice that becomes dissonant.
least two tension classes below the sharp event.
reads as a failure to close. The cadence takes the phrase's LOWEST tension.
Volume 1 section 6.2 gives four discriminators — consistent approach and departure, resolution or
sufficient duration, placement rather than sprinkling, and spacing for audibility. Three more are
specific to the counter-melody case and are not on that list.
because both are audible as themselves; the same semitone inside a string mass is heard as tuning.
McAdams's timbral heterogeneity as an expressive licence: schedule a clash onto the two most
timbrally distinct active parts.
signature; one landing somewhere different each time is noise. Testable from the plan as the
variance of the maximum-tension simultaneity's position, normalised within each phrase.
mud and at C5 is shimmer. Choosing the register is choosing the meaning, and the number is in the
table.
computable before render; the tension value is a table lookup by interval and register slot; the
phrase-position check is arithmetic on the section map. MEASURE-ONLY partially: a
roughness-over-time curve computed from the render's STFT via the Sethares or Vassilakis model needs
no transcription and would tell us where a cue clashes and how hard — volume 1 section 9 already
names this as the cheapest high-value measurement addition, and this section supplies the reference
values it would be read against.
The standard orchestration-analysis taxonomy assigns every part to one of seven roles: primary melody,
secondary melody, parallel supporting melody, static support, harmonic support, rhythmic support, and
combined harmonic-and-rhythmic support. Homophonic textures usually contain one primary melody;
polyphonic textures may contain several. Parallel supporting melodies double or parallel the primary
melody they support; static support is pedal tones and ostinati
(https://chromatone.center/theory/composition/texture/; the taxonomy is the one taught alongside
Blatter, Instrumentation and Orchestration, 2nd ed., though the page carries no citation of its own
and that weakness is declared).
Read our manifest against that roster and the diagnosis is immediate. Our fifteen role values —
carrier, carrier_return, carrier_close, statement, counterline, pokok, elaboration_polos,
elaboration_sangsih, figuration, bass_line, pad, drone, pulse, cycle_marker, colour — map almost
entirely onto primary melody and the support classes. Only counterline maps to secondary melody, and
section 1 showed what those two layers contain. What we are missing by name is SECONDARY MELODY as a
first-class, populated, derived role.
In a well-orchestrated cue, melodic interest MOVES. Bolero is the extreme case and
ORCHESTRATION_TEXTURE_AND_DENSITY.md section 5.1 already records its promotion law: the instrument
that has just carried the melody joins the accompaniment for the next statement, so the texture
thickens by promotion rather than by piling on. What this lane adds is that the idea generalises past
the carrier to melodic INTEREST as a quantity distributed across the strata.
Stated as a schedule, per section: which layer holds primary melody, which holds secondary melody, and
which layer just relinquished one of those roles. Three fields, and once they exist four properties
become checkable before render.
section's.
than exiting, which is the promotion law made explicit and countable.
Volume 1 section 5 and ORCHESTRATION_TEXTURE_AND_DENSITY.md section 6.1 both carry Huron's voice
denumerability result — accuracy at counting concurrent voices drops markedly from three to four and
more than about four may be impossible. This lane does not restate the ceiling; it converts it into a
target, which is a different quantity. The ceiling says how many voices a listener CAN track; the
composer's question is how many should be melodically ACTIVE, and the answer is smaller, because a
texture running at the ceiling has no headroom for an event. The working rule proposed here, CRAFT
CONSENSUS, to be calibrated against our own corpus once the measure of section 9 exists:
timbrally distinct of the set.
texture becomes a surface — McAdams's surface-texture integration, the honest name for the wall.
Against that scale our current cue runs one voice in fourteen of twenty-four cycles and never exceeds
three, and the three-voice cycles are two carriers of the same tune plus a scale. The fix is not more
layers. It is second and third IDEAS in the cycles that already have layers to spare.
melodic_role field per layer valued primary, secondary, parallel_support or none, and a per-section melodic_schedule block naming the primary and
secondary holders. Predicates: the count of simultaneously primary-or-secondary layers stays
inside the band above per section; no layer holds primary across more than N consecutive sections;
a relinquishing layer takes a support role rather than exiting. All three are arithmetic on fields
the plan would then carry, and all three are refusals rather than warnings.
Our slice cultures are mostly not Western, and Western counterpoint is not the only way to get melodic
richness. Heterophony is the alternative and it is a FIRST-CLASS system, not a simplification of
counterpoint. Volume 1 section 4 defined it in five lines as simultaneous variation of the same
melody; this section gives the three worked traditions and the operation.
Gamelan music is heterophonic, and the balungan is the skeletal core melody the elaborating instruments
elaborate (https://en.wikipedia.org/wiki/Balungan). The system governing HOW MUCH elaboration is irama:
melodic tempo and the density relationships between balungan, elaborating instruments and gong
structure, such that as the balungan's notes spread further apart in time the elaborating instruments
subdivide more finely, at two or four times the rate (https://en.wikipedia.org/wiki/Irama; Sumarsam,
Temporal and Density Flow in Javanese Gamelan,
https://sumarsam.faculty.wesleyan.edu/files/2023/01/4_Temporal_and_Density_Flow.pdf).
Our pipeline already has half of the transferable operation: it compiles kotekan splits of a pokok at a
FIXED subdivision factor. Irama says that factor is a VARIABLE moving with the structural tempo, and
that moving it is a formal event. That one change turns a static elaboration layer into a developmental
one at no cost in new material.
In Carnatic performance the violin shadows the vocalist's every curve and pause, mirroring the voice so
closely that the boundary blurs, and in improvisational sections the two TAKE TURNS elaborating
(https://darbar.org/the-violin-a-western-instrument-takes-centre-stage-in-carnatic-classical/,
https://www.hclconcerts.com/blogs/role-of-violin-in-carnatic-music/,
https://en.wikipedia.org/wiki/Performances_of_Carnatic_music). Two devices in one tradition: continuous
heterophonic shadowing, and discrete call-and-response at section boundaries. That pairing answers
section 6.2's schedule question in a non-Western idiom — the secondary voice is a shadow during
statements and a respondent during elaborations.
Volume 1 section 7.1 documents the four named constituent parts of the Aka system and their fixed
identities and variation spaces, from Furniss's analysis. What this lane adds is the classification:
because each singer varies a FIXED part rather than inventing a free line, the system is heterophony
over four models simultaneously rather than counterpoint between four independent lines. That matters
operationally, because the richness then comes from the VARIATION OPERATOR rather than from the number
of ideas — cheaper than four inventions, and a different mechanism from Western counterpoint.
Heterophony is a per-part ornamentation function over shared material. Given a line L and a part P,
produce P's realisation as L with a declared ornament density, a declared ornament vocabulary
(neighbour, passing, anticipation, delayed attack, grace), and a declared timing offset. Parts differ
in density and vocabulary, never in identity. Sharing identity they cannot fight; differing in detail
they cannot fuse — volume 1's framing of heterophony as the middle setting between unison and
counterpoint, now with the three parameters that make it executable.
heterophony pattern kind taking from_layer, ornament_density, ornament_set and offset_beats, compiling deterministically from the referenced layer's compiled
notes. This is the cheapest melodic-richness operation in this document: it needs no new invention
at all, it produces genuinely different lines, and it is idiomatically CORRECT for most of our
regions rather than an import.
irama_level on the section, multiplying the kotekan subdivision factor. A section-to-section change in irama level is a colour_change device that no existing
entry in the enumerated list covers.
Each with its mechanism and its detection.
timbre. The listener switches attention rather than integrating and the cue reads as busy rather than
rich. Mechanism: two primary melodies where the texture supports one. Detection: two layers marked
primary with overlapping arcs and comparable gain.
grid, no arc, no climax. What our manifest contains; section 1 gives the two measured signatures,
onset coincidence 1.00 and exclusive-sounding fraction 0.000, so detection is already specified.
rises and the melodic interest does not. Mechanism: figuration promoted in the mix without being
promoted in the schedule. Detection: high onset rate with a melodic_schedule showing one primary
and no secondary.
section 2.2 gives the mechanism — parallel perfects reduce the effective voice count — and section
5.2 gives the number: a fifth reads 0.180 roughness at C4 and an octave 0.026, the smoothest and
most fusion-prone intervals available, which is why a doubled wall sounds thick and empty at once.
asymmetry law is the fix; detection is that the answering layer's phrase length equals its caller's.
and becomes a texture, and the layer system can no longer use it as an event. Detection: an
exit-minus-entry span approaching the full cue with no rest_spans — the existing V11 predicate,
extended from ostinati to melodic roles.
A parallel lane is building these axes. Proposed definitions, each split into the symbolic form (which
is exact) and the audio form (which is a proxy and must be labelled as one).
The question: does this cue contain a real second melodic voice, or one tune and accompaniment.
From the SYMBOLIC score, exactly. For each pair of simultaneously active pitched layers (A, B) where A
holds primary melody, compute four component scores in the unit interval and take their product,
then take the maximum over pairs and the time-weighted mean over the cue.
measured only over spans where A sounds. Our current counterlines score 0.125.
onset. Our current counterlines score 0.000.
opposite directions, floored at zero after subtracting the 0.5 expected by chance and rescaling.
interval-class variety divided by A's, and an indicator that B has exactly one contour maximum.
The product is the right combinator rather than the mean, because a zero on any component means there
is no counter-melody, and a mean would let three good components hide a fatal one. Our current cue
scores zero on onset independence and therefore zero overall, which is the correct verdict.
From RENDERED AUDIO, as a proxy and labelled as such. The card's melodic tracker is explicitly a
PREDOMINANT-pitch tracker and to_midi.py declares that polyphony below the predominant voice is
dropped, so a true second line is not extractable today. Two honest proxies: the stem-wise route, in
which per-stem monophonic tracking on the rendered per-layer stems — which we keep on disk — gives an
exact answer for the authored lane and is the recommended path; and the mix-wise route, a
predominant-melody tracker run twice with the first melody's harmonic comb notched out, whose
confidence must be reported and which will fail on dense mixes. The stem route is not available for a
generated bed, and the measure must return no verdict rather than a number in that case.
The question: does melodic interest MOVE around the ensemble, or does one part carry it.
From the SYMBOLIC score, exactly. Four components, reported separately and as a single index.
the number of sections. Our current cue: seven layers over nineteen sections, though six of the
seven state the same cell.
number available in the plan's instrumentation.
layer takes a support role in the next section rather than exiting. This is Bolero's law as a
percentage.
only if the two layers share less than a declared fraction of sounding time and the answer's phrase
length differs from the call's.
From RENDERED AUDIO, as a proxy. Melodic salience already exists on the card; what dialogue needs
additionally is a timbral-identity trajectory of the predominant voice — MFCC centroid of the frames
carrying the melody, segmented, with a change count. A cue where the melody's timbre changes N times
has at least N carrier changes; it cannot say whether the material changed, only that the voice did.
That is a weaker claim than the symbolic measure and must be reported as such, in the same register
the existing lane_analysis.confidence field uses.
Both measures should be added as axes to the grading criteria Josh asked to be expanded, and both
should be calibrated on the exemplar corpus before any threshold is set, because a target derived from
our own generator's output is not a bar.
Every principle in this document, marked and ordered by value.
| # | operation | where | tier |
|---|---|---|---|
| 1 | counter_melody pattern kind with the five-step derivation rule of section 2.4 | pass2_realise.py compile branch plus pass2_plan.py PATTERN_KINDS | CODE-ABLE |
| 2 | melodic_role per layer plus a per-section melodic_schedule | plan schema plus three predicates | CODE-ABLE |
| 3 | heterophony pattern kind deriving from another layer | compile branch | CODE-ABLE |
| 4 | invertible_pair block plus a swap transform | plan schema plus validator plus compiler | CODE-ABLE |
| 5 | answers field with delay_beats and the shared-sounding-time predicate | plan schema plus validator | CODE-ABLE |
| 6 | tension_profile per idiom, defaulting from the mode's own interval content | plan schema plus per-simultaneity check | CODE-ABLE |
| 7 | irama_level per section multiplying the kotekan subdivision | section schema | CODE-ABLE |
| 8 | Counter-melody presence and orchestrational dialogue, symbolic form | a new instrument beside floor_instruments.py | CODE-ABLE |
| 9 | Roughness-over-time curve from the render's STFT | feature rig | MEASURE-ONLY |
| 10 | Per-stem monophonic tracking for the audio form of measure 9.1 | measurement rig, authored lane only | MEASURE-ONLY |
| 11 | Melody-timbre trajectory change count | feature rig | MEASURE-ONLY |
| 12 | Hook breathability — the head cell must contain rests | a predicate on hook.head_cell | CODE-ABLE |
The asymmetry volume 1 identified holds here and is sharper. Eight of these twelve are constraints on
authored symbolic material, enforceable before a single sample is rendered, and the four measurements
are all weaker than their symbolic counterparts. A generated bed can be graded for roughness and for
melodic-timbre travel and for nothing else in this document. The counter-melody, the dialogue
schedule, the invertible pair and the tension profile are the authored lane's to carry — one more
technical reason the current split between the authored and generative lanes is correct rather than
merely convenient.