systems/ENGINE_OPTIMIZATION_DOCTRINE.md
Applied: opt_DOCTRINE_critique.md (fresh-context critic, 2026-07-27; verdict FIXES —
5 HIGH · 6 MEDIUM · 5 LOW, plus the three anchor corrections in its §5). Revised: 2026-07-27.
All sixteen findings are closed in place. The critic's core charge is accepted without argument:
*the doctrine applied its own evidence rules to Epic's numbers rigorously and to everyone else's
loosely.* Every HIGH finding was an instance of that, and the corrections below are all in the
direction of holding non-Epic evidence to the same bar.
CLAIMS WITHDRAWN — stated plainly rather than quietly overwritten:
1. "8 GB is still the largest VRAM slice at 35.74%" — WITHDRAWN. Wrong by 10.1 points. Valve's
own survey page, June 2026, read direct by the critic 2026-07-27: 8 GB = 25.64% (falling
~0.25 pts m/m), 16 GB = 24.50% (rising ~0.45). The 35.74% figure came from secondary
aggregators and is the likely residue of the VRAM-reporting error Valve publicly corrected.
2. "1080p is still the modal primary resolution at ~57.7%" — WITHDRAWN. Overstated by 6.6
points. Live figure: 1920×1080 = 51.12% (1440p 21.44%). The *conclusion* survives — 1080p
leads by 2.4× — but the number did not.
3. **The [VERIFIED, June 2026 Steam Hardware Survey] tag on both figures — WITHDRAWN as a tier
upgrade.** W1's own unknowns register (line 824) had marked them "Partially VERIFIED — from
secondary aggregators, not Valve's page directly." Promoting that to flat VERIFIED broke this
document's own inheritance rule, in the section arguing its most contested decision.
4. The audience-share justification for the 7.0 GB VRAM ceiling — WITHDRAWN. The ceiling is
KEPT and is now justified as a *memory-discipline* constraint (the AC Shadows lesson), because
the audience slice it was argued from is shrinking under it. It now carries a scheduled
falsification test alongside the GPU anchor (§1.4).
5. **§2.2 row 8 as a single 1.15 ms "post-process + temporal upscale" row — WITHDRAWN as
arithmetically unmeetable.** §2.3's own 1.0–1.5 ms upscaler range consumed the entire row. Split
into 8a/8b with a named donor and the arithmetic shown (§2.2).
6. The "~75% of minimum-spec VRAM" texture rule in R-29 — WITHDRAWN. It was a U-27
unverified-blog heuristic written into the requirements section, and it contradicted §7.3's own
pool numbers by 2.4×. Replaced with §7.3's per-class pools.
7. D2's "~30 emitters" as a DON'T with a check — WITHDRAWN as a gate. Same U-27 origin.
Demoted to a spike-investigation trigger.
8. The bare Reflex 2 "75% / 56 ms → 14 ms" citation — WITHDRAWN as stated. Numbers confirmed,
but the capture is 4K/max/RTX 5070 and the feature debuts RTX-50-only. Both qualifiers are now
on the row, and latency is explicitly barred from being a combat-tuning assumption (§4.7 R-49).
9. "The content does not get simpler here" (§2.5) — WITHDRAWN. It gets more expensive on three
axes at once. §2.5 is now marked as the least-evidenced table in the document.
10. Three references to a non-existent "§3.6" — WITHDRAWN. Repointed to §7.3 (§2.4, §5.4 D5,
§7.2). One of them carried the whole 5090-proxy discipline.
Re-checked and CONFIRMED (no change needed): the UE 5.8 release date 2026-06-17 — Epic
released 5.8 at State of Unreal 2026 in Chicago on June 17, 2026 (re-verified live 2026-07-27;
Epic,
GamesBeat). F-14's second half is
closed; the date stands as W4 gave it.
AMENDMENT, 2026-07-27 — §2.7 THE MODEL-INFERENCE BUDGET. Josh asked whether this doctrine
considers the extra GPU usage for AI guided by deterministic code. It did not — rev 2 budgets
rendering only and mis-filed the runtime-inference lane as game-thread work in §2.1 (corrected in
place). §2.7 is the budget: the deterministic-first shield as a scheduling-class taxonomy with
model inference banned from the frame's critical path at every class (AI-1); the
no-hardware-tiered-intelligence rule (AI-2, director ruling); the per-spec-class VRAM policy
with the carve arithmetic shown (AI-3); and the zero-displacement bound that keeps §2.2's
15.00 + 1.67 closure intact without a new ms row (AI-4). It integrates the pre-registered KT-5 bars
rather than inventing parallel ones, names the fixtures that do not exist rather than claiming
coverage, adds §6.5's inference-off validation pass, and opens U-37 … U-42 (extended to U-43
by the 2026-07-28 canon-residue batch, registering the game lane's unruled PC-class latency bar). It was itself put
through a fresh-context adversarial arithmetic critic before landing; three HIGH findings were
applied, and the two claims that failed — a false "independent corroboration" of canon's VRAM
ladder, and a closure argument that had quietly ignored the ≥12 GB recommended configuration — are
withdrawn in place and named in §2.7.3 and §2.7.4 rather than silently overwritten, per this
document's own rev-2 discipline.
Confirmed unchanged by the critic and deliberately left alone: all §2.2/§2.4/§2.6 arithmetic
(recomputed independently, clean); the Lumen 4.0 ms reconciliation; RULE PROF-5; §7.2's
scalability-preview rule; §8's unknowns register; the hype-resistance verdicts (not one W4-IGNORE
item is adopted); and §1.4's "hypothesis with a scheduled falsification test" framing.
---
The build-it-properly-from-the-start ruleset, synthesized from four live-web research lanes
(opt_W1_state_of_art.md · opt_W2_aaa_techniques.md · opt_W3_our_stack.md ·
opt_W4_upcoming.md, all compiled 2026-07-27). Josh RULED the frame-rate canon 2026-07-27;
this document is the engineering consequence of that ruling.
The directive it serves: "we have to build things properly so it is efficient." Every rule
below is chosen because it costs approximately nothing today and is expensive-to-impossible to
retrofit into a 79-node, ~30M-word world. Nothing here is a tuning knob. Tuning knobs are what
you reach for when the doctrine was never written.
---
Every load-bearing claim carries its lane citation and its evidence tier, inherited unchanged
from the lane that sourced it:
belt is the instrument that moves it. Never presented as measured.
Citations read [W3 §3.1 — EPIC], meaning: lane W3, section 3.1, Epic first-party tier. The
lane reports carry the URLs and retrieval dates; this doctrine carries the rulings.
TWO SELF-CHECKS, added in rev 2 because the first draft failed both (critique F-1, F-3):
"partially verified — secondary aggregator," this document may not print [VERIFIED] and may not
attach a first-party attribution the lane withheld. Downgrades are always allowed; upgrades
require a named live re-verification with its own retrieval date.
§5.** Those numbers size *spikes*, never *budgets*. Two of them had leaked into R-29 and D2; both
are corrected in rev 2. This check is repeated at §8.3 so it is enforceable from either end.
A house rule inherited from the lanes and made binding here: all four lanes independently
recorded that searches for UE performance topics in 2026 surface a large volume of SEO/AI-generated
blog content carrying very specific-sounding millisecond figures that appear nowhere in primary
sources [W1 §"How to read"]. **Treat any ms figure arriving from outside this document as
fabricated until an Epic doc, a GDC/SIGGRAPH talk, or our own capture confirms it.**
---
| Class | Target | Frame budget | Status |
|---|---|---|---|
| minimum spec | 30 fps floor | 33.33 ms | ruled canon |
| recommended spec | 60 fps | 16.67 ms | ruled canon — the class the factory authors against |
| high-end | uncapped / 120 target | 8.33 ms | ruled canon |
Milliseconds, not fps, is the unit of all budgeting. Ms is additive; fps is not
(30 = 33.33 · 60 = 16.67 · 90 = 11.11 · 120 = 8.33) [W3 §0 — COMMUNITY, Unreal Art Optimization].
Any conversation that proceeds in fps rather than ms is a conversation that cannot be summed.
A ms budget is meaningless without the machine it is measured on. W1 derived the anchor from the
June 2026 Steam Hardware Survey plus Epic's own shipping configuration, and W3 independently
raised the same gap as FINDING W3-2 ("the whole budget is unanchored until it exists"). The
anchor is adopted as stated:
RECOMMENDED SPEC — the machine the 60 fps number derives against:
**RTX 5060 / RTX 4060 Ti / RX 7600 XT class (mobile: RTX 4070 / 5070 Laptop class) · 1080p output
· High preset · temporal upscaler in Quality mode · NO frame generation · under a hard 7.0 GB VRAM
working-set ceiling · 6-core/12-thread CPU · 16 GB system RAM · NVMe SSD.**
Full ladder:
| Class | GPU anchor (desktop) | Laptop equivalent | Output | GI rung | Lighting | Upscaler | Frame gen |
|---|---|---|---|---|---|---|---|
| minimum (30 fps floor) | RTX 2060 SUPER / RX 6600 | RTX 4060 Laptop class [DERIVED — provisional] | 1080p, Low-Medium | Lumen Lite (Medium GI/Reflections) | MegaLights OFF, deferred fallback | Balanced/Performance | never |
| recommended (60 fps) | RTX 5060 / 4060 Ti / RX 7600 XT | RTX 4070 / 5070 Laptop class [DERIVED — provisional] | 1080p, High | Lumen High | MegaLights ON | TSR/DLSS/FSR/XeSS Quality | OFF — the 60 is real frames |
| high-end (120 target) | RTX 5070 Ti+ / RX 9070 XT-class | RTX 5080 / 5090 Laptop class [DERIVED — provisional] | 1440p–4K, Epic | Lumen Epic + HWRT | MegaLights full | Quality/DLAA | allowed here and only here, on a real 60–90 base |
Why the laptop column exists at all, and why it is marked provisional (critique §5 correction 2):
as of the June 2026 Steam survey a laptop GPU is the #1 GPU on Steam outright — RTX 4060 Laptop
at 3.81%, ahead of the RTX 3060 desktop at 3.73% and the RTX 5060 at 2.66%
[VERIFIED — Valve, Steam Hardware & Software Survey, June 2026, read direct 2026-07-27 (critic
re-verification V-6)]. A published recommended spec that names desktop parts only leaves the single
most common configuration on Steam with no stated target. The mobile parts above are **[DERIVED]
placeholders and must not be printed on a store page until they are measured** — §7.4 correctly
records that laptop thermal and power behaviour is not reproducible by downclocking a desktop part,
so the real question ("which class does the RTX 4060 Laptop actually land in?") is a benchmark-day
item, filed as U-35.
1. Epic publishes the 60 fps arithmetic. The Lumen Performance Guide states the shipping
configurations target "30 and 60 frames per second on consoles with 8ms and 4ms frame budgets
at 1080p" — i.e. Lumen's *share* is 4 ms of a 16.67 ms frame, 24%
[W1 §0.2, §2 — VERIFIED, Epic Lumen Performance Guide]. This is the single most useful published
number in the entire research fan-out and it anchors §2's table.
2. Epic's quality ladder already maps onto our three tiers. Lumen *Epic* targets 30 fps;
Lumen *High* targets 60 fps on current-gen consoles; Lumen *Medium* (Lumen Lite, new in 5.8)
targets Switch 2 and lower-end PCs at 60 fps [W1 §3.1 — VERIFIED]. **We do not invent a custom
ladder.** Epic already tuned these against real console budgets.
3. 60 fps is a 1080p-INTERNAL number, never a native-resolution number. Epic's own shipping
pattern renders Lumen and other features at 1080p internal and upsamples with TSR, and states
this beats native 4K at reduced settings [W1 §3.4 — VERIFIED]. Every ms figure in this doctrine
is per-1080p-internal-pixel. Doubling internal resolution roughly doubles the GPU-bound frame.
4. Frame generation never counts toward the 60. Vendors converge on ~60 fps as the ideal base
before enabling FG (AMD), 40 fps as the floor (Intel); Digital Foundry measured 40–60 base as
responsive and 30–40 as poor; generated frames carry no game state, so input latency stays
pinned to the base rate [W1 §8.4 — VERIFIED]. **A shipped configuration that only reaches 60
*with* FG is a failed configuration.**
5. RT-capable minimum spec is now industry-normal, not a risk. DOOM: The Dark Ages and
Indiana Jones and the Great Circle both refuse to launch without RT hardware; DOOM's minimum is
RTX 2060 SUPER / RX 6600, explicitly excluding GTX, Polaris, Vega and non-RT integrated graphics
[W1 §0.14, §9.1 — VERIFIED]. That precedent is what makes a MegaLights + Lumen-HWRT architecture
safe to commit to at the start, and it is where our 30 fps floor's GPU comes from.
Objection: the recommended tier should be RTX 4070/5070-class, because a photorealistic AAAAA
open world with Lumen, Nanite, MegaLights and Substrate will not hold 60 fps at High on an xx60
card with 8 GB — setting the anchor too low forces either a compromised look or a broken promise.
Answer: correct about the *risk*, wrong about the *response*. A 4070-class anchor lets the
budget drift upward silently for three years, after which the shipped game excludes most of its
audience and the choice is unrecoverable. An xx60 anchor makes the 16.67 ms / 7.0 GB budget an
enforced constraint from the first asset, which is exactly the directive. The honest hedge is
*resolution*, not GPU class: 1080p at recommended, 1440p/4K positioned as high-end. Commercially
this matters too — the three most common GPUs on Steam are all xx60-class (RTX 4060 Laptop
3.81%, RTX 3060 3.73%, RTX 5060 2.66%) and **1080p remains the modal primary resolution at 51.12%
against 1440p's 21.44%, a 2.4× lead** [VERIFIED — Valve, Steam Hardware & Software Survey,
June 2026, read direct 2026-07-27 (critic re-verification V-6)]. For a permanently-discounted Early
Access title (ship-rulings-vo-steam-pricing), pricing the audience out is a commercial failure
mode, not a technical one.
A correction, and the conflict it resolves. Rev 1 of this section printed "8 GB is still the
largest VRAM slice at 35.74%" and "1080p … ~57.7%" tagged `[VERIFIED — June 2026 Steam Hardware
Survey]`. Both are withdrawn: the live figures are 1080p 51.12% and **8 GB 25.64% against
16 GB 24.50%**. Two of our own lanes had disagreed about the VRAM row — W1 said 8 GB at 35.74%,
W4 said 8 GB at 26.76% / 16 GB at 23.51% — and rev 1 resolved that silently, in the direction that
flattered the argument. **Resolved here the same way §2.2 resolves the Lumen disagreement: by
reading the primary source.** Valve's page wins; W1's number was an aggregator artifact, plausibly
downstream of the VRAM-reporting error Valve publicly acknowledged. The lesson is procedural and
now sits in "How to read": *a conflict between two of our own lanes is named and adjudicated in the
open, never averaged away.*
The corrected data weakens the ceiling's *stated justification* and not the ceiling itself, and the
honest response is to change the argument rather than the number:
| VRAM tier | Apr 2026 (W4) | Jun 2026 (verified live) | Direction |
|---|---|---|---|
| 8 GB | 26.76% | 25.64% | falling (−0.25 pts m/m) |
| 16 GB | 23.51% | 24.50% | rising (+0.45 pts m/m) |
The two tiers are ~1.1 points apart and converging; on this trend 16 GB is modal within the year
and comfortably modal by any plausible ship date. Anchoring a multi-year, expensive-to-relax
constraint to an audience slice that is shrinking under it is a bad argument for a good rule.
The ceiling stands on discipline, not on audience share. §3.2 item 3 is the real reason: the
minimum spec is defined first *because that is what forces the memory and streaming architecture*,
and the recommended and high-end tiers then fall out as quality dials rather than as separate
engineering. That is the Assassin's Creed Shadows lesson verbatim [W2 §A.4.1 — TRADE]. The
constraint is worth more than the audience it was originally justified by, and it would be worth
keeping even if 8 GB fell to single digits — a 7.0 GB working set is what makes a 76-region world
stream at all.
And it is now on the falsification schedule, which rev 1 let it escape.
THE ANCHOR AND THE CEILING ARE ONE HYPOTHESIS WITH ONE SCHEDULED FALSIFICATION TEST. The test:
build one complete region; profile it on an RTX 5060-class card at 1080p High with the budget
enforced; check the frame against §2 **and the VRAM working set against 7.0 GB, as two independent
pass/fail axes** (vram_working_set_mb in the §6.2 record is the enforcing metric, from the first
region rather than at the end). If either fails, that axis is revised once, with evidence, at
month six — not assumed at month zero. That benchmark day belongs on the 5090-window plan, and both
axes appear in §8.5's re-check triggers.
generation off."** Not "60 fps." The unqualified number is unfalsifiable.
seconds reads to a player as a broken game [W1 §11]. §4 is therefore not an appendix — it is
half the canon.
---
RULE A — design GPU-bound, on purpose. The Witcher 4 UE5 tech demo (the closest public
analogue to our shape) held ~16.666 ms frame time with GPU render time at essentially the same
~16 ms — entirely GPU-bound, with the game thread carrying gaps deliberately left for gameplay
features to fill later; at 300 individually animated NPCs the CPU sat at ~87% utilization, at 200
NPCs ~60% [W3 §1.1 — COMMUNITY, summary of Unreal Fest Orlando 2025 talk by Kevin Ortegren (Epic)
and Jarosław Rudzki (CDPR)]. This names the *shape* of a healthy 60 fps UE5 open world: the GPU is
saturated and the CPU has headroom.
The counterweight, and it is the sharpest finding in the whole fan-out: **in a dense simulated
open world the frame-rate ceiling is set by simulation cost on the CPU, not by the renderer.**
Digital Foundry's June 2026 position on GTA 6 is that the 60 fps blocker is the CPU — traffic,
crowd routines, daily NPC schedules, vehicle physics, world simulation [W2 §A.2 — DF-SEC].
A photorealistic mandate does not automatically make you GPU-bound. Our stack adds a
runtime-inference lane, a foundry, generative-entity logic, personas on every entity, companion
parties, Tier-C raid roles and a 100-player World Boss — all game-thread work. If we let the
game thread become the bottleneck, every one of those systems becomes a frame-rate negotiation
instead of a design decision [W3 §1.1].
One correction to that sentence, made in the amendment of 2026-07-27 (§2.7). The
runtime-inference lane is NOT game-thread work and listing it there is how it escaped being
budgeted at all. Its 60 Hz half — the deterministic controller executing a strategy token — is
game-thread work and is budgeted in §2.6. Its *model* half is async-worker work with an optional
GPU/VRAM or NPU residency that no table in this document carved. §2.7 is that budget. Everything
else in the sentence is game-thread work and the warning attached to it stands unchanged.
RULE B — budget in percentages, instantiate in milliseconds. Content authored to a percentage
allocation survives a change of target frame rate; content authored to an absolute ms number does
not [W3 §1.1].
W1 and W3 produced two independent tables. They agree on shape and disagree on one row: W1 holds
Lumen at Epic's published 4.0 ms; W3 derived 3.33 ms (20% of frame). **This doctrine takes
the published number.** A derived figure below a published figure is optimism, not discipline —
and Lumen is the one row where Epic has told us the answer. The remaining rows are rebalanced
around it, keeping W3's percentage architecture and its stat gpu measurement handles, which are
what make the table a gate rather than a wish.
| # | Pass group | Budget @ 60 fps | % of frame | Measurement handle | Tier |
|---|---|---|---|---|---|
| 1 | PrePass + HZB (depth prepass, occlusion) | 0.75 ms | 4.5% | PrePass, HZB | [DERIVED] |
| 2 | Nanite VisBuffer (cull + raster) | 2.25 ms | 13.5% | Nanite VisBuffer, r.Nanite.Visualize | [DERIVED] |
| 3 | BasePass / Substrate material resolve | 2.00 ms | 12.0% | BasePass, Substrate Stats Panel | [DERIVED] |
| 4 | Shadows — VSM depths + projection (sun/directional clipmaps only) | 1.75 ms | 10.5% | ShadowDepths, VirtualShadowMapProjection, r.Shadow.Virtual.Stats | [DERIVED] |
| 5 | Lumen GI + reflections | 4.00 ms | 24.0% | Lumen, LumenSceneUpdate, Reflections | [VERIFIED — Epic Lumen Performance Guide, W1 §2] |
| 6 | Direct lighting (MegaLights, incl. its RT shadows) | 1.25 ms | 7.5% | Lights, r.MegaLights.Visualize.LightComplexity | [DERIVED] |
| 7 | Translucency + Niagara + volumetrics + water | 1.10 ms | 6.6% | Translucency, stat Niagara | [DERIVED] |
| 8a | Post-process (tonemap, bloom, DOF, exposure, motion blur) | 0.65 ms | 3.9% | PostProcessing | [DERIVED] |
| 8b | Temporal upscale (TSR / DLSS / FSR / XeSS, Quality) | 1.25 ms | 7.5% | TemporalSuperResolution / vendor pass; verify per preset within 1.0–1.5 ms | [DERIVED — §2.3] |
| — | GPU subtotal | 15.00 ms | 90% | — | — |
| — | RESERVE — hitch and streaming absorption | 1.67 ms | 10% | never allocated to content | non-negotiable |
The arithmetic, shown rather than asserted (critique F-2 required this row to be recomputed):
0.75 + 2.25 + 2.00 + 1.75 + 4.00 + 1.25 + 1.10 + 0.65 + 1.25 = 15.00 ms (GPU subtotal) 15.00 + 1.67 = 16.67 ms (the 60 fps frame) percentages @ /16.67: 4.5 + 13.5 + 12.0 + 10.5 + 24.0 + 7.5 + 6.6 + 3.9 + 7.5 = 90.0% reserve: 1.67 / 16.67 = 10.0% — the non-negotiable floor, unchanged
**This table budgets RENDERING ONLY. Model inference has no row here and must not acquire one:
its carve is VRAM + async compute, and at the recommended-class anchor card it is OFF — see §2.7
(AI-1 through AI-4b) for the arithmetic, the per-spec-class policy, and the pre-named donor if
benchmark day ever forces a row.**
What changed in rev 2, and which rows paid for it. Rev 1 carried a single row 8 at 1.15 ms
covering post-process *and* temporal upscale — while §2.3, four lines later, budgeted **1.0–1.5 ms
for the upscaler alone. At the top of that range post-processing was left with −0.35 ms**; even
at the midpoint the row was already over. The table could not be met as written, and the document
was violating its own governance rule ("name which row it draws from and what it gives back") in
the table that rule is attached to. The split is 8a 0.65 + 8b 1.25 = 1.90 ms, up 0.75 ms, and
the donors are named, with a reason each, rather than taken from the reserve:
every local light onto ray-traced shadows *because Epic defaults it that way, they being cheaper
to generate than VSMs* (§3.1), which removes the per-light VSM tax entirely. Row 4 therefore
carries the sun/directional clipmap set and a justified handful of exceptions — not the scene's
light list. The row header now says so.
bone hierarchies, tightening cluster bounds onto the fixed-function raster path, and B8 bans
Nanite on single-plane cards and trivially low-poly props. Both shrink cluster counts against the
15% rev-1 figure, which was itself derived.
rule, and Epic states one slab is at parity with the legacy path while each additional slab adds
almost linearly. The saving is bought by an enforced content rule, not by hope.
Row 5 is still the only VERIFIED row, and it did not donate: a published number is not a
negotiating position. Everything else is a derived allocation whose job is to be falsified by
measurement. That is the correct epistemic status for a budget table and it must be stated on the
face of the artifact so no one later mistakes it for data. **The three donations above are
themselves the first three predictions the §1.4 benchmark day should try to break.**
THE RULE THIS TABLE PRODUCES — the single most important governance line in this doctrine:
**Every new rendering feature proposed after this point must name which row it draws from and
what it gives back. A feature with no budget row does not get built.** [W1 §2]
Corroboration of SHAPE only, and explicitly not of the per-row percentages: a publicly circulated
stat gpu breakdown of a UE5 scene (12.4 ms total: PrePass 1.2, BasePass 3.1 with Nanite raster
1.8, Lumen 4.2 with screen-probe-gather 2.1 / reflections 1.4 / scene update 0.7, Shadows 1.8,
Translucency 0.8, PostProcess 1.3) puts **Lumen as the single largest row at roughly a quarter to a
third of the frame, with BasePass + Nanite raster next** — which is the shape above. It is a 12.4 ms
captured scene, not a 16.67 ms budget, so row-by-row percentage agreement is not claimed and
would be over-reading a community summary [W3 §1.2 — COMMUNITY, summary only, direct fetch blocked].
Row 8b exists because of this section — rev 1 folded the upscaler into post-processing and the
row became unmeetable (see the split above). Row 8b is easy to under-budget because "enable DLSS,
gain 40%" is folk wisdom. The 2026 reality:
DLSS 4.5's 2nd-generation transformer is more computationally intensive than 4.0's, offset by
FP8 acceleration only on RTX 40/50 Tensor Cores — measured costs of ~16% on an RTX 4060 Laptop,
~15% on an RTX 4060 Ti moving Preset K→L, and 7–30% across RTX 30/20-class cards
[W1 §8.3 — VERIFIED secondary; NVIDIA's own framing confirms the increased demand]. The circulating
"~5× the compute of DLSS 4.0" figure is [UNVERIFIED] and excluded. **Row 8b instantiates this
range at 1.25 ms — its midpoint — and the 1.0–1.5 ms band is the verification tolerance, not a
licence to overrun the row.** A preset measuring above 1.5 ms on the anchor class is a row failure
and goes back to §2.2 for a named donor, exactly like any other feature.
This class does not get "everything, doubled." Its shape changes because features switch off, and
that is the entire point of the three-level scalability architecture in §7.3 [W3 §1.3].
EVERY ROW OF THIS TABLE IS [DERIVED]. Rev 1 shipped it with no tier column at all, so eight
derived numbers read with the same authority as §2.2's one published row. They do not have it.
| Pass group | Budget @ 30 fps (33.33 ms) | Tier | Feature delta at this tier |
|---|---|---|---|
| PrePass + HZB | 1.7 ms | [DERIVED] | — |
| Nanite VisBuffer | 6.0 ms | [DERIVED] | r.Nanite.MaxPixelsPerEdge raised |
| BasePass | 5.5 ms | [DERIVED] | reduced material complexity via quality switch |
| Shadows (VSM) | 4.0 ms | [DERIVED] | resolution LOD bias, reduced ray/sample counts |
| GI + reflections | 5.0 ms | [DERIVED — see the reconciliation below] | Lumen Lite — "twice as fast as Lumen high quality" [W1 §3.2, W4 §1 — VERIFIED, Epic 5.8 release notes + Daniel Wright, Epic] |
| Direct lighting | 3.5 ms | [DERIVED] | MegaLights OFF — architectural, not a slider (§3.3) |
| Translucency + Niagara | 3.0 ms | [DERIVED] | Niagara significance culling aggressive |
| Post-process | 0.5 ms | [DERIVED] | reduced post chain |
| Temporal upscale | 0.8 ms | [DERIVED] | TSR Low / FSR Performance |
| GPU subtotal | 30.0 ms | 1.7+6.0+5.5+4.0+5.0+3.5+3.0+0.5+0.8 = 30.0 | |
| Reserve | 3.33 ms | 10%, as at every class |
The GI row reconciled against Epic's OTHER published number. §1.3 fact 1 quotes Epic's shipping
configurations as "30 and 60 frames per second on consoles with 8 ms and 4 ms frame budgets at
1080p." §2.2 takes the 4 ms half as VERIFIED and refuses to derive below it. The 30 fps half of the
same sentence is 8 ms, and this table allocates 5.0 ms — so the asymmetry has to be argued,
not assumed, or §2.2's own standard ("a derived figure below a published figure is optimism, not
discipline") is being applied selectively. The argument, stated so it can be attacked:
Lite**, which Epic states is ~2× faster than Lumen High — itself the rung below Epic.
RX 6600 desktop floor.
why it is a benchmark-day item rather than a settled number. If Lumen Lite measures above 5.0 ms
on the min-spec card, this row is where the 30 fps floor breaks first.
**A caution that must travel with these tables: ms budgets are per-hardware-class and are NOT
portable between classes.** Lumen Lite is ~2× faster than Lumen High, so on identical hardware it
would be ~2.0 ms where High is 4.0 — but the minimum-spec class runs on an RTX 2060 SUPER, not on
the anchor card, so 5.0 ms is the correct budget there. A ms number without its hardware class
attached is not a budget, it is a rumor.
THE LUMEN LITE RETREAT (rev 2, critique F-8). Lumen Lite is Beta in 5.8, and 5.8 is the
terminal UE5 feature release — so a Beta→Production promotion may never arrive on our engine line.
The entire minimum-spec GI rung rests on it, and rev 1 stated no retreat. The retreat is: **if
Lumen Lite regresses or is withdrawn, the minimum class falls back to Lumen High at reduced quality
settings and a lower internal resolution (Balanced/Performance upscaling), and the GI row is
re-derived against the 4.0 ms published High figure rather than against Lite.** That path costs
image quality at the floor; it does not cost an architecture, because both rungs are the same
dynamic solve and both are reached through sg.GlobalIlluminationQuality. **This is why §3.1 bans
baked lighting rather than holding it as the min-spec fallback** — a baked fallback would be an
architecture change, and it is authorized only under §3.3's single ruled exception. "Lumen Lite
stabilizes or regresses in a 5.8 hotfix" is now a §8.5 re-check trigger.
Same percentages, halved absolutes (3 decimals, so the arithmetic closes exactly rather than
appearing to miss by 0.02):
0.375 + 1.125 + 1.000 + 0.875 + 2.000 + 0.625 + 0.550 + 0.325 + 0.625 = 7.500 ms (GPU subtotal) 7.500 + 0.833 = 8.333 ms (the 120 frame)
**Read the §2.4 caution again before using any number above: ms budgets are per-hardware-class and
are NOT portable between classes.** §2.4 did that hardware-class reasoning explicitly for the 30 fps
table. This table does not do it at all — it halves the recommended-class absolutes and asserts that
a 5070 Ti-class part absorbs the difference. That may well be true; it is asserted, not argued,
and this heading says so.
Correction to rev 1, which said "the content does not get simpler here." That sentence was
wrong in a way that hides the risk. The asset set does not get simpler — but the *frame* gets more
expensive on three axes simultaneously, per §1.2's own ladder:
1. A higher GI rung — Lumen Epic, not Lumen High.
2. Hardware ray tracing, not the SWRT default (§3.1 RULE GI-1 exists precisely because SWRT is
the performant path in scenes of many overlapping instances — which is our construction).
3. A higher internal resolution — 1440p–4K, not 1080p; and §1.3 fact 3 states doubling internal
resolution roughly doubles the GPU-bound frame.
The sharpest consequence, stated so it cannot be missed: **this table budgets Lumen at 2.000 ms for
the Epic rung, which is the very configuration Epic's published guide describes at 8 ms / 30 fps** —
a 4× gap on the one row where Epic has told us the answer. It is not impossible (a 5070 Ti-class GPU
is far above the console parts Epic tuned against, and 5090-class further still), but nothing in
this document evidences the factor. **The high-end class therefore gets its own benchmark day, filed
as U-34** — it is not a free consequence of the recommended-class measurement.
Uncapped is a mode; 120 is the bar we measure. Multi-frame generation is available in this class and
only this class, layered on a real 60–90 fps base [W1 §9.3, W3 §1.3].
CPU threads run concurrently with the GPU and each other; each has the *same* wall-clock budget and
the frame is bound by the worst of them. Per RULE A we deliberately under-fill [W3 §1.4 — DERIVED]:
| Thread | Budget @ 60 fps | Design target | Sub-allocation |
|---|---|---|---|
| Game thread | 16.67 ms | ≤ 13.5 ms (80%) | gameplay/AI/foundry 6.0 · animation 3.0 · PCG runtime 1.5 · UI 1.0 · tick/misc 2.0 |
| Render thread (Draw) | 16.67 ms | ≤ 11.0 ms | draw submission; stat rhi draw calls |
| RHI thread | 16.67 ms | ≤ 9.0 ms | — |
The animation row is not a hope — UE ships a system that enforces it. The **Animation Budget
Allocator** takes a fixed per-platform game-thread ms budget and dynamically throttles skeletal-mesh
ticking (Update Rate Optimisation / frame skipping) by significance, so hero characters animate
at full rate and distant ones degrade, with the total held to the declared budget [W3 §1.4 — EPIC,
Animation Budget Allocator doc; corroborated by Coconut Lizard].
This is the single most important CPU decision for our shape, because canon carries Tier-C raid
roles and a 100-player World Boss (docs/COMBAT_PROGRAM_ADDENDUM.md §12) plus companion/familiar
parties with their own stat builds and AI-intelligence tiers. Wiring ABA + Significance Manager
now, with the significance function keyed off (party membership · boss role · screen size ·
distance), is the difference between "the raid scales" and "the raid ships at 30 fps." The Witcher
4 demo's 300-NPC market scene is the public proof this holds at crowd scale [W2 §B.10 — TRADE].
---
Josh, 2026-07-27, verbatim: *"Does that doctrine consider the extra gpu usage for any AI guided
by deterministic code?"*
It did not. Rev 2 budgets rendering and nothing else. The one place the runtime-inference lane
appears at all is §2.1's RULE A counterweight, which lists it among "a foundry, generative-entity
logic, personas on every entity, companion parties, Tier-C raid roles and a 100-player World Boss —
all game-thread work." That classification is WITHDRAWN as applied to model inference
(§2.1 now carries the correction and points here). The rest of that list really is game-thread work
and the warning attached to it stands. Model inference is not: it is async-worker work with an
optional GPU/VRAM or NPU residency that no table in this document had carved. A stack that ships
a local model co-resident with the renderer and budgets only the renderer has an unbudgeted
consumer sitting on the same bus — which is precisely the class of omission §2.2's governance line
exists to forbid ("a feature with no budget row does not get built").
This section is that budget. It is written to the same standard as the rest of §2: rules numbered
so a gate can cite them, arithmetic shown rather than asserted, and every number carrying the tier
its source assigned it.
A tier note, because this section imports a NEW source family. The engineering claims below
inherit the doctrine's tiers ([VERIFIED] / [DERIVED] / [COMMUNITY] / [UNVERIFIED]). The *canon*
sources carry their own declared provenance and it is inherited unchanged, exactly as the lane
tiers are:
canon_ruled). load-bearing sources say so on their face: docs/RUNTIME_GENERATIVE_LAYER.md §B.3's memory table
is headed "Proposed inference-layer memory budget," and
docs/RUNTIME_GUARDRAIL_ARCHITECTURE.md §8.2 is headed "the efficiency budget (**illustrative
shape; locked at a runtime perf audit**)."
Tools/bench/bench_criteria.json(criteria_version 1.1.0, pre-registered 2026-07-27T19:54Z, sha256-locked, moved only through
run_all_benchmarks.py --relock).
No inference VRAM or latency figure in this section may be printed as [VERIFIED]. None of them
is measured yet; the runtime perf audit is what measures them, and it is KT-5.
The architecture that answers Josh's question already exists in canon and predates this doctrine.
It is worth restating in budget terms because that is the form in which it becomes checkable:
Per-frame AI is deterministic code. The LLM emits only a schema-constrained strategy token
({goal ∈ enum, plan ∈ legal-set, knobs ∈ ranges, line ≤ N tok}) — never an action. A
deterministic controller (utility AI + StateTree) executes that token at 60 Hz. The token is a
read-only leaf with no write path into the sim, enforced at the API boundary, not by policy
[CANON-RULED — docs/translation/T99_Translation_Combat.md §6.1;
docs/RUNTIME_GUARDRAIL_ARCHITECTURE.md §1.1 (the strategic/tactical wall), §7 F10].
The budget consequence: the 60 Hz half of that architecture is ordinary game-thread CPU work and
is already budgeted — it lives inside §2.6's gameplay/AI/foundry 6.0 ms sub-allocation of the
13.5 ms game-thread target. It needs no new row because it was never the new thing. The *model* half
is the thing this section budgets, and the rule that governs it is:
AI-1. MODEL INFERENCE IS BANNED FROM THE FRAME'S CRITICAL PATH AT EVERY SPEC CLASS. No frame,
at any class, on any hardware, may block on a model. The render/sim thread posts and polls a
double-buffered mailbox and reads only the last completed result; if the answer is not ready it
uses the previous token or the authored fallback [CANON-RULED — RUNTIME_GENERATIVE_LAYER.md
§B.4]. This is not a performance preference that a fast machine may relax — it is the property
that makes a late or dropped generation structurally incapable of stalling a mechanic.
THE SCHEDULING-CLASS TAXONOMY. Four classes, and the class is a property of the *work*, not of
the hardware it lands on:
| Class | What runs there | Where it runs | What preempts it | On miss |
|---|---|---|---|---|
| Frame-critical (every frame, inside 16.67 ms) | NOTHING. No model call of any kind, at any spec class. Deterministic controller only — StateTree/utility AI executing the last token | Game thread, inside §2.6's 6.0 ms gameplay/AI/foundry row | n/a — it *is* the frame | n/a — it cannot miss; it is deterministic code |
| Sub-second amortized — the strategy-token refresh | Strategy-token re-issue on discrete events (phase change, aggro shift), 2–10 Hz strategic tick; combat barks (soft P99 120 ms, hard cap 250 ms) | A dedicated async worker; on GPU-shared platforms submitted to a low-priority async-compute queue | The graphics queue preempts it. A tick spilling across several frames is a normal outcome the mailbox tolerates, not a defect | Authored fallback / previous token, dropped silently [RUNTIME_GEN §B.4] |
| Seconds-scale | Player-facing dialogue probe (≤ 700 ms, "thinking" beat permitted, then fallback); persona/dialogue generation; faction, disposition and environment plans (seconds, no player waits) | Same worker, lower priority band (queue order: combat bark > player probe > ambient > background disposition) | Graphics queue, and every higher band in the priority micro-batch queue | Authored fallback; for strategic output the fallback is the deterministic default behaviour the sim would run absent the model [RUNTIME_GUARD §6] |
| Idle / loading / background | Ambient response-cache pre-warm per (entity, state_bucket); per-entity persona prefix-cache computation at entity load; foundry work | Idle frames and load screens on the player's machine; all asset/audio/VO foundry work runs off the player's machine entirely | Evicted on any load or streaming spike | Cache miss → generate in the sub-second class, or fall back |
The latency figures in that table are canon's, not this document's
[CANON-PROPOSED — RUNTIME_GENERATIVE_LAYER.md §B.4 latency table, §3; RUNTIME_GUARD §8.2]. They
are reproduced rather than re-derived so there is exactly one set of numbers in the house.
Check (AI-1): the §6.2 record gains inference_scheduling_class and inference_queue_priority;
a build in which any model call is reachable from the game thread's synchronous path fails the
frame-critical scan (the fixture for this does not exist yet — §2.7.5 lists it).
**AI-2. Gameplay-affecting AI DECISIONS are IDENTICAL at every spec class — and at any future
console tier — and are never scaled by available inference hardware.** A stronger GPU buys a
richer *texture* over the same decisions. It never buys a better decision.
Four grounds, each cited, because this rule will be argued with (the argument is always "but the
5090 could run the 8B, why not let it think harder?"):
1. INT is a BUILD STAT, and build viability is a certified property. The companion/familiar
AI-INTELLIGENCE axis makes an ally's brain something the *player* invests in — "if they're more
intelligent their AI will start to detect more patterns and be more proactive and aware," with a
mandatory tradeoff against combat stats [CANON-RULED — docs/COMBAT_PROGRAM_ADDENDUM.md §12,
Josh 2026-07-26, verbatim-seeded]. That axis is explicitly enrolled in the build-space sweeps
of the AAAAA balance doctrine, where a never-picked pool option is a balance *defect*
[COMBAT_PROGRAM_ADDENDUM.md §10, §12]. Hardware-variable intelligence makes the *same build*
play differently on two machines, which does not merely disadvantage a player — it makes the
sweep unmeasurable, because "is this spread viable?" stops having an answer.
2. Multiplayer fairness. The same architecture scales into Tier-C group/raid roles and the
canonical 100-player World Boss [CANON-RULED — COMBAT_PROGRAM_ADDENDUM.md §12;
T1_Tier_C_MMO_Spec §8; CVD §13.1]. Hardware-tiered ally intelligence in that arena is
pay-to-win via GPU. Decisions are identical across clients or server-authoritative; there is
no third option.
3. QA reproducibility. Soaks, kill tests and the build-space sweeps certify *reproducible*
behaviour. Canon already commits the sim to determinism and generated output to a seeded,
cosmetic leaf, with a "deterministic text" accessibility mode that is bit-identical to the no-LLM
configuration [CANON-RULED — RUNTIME_GUARD §8.3; RUNTIME_GEN §B.4, §3]. A decision layer that
varies with the tester's GPU makes every soak result conditional on the machine that produced it.
4. The inference-off pass already settles it. §6.5 item 4 (added by this amendment) requires the
game to run and play correctly with all model inference disabled. If decisions were
hardware-tiered, that pass would be measuring a *different game* rather than the same game
without ornament — and the pass is the thing that proves the deterministic layer IS the game.
WHAT MAY TIER BY HARDWARE — the DLSS-of-AI framing. Same decisions, richer texture. Exactly the
sub-second-amortized and seconds-scale enrichment work of §2.7.1's taxonomy, and nothing else:
| May tier | May NOT tier |
|---|---|
| Ambient variety — generated bark / persona-flavor pool size and refresh rate | Any tactic, target, plan, disposition, ability choice, aggro decision, or timing |
| Local TTS quality for non-VO lines (see the note below) | Companion/familiar behaviour at a given INT tier |
| Seconds-scale generation latency (a faster machine hears a probe answer sooner) | Boss phase logic, tells, punish windows, win conditions |
| Cosmetic reactivity — how often a line is freshly generated vs drawn from the authored bank | Anything a build-space sweep, a soak, or a raid encounter measures |
AI-2 bans HARDWARE-tiered intelligence. It does not ban tiered intelligence. The distinction
matters because canon already ships a system that scales antagonist intelligence deliberately:
higher difficulty means "smarter antagonists (deeper strategic reasoning, longer faction memory,
better use of the player's tells) — *smarter, not bigger numbers*," scaling with the vril-density
dial and never with stat buffs [CANON-RULED — RUNTIME_GEN "DIRECTIVES FOLDED IN," Josh 2026-07-04;
docs/proposals/DIFFICULTY_SYSTEM.md]. That is a player-chosen, machine-independent axis: two
players on the same difficulty face identically-smart antagonists on a 2060 and a 5090. AI-2 is
satisfied, and the difficulty axis is enforced by the same decision-parity check — the trace must
match across inference states *at a fixed difficulty*, not across difficulties.
The player-facing statement of AI-2, which is also canon's own: a weaker machine gets the
*identical fight* and the *identical companion tactics* with a smaller ornamental vocabulary.
Canon already states this as a hardware policy — "every tier plays the identical, canon-correct
game… a Series S player and an RTX-4090 player fight the same boss with the same tells and the same
win condition; one just hears more varied taunts" [CANON-RULED — RUNTIME_GEN §B.9]. AI-2 promotes
that from a property of the current design to a rule the design may not later trade away.
Note on local TTS, marked honestly because it is the one row above that canon does not yet carry.
The ruled VO policy is PARTIAL VO — generated hero/canonical lines only, bulk and ambient stay
subtitle, high-care cultural roles subtitle-or-synthetic
[CANON-RULED — docs/ROADMAP_POST_5090_TO_SHIP.md D-VO-POLICY, Josh 2026-07-24]. All of that is
bake-time: the audio and VO foundry runs on the factory's machines, not the player's, and the
audio stack's local models are production tools [docs/pipeline_review/tech_research/AUDIO_STACK.md].
UE's runtime audio synthesis is MetaSounds DSP, not model inference. **So there is no runtime TTS
lane in canon today, and this doctrine does not create one.** AI-2 permits a local TTS lane as a
tierable enrichment; §2.7.3 constrains it: **if it is ever built it enters the SAME carve as the
language model — never a second one — and it is barred at minimum and 8 GB-recommended by the same
arithmetic that bars the language model there. Filed as U-39**.
Check (AI-2): decision-parity is proven by a paired capture — the same route_id and the same
seed, run inference-on and inference-off, asserting an identical decision trace (token
sequence, ability selections, phase transitions, ally tactics). Divergence in the decision trace is
a failure; divergence in the *line text* is the feature working. No fixture enforces this today
— §2.7.5 lists it as the highest-priority one to add.
This is where Josh's question gets its number. The doctrine's own two constraints — the **8 GB
anchor card and the 7.0 GB working-set ceiling** — turn out to decide the policy almost
entirely, and they decide it more sharply than the design expected.
**FIRST, WHICH MODEL — because canon carries three answers and they are dated, not contradictory
by accident.** This must be stated before any arithmetic, or the arithmetic silently picks one:
| Source | Date | Model target | Scope |
|---|---|---|---|
RUNTIME_GEN §B.1/§B.10 | design workflow | 3–4B Q4_K_M as "the single canonical target"; 8B a PC-only prestige option | Proposed. §6's own open items still carry "[DECISION NEEDED - JOSH] Target local model + quant (3B vs 8B)" — so §B was never closed |
RUNTIME_GEN "DIRECTIVES FOLDED IN" | Josh, 2026-07-04 — RULED | 28–32B at a ~2030 ship window on next-gen PC/console NPUs; dev/test on the local 14B now | The long-window ruling. Also rules the escape hatch used below: *"Tier DOWN to min-spec via smaller distilled models + a larger authored-fallback surface; the deterministic sim is always fully playable with zero model."* |
T99_Translation_Combat §6.4 + bench_criteria.json v1.1.0 | 2026-07-19 / 2026-07-27 | 3–4B = the SHIP target sized for console shared memory; 14B = a dev-side reference tier only | The near-window lane, and the newest. KT-5's ship / reference tiers are exactly these two |
**This amendment budgets against the near-window lane (3–4B ship / 14B reference), because that is
the newest statement and the one the pre-registered KT-5 bars encode** — and it flags the rest
honestly rather than resolving it here. Two consequences worth stating plainly:
is further from the 14B and nowhere near 28–32B. A larger target only widens the gap, and the
strategy-token architecture is explicitly model-size-agnostic (Josh 2026-07-04: "scaling the
model up is a quality dial, not an architecture change"), so nothing structural rides on the pick.
GPU-carve question this section answers is properly the interim, no-NPU case — which is
exactly the case an Early-Access ship on 2026-class hardware is in. The doctrine's spec ladder is
2026-anchored; canon's full generative target is ~2030-anchored. Filed as U-42, together with
the un-applied ACTION in the same 2026-07-04 block that tells §B.1/§B.3/§B.9/§B.10 to be
re-derived — those tables are the ones this section cites, and they are stale on their face.
THE CARVE ARITHMETIC AT THE RECOMMENDED ANCHOR (the 8 GB floor of the anchor class):
8.00 GB anchor-class FLOOR VRAM (the 8 GB xx60 card) [§1.4 "an xx60 card with 8 GB";
§7.1 "the 8 GB minimum card"]
-7.00 GB game working-set ceiling, enforced as a gate [§1.2, §1.4b, §7.2]
=1.00 GB nominal remainder
- ? OS / desktop-compositor / driver VRAM reserve [U-38 — UNMEASURED]
=<1.00 GB actually available to an inference lane
Note the anchor class is not uniform: §1.2 names RTX 5060 / RTX 4060 Ti / RX 7600 XT, which
spans 8 GB and 16 GB parts. The arithmetic above is the floor of that class, and it is the case
that binds, because §3.2 item 3 defines the memory architecture from the bottom.
Against what the near-window ship model needs resident:
~2.6-3.2 GB 3-4B Q4_K_M weights + KV cache [CANON-PROPOSED - RUNTIME_GEN B.1] +~0.4 GB draft model + grammar FSM + scratch [CANON-PROPOSED - RUNTIME_GEN B.3] =~3.0-3.6 GB inference total [CANON-PROPOSED - RUNTIME_GEN B.3] 3.00 GB required (low end) vs <1.00 GB available -> short by AT LEAST 2.00 GB
The 8 GB card cannot host the near-window ship model at all. This is not a tuning outcome and no
quantization step closes a 2 GB gap on a 3 GB model. It is arithmetic, and it is the direct answer
to the question: the extra GPU usage was never considered, and once it is, the recommended-class
anchor turns out to have no room for it.
*(Canon carries an internal tension in that stack worth naming, since it does not change the result:
§B.3 budgets "draft + FSM + scratch" at ~0.4 GB while §B.1 sizes the 0.5–1B draft tier alone at
~0.6–0.9 GB resident. Taking §B.1's higher figure moves the threshold below from 11.10 to ~11.60 GB
— still 12 GB.)*
Where the threshold actually falls, derived the same way:
7.00 GB game working set (the ceiling) +3.60 GB inference total, worst case [CANON-PROPOSED - RUNTIME_GEN B.3] =10.60 GB +~0.50 GB a conservative OS/compositor reserve [U-38 - UNMEASURED, illustrative only] =~11.10 GB -> the smallest COMMON VRAM tier that clears it is 12 GB
**Two honest qualifications on that "12 GB," because rev 2's whole lesson was about numbers that
acquire authority by being written confidently. First, an 11 GB** tier exists (RTX 2080 Ti /
GTX 1080 Ti) and clears the pre-reserve 10.60 — so the step to 12 GB leans on the 0.50 GB term,
which is itself admittedly illustrative and is U-38. Second, **7.00 GB is the recommended class's
working set**; the threshold is class-conditional, and at the high-end's larger working set a 12 GB
card would leave far less room (see the high-end block below).
AND A DIRECT CONTRADICTION WITH CANON, NAMED RATHER THAN AVERAGED AWAY (§1.4's own procedural
lesson, applied to a canon source instead of a lane): RUNTIME_GEN §B.9's degradation ladder places
the 3–4B baseline at "mid PC (8–12 GB)" and still runs "1–3B … generative on bosses + key NPCs
only" at "entry PC (≤8 GB)". That ladder does not corroborate the conclusion above — it
contradicts it. An earlier draft of this section claimed §B.9's "PC ≥12 GB" row as independent
corroboration; that claim is WITHDRAWN and was wrong twice over: §B.9's ≥12 GB row is the *8B
prestige* row, a different model from the one the 11.10 GB was derived for, and the 3.60 GB term is
cited from §B.3 — the same document — so the derivation was never independent of it.
The resolution, and why this doctrine's number wins here: §B.9's ladder is keyed to **raw card
VRAM** and never nets out the game's working set. A "mid PC 8–12 GB" running a 3–4B at ~3.0–3.6 GB
implicitly assumes the renderer fits in ~5 GB. This doctrine rules the working set at 7.0 GB
(§1.2, §1.4b) and enforces it as a gate (§7.2), so the same ladder re-derived against our actual
ceiling shifts up by roughly one tier. §2.7.3 therefore tightens canon's ladder rather than
reproducing it, and RUNTIME_GEN §B.9's tier table is exactly the kind of absolute-number table its
own §6 flags as "illustrative … deferred to the Phase-5M balance audit." **The re-derivation is a
proposal against a canon doc that invites it, not a silent override** — logged in U-42 for the canon
lane to land or reject.
THE POLICY:
| Spec class | GPU inference | What runs | Player-facing difference |
|---|---|---|---|
| minimum (30 fps) — RTX 2060 SUPER / RX 6600, 8 GB | ZERO. None, of any kind. | Authored-fallback-only — the model never loads. This is canon's declared floor tier and RUNTIME_GUARD §6 pre-authorises it as a valid ship configuration with no code-path change. Read this cell as a RULING, not as an arithmetic consequence: the arithmetic above excludes only the 3.0–3.6 GB baseline tier, and a 0.5–1B floor model (~0.6–0.9 GB) against a <1.00 GB remainder is genuinely open until U-38 is measured. The class is ruled to zero anyway, on the §3.2-item-3 ground that the minimum spec is what forces the memory architecture — a sub-1 GB lane on the class with the least margin buys prose variety at the cost of the one budget the whole streaming design rests on. Josh's own 2026-07-04 escape hatch is exactly this: *"tier DOWN to min-spec via smaller distilled models + a larger authored-fallback surface"* — and at this class the surface is the whole of it | None that touches the game. Identical bosses, tells, win conditions, companion tactics, facts and progression (AI-2). Barks come from the authored 3–6-line-per-bucket bank instead of being generated: less prose variety, nothing else [CANON-RULED for the property — RUNTIME_GUARD §6; RUNTIME_GEN §B.9] |
| recommended (60 fps) — the 8 GB anchor | ZERO on the 8 GB anchor — arithmetically forced above | Authored-fallback-only, exactly as at minimum. A CPU-side small-model lane is NOT authorised here — it is filed as U-41 and is INCONCLUSIVE until measured against §2.6's game-thread budget on the ruled 6c/12t part | As above: identical game, smaller ornamental vocabulary |
| recommended (60 fps) — ≥12 GB variants of the anchor class (16 GB RTX 4060 Ti, RX 7600 XT) | A BOUNDED CARVE, probe-gated | The canon baseline tier: one shared 3–4B Q4_K_M instance + 0.5–1B draft, speculation on, batching, ~3.0–3.6 GB [CANON-PROPOSED — RUNTIME_GEN §B.9, §B.10]. This is the one configuration where the lane runs INSIDE §2.2's 16.67 ms frame, so AI-4b's zero-displacement bound binds here first and hardest — see §2.7.4 | Generated barks and persona flavour appear. No decision changes (AI-2) |
| high-end (120 target) — RTX 5070 Ti+ / RX 9070 XT, 16 GB | The full local tier at the 3–4B baseline; the 8B prestige rung is probe-gated and NOT yet cleared | See the high-end arithmetic below | Widest ambient variety, lowest probe latency. Same fight |
| any class with an NPU (Copilot+, Ryzen AI, future consoles) | Off the GPU entirely | Canon: an NPU "runs the LLM without touching the GPU's bandwidth; treat as a bonus, never a requirement" [CANON-PROPOSED — RUNTIME_GEN §B.3]. This is the one path that puts generated barks on an 8 GB recommended machine, and it costs zero VRAM and zero graphics ms | Baseline tier behaviour on a machine the VRAM arithmetic otherwise excludes. Breadth of NPU presence in our audience is [UNVERIFIED] — the Steam survey does not report it |
The high-end arithmetic, and the one place it does not close:
16.00 GB high-end anchor (RTX 5070 Ti / RX 9070 XT)
- 9.50 GB working set at high-end [DERIVED: 7.00 + the 7.3 texture-pool delta 5.0 - 2.5]
= 6.50 GB remainder before any OS reserve
3-4B baseline ~3.0-3.6 GB -> CLEARS with ~2.9-3.5 GB of margin
8B prestige ~5.9-6.9 GB -> DOES NOT CLEAR at the top of its range (6.90 > 6.50), and at
the bottom leaves 6.50 - 5.90 = 0.60 GB - which is INSIDE the
~0.50 GB illustrative OS reserve, so it is not a clearance
So: **the high-end class gets the full local tier at the baseline model, and the 8B prestige rung is
[DECISION AT BENCHMARK DAY]. The blocking term is that this doctrine states no working-set
ceiling for the high-end class at all** — the 9.50 GB above is [DERIVED] from §7.3's pool delta,
not ruled. Filed as U-40. Refusing to invent that ceiling here is the same discipline §2.4
applies to the Lumen Lite row.
**AI-3. The inference lane is admitted by a LAUNCH-TIME VRAM/NPU PROBE, never by a spec-class
label.** The probe measures free VRAM *after* the game's working set is established and enables
the highest canon tier that fits with margin; where no tier fits, the lane does not load and the
authored floor is the configuration. The probe exists in canon already [RUNTIME_GEN §B.9]; AI-3
supplies its numbers and makes the ceiling — never the feature — the thing that wins a tie.
Why the ceiling wins. §1.4b establishes the 7.0 GB ceiling as a *memory-discipline* constraint
that the entire 76-region streaming architecture rests on — "worth keeping even if 8 GB fell to
single digits." The inference lane is ruled "upside, never a dependency" on its own face
[CANON-RULED — RUNTIME_GUARD §6]. Lowering an architecture constraint to fit an upside feature
inverts the priority order both documents already declare, so the option is named and rejected here
rather than left available to be discovered as a compromise later.
AI-3b. The validator's classifier is inside the carve, not beside it. RUNTIME_GUARD §5's
V1–V5 validator is described as stateless and cheap — regexes, set membership, a length check, an
edit distance — plus "one small intent classifier" (V2). That classifier is a model. It
inherits the scheduling class of the generation it validates (it runs in the same post-generation
stage on the same worker) and its residency is inside the §2.7.3 carve. It is never an
always-resident second consumer, and a build in which it loads while the language model is
tier-disabled is a defect — because that is the configuration the min-spec and 8 GB-recommended
classes ship in.
The closure claim, stated so it can be attacked:
0.75 + 2.25 + 2.00 + 1.75 + 4.00 + 1.25 + 1.10 + 0.65 + 1.25 = 15.00 ms (2.2 GPU subtotal) 15.00 + 1.67 = 16.67 ms (the 60 fps frame)
**Unchanged by this amendment — but the reason is narrower than it first looks, and the narrow
version is the honest one.** §2.2 is instantiated *at the recommended class, on the anchor card*,
and §2.7.3 shows the 8 GB anchor runs zero GPU inference. So §2.2 as measured — on the card the
budget is defined against, in the configuration §7.2's proxy rule pins — measures a frame with the
inference lane off by construction. §2.4's minimum-spec table (30.0 + 3.33 = 33.33 ms) closes the
same way, and there the zero is a ruling as well as arithmetic.
**What that does NOT cover, stated because an earlier draft of this section quietly assumed it
did.** The recommended *class* is not only the 8 GB card. §2.7.3's policy table enables a bounded
carve on ≥12 GB variants of the same class — 16 GB RTX 4060 Ti, RX 7600 XT — and those run the
lane at 60 fps, inside §2.2's 16.67 ms frame. So the recommended class has two instantiations
with different inference states, and only one of them is the one §2.2 was measured on. The rule
below is what governs the other, and it binds at recommended-≥12 GB first, not only at the 120
target:
**AI-4. The inference carve is VRAM + async compute. It is NOT a row on the frame's critical
path, and it may NOT be paid for out of the §2.2 reserve.** The 1.67 ms reserve is declared for
hitch and streaming absorption and is non-negotiable; an inference lane is neither hitch nor
streaming, so spending reserve on it would be exactly the quiet raid rev 2 was written to stop.
Where the ms question bites: every configuration with the lane ON. Two of them exist —
recommended-≥12 GB (16.67 ms) and the 120 target (8.33 ms) — and the second is the sharper
case only because an 8.33 ms frame makes the same absolute displacement proportionally twice as
large. The first is the more *common* case and is the one a store-page "recommended spec" claim
rests on. The rule covers both:
AI-4b — THE ZERO-DISPLACEMENT BOUND. On any class where the inference lane is enabled, the
inference-on GPU subtotal must equal the inference-off GPU subtotal for the same
route_id within measurement noise. Any displacement above the noise floor is a **duty-cycle
defect, throttled at the scheduler** — never a donation from a graphics row. The lever is the
strategic tick's rate and batch size, which the mailbox architecture is explicitly built to
tolerate ("a strategic tick spilling across several frames is fine"). Check: a paired A/B
capture on the same route_id, both records complete under RULE PROF-5; the delta is the metric.
**The bound's value is [DECISION AT BENCHMARK DAY], and neither configuration needs a benchmark day
of its own. The 120-target A/B rides on U-34** ("whether the high-end class actually holds
8.33 ms"); the recommended-≥12 GB A/B rides on U-07 / §1.4's falsification test, which already
builds one complete region and profiles it against §2.2 — it simply runs the route twice.
Deliberately not invented here: any specific ms figure for async-compute displacement would be
exactly the kind of number the "How to read" house rule tells us to treat as fabricated until our
own capture confirms it.
THE CONTINGENCY, pre-named so benchmark day is turnkey. If U-37 measures that displacement
cannot be throttled to the noise floor, or if the director later rules generated barks a
*requirement* at recommended spec rather than upside (which would force the lane onto a class §2.2
budgets), then a ms row IS forced and the donor must be named in advance rather than taken from the
reserve under pressure:
already rules that "the honest hedge is *resolution*, not GPU class," and every content row in
§2.2 is either published (row 5), already a donor (rows 2, 3, 4), or canon-bound (row 7 via D1 and
the presentation ladder). Dropping the upscaler from Quality to Balanced at that class reduces
internal pixel count on every pixel-bound row at once.
§1.3 fact 3 — doubling internal resolution roughly doubles the GPU-bound frame — which gives the
direction and not the coefficient. **This contingency is a prediction to be broken, in the same
register as §2.2's three rev-2 donations, and it is filed as U-37.**
Tools/bench/kt5_runtime_inference.py is the runtime perf audit this section defers its numbers
to (the Phase-5M analog; T99_Translation_Combat §6.4, spike C-11, BENCHMARK-DAY). Its bars are
pre-registered in bench_criteria.json v1.1.0 and are reproduced here **verbatim rather than
re-derived** — this amendment introduces no parallel numbers:
| Bar | Value | Scope |
|---|---|---|
escaped_invalid_max | 0 | the safety tooth; outranks every perf number in the run |
min_gen_per_s_median / kill_gen_per_s_median | 8.0 / 4.0 gen/s | the ~8–15 gen/s busy-scene target |
ship tier: max_vram_mb / max_p99_total_ms / max_p99_ttft_ms | 6144 MB / 1200 ms / 400 ms | sized for console shared memory, not for a PC spec class |
reference tier | 20480 MB / 2500 ms / 900 ms | the 14B dev-side reference tier |
min_requests / min_adversarial_requests | 50 / 10 | sample adequacy |
The mapping, stated so a "covered" claim can be checked:
**COVERAGE REFRESH (2026-07-28, game systems wave 2 — bench criteria relocked v1.2.0 with the
dated audit row; director-verified: X9.5 self-test 59 ok / mutation-test 25 armed, both exit
0).** The five follow-up fixtures below were BUILT: AI-2 decision-parity KILL (the
no-hardware-tiered-intelligence rule's first enforcement — a decision_trace divergence
across inference-on/off is now a KILL), the AI-1 frame-critical ban (+ a missing-declaration
INCONCLUSIVE), AI-3pc_minimumzero-VRAM + thepc_recommended/pc_highco-residency
pairs (PC bars null-with-co-residency-evidence-mandatory — no number invented pre-benchmark),
and AI-4b's A/B in KT-3 where this section rules it; the in-engine producer wired tri-state
(undeclared/off/on) so real captures can verdict. The NO-COVERAGE rows below stand as the
HISTORICAL record of the gap this amendment refused to paper over.
| Rule | Verdict class | Enforcing fixture | Status |
|---|---|---|---|
| AI-1 frame-critical ban | — | NONE | NO COVERAGE. KT-5 carries no frame-time field at all; a synchronous per-frame call would surface only as latency, well inside the 1200 ms ship p99. The bench cannot see this defect class |
| AI-2 decision parity across inference-on/off | — | NONE | NO COVERAGE, and this is the highest-priority gap. Nothing in the record expresses a decision trace; escaped_invalid catches a *line* leaking, never a *decision* diverging |
| AI-3 per-spec-class VRAM policy | OVER_BUDGET | kt5/overbudget_latency_vram.json — partial | PARTIAL. The VRAM tooth is real and proven to fire (7000 MB > the 6144 MB ship bar). But the tiers are ship/reference — console and dev tiers, not PC spec classes — and no field carries the game's concurrent working set, so KT-5 cannot express "3.0 GB of model + 7.0 GB of game exceeds an 8 GB card" |
| AI-3 min-spec zero-inference | INCONCLUSIVE | kt5/inconclusive_unknown_tier.json — wrong shape | NO COVERAGE. An unknown tier lands INCONCLUSIVE, which is the right instinct for a missing budget and the wrong tooth for a *forbidden* one |
| AI-3b validator classifier inside the carve | OVER_BUDGET | kt5/overbudget_latency_vram.json — incidentally | PARTIAL. vram_peak_mb is a whole-lane peak, so a classifier loaded beside the model inflates it — but nothing attributes the excess, and nothing fires when the classifier loads in a tier where the model is disabled |
| AI-4 / AI-4b displacement bound | — | NONE | NO COVERAGE. No frame-time field; this is a KT-3 measurement, not a KT-5 one |
| The authored floor as a valid ship configuration | KILL | kt5/kill_escaped_invalid.json + kt5/kill_throughput_collapse.json | FULL — and the mapping is exact. Both KILL verdicts prescribe "ship fallback-only for this tier (canon-pre-authorised, no code path changes)," which is *the same configuration* as §2.7.3's minimum/8 GB-recommended policy and as §6.5's new inference-off pass. Three routes, one build |
| Throughput and latency bars | PASS / OVER_BUDGET / KILL | kt5/pass_ship_tier.json (must-not-fire) + the three above | FULL at the ship and reference tiers |
| Sample adequacy → INCONCLUSIVE, not a pass | INCONCLUSIVE | kt5/inconclusive_no_adversarial.json, kt5/inconclusive_too_few_requests.json | FULL — and it is this doctrine's RULE PROF-5 in another lane. KT-5's "a bench that never stresses the guardrail cannot clear it; zero stress is no evidence, not zero risk" is *verbatim the same discipline* as §6.3's "a record missing any declaration is INCONCLUSIVE, not a pass" and the house positive-control rule. The two harnesses independently converged on it |
FIXTURES TO ADD — listed as follow-up work, NOT claimed as coverage. Ranked; each is a criteria
relock plus (where noted) a record-schema field. Nothing here is applied by this amendment.
1. kt5/kill_decision_divergence.json → KILL (AI-2). Requires a new decision_trace field
(seeded token/ability/phase sequence) and a paired inference-off trace; any divergence KILLs the
configuration, on the same "outranks perf" footing as the escaped-line tooth. **Highest priority:
AI-2 is a director ruling with zero enforcement today.**
2. A pc_minimum tier with max_vram_mb: 0 → OVER_BUDGET (AI-3). The cleanest tooth
available: kt5_runtime_inference.py's existing check is vram > tier_crit["max_vram_mb"]
(line 142), so a zero budget makes any GPU-resident model fail with no code change — a
criteria relock only. Fixture: kt5/overbudget_min_spec_gpu_inference.json. **Two scope limits,
stated so this is not over-claimed:** the tooth is reached only after the record clears sample
adequacy (≥50 requests / ≥10 adversarial) and the 4.0 gen/s kill bar, since both return earlier
in evaluate(); and it keys on vram_peak_mb, so a CPU-side lane declaring 0 passes it. It
closes the *GPU-resident* half of AI-3's min-spec zero, not the whole ruling — U-41's CPU lane
needs its own field.
3. pc_recommended / pc_high tiers + a game_working_set_mb field → OVER_BUDGET (AI-3).
The co-residency check game_working_set_mb + vram_peak_mb > card_vram_mb is the one that would
have caught the 8 GB gap by measurement instead of by arithmetic. Needs a small evaluator change,
so it is ranked below (2).
4. kt5/inconclusive_no_scheduling_class.json → INCONCLUSIVE (AI-1). A bench that does not
declare inference_scheduling_class and inference_queue_priority cannot clear the
frame-critical ban — the direct KT-5 analogue of RULE PROF-5.
5. The AI-4b displacement A/B belongs to KT-3, not KT-5 — it is a gpu_pass_ms delta between
two §6.2 records, and it lands as a §6.2 schema addition (inference_enabled, bool) rather than
as a KT-5 fixture. Recorded here so it is not lost between the two harnesses.
One honest note on the ship tier's 6144 MB bar. It is not wrong; it is **out of scope for the
PC lane**. It was pre-registered against console shared memory (RUNTIME_GUARD §8.2's console-parity
bullet, the 16 GB unified pool), where RUNTIME_GEN §B.3 budgets ~12.5–13 GB to the game and
~3.0–3.6 GB to inference. Read as a PC bar it is **~2.5 GB looser than canon's own inference
total**, ~1 GB looser than the "~2–5 GB VRAM" that RUNTIME_GUARD §8.2 — the very bullet the bar
cites as its origin — states for the model, and ~5 GB looser than the 8 GB anchor can survive. So a
configuration could PASS KT-5 at 6.0 GB and still be unshippable at recommended spec. That is precisely the gap fixture (3) closes, and it is
named here so nobody reads a KT-5 PASS as a PC-class clearance.
---
| Technique | Status 2026-07-27 | Verdict | Concrete commitment | Citation |
|---|---|---|---|---|
| UE 5.8 as ship engine | released 2026-06-17; terminal UE5 feature release | ADOPT | ship on 5.8; 5.9 only if it lands as stability. No engine version arrives mid-build that changes the cost model | W1 §1, W4 §1 — VERIFIED |
| UE6 | EA end-2027, stable ~2029 | IGNORE for EA; post-1.0 replatform | Niagara-only VFX (Cascade is removed); gameplay logic out of monolithic Blueprint graphs and behind plugin/module boundaries so a Verse-era port is mechanical | W4 §1 — VERIFIED timing / secondary on rendering |
| Lumen High (recommended tier) | Production | ADOPT — the 60 fps GI rung | sg.GlobalIlluminationQuality / sg.ReflectionQuality; Epic's own stated 60 fps open-world configuration | W1 §3.1, W2 §B.6, W3 §5.2 — EPIC |
| Lumen Lite (minimum tier) | Beta in 5.8 — and 5.8 is terminal, so Beta may be its final state on our engine line | ADOPT — the 30 fps GI rung, WITH THE §2.4 RETREAT STATED | r.Lumen.FinalGatherMethod 0; irradiance fields + probe occlusion + automatic probe placement; ~2× Lumen High; hits 60 on Switch 2. Retreat if it regresses or is withdrawn: min-spec falls back to Lumen High at reduced quality + lower internal resolution (§2.4) — same dynamic solve, same scalability group, no architecture change. §8.5 re-check trigger | W1 §3.2, W4 §1 — VERIFIED (Daniel Wright, Epic); Beta status re-confirmed live (critic V-3) |
| Lumen Epic + HWRT (high-end) | Production | ADOPT as the high-end rung only | HWRT has the highest setup cost in large scenes and heavy deforming-mesh cost | W3 §5.2 — EPIC |
| Software ray tracing as the DEFAULT | Production | ADOPT | SWRT is "the only performant option in scenes with many overlapping instances," and our worlds are by construction many overlapping instances (PCG scatter over Nanite architecture + skinned parties). HWRT is the high-end lane | W3 RULE GI-1 — EPIC |
| Baked lightmaps | — | BANNED | Every GI future in the industry (ReSTIR GI, ORCA, NRC, FSR Radiance Caching, UE6 volumetric Lumen) assumes a dynamic solve. Baking is the one decision that permanently closes that door. Baking survives only as a surgical per-asset optimization on a measured over-budget row, never as a tier | W4 §4 ruling, W1 §3.5 |
| MegaLights | Production-Ready in 5.8 | ADOPT for local lights, recommended + high tiers | Constant GPU cost, quality degrades with per-pixel light complexity — the inverse of deferred. Requires HWRT. r.MegaLights.DirectionalLights stays off (sun remains on the VSM clipmap path). Incompatible with the Forward Renderer → we are deferred-only → temporal AA/upscaling is mandatory, not optional | W1 §4, W2 §B.7, W3 §5.3 — EPIC |
| MegaLights at minimum spec | unsupported below current-gen / non-RT | DESIGNED FALLBACK REQUIRED | FINDING W3-4: every scene using MegaLights needs an authored deferred fallback lighting state. r.MegaLights.Allow 0 is a scalability row. The fallback is authored, never discovered when a min-spec capture goes dark | W3 §5.3, W2 §B.7 — EPIC |
| Virtual Shadow Maps | Production | ADOPT — sun/directional on clipmaps | Narrowed clipmap range (default levels 6–22 = 64 cm to ~40 km is more than we need); quantized/stepped sun movement because any light rotation invalidates ALL cached pages for that light | W1 §5, W2 §B.5, W3 §5.1 — EPIC |
| Ray-traced shadows for local lights | via MegaLights, default | ADOPT | Epic defaults MegaLights to RT shadows *because they are cheaper to generate* than VSMs, removing the per-light VSM CPU/memory/GPU tax. Adopting MegaLights is also a shadow-architecture decision | W1 §4.1 — EPIC |
| Contact shadows for grass/high-motion foliage | Production | ADOPT | Epic names contact shadows alone "a sufficient substitute for high resolution shadow maps" for high-motion geometry such as grass and foliage | W1 §5.5 — EPIC |
| Nanite (static + skeletal) | Production | ADOPT — the geometry architecture | Nanite-first for all static world geometry and architecture. Opaque/Masked only; no translucency, no two-sided. Hard cap 16M instances. Do NOT Nanite single-plane cards or trivially low-poly props (Epic's own counter-example is a sky sphere) | W1 §6.1, W2 §B.1 — EPIC |
| Nanite Foliage | Experimental in 5.8, and 5.8 is terminal | ADOPT HYBRID, KEEP THE RETREAT REACHABLE | Nanite Foliage for hero vegetation (trees, large plants) where wins are largest and asset count lowest; traditional instanced meshes + contact shadows for grass and ground cover. Keep the source of truth in the Procedural Vegetation Editor / source meshes so a forced retreat is a re-export, not a re-author | W1 §6.5 (option c), W2 §B.2, W4 §5 |
| Nanite Skinning (bone-driven wind) | Experimental, part of Nanite Foliage | ADOPT wherever Nanite Foliage is used | Replaces per-vertex WPO with a bone hierarchy → tight cluster bounds → fixed-function raster and removal of continuous WPO-driven VSM page invalidation. Simultaneously a Nanite win and a VSM win. 100k bones ≈ 0.1 ms GPU; one demo tree 3.5 GB → ~29 MB on disk, ~36 MB → ~2.7 MB streaming per view | W1 §6.2–6.3, W2 §B.2 — EPIC |
| Substrate — Blendable GBuffer | default-on for new 5.8 projects, not declared production-ready | ADOPT AND LOCK | Epic: "Performance is in parity with the non-Substrate path," "targets performance-oriented project (60Hz)," no cooking overhead. The cheapest build-correctly win in the entire fan-out — one project setting today, a multi-year re-authoring cost to reverse | W1 §7 — EPIC |
| Substrate — Adaptive GBuffer | not production-labelled | GATE BEHIND BUDGET REVIEW | Cost scales with on-screen material complexity; +~15% cook time. A documented per-scene exception, never a project-wide default | W1 §7 — EPIC |
| World Partition + FastGeo Streaming | FastGeo Experimental in 5.6→5.8 | ADOPT CONDITIONALLY — A/B ON OUR OWN AOI FIRST; R-05 is gated on that result | FastGeo swaps static geometry onto a lightweight non-actor registration path, killing the game-thread actor/component registration spike that *is* traversal stutter. Built by CDPR × Epic for the Witcher 4 demo. 5.8 adds a PCG FastGeo Interop plugin. Retreat: conventional World Partition actor streaming remains buildable and is the fallback — U-15 flags the supported/unsupported component matrix and the measured gain as UNVERIFIED, and "pin the engine version" is a lock-in, not a retreat | W2 §B.9.2, W3 RULE WP-5 — EPIC/TRADE |
| PCG (Production-Ready since 5.7) | 5.8 adds GPU scatter at landscape-grass parity + actor/component-less runtime generation | ADOPT | One graph per biome archetype, parameterized by registry row — never one graph per region (76 regions × bespoke graphs is 76 things to re-profile) | W3 §3.3, §3.4 — EPIC |
| HLOD (Instancing / Merged / Simplified) | Production | ADOPT | Overnight CI job, never an interactive step. 5.8's perceptual difference heuristic rebuilds HLODs only when needed — a direct factory-throughput win across 76 regions. Disable Allow Distance Fields on distance-only HLODs | W2 §B.9.3, W3 RULE WP-6 — EPIC |
| Mass + Instanced Skinned Meshes for crowds | MetaHuman Crowds Experimental in 5.8 | ADOPT the pattern; treat MetaHuman Crowds as experimental | Proximity-graded fidelity ladder authored once. Our persona system must be data-oriented from the start — persona data in Mass fragments / registries, never per-actor Blueprints, or density becomes the GTA-6-class CPU wall | W2 §B.10 — EPIC/TRADE |
| Animation Budget Allocator + Significance Manager | Production | ADOPT NOW | The enforcement mechanism for the 3.0 ms animation row (§2.6). Every skeletal mesh registers with a significance function | W3 §1.4 — EPIC |
| Ray Tracing Proxies + Reference-Based Residency | shipped, documented | ADOPT NOW | Disable "Generate Nanite Fallback Meshes" project-wide where fallbacks exist only for RT; enable Generate Ray Tracing Proxies (stream on demand) — significant memory saving. Pool via r.RayTracing.ResidentGeometryMemoryPoolSizeInMB | W4 §5 — EPIC |
| RT instance budget as a LEVEL-DESIGN rule | Epic-documented | ADOPT NOW AS A DESIGN RULE | ≤ 100,000 RT-scene instances after culling at 30 fps; far below at 60. Dynamic geometry 10k–30k tri/frame (20k at the 60 fps tier); SkeletalMeshes.LODBias=2; landscape update threshold 0.3. BVH health: 50–100 traversal iterations typical, 150+ = redesign the overlap, not the setting | W2 §B.8, W4 §5 — EPIC |
Visible in Ray Tracing = OFF by default | — | ADOPT NOW | Grass, sky and decorative layers opt in, never out. Skyboxes and grass are Epic's named culprits | W2 §B.8 — EPIC |
| TSR | ships with 5.8 | ADOPT — the vendor-neutral baseline | The 60 fps canon is measured against TSR Quality so the canon does not depend on a vendor. Epic's own 1080p→4K console path; City Sample measured near-half GPU frame time from it | W1 §8.6, W2 §B.11 — EPIC |
| DLSS 4.5 / FSR (Redstone) / XeSS 3 via Streamline-class plugins | shipping | ADOPT ALL | The real engineering cost is getting the temporal inputs correct once — motion vectors on everything that moves (including WPO/skinned/particle/translucency), correct jitter, depth, exposure and UI composition order. One setup unlocks DLSS + FSR + XeSS + DLSS 5 simultaneously, and broken motion vectors are invisible until any upscaler is switched on | W1 §8.6, W4 §9 rule 6 — VERIFIED |
| Dynamic resolution | off by default | ADOPT with an explicit per-class budget | r.DynamicRes.FrameTimeBudget defaults to 33.3 ms — a 30 fps target. Set 33.3 / 16.7 / 8.3 per class; always paired with a temporal upscaler, never run bare; enable the panic path (MaxConsecutiveOverbudgetGPUFrameCount) — our combat is full of exactly the boss-phase/VFX-burst/camera-cut events it exists for | W3 §5.4 — EPIC |
| Reflex 2 / Anti-Lag / XeLL | shipping | ADOPT FROM THE FIRST PLAYABLE — as additive headroom, never as a tuning assumption (R-49) | NVIDIA reports Reflex 2 Frame Warp cutting PC latency up to 75% (THE FINALS: 56 ms → 14 ms) — declared capture: RTX 5070, 4K, max settings, and NVIDIA's own announcement states Frame Warp debuts on RTX 50-series only, with other RTX GPUs "in a future update" (current breadth UNVERIFIED, U-33); Reflex is NVIDIA-only in any case. Rev 1 quoted the figure bare, which broke §6.5/§7.2's own undeclared-capture rule on the document's headline latency number. The reasoning stands and is unaffected: for an action RPG with parry as a ruled core primitive, end-to-end latency is a GAMEPLAY system, not a graphics setting — a parry window measured in frames is meaningless if 40 ms of it is pipeline latency. Latency belongs in the combat design spec with a target number, and that number is authored on the no-Reflex path (R-49) | W1 §8.5 — VERIFIED numbers, config + hardware gate per critic V-7 |
| Frame generation | shipping (DLSS 4.5 MFG up to 6X, FSR FG RX 9000+, XeSS 3 MFG) | HIGH-END TIER ONLY | Never counts toward 60. Never on in a recommended-class capture | W1 §8.4, W3 RULE DR-4 |
| Advanced Shader Delivery (ASD) | Agility SDK retail; UE support "in progress"; NVIDIA consumer support later 2026 | ADOPT THE PLUMBING NOW, not the feature | ASD's SODB is produced by tracing the game. Our AI QA loop is already an automated player — wiring "record a PSO trace on every automated playthrough" into it costs almost nothing today and *is* the entire ASD deliverable later. This is a genuine structural advantage of the factory model: most studios must hand-build trace coverage | W4 §6 — VERIFIED (Microsoft DirectX devblog 2026-03-12) |
| DLSS 5 neural rendering | Fall 2026, RTX 50-optimized, Streamline pipe | STAY COMPATIBLE | Additive skin only. Never build art direction that only reads correctly under neural relight — if the lighting only looks AAAAA with a Blackwell-only model on, we shipped a game that looks broken to most of Steam | W4 §2 — VERIFIED announcement, hardware gating UNVERIFIED |
| NVIDIA NRC / FSR Radiance Caching / ORCA-class radiance caches | experimental / dev-only / research | STAY COMPATIBLE, INTEGRATE NOTHING | Every vendor is converging here and every one of them replaces the *indirect* solve. Keeping GI fully dynamic and probe-agnostic is the entire compatibility action | W4 §4 |
| Neural texture compression (NVIDIA NTC / AMD NTBC / Intel TSNC) | SDK/beta/research; zero shipped games | STAY COMPATIBLE — architecture only | Texture format is a COOK-TIME decision, never an authoring decision. Keep sources unbaked at full authoring resolution, keep one swappable per-platform encoder stage, and never hand-pack channels in a way that assumes a specific block format. AMD NTBC is the safest forward bet (outputs standard BC, no shader changes, cuts install size too) | W4 §3 — VERIFIED SDKs, shipped-title claims UNVERIFIED |
| DXR 2.0 (CLAS / partitioned TLAS / Opacity Micromaps, SM 6.10) | preview planned late summer 2026 | STAY COMPATIBLE | Keep HWRT a live scalability rung in every region (never "RT off for perf" as a permanent state) and keep Nanite as the authoring source of truth, so a CLAS-era upgrade converts geometry automatically. CLAS is the standardized form of RTX Mega Geometry and directly attacks our Nanite+HWRT bottleneck | W4 §5 — VERIFIED announcement |
| Cooperative Vectors / SM6.9 / DX Linear Algebra | retail Agility SDK 1.619 (2026-02-26); no UE authoring path found | WATCH | Nothing to do. Re-check at 5.8 hotfixes and UE6 EA | W4 §2 — UNVERIFIED GAP |
| Shader Execution Reordering | required in SM6.9 | FREE | Inherited via engine/driver | W4 §5 |
| Incremental Garbage Collection | 5.8, not fully thread-safe | DO NOT PLAN ON IT | Epic recommends enabling it only in single-threaded builds (e.g. dedicated servers). The client-side cure is architectural (§4.3). Use TObjectPtr everywhere anyway — it costs nothing and future-proofs | W2 §C.3 — EPIC, and the highest-consequence trap in the fan-out |
| DirectStorage | no Epic documentation confirming UE ships it (as of Feb 2026); DS 1.4 + Zstd retail at the OS level | DO NOT ARCHITECT AROUND IT | Close the I/O budget on conventional async loading. Any DirectStorage benefit is upside, never a plan | W2 §B.9.5, W4 §6 — UNVERIFIED/assume absent |
| GPU work graphs / mesh nodes | no documented shipped-game adoption; no UE authoring path | IGNORE | Engine-internals bet. If Epic adopts it we inherit it free. Zero architecture cost today | W4 §5 |
| Hand-written mesh shaders | not broadly adopted as an engine-level architecture | IGNORE | Nanite *is* our geometry architecture; its raster path is Epic's concern | W4 §5 |
| Mesh Terrain | Experimental in 5.8; overlaps the researched Gaea/heightmap chain | EVALUATE, DON'T DEPEND | Prototype only; heightfield + Nanite is the safe path. Open decision against RESEARCHED_STACK.md Chain 2 step 3 | W1 §1.1, W4 §1 |
| Substrate NPR / X-Rite AXF / Fog SSS | Experimental | IGNORE / WATCH | Photoreal mandate makes NPR irrelevant; FSSS is atmosphere polish, not architecture | W1 §1.1 |
| PS6 / next Xbox / RTX 60 specs | unannounced; rumor-only; memory-crisis pressure | IGNORE AS A PLANNING INPUT | Neither will be a meaningful share of the *PC* install base before 2029. Do not assume a generational uplift rescues our frame rate | W4 §7 |
1. **MegaLights is adopted → minimum spec becomes RT-capable → the shadow architecture changes →
and artists light differently.** Under deferred + VSM the discipline is "few lights, each
cheap." Under MegaLights it is "**as many lights as you like, but keep per-pixel light overlap
low**," because *quality*, not cost, is what degrades. That is a completely different
lighting-authoring rule set and **it must be written into the art bible before the first region
is lit** [W1 §4.4 — EPIC]. This is a design *unlock* for a canon dense with lantern-lit
settlements, ritual grounds, vril sites and cave systems [W3 §5.3].
2. Substrate stays on Blendable GBuffer. One project setting, made today, worth years [W1 §7].
3. **The minimum spec is defined FIRST and it forces the memory and streaming architecture; the
recommended and high-end tiers then fall out as quality dials rather than as separate
engineering.** This is the explicit lesson from Assassin's Creed Shadows, whose technical
director states that targeting Xbox Series S at 1080p is what forced the memory discipline that
then made broad PC scaling work [W2 §A.4.1 — TRADE, GameDeveloper 2025-03-19].
The best-performing UE5 title of the 2025 cohort got there by **subtracting the three flagship
features** — Embark shipped without Lumen, Nanite or VSM, using probe-based RTXGI instead, and
Digital Foundry found shader-compilation stutter "simply doesn't exist" with reportedly ~3× the
frame rate of contemporary UE5 releases [W2 §E — DF-SEC]. That is a real shipped datapoint and it
deserves to sit next to the photorealistic mandate.
It is not our answer, for a specific reason: ARC Raiders paid for its frame rate with
low-resolution shadows, dated screen-space reflections and light-leaking probe GI — compromises the
art direction has already ruled out, in a 120-fps-competitive genre we are not in.
The correct synthesis is the AC Shadows posture, not the ARC Raiders posture: keep the flagship
features, tier them against the spec ladder, and hybridize where the cost curve is bad. But
take ARC Raiders' *actual* transferable lesson, which is not "disable everything":
**Every enabled feature must be one you can name a reason for. A feature that is on by default
and unjustified is a budget leak.** [W2 §E.1]
And keep AC Shadows' fallback in the pocket: when full RTGI would not hold 60 fps, a shipped AAAA
open world chose hybrid baked-GI × ray-traced GI [W2 §A.4.1 — GDC 2025 + 80.lv]. That is the
documented escape hatch if §2 row 5 will not close — and it is the *only* circumstance under which
the "no baked lighting" ban in §3.1 is revisited, by director ruling, with capture evidence.
grep -i licen returned zero hits across the rev-1 draft *and* all four lane reports. That is a
gap, not a position: the house carries a Josh ruling — **"unrestricted-components-by-default in the
shipped-asset path … licensing never constrains where the game ships or who owns it"**
(docs/translation/T99_Translation_AssetGen.md, "Four constraints are ruled"; enforced at
docs/BUILD_PLAN_END_TO_END.md:478, "No engineering closes a license"). Read carefully, that ruling
scopes to the shipped-asset path while this doctrine governs the shipped-binary path, so
§3.1 is not in violation. But three of its rows depend on the answer, and one of them is now a
gameplay dependency:
redistributables shipped inside the game binary.
latency dependency touching a ruled core primitive.
The posture, ruled here because the architecture already implies it:
1. **TSR and the fully vendor-neutral path are the unrestricted default and the ship-viable
baseline.** §3.1 already says "the 60 fps canon is measured against TSR Quality so the canon does
not depend on a vendor" — rev 1 reached that for vendor-neutrality reasons and never connected it
to the licensing ruling. It is now doing both jobs.
2. Every vendor SDK is an optional, removable plugin behind a build flag. DLSS/Streamline, FSR,
XeSS, Reflex, and any NvRTX-branch dependency. **A build with all of them stripped must be
shippable and on-budget at every spec class** — that is the enforceable form of the rule, and it
is falsifiable in one CI configuration.
3. The validation already exists; it is just renamed. §6.5 item 3's neural-features-off pass
and this licensing-removability test are the same pass — one capture set, two guarantees. If
the all-stripped build looks correct and holds its budget, both the art-direction guard and the
licensing guard are green. Gate: the stripped configuration is a named route_id bucket in the
§6.2 record schema, run at every milestone, not a one-off.
4. The scope question goes to the director, not to this document. *Does the asset-path
unrestricted-by-default ruling extend to runtime redistributables shipped in the binary?* Filed
as U-36. This doctrine does not assume either answer; it simply guarantees that the answer
cannot become expensive, because rule 2 keeps the vendor-free build permanently shippable.
---
Epic's own taxonomy names seven hitch classes: level streaming, physics, actor spawning,
PSO/shader compilation, garbage collection, asset loading, Blueprint overload [W2 §C.1 —
EPIC-TALK, Ari Arnbjörnsson, "The Great Hitch Hunt," Unreal Fest 2025]. Each requirement below is
a decision made before content exists, not a fix applied after. They are numbered so a gate can
cite them.
of distance-culling and min-screen-radius filtering and **control density through placement rules
(PCG + HLOD) instead** — Nanite forbids those cull knobs, so density that is not authored cannot
be recovered later [W2 §B.1, §D.1 — EPIC].
assemblies + voxelization + bone-driven wind. **Alpha-card foliage and WPO-driven foliage are
banned as INPUTS to the Nanite-foliage path** in the pipeline spec. Retrofitting card foliage into
Nanite foliage is a full re-author [W2 §B.2, §D.1 — EPIC]. (Grass/ground cover remains
traditional instanced + contact shadows per §3.1.)
vegetation asset — W4's verdict on Nanite Foliage is "ADOPT WITH FALLBACK** | author
vegetation so conventional instancing still works," and rev 1 kept the ban half of that while
dropping the fallback half. Nanite Foliage is Experimental in the terminal UE5 release; an
untested retreat is not a retreat. Check (gate-tier): the asset validator asserts that every
asset tagged hero_vegetation resolves both a Nanite-foliage route and a conventional
instanced-mesh route, and the conventional route is built and captured on **one biome per
milestone**, not merely declared to exist [W4 §5, critique F-8].
UPROPERTY uses TObjectPtr from the first line of C++ [W2 §D.1 — EPIC].[W2 §B.10, §D.1].
first region ships.** Actors are reserved for things that genuinely need actor semantics
(that half is unconditional and is the part that cannot be retrofitted). **Retreat: conventional
World Partition actor streaming remains buildable**, and the A/B result is what promotes FastGeo
from conditional to architectural. Rev 1 stated R-05 unconditionally while §3.1 stated it
conditionally, and offered "pin the engine version" as the mitigation — which is a lock-in, not a
retreat. Pinning the version is still correct; it is just not the fallback [W2 §D.1,
W3 RULE WP-5, U-15, critique F-8].
becomes the GC problem that incremental GC cannot fix on a client build [W2 §C.3, §D.1].
spawning is Epic hitch class #3 and has no engine-side fix [W2 §C.1, §D.1].
MegaLights-free path at minimum spec [W2 §D.1, W3 FINDING W3-4].
FSceneView, the documentedprecondition for mesh-draw-command caching — otherwise dynamic instancing silently stops merging
[W2 §B.4, §D.1 — EPIC]. Related: loose shader parameters prevent instancing; exceeding the inline
allocators (2 shader frequencies, 10 bindings, 4 vertex streams) triggers a heap allocation per
draw.
Visible in Ray Tracing defaults to OFF for grass, sky and decorative layers; opt-inonly [W2 §D.1, W4 §5 — EPIC].
Nanite cluster explosion and Lumen cost all key off this one property [W1 §5.4, W2 §D.1,
W3 RULE VSM-1 — EPIC].
internal render resolution + upscaler, never at native [W2 §D.1, W1 §3.4].
DESIGN DEFECT, reviewed at design time** — GPUScene cache invalidation during gameplay causes a
hitch in larger scenes [W2 §B.4 — EPIC].
TSoftObjectPtr / TSoftClassPtr) across every cell boundary. A hardreference from an actor in cell A to an actor in cell B **silently promotes both into a coarser
or persistent grid** — the world still works, streaming quietly stops working, and no gate fires
[W3 RULE WP-4 — COMMUNITY]. This one is ours specifically to worry about: the
registries → DataTables → PCG-placed-actors chain can emit hard references *mechanically, at
scale, in a single import run*.
Mechanism: modern explicit APIs require a fully compiled Pipeline State Object before a draw; PC
cannot ship every permutation; compilation happens on the render thread, so a miss is a visible
freeze [W2 §C.2]. Three tiers exist — PSO precaching (5.1+, automatic), the bundled .spc cache
(UE4-era, still needed), and the driver cache (which will mask your bugs during testing). Epic
intends precaching to eventually replace bundled caches; **coverage was still incomplete as of 5.6,
so shipping titles run both** [W2 §C.2 — EPIC-DOC + Tom Looman, updated 2026-04-10].
.spc cache.r.PSOPrecache.Validation=2 in dev builds; =1 is shipping-safe for telemetry. stat PSOPrecache reports Missed / Too late / Hit / Untracked.
-clearPSODriverCache. A test without it is a FALSEGREEN** — the driver cache hides every miss [W2 §D.2].
which PSOs are generated; scalability level does [W2 §C.2].
r.PSOPrecache.ProxyCreationDelayStrategy as an ART/DESIGN decision(0 = skip rendering until compiled; 1 = swap in the engine default material). You are choosing
between a hole and a grey blob — never a hitch — and it is a visible artifact either way
[W2 §C.2, §D.2 — EPIC].
as content scales [W2 §D.2].
IPSOCollector work as ENGINEERING, not configuration. Precaching does notcover hair, Nanite, or ray-tracing dynamic geometry; render-target information is only known at
draw time [W2 §C.2 — EPIC].
rate is a standing reported metric (PIX exposes it as a real-time counter as of May 2026). This
is simultaneously the PSO gate and the Advanced Shader Delivery deliverable [W4 §6 — VERIFIED].
Tailwind, not a substitute: Epic's 5.8 shader-compiler and deduplication work **cut Fortnite's
shader count by 68%**, reducing cook time, load time and memory, and 5.8 improves PSO-precaching
fallback rendering [W2 §C.2 — EPIC-TALK, Unreal Fest Chicago 2026].
incremental reachability is not fully thread-safe — an object manipulated on a worker thread
may not be marked reachable — and **recommends enabling it only in single-threaded builds (for
example, dedicated servers)** [W2 §C.3 — EPIC]. The client-side cure is architectural: keep the
live UObject count low by construction (R-04, R-06, R-07, R-05). Set
gc.IncrementalReachabilityTimeLimit and friends knowingly, but never as the plan.
activation-cost budget. New in 5.8, and the first Epic-native tool that answers "which cell**
caused that hitch" instead of "there was a hitch" [W2 §D.3, W3 RULE WP-7 — EPIC].
speed (mounts, vehicles, fast travel, the 78 typed seams). Epic publishes no recommended
value**; City Sample's Big City is the only reference setup [W2 §B.9.1, §D.3 — EPIC].
The cost law that makes this urgent: **doubling the loading range costs roughly 4× the loaded set
on a 2D grid** [W3 §2.1 — COMMUNITY]. Our derived starting point, one runtime grid per map, cell
size varying by mpv tier while the loaded-cell count is held roughly constant (~25 cells):
hero (1 m) 128 m cells / 256 m range · standard (2 m) 256 m / 512 m · ambient (5–10 m) 512 m /
768 m [W3 RULE WP-2 — DERIVED].
performance" [W3 RULE WP-1 — EPIC].
vehicle accelerating — not merely on player position [W2 §D.3 — EPIC mechanism].
paths** [W2 §D.3].
sg.TextureQuality values — ~1.0 GB minimum / ~2.5 GB recommended / ~5.0 GB high — with r.Streaming.LimitPoolSizeToVRAM=1;
never set the pool to 0 (0 = unlimited; the engine then assumes infinite VRAM). Raising the
pool is the LAST fix, never the first** — fix texture sizes and compression, pack masks into
channels, set Mip Gen Settings from the texture group, *then* consider the pool
[W2 §B.9.4, W3 RULE TS-1/TS-2 — EPIC + COMMUNITY]. Check: texture_pool_mb as configured in
the §6.2 record, compared against the class's declared value; a capture whose configured pool
exceeds its class value is INCONCLUSIVE, not a pass.
WITHDRAWN in rev 2: the "keep texture memory under ~75% of the MINIMUM-SPEC VRAM" clause. It
was a §8.3 U-27 practitioner-blog heuristic written into the requirements section, which this
document's own rule forbids ("use them to size SPIKES, never to size BUDGETS") — and it was
incoherent besides: 75% of the 8 GB minimum card is 6.0 GB, i.e. a texture pool 2.4× the
§7.3 recommended-class value and 86% of the entire 7.0 GB working-set ceiling, before render
targets, the Nanite streaming pool, RT BLAS or geometry. §7.3's numbers are ours and internally
consistent; they are now the only ones stated [critique F-3].
[W2 §B.9.5, §D.3].
Misalignment inflates the texture-streaming pool requirement for no visual gain
[W3 RULE WP-3 — COMMUNITY, mechanism-consistent].
r.Shadow.Virtual.Stats "Invalidated" is a FIRST-CLASS metric in the perf harness.Rising invalidation is the leading indicator of a shadow blowout [W2 §D.4 — EPIC]. Targets:
Allocated stays below Max Pages, Cached is high, Invalidated is near zero for static geometry.
— architecture, ruins, vril sites, props that never deform [W3 RULE VSM-2 — EPIC].
Accidental invalidators — Blueprint property writes with no visible movement — are a named cause
[W2 §B.5, §D.4 — EPIC].
rotation invalidates all cached pages for that light, so a continuously-moving day/night sun
invalidates the entire directional clipmap set every frame it moves. **Mitigation: stepped /
quantized sun movement** so invalidation happens on discrete ticks [W1 §5.3 — INFERRED mitigation
from VERIFIED behaviour]. Also narrow the clipmap range (Clipmap.FirstLevel/LastLevel) — the
default 64 cm-to-40 km throw is more than our regions need — and use the *Moving resolution-bias
variants, which exist precisely so moving lights shadow at lower resolution.
is much more expensive to render into VSMs than Nanite geometry" [W2 §B.5 — EPIC]. Distant
foliage uses contact shadows.
r.Shadow.Virtual.MaxPhysicalPages for the WORST region and treat page-pooloverflow as a HARD FAILURE — it produces checkerboard corruption and missing shadows; it does
not degrade gracefully** [W2 §B.5, §D.4 — EPIC].
six 16k VSMs — six times the shadow surface of a spot [W1 §5.1 — EPIC]. Reduce **Source Radius /
Source Angle before** cutting ray or sample counts (Epic's stated ordering) [W3 RULE VSM-3].
dramatic many-light set-pieces through MegaLights instead [W3 RULE VSM-4 — EPIC].
under that at 60** [W2 §D.5, W4 §5 — EPIC].
r.RayTracing.Geometry.SkeletalMeshes.LODBias=2.** "High polygon Skeletal Meshes are the most
common cause of high BLAS build costs," and BLAS rebuild scales linearly with triangle count
every frame for deforming geometry [W2 §B.8 — EPIC].
setting** [W2 §D.5 — EPIC].
r.DistanceFields.SupportEvenIfHardwareRayTracingSupported=0** to avoid paying for both distance
fields and RT acceleration structures at once [W3 RULE GI-2 — EPIC].
The GTA-6 lesson is that a dense simulated world hits the *game thread* wall first [W2 §D.6].
r.RDG.AsyncCompute 0 for attributable timings, then re-enable [W2 §D.6, W3 — EPIC]. r.Lumen.AsyncCompute is likewise disabled only for profiling.
alters the numbers [W3 RULE PROF-1 — COMMUNITY].
our cost lives [W3 RULE PROF-2].
visual-verification directive [W2 §D.6; memory: visual-verification-standing-directive].
NO-REFLEX PATH. Parry and dodge windows are authored and validated at the minimum spec, with
Reflex/Anti-Lag/XeLL OFF and frame generation OFF**; vendor latency reduction is additive headroom
and is never a tuning assumption. This is the exact discipline §3.1's DLSS 5 row already applies to
*visuals* ("never build art direction that only reads correctly under neural relight") and §6.5
item 3 already applies to *lighting* — rev 1 simply failed to extend it to *latency*, which is the
axis that touches a ruled core primitive (puzzle-boss-chain-meter-design: parry is a new
primitive to build). A window tuned against Reflex-2-class latency is invalid for the ruled
minimum spec, for every AMD and Intel player, and for the RTX 4060 Laptop that is now the #1 GPU
on Steam. Check: the §6.2 record carries reflex_mode and a measured end-to-end latency in ms;
the combat spec's window numbers cite a capture with reflex_mode = off at
scalability_level = minimum. A window citing any other capture is INCONCLUSIVE, not a pass
[critique F-4; W1 §8.5].
If nothing else from this doctrine is adopted, adopt these [W3 §10 — all EPIC]:
1. pcg.FrameTime = 16.667 ms by default — PCG is permitted to consume 100% of a 60 fps frame.
Every "PCG causes hitches" report is, at root, this default. Set 1.5 ms at recommended,
3.0 at minimum, 0.75 at the 120 target, and halve
pcg.RuntimeGeneration.NumGeneratingComponents from 16 to 8.
2. r.DynamicRes.FrameTimeBudget = 33.3 ms — dynamic resolution targets 30 fps by default.
Set 33.3 / 16.7 / 8.3 per class, and set r.DynamicRes.OperationMode (which defaults to off).
3. r.Streaming.PoolSize = 0 means UNLIMITED. On a 32 GB dev box this hides a minimum-spec hard
failure indefinitely and makes the harness's pool metric meaningless. Declare per class; set
LimitPoolSizeToVRAM=1.
4. WPO Disable Distance is unset by default on every asset → constant VSM page invalidation
across every PCG-scattered plant in 76 regions. Make it a factory-emitted property and a gate row
(R-11).
And the fifth, which is not a cvar but a decision: **the runtime-grid cell size and loading range
are chosen once, per mpv tier, before content exists** — because doubling the loading range costs
~4× the loaded set, and discovering that after 76 regions of placement is a re-lay-out, not a
tuning pass.
---
Status marking, per the task frame: every row below is a FUTURE FACTORY-CONTRACT CONSUME ROW.
They are written as checkable triples (rule → grounding → check) so they become gate teeth rather
than advice — **and rev 2 adds the Check column to §5.5 and §5.6, which shipped as pairs and
therefore as advice, §5.6 being the worst place for that since it is the region-gate surface**
[critique F-7].
The first wave — deterministic, zero-token, lands in docs/FACTORY_CONTRACT.md first
[W3 §11 item 5], widened in rev 2: A3, A4, A5, B1, B2, B4, C1, C2, C3, plus **E6, F6, F7, F8,
F9** from the newly-checked tables. A5 (asset memory report vs cap), C1 (tri count per LOD) and
C2 (LOD count) are exactly as deterministic as the original six and were omitted for no reason.
Gate-tier vs critic-tier is now marked on the face of every row. The doctrine is stronger for
naming the rows a machine cannot check than for implying all of them are checkable — a row whose
"check" is a judgement call is a critic instruction, and calling it a gate is how a gate that
cannot fail gets built (the §6.1 lesson, applied to §5).
Nothing here is applied to the repo by this doctrine; the director adjudicates the landing.
Note on §1.7: every reference below is to T1_Build_Pipeline_Contracts §1.7 (the build
pipeline contract's spec-class enum + D-POLY-BUDGETS table), never to a section of this document —
this doctrine's own §1 stops at §1.5.
| # | Rule | Grounding | Check |
|---|---|---|---|
| A1 | DO author solid, closed, smooth-normal geometry. DON'T ship fully-faceted dense meshes | vertex:tri ratio >2:1 inflates Nanite data and cost; "100% faceted very dense meshes are far more expensive than intended" [W3 §9.1 — EPIC Nanite content doc] | GATE-TIER: mesh-import asserts vertex:tri ratio ≤ 2:1 — the grounding already names a hard, deterministic threshold; rev 1 left it in the prose and put a vague "normal check" in the check column |
| A2 | DON'T stack surfaces closer than a pixel apart (decals-as-geometry, double-sided sheets) | Nanite cannot order sub-pixel layers → both render [EPIC] | CRITIC-TIER, not gate-tier: authoring review + r.Nanite.Visualize overdraw capture attached to the asset record. No deterministic threshold exists; the capture is the evidence, a human/critic is the judge |
| A3 | DO hold the material-slot cap: 4 / 6 / 8 / ∞ by spec class | draw calls = unique meshes × unique material IDs; T1_Build_Pipeline_Contracts §1.7 D-POLY-BUDGETS [EPIC optimizing-rendering guidelines] | GATE-TIER: automated slot count per asset |
| A4 | DO ship simple collision only — convex / capsule / box, at every spec class | T1_Build_Pipeline_Contracts §1.7 ("collision is ALWAYS simple") | GATE-TIER: collision-primitive type check |
| A5 | DON'T exceed the per-prop memory-sanity cap: 8 / 16 / 32 MB / ∞ | T1_Build_Pipeline_Contracts §1.7 | GATE-TIER: asset memory report vs the class cap (first-wave, per §5 preamble) |
| A6 | DO prefer Opaque; DON'T use Masked where Opaque will read | "Masked materials are fairly expensive compared to Opaque ones" [EPIC] | CRITIC-TIER, not gate-tier: "where Opaque will read" is irreducibly a judgement call. The deterministic half that IS gated is A7 (blend mode vs the Nanite routing table); A6 is the material-review rubric's question, and the reviewable artifact is a Masked-material count per region, trending down |
| A7 | DON'T author translucent or two-sided materials on anything that must stay on the Nanite path | Nanite supports Opaque and Masked only; anything else silently falls back to the default material and pays the far more expensive non-Nanite VSM depth path [W1 §6.1, W2 §B.1 — EPIC] | blend-mode audit against the Nanite routing table |
| A8 | DO keep material slabs to one. A second slab is an EXCEPTION that names its reason in the asset record; three or more is a rejection | Substrate: one slab ≈ legacy cost; "the second slab will be more expensive, and following slabs will increase the cost almost linearly" [W1 §7 — EPIC]. §2.2's BasePass donation (2.25 → 2.00 ms) is bought by this rule, so it is load-bearing | GATE-TIER as restated: Substrate Stats Panel slab count per material — 1 passes, 2 passes only with a recorded reason string, ≥3 fails. Rev 1's "wherever possible" was unenforceable and is withdrawn; the exception mechanism replaces it |
| # | Rule | Grounding | Check |
|---|---|---|---|
| B1 | DO set World Position Offset Disable Distance on every WPO asset. No exceptions | constant VSM page invalidation otherwise [W3 §9.2 — EPIC VSM doc] | property presence = gate row |
| B2 | DO clamp with Max World Position Offset Displacement | WPO cost scales with vertex count [EPIC Nanite doc] | property presence |
| B3 | DO split trees: Nanite trunk/structural branches + card leaves (or full Nanite Foliage for hero vegetation) | aggregate geometry — "disjointed parts that become a volume in the distance, such as hair, leaves on trees, and grass" — defeats Nanite: it cannot simplify a partially-opaque volume aggressively and cannot occlusion-cull through layered geometry [W3 §4 — EPIC] | asset-class routing table |
| B4 | DO enable Preserve Area on foliage | prevents canopy thinning at distance [EPIC Nanite doc] | property check |
| B5 | DO spawn via ISM/HISM; DON'T spawn individual static mesh actors | per-actor overhead at scale; 5.8's "actor/component-less runtime generation" is the intended path [W3 §9.2 — EPIC + COMMUNITY] | PCG graph node audit |
| B6 | DO capture Nanite overdraw at the densest authored radius before a biome graph ships | overdraw drives cull + raster cost; "excessive amounts result in higher Nanite culling and rasterization costs" [EPIC] | Today CRITIC-TIER, gate-tier from the first biome: rev 1's check proved only that an artifact exists, because Epic publishes no acceptable overdraw value and this document refuses to invent one. The bar is set by measurement, once: the first biome that closes §2.2 row 2 (Nanite VisBuffer ≤ 2.25 ms) at its densest authored radius becomes the reference overdraw capture, and its measured value becomes the numeric bar for every subsequent biome. Until that capture exists the check is "capture attached + critic read"; after it, it is a number |
| B7 | DO set PCG Generation Radii monotonically increasing with grid size | Epic: "avoid small grid sizes having a larger generation radius than larger grid sizes" — otherwise the hierarchy generates detail beyond where its own proxy exists [EPIC PCG generation-modes doc] | graph-settings lint |
| B8 | DON'T Nanite single-plane cards or trivially low-poly props | Nanite's fixed overhead is not repaid; Epic's own counter-example is the sky sphere [W1 §6.1, W2 §B.1 — EPIC] | asset-class routing table |
| B9 | DO keep Nanite displacement magnitude as small as possible; static-displacement Trim Relative Error stays at the 0.04 default and never below 0.02; landscape materials use a magnitude roughly 64× smaller than standard meshes | displacement costs culling and VSM [W3 RULE NP-6 — EPIC] | material-parameter lint |
| # | Rule | Grounding | Check |
|---|---|---|---|
| C1 | DO hold LOD0 triangle budgets: Creatures 24k/60k/120k, Familiars 18k/45k/90k | T1_Build_Pipeline_Contracts §1.7 D-POLY-BUDGETS | GATE-TIER: tri count per LOD (first-wave) |
| C2 | DO ship the LOD chain: 4/4/3/2–3 steps | T1_Build_Pipeline_Contracts §1.7 | GATE-TIER: LOD count (first-wave) |
| C3 | DO register every skeletal mesh with the Animation Budget Allocator and give it a Significance function | fixed game-thread ms budget with URO throttling by significance [W3 §9.3 — EPIC ABA doc] | component-type audit |
| C4 | DON'T rely on hardware-RT Lumen in skinned-heavy scenes | "dynamically deforming meshes… incur a large cost to update the Ray Tracing acceleration structures each frame" [EPIC Lumen doc] | scene lighting-path declaration |
| C5 | DO budget MetaHuman NPCs by their native LOD system, not this table | T1_Build_Pipeline_Contracts §1.7 | asset-class routing |
| C6 | DO treat crowds as Mass + ISKM with a proximity-graded fidelity ladder; DON'T place hero-fidelity actors at distance | City Sample uses MetaHumans up close and custom vertex-animated static meshes for distant crowds [W2 §B.10 — EPIC] | crowd-spawner audit |
Realm-class skinned assets (generation_tier='realm') inherit this table as their FLOOR until §8.2
U-44 closes — see the realm-class block at §5.6.
The presentation ladder demands mastery-tier animation and VFX variants — "every spell and ability
and movement needs to be beautiful and fluid art" (Josh, 2026-07-26, COMBAT_PROGRAM_ADDENDUM.md).
That multiplies VFX asset count by the tier ladder, making VFX budget discipline a scaling
problem rather than a per-effect one. (The ruling is now sized: FULL per-tier variants — 248 live
variant keys, 372 at the 12-tier clamp — Josh, 2026-07-28 fourteenth sitting, the band collapse
OVERRULED; ASSET_DROPIN_CONTRACT.md §7 item 7.) Realm-class VFX surfaces inherit this table as
their FLOOR until §8.2 U-44 closes — see the realm-class block at §5.6.
| # | Rule | Grounding | Check |
|---|---|---|---|
| D1 | DO hold the frame's translucency + Niagara budget: 1.10 ms @ 60 fps | §2.2 row 7 [DERIVED] | stat gpu Translucency + stat Niagara |
| D2 | SPIKE-INVESTIGATION TRIGGER, NOT A LIMIT: emitter counts above ~30 actively processing in a scene TRIGGER A PROFILE, never a rejection. The binding limit is D1's 1.10 ms row, which is measured | frame time inflates beyond that, especially with GPU particles [W3 §9.4 — COMMUNITY, summary only] — this is a §8.3 U-27 practitioner number: "use them to size SPIKES, never to size BUDGETS" | [UNVERIFIED — NOT A GATE]. stat Niagara emitter count is *reported* alongside the D1 capture and opens an investigation; it never fails a build on its own. Rev 1 wrote it as a DON'T with a named check, which is a gate built on an unverified number — the one thing "How to read" forbids outright [critique F-3] |
| D3 | DO use cutout textures / alpha erosion instead of full rectangular sprites | kills corner overdraw [COMMUNITY] | material audit |
| D4 | DO keep big-effect materials to ~100–150 instructions | [COMMUNITY] | Material Editor Stats |
| D5 | DO give every ability VFX a Significance/scalability entry so sg.EffectsQuality can scale it — within the D10 limit: cost scales, IDENTITY does not | §7.3 (the three-level scalability architecture; rev 1 cited a non-existent "§3.6") [DERIVED] | GATE-TIER: scalability entry presence, plus the D10 pairing |
| D6 | DO treat mastery-tier variants as scalability TIERS, not unconditional upgrades | a tier-12 VFX must still fit 1.10 ms at recommended spec — the ladder is a fidelity ladder, not a budget exemption [DERIVED] | per-tier capture |
| D7 | DO route many-small-lights effects through MegaLights, with an authored deferred fallback | MegaLights = constant cost, but unsupported at the minimum tier [EPIC MegaLights doc] | fallback-state presence |
| D8 | DO author correct motion vectors on particles, translucency, WPO and skinned geometry | broken motion vectors are invisible until any upscaler is on, then they are everywhere — and they block TSR, DLSS, FSR, XeSS and DLSS 5 simultaneously [W1 §8.6, W4 §9 rule 6] | upscaler-on A/B capture |
| D9 | DO author VFX with Niagara only | Cascade is fully removed in UE6; this is the one cheap forward-compat action [W4 §1] | asset-type audit |
| D10 | THE MASTERY-TIER DISTINCTION IS PRESERVED AT EVERY SPEC CLASS; ONLY ITS COST SCALES. A tier-12 ability must remain visually distinguishable from its tier-1 form at minimum spec. What degrades under sg.EffectsQuality is layer count, particle count, resolution and secondary detail — never the identity of the variant. Serving the tier-1 effect for a tier-12 ability is a canon defect, not a scalability decision | Josh's ruled presentation ladder — mastery-tier animation/VFX variants exist and "every spell and ability and movement needs to be beautiful and fluid art" (COMBAT_PROGRAM_ADDENDUM.md, 2026-07-26). D5 + D6 as written in rev 1 permitted a *compliant* implementation to cull the tier-12 variant down to tier-1 for the recommended-spec majority — silently deleting an earned reward, and the perf gates would have scored it as a pass [critique F-6] | CRITIC-TIER with a defined pass/fail: a per-tier paired capture at minimum AND recommended spec, judged by the image critic on one question — *is the tier-12 variant identifiable as tier-12 at this class?* Fail = the scalability entry is re-authored. This is also the pass/fail criterion §6.5's presentation-ladder tier pass previously lacked |
D6 is the rule most likely to be violated, because the ladder's whole emotional point is "the
tier-12 fireball is spectacular." The discipline that makes it survivable: **spectacle comes from
motion, timing, silhouette and light — which are nearly free — rather than from more overlapping
translucent layers, which is the single most expensive thing a VFX author can add** [W3 §9.4].
D10 is why that discipline is not merely a cost trick but the design answer: motion, timing,
silhouette and light survive every scalability level intact, so an effect whose spectacle is built
from them keeps its tier identity on the minimum-spec card — whereas an effect whose spectacle is
built from stacked translucent layers loses exactly the thing the player earned the moment
sg.EffectsQuality drops. **D6 protects the budget; D10 protects the reward; the same authoring
discipline satisfies both.**
| # | Rule | Grounding | Check |
|---|---|---|---|
| E1 | DON'T run logic in On Tick / On Paint; use events and dispatchers | [W3 §9.5 — EPIC, Optimization Guidelines for UMG] | GATE-TIER: widget-Blueprint graph scan — any node graph rooted at Tick/OnPaint containing non-trivial logic fails. Deterministic; the asset graph is machine-readable |
| E2 | DO wrap rarely-changing widget groups in an Invalidation Box first; add Retainers only if draw calls still need cutting | [EPIC] | CRITIC-TIER: UI review — "rarely-changing" is a judgement. The reviewable artifact is a Retainer-count report; a Retainer added without a preceding Invalidation Box is the flag |
| E3 | DO mark frequently-changing widgets Volatile — caching them costs more than repainting | [EPIC] | CRITIC-TIER: same review, opposite direction — a widget inside an Invalidation Box that invalidates every frame is the flag |
| E4 | DON'T nest Canvas Panels; DO use Spacers over Size Boxes; DON'T combine Scale Box with Size Box | Canvas Panels increment Layer IDs → extra draw calls; Scale+Size causes layout thrash [EPIC] | GATE-TIER: widget-hierarchy lint — nested CanvasPanel, and ScaleBox with a SizeBox parent/child, are both structural patterns and both deterministically detectable |
| E5 | DO prefer material-driven UI animation (GPU, zero CPU) over Sequencer or layout-invalidating animation | Epic's own cost ranking [EPIC] | CRITIC-TIER: animation-source audit; the reviewable artifact is a count of Sequencer-driven UI animations per screen, trending down |
| E6 | DON'T build 1000+-child widgets; segment into always-loaded / preloaded / async-on-demand | [EPIC] | GATE-TIER (first-wave): child count per widget tree — a hard number, fails above 1000 |
| E7 | DON'T ship the debug HUD's VERB KIT | canon: no mode taxonomy on player surfaces (Josh 2026-07-24) — a *canon* rule with a perf side-benefit | GATE-TIER: string/asset scan of shipping-config UI for the verb-taxonomy tokens, positive-controlled (the scan must be proven able to find the token in the debug HUD before its zero on the shipping build means anything — memory: a-search-that-cannot-match-reports-zero) |
This is the region-gate surface, so it is the worst table in the document to ship without checks.
Rev 1 did exactly that. Most of these rows already had a deterministic check available through their
R-number; it simply was not written down [critique F-7].
| # | Rule | Grounding | Check |
|---|---|---|---|
| F1 | DO use soft references across cell boundaries | R-14 — silent grid promotion, no gate fires | GATE-TIER: the hard-cross-cell-reference scan (§9 item 6) — every hard UPROPERTY reference whose target actor lives in a different cell fails. Deterministic, and it catches a defect class frame captures misattribute |
| F2 | DO set Shadow Cache Invalidation Behavior Static/Rigid on non-deforming geometry | R-33 | GATE-TIER: property audit — non-deforming primitives with invalidation behaviour left at the default are listed; the list must be empty or annotated |
| F3 | DON'T place multiple large shadow-casting lights covering most of the screen | R-39 [EPIC VSM doc] | CRITIC-TIER with a reported metric: count of shadow-casting lights whose attenuation volume exceeds the region's screen-coverage threshold, read off the same capture as R-32's Invalidated metric. Screen coverage is view-dependent, so the number opens a review rather than failing a build |
| F4 | DON'T place lights inside geometry; DO narrow attenuation ranges, spot cones, barn doors | [EPIC MegaLights doc] | GATE-TIER for the first half: a light-placement lint — a light origin inside a collision/Nanite volume is deterministic and fails. CRITIC-TIER for the second half (how narrow is narrow enough is a lighting judgement) |
| F5 | DO merge many tiny lights into fewer area lights | [EPIC MegaLights doc] | CRITIC-TIER: light count per volume reported per region, with r.MegaLights.Visualize.LightComplexity attached; the metric that actually gates is §2.2 row 6 (1.25 ms) |
| F6 | DO keep to ONE runtime grid per map | R-26 [EPIC WP doc] | GATE-TIER (first-wave): runtime grid count per map — must equal 1 |
| F7 | DO align streaming grid cell size to landscape component resolution | R-31 | GATE-TIER (first-wave): cell size modulo landscape component size = 0 |
| F8 | DO hold the per-region RT-instance budget (≤100k post-cull at the floor; far below at 60) | R-40 [EPIC RT Performance Guide] | GATE-TIER (first-wave): rt_instances_post_cull from the §6.2 record vs the class budget |
| F9 | DO attach a KT-3-shaped capture record to every region before it is called done | §6 + the visual-verification directive | GATE-TIER (first-wave): record presence + schema completeness — a record missing any §6.2 declaration field is INCONCLUSIVE, not a pass (RULE PROF-5) |
| F10 | DO keep hardware RT a LIVE scalability rung in every region — never "RT off for perf" as a permanent state | DXR 2.0 / CLAS forward-compat [W4 §5] | GATE-TIER: an automated load-and-capture smoke of every region at the high-end rung with HWRT enabled; a region that will not load or render on that rung fails, which is what stops "RT off for perf" from silently becoming permanent |
**THE REALM CLASS (reservation clause 6 — ASSET_DROPIN_CONTRACT.md §7 item 6; register v1.4
entry 95).** The generation_tier enum reserves realm above region_hero at the DR-2 freeze, and
this document's budget tables are where that class must not be silent: **realm surfaces inherit the
region budgets of §5.3, §5.4 and §5.6 as their FLOOR ruleset** — a realm asset passes the same C/D/F
checks as a region asset — until benchmark day mints realm-specific numbers (§8.2 U-44). F8's
"class budget" includes the realm class explicitly: a realm row holds the same ≤100k post-cull bound
until U-44 closes. What a realm may spend ABOVE the region floor — the ceiling the reservation
exists to protect — is a benchmark-day derivation, never an assumed headroom; per this document's
own reading rule, no gate is built on an unmeasured realm budget.
---
FINDING W3-1 (HIGH) — the pre-registered KT-3 bars CONTRADICT the ruling.
Tools/bench/bench_criteria.json (criteria_version 1.1.0, locked 2026-07-27T20:18Z) declares four
spec classes and three of them target 60 fps [W3 §0]:
steam_deck_minimum : target_fps 30, p99 50.0 ms, draws 2000, pool 1024 MB mid_range_pc : target_fps 60, p99 33.3 ms, draws 5000, pool 2560 MB high_end_pc : target_fps 60, p99 25.0 ms, draws 10000, pool 5120 MB ultra : target_fps 60, p99 25.0 ms, draws 0(∞), pool 0(∞)
Reading note, because these two zeros mean opposite things (critique F-15): ultra's
draws 0(∞) and pool 0(∞) are bench-criteria sentinels meaning "this class declares no cap,"
which is a deliberate and correct statement about a fidelity ceiling. They are not
r.Streaming.PoolSize = 0, which §4.8-3 identifies as a defect because there 0 tells the *engine*
to assume infinite VRAM. Same glyph, different object: one is a bar that declines to constrain, the
other is a runtime setting that conceals a hard failure. **ultra still declares a real texture
pool cvar per §7.3; only its bench bar is uncapped.**
Under the ruled canon, high_end_pc and ultra are the 120-target class. Leaving them at 60
means KT-3 will pass a build that misses the ruled high-end target by 2× — and a bar that
cannot fail is not a bar. The harness's own doc block says these are canon_illustrative and that
"the bar or the build must move — director adjudicates." This is the director's move to make,
through run_all_benchmarks.py --relock so the change is **dated and attributed, never a silent
edit**.
The recommended relock (three ruled classes onto the four-class §1.7 enum, which stays):
| §1.7 enum | Ruled class | target_fps | p99 max | Rationale |
|---|---|---|---|---|
steam_deck_minimum | minimum spec | 30 | 50.0 ms | ship-viability floor, unchanged |
mid_range_pc | recommended spec | 60 | 33.3 ms | the class the whole doctrine is written to |
high_end_pc | high-end | 120 | 16.7 ms | ruled uncapped/120 |
ultra | high-end, fidelity ceiling | 120 | 16.7 ms | budgets unconstrained; the frame target is not |
Second-order consequence, to be adopted deliberately rather than discovered on benchmark day:
with high_end_pc at 120, the existing kill_margin: 2.0 on that class now fires at **16.7 ms p50
— exactly the recommended-spec budget**. That is the correct semantics (a high-end build that only
reaches the recommended target has lost its class), but it should be a deliberate re-lock.
kt3_wp_pcg_runtime.py consumes a record produced by UHumanityBenchSamplerSubsystem and pinned
by the Humanity.Bench.SamplerRecordContract test, so this is a coordinated two-file change
[W3 §8.3]:
| Field | Why it must exist |
|---|---|
route_id | a scenario must be a fixed, replayable path registered by id — otherwise "the frame got worse" and "the tester walked somewhere else" are indistinguishable |
scalability_level | a capture at "high" scored against the recommended budget is meaningless |
internal_resolution + upscaler + upscaler_mode | a 50%-screen-percentage capture is not a native one; the canon is defined at internal resolution (§1.5) |
frame_generation (bool) | must be false for any recommended-class verdict (§1.3 fact 4) |
gpu_pass_ms (map: pass → ms) | promotes KT-3 from pass/fail on total frame time to attribution against the §2.2 table — without it the budget table can never be falsified |
game_ms / draw_ms / rhi_ms, p50 + p99 | RULE A is unverifiable without them; today the record cannot prove we are GPU-bound |
wp_cell_size_cm / wp_loading_range_cm | the §4.4 settings under test |
pcg_frame_time_cvar | a KT-3 run taken with pcg.FrameTime at its 16.667 ms default measures a configuration we would never ship — and today nothing in the record would reveal it |
vram_working_set_mb | the 7.0 GB recommended-spec ceiling (§1.2) needs an enforcing metric from the first region, not a measurement at the end — and per §1.4b it is now a pass/fail axis of the falsification test in its own right, not a note |
texture_pool_mb (as configured) | distinct from vram_working_set_mb, which is measured. R-29 and §7.3 constrain the *configured* pool (~1.0 / ~2.5 / ~5.0 GB per class); without this field a capture taken with the dev box's oversized pool is indistinguishable from a compliant one, and §4.8-3's whole hazard is invisible |
reflex_mode (off / on / on+boost) + end_to_end_latency_ms | R-49. Parry and dodge windows are authored on the no-Reflex path at minimum spec; without these two fields that rule cannot be enforced and the latency claim in §3.1 cannot be reproduced. NVIDIA's own headline figure is a 4K/max/RTX-5070 capture, which is precisely the kind of undeclared configuration RULE PROF-5 exists to reject |
pso_misses / shader_cache_hit_rate | R-16, R-20, R-22 |
vsm_invalidated_pages | R-32 |
rt_instances_post_cull | R-40 |
frame budget, p99 vs a separate ceiling, hitches >50 ms capped at 2/min) is the right shape and
becomes the house standard for every perf gate, not just KT-3 [W3 §8.2].
capture taken with PCG disabled or at the wrong mpv tier. **Generalize it: every perf record
declares the feature state it measured** (PCG, Nanite, MegaLights, Lumen rung, upscaler + mode,
frame generation, scalability level, internal resolution, driver-cache-cleared) — and **a record
missing any of those is INCONCLUSIVE, not a pass.** This is the direct analogue of the
positive-control rule that already binds our search claims [W3 §8.2; memory:
a-search-that-cannot-match-reports-zero].
-clearPSODriverCache. A testwithout it is a false green.
streaming cost lives in the tail: a 5-second sample has no p99" — that clause is correct and is
quoted into the doctrine rather than re-derived [W3 §8.2].
The soak is the long-form run that catches what a 60-second scenario cannot: GC pauses, streaming
churn, VSM cache drift, PSO misses in rarely-visited content, memory creep.
| Instrument | What it must capture | Source |
|---|---|---|
Unreal Insights .utrace, emitted alongside the CSV | the full timeline; "which call grew" — the only instrument that attributes a hitch to a callsite | W3 §8.1 — EPIC |
| World Partition Insights (new in 5.8) | per-cell streaming analysis with session playback — "which cell caused that hitch" instead of "there was a hitch" | W3 §2, W2 §B.9.1 — EPIC |
CSV Profiler → PerfReportTool.exe -csvdir <path> -summarytable all -reporttype Default60fps | per-frame metric trend + charts. Default60fps is a stock report type — literally our target | W3 §8.1 — EPIC |
stat gpu per-pass series | the §2.2 table, measured rather than asserted | W3 §8.1 |
r.Shadow.Virtual.Stats Invalidated / Allocated / Cached over time | shadow-cache drift, the leading indicator of a blowout (R-32, R-37) | W2 §D.4 |
stat PSOPrecache Missed / Too late / Hit / Untracked + PIX shader-cache hit rate | R-16, R-20, R-22; also the ASD SODB harvest | W2 §C.2, W4 §6 |
| PSO/SODB trace on every automated playthrough | R-22 — the QA loop *is* the trace-collection mechanism | W4 §6 |
| Hitch histogram, not a hitch average | the canon is "60 fps HELD" (§1.5). Proposed gate: zero frames over 33 ms in a scripted traversal run | W1 §11 |
VRAM working set + stat streaming + stat TextureGroup over the full soak | the 7.0 GB ceiling and the pool declaration (R-29) | W1 §9.3, W3 §6.1 |
stat WorldPartition / stat levels / wp.Runtime.ToggleDrawRuntimeHash2D | part of every capture, not a thing reached for when it breaks | W3 RULE WP-7 |
| Gauntlet as the harness | headless execution that drives the game, triggers the scenario, invokes PerfReportTool | W3 §8.1 — COMMUNITY |
Soak duration and route: at least one full traversal of the largest authored region at maximum
traversal speed (mount / fast-travel / vehicle), plus one combat set-piece with the densest
authored light and VFX load, plus one dwell in the densest settlement — each registered by
route_id so week-over-week comparison is real.
The QA loop is the sole player until the Josh Gate, so the visual critics carry the entire
play-quality claim. Under this doctrine they need three things they do not have today:
1. Every screenshot and clip declares its spec class and its full feature state. A critic
reading a frame captured on the 5090 at Epic/4K/DLAA is reading a frame **no player in the
recommended class will ever see**. The image record carries the same declaration block as the
perf record (§6.2): scalability level, internal resolution, upscaler + mode, frame generation,
GI rung, MegaLights on/off. **An undeclared capture is INCONCLUSIVE for a look claim exactly as
it is for a perf claim.**
2. Paired captures at recommended AND minimum spec for every look-bearing verdict. The
MegaLights-off deferred fallback (W3 FINDING W3-4) is an *authored lighting state*, not a
degradation — so it needs its own critic pass. The failure mode this catches is precise: a
region that reads AAAAA under MegaLights and reads flat, dark or wrong with the fallback on,
discovered at cert instead of at authoring time.
3. **A neural-features-off validation pass — which is ALSO the licensing-removability test
(§3.4) and ALSO the R-49 latency baseline.** Every region must look correct, hold its budget,
and play correctly with zero neural features and zero vendor SDKs enabled — that single
configuration is simultaneously the min-spec rung, the AMD/Intel/older-NVIDIA path, the
all-vendor-plugins-stripped build that §3.4 rule 2 requires to stay shippable, the no-Reflex
path on which combat timing windows are authored, and the guard against art direction that only
reads under DLSS 5's neural relight [W4 §2, §9 rule 7]. One pass, four guarantees — which is
the argument for running it at every milestone rather than once. If the lighting only looks right
with a Blackwell-only model on, we have shipped a game that looks broken to most of Steam.
4. AN INFERENCE-OFF VALIDATION PASS — the deterministic layer IS the game (§2.7). Every region,
every boss and every companion encounter must **run, hold its budget, and PLAY CORRECTLY with all
model inference disabled**. Not "degrade acceptably" — *play correctly*: identical bosses, tells,
punish windows, win conditions, companion tactics, facts and progression, with the authored
fallback bank supplying lines in place of generated ones. Models enrich; they never carry.
Item 3 strips *vendor* neural features and asks whether the image is still correct. This one
strips *our own* model layer and asks whether the game is still there.
the argument for running it at every milestone: it is canon's ruled authored-floor ship
configuration (RUNTIME_GUARD §6 — "the floor is already the whole game," with no code-path
change), it is the configuration both of KT-5's KILL verdicts prescribe, and it is the shipped
configuration for the minimum class and the 8 GB recommended anchor per §2.7.3. **One pass,
four guarantees: the AI-2 decision-parity guarantee, the min-spec ship configuration, the
accessibility "deterministic text" mode, and the replay/speedrun-legit mode**
[RUNTIME_GUARD §8.3].
inference disabled — and that union is exactly what a minimum-spec player runs. It is
therefore the one capture set that must never be skipped.
capture on the same route_id at the same seed, inference-on and inference-off,
asserting an identical decision trace (AI-2's check). Divergence in the decision trace is a
fail; divergence in the line text is the feature working. The enforcing fixture does not exist
yet and is listed first in §2.7.5.
Two additional critic hooks fall out of the doctrine:
the critics must evaluate the *upscaled* image — ghosting, disocclusion trails, foliage shimmer —
because that is the shipped image. Broken motion vectors (D8) are invisible in a native capture
and obvious in an upscaled one.
captures: the tier-12 variant is judged both on spectacle and on whether it held the 1.10 ms
translucency row. D10 supplies the criterion rev 1 lacked: the captures are **paired at
minimum AND recommended spec**, and the question the critic answers is binary — *is the tier-12
variant identifiable as tier-12 at this class?* A "yes" on the 5090 and a "no" at minimum spec is
a fail, because it means the scalability entry culled the reward rather than its cost.
---
Our dev box is an RTX 5090 — the fastest consumer GPU available, and 0.41% of Steam systems.
Meanwhile a laptop GPU took the #1 spot in the survey for the first time in history (RTX 4060
Laptop, 3.81%), 1080p remains the modal resolution at 51.12% (1440p 21.44%), and **8 GB is
still the largest VRAM tier at 25.64% — but only just, with 16 GB at 24.50% and closing** (8 GB
−0.25 pts m/m, 16 GB +0.45; see §1.4b) [VERIFIED — Valve, Steam Hardware & Software Survey,
June 2026, read direct 2026-07-27 (critic re-verification V-6)]. Rev 1 attributed the 0.41%
figure to the April 2026 survey; it is also the June figure, and June is what §1.4 cites, so the
whole document now reads off one survey month.
**Every frame measured on the 5090 is a lie about the recommended class unless a declared proxy
converts it** [W3 FINDING W3-2]. Worse, three of the four wrong-by-default settings (§4.8) are
specifically the ones a 32 GB dev box conceals: an unlimited r.Streaming.PoolSize will happily
hide a defect that hard-fails on the 8 GB minimum card, and the KT-3 streaming_pool_mb_peak
metric will report a number that means nothing.
**No capture taken at the dev box's own settings is ever evidence for a spec-class claim. Every
claim about a spec class is made from a capture taken in that class's declared configuration —
either on the class's hardware, or on the 5090 under a PRE-REGISTERED PROXY whose extrapolation
is recorded in the record rather than assumed.**
The pre-registered proxy for the recommended class, until a real xx60-class machine exists in the
lane [W3 §0 FINDING W3-2 — DERIVED]:
(§7.3, immediately below — rev 1 pointed this at a non-existent "§3.6", which meant the single
rule the whole 5090-proxy discipline rests on cited a void), never the editor default and never
Epic/Cinematic.
t.MaxFPS 0 so the frame time is the measurement, not the cap.r.Streaming.PoolSize at the recommended-class value, LimitPoolSizeToVRAM=1, and the **7.0 GB VRAM working-set ceiling
enforced as a gate**, not observed.
*comparative* verdict against a stored 5090 baseline for the same route_id; it produces an
*absolute* 60 fps verdict only on the anchor hardware.
Build Config/DefaultScalability.ini from BaseScalability.ini with exactly **three named levels
mapping 1:1 onto the ruled classes** — not four, not five [W3 §7.2]:
| Scalability group | minimum (30) | recommended (60) | high (120) |
|---|---|---|---|
sg.ViewDistanceQuality | foliage/grass density reduced; WP loading range reduced | baseline | extended |
sg.ShadowQuality | VSM resolution LOD bias +, SMRT samples 4 | bias 0, samples 8 | bias −, samples 8 |
sg.GlobalIlluminationQuality | Lumen Lite | Lumen High (SWRT) | Lumen Epic (HWRT) |
sg.ReflectionQuality | screen-space / Lite | Lumen reflections | Lumen HWRT reflections |
sg.PostProcessQuality | TSR Low | TSR/DLSS Quality | native or DLAA |
sg.TextureQuality | pool ~1.0 GB, mip bias + | pool ~2.5 GB | pool ~5.0 GB |
sg.EffectsQuality | Niagara significance culling aggressive | baseline | full |
custom sg.NaniteQuality | r.Nanite.MaxPixelsPerEdge raised; small streaming pool | baseline | detail-max |
custom sg.PCGQuality | pcg.FrameTime=3.0, radii reduced, density scaled | pcg.FrameTime=1.5 | pcg.FrameTime=0.75, radii extended |
custom sg.LightingQuality | r.MegaLights.Allow 0 (authored deferred fallback) | MegaLights on, DownsampleMode conservative | MegaLights full |
A number set in a map, a Blueprint, or a level-specific ini cannot scale, and the minimum-spec
class will silently inherit the recommended-spec value.
sg.PCGQuality and the Nanite/lighting group) are OURS tocreate.** The engine ships no PCG scalability group, and PCG is the single largest runtime-tunable
cost in our stack. Creating it now is cheap; retrofitting after 76 regions of content is not.
Config/<Platform>/ and -CustomConfig=<Directory> absorb console/handheld classeslater** with a MINOR schema bump rather than a rework.
win**; Epic's strong recommendation is to change parameters inside *Scalability.ini and observe
the groups consistently rather than scattering small changes into device profiles
[W3 §7.1 — EPIC, Lyra scalability doc].
The proxy is a bridge, not a substitute. Two reference machines close the gap permanently and both
are cheap relative to the risk they retire:
1. An xx60-class desktop reference (RTX 5060 / 4060 Ti / RX 7600 XT, 8 GB) — the anchor
hardware for the §1.4 falsification test. Without it the 60 fps canon is a hypothesis
indefinitely.
2. A laptop-class reference, because **a laptop GPU is now the single most common GPU on
Steam** and thermal/power behaviour is not reproducible by downclocking a desktop part
[W4 §7, §9 rule 9].
Until they exist: every recommended-class claim in the repo carries the words "5090 proxy capture"
and the recorded extrapolation, and the §1.4 falsification test stays open on the 5090-window
plan rather than silently assumed closed.
Not everything about the dev box is a hazard. HLOD material-baking wants ≥24 GB VRAM and the 5090's
32 GB clears it comfortably [W3 RULE WP-6 — COMMUNITY]; DDC warming (-run=DerivedDataCache -fill)
and overnight HLOD builds are zero-token local compute that never waits on a model budget
(standing efficiency policy). Use the 5090 for throughput, not for verdicts.
---
| # | Item | Status | How to close it |
|---|---|---|---|
| U-01 | "Hitting 60 FPS in UE5 and Keeping It There," Matt Oztalay, Epic, GDC 2026 | abstract only; GDC Vault body not reachable | Watch it. The abstract is almost verbatim this doctrine's thesis — *start projects at the target frame rate rather than optimizing afterward* — so the talk is either our strongest corroboration or our sharpest correction [W1 §11.1] |
| U-02 | "MegaLights: Stochastic Direct Lighting in UE5," SIGGRAPH 2025 Advances | PDF exceeded the fetch size limit | Download manually. The only public document likely to contain per-platform MegaLights millisecond costs — §2.2 row 6 is pure [DERIVED] until it is read [W1 §4.3] |
| U-03 | "The Road to 60 fps in The Witcher 4 UE5 Tech Demo," Unreal Fest Orlando 2025 (Ortegren/Rudzki) | per-pass numbers via secondary summary only (PC Gamer 403'd) | Watch the talk. It is the closest public analogue to our exact shape [W3 U3] |
| U-04 | "A Tech Artist's Guide to Automated Performance Testing," Unreal Fest 2025 | page metadata fetched, body did not | Watch before finalizing the CI belt [W3 U2] |
| U-05 | GDC 2025 AC Shadows talks — *Rendering 'Assassin's Creed Shadows'* (Lopez) and *Micropolygon Rendering in 'Anvil'* (Berenguier/Minnetian) | listings + speakers + scope confirmed; slide-level numbers NOT-READ | GDC Vault access. Harvest: frame-time breakdowns, RTGI ms costs, the hybrid baked×RT GI split, the GPU-driven streaming architecture [W2 §F.6] |
| U-06 | Epic's FastGeo Streaming guide | fetch returned only a table of contents | Read on the dev box. The mechanism description (lightweight non-actor registration replacing AActor/UActorComponent registration) is consensus, not Epic-confirmed [W2 §F.3, W3 U5] |
| # | Item | Status |
|---|---|---|
| U-07 | Whether the xx60 class actually holds 60 fps at 1080p High with OUR content | THE HYPOTHESIS. §1.4's falsification test. Build one complete region, profile on RTX 5060-class |
| U-08 | Nanite's fixed GPU overhead in ms | UNVERIFIED — Epic publishes none; only SEO blogs claim figures |
| U-09 | VSM cost in ms for our content | UNVERIFIED by design — cost is content-dependent (invalidation-driven) |
| U-10 | TSR vs DLSS ms cost at 1080p/1440p on 2026 builds | UNVERIFIED — no 2026 primary comparison found |
| U-11 | Nanite skinned-mesh costs (~0.6 ms vs 0.4 ms per PS5 MetaHuman; +40–60% memory) | [BLOG] tier, NOT Epic-stated. Verify on our own hardware before any character budget depends on it [W2 §F.2] |
| U-12 | r.Nanite.Streaming.StreamingPoolSize default (512 MB widely quoted) | UNVERIFIED — read it off the running build before writing it into any ini [W3 U1] |
| U-13 | The exact stat gpu pass-name set in 5.8 | UNVERIFIED — names shift between versions. Capture one frame on our build and pin the names before §2.2 becomes a gate [W3 U7] |
| U-14 | World Partition cell-size / loading-range recommendation | Epic publishes none. Our §4.4 numbers are DERIVED; City Sample's Big City is the only reference [W2 §F.4] |
| U-15 | FastGeo's supported/unsupported component matrix and measured gain | UNVERIFIED — Experimental; A/B on our own AOI [W3 U5] |
| U-16 | Nanite per-section hybrid (Nanite trunk + card leaves on one mesh) from 5.7 | [COMMUNITY summary only] — verify against Epic docs before it becomes a factory rule (it is currently B3) [W3 U4] |
| U-33 | Reflex 2 breadth beyond RTX 50-series, AND our own measured end-to-end latency on the no-Reflex path | NVIDIA's announcement says Frame Warp debuts RTX-50-only with other RTX GPUs "in a future update"; whether that has broadened is UNVERIFIED as of 2026-07-27. The gating number is ours, not NVIDIA's: R-49 needs measured end-to-end latency at minimum spec with Reflex off, because that is the configuration the parry window is authored against [critique F-4] |
| U-34 | Whether the high-end class actually holds 8.33 ms at Lumen Epic + HWRT + 1440p–4K | ITS OWN BENCHMARK DAY. §2.5 is the least-evidenced table in the document and budgets Lumen at 2.000 ms for the rung Epic publishes at 8 ms / 30 fps. The recommended-class measurement does not settle this by halving [critique F-9] |
| U-35 | Which spec class the RTX 4060 Laptop — the #1 GPU on Steam — actually lands in | §1.2's laptop column is [DERIVED] placeholder. §7.4 records that laptop thermal/power behaviour is not reproducible by downclocking a desktop part, so this closes only on the laptop reference machine. Blocks publishing a mobile part on a store page [critique §5 correction 2] |
| U-37 | How much GPU frame time the inference lane actually DISPLACES at its worst-case duty cycle — measured as an inference-on/inference-off A/B on the same route_id, at BOTH configurations where the lane is enabled: recommended-≥12 GB (16.67 ms) and the 120 target (8.33 ms) | UNMEASURED. The number AI-4b's zero-displacement bound is set from, and the number that decides whether §2.2 ever needs an inference row at all. Async compute is not free: a low-priority queue is preemptible, not costless, and no capture exists. Needs no benchmark day of its own — the recommended-≥12 GB A/B rides on U-07/§1.4's region profile (run the route twice), the 120-target A/B on U-34. If displacement cannot be throttled to the noise floor, §2.7.4's contingency fires and the pre-named donor is internal resolution at that class [§2.7.4] |
| U-42 | Which local-model target the ship configuration is actually built against — and the un-applied ACTION that follows from it | A DATED CANON TENSION, routed rather than resolved by this doctrine. Three statements: RUNTIME_GEN §B proposes 3–4B and its own §6 still carries "[DECISION NEEDED - JOSH] Target local model + quant (3B vs 8B)"; Josh's 2026-07-04 directive block RULES 28–32B at a ~2030 window on next-gen NPUs with dev/test on the local 14B; T99_Translation_Combat §6.4 + bench_criteria.json v1.1.0 (the newest) pin 3–4B as the SHIP target and 14B as a dev-side reference. §2.7.3 budgets against the newest and shows the conclusion is invariant across all three. Two live items: (a) the 2026-07-04 block's own ACTION to re-derive §B.1/§B.3/§B.9/§B.10 is un-applied, so the tables §2.7.3 cites are stale on their face; (b) Josh's window assumes NPUs, while this doctrine's ladder is 2026-anchored — the GPU-carve question is properly the interim no-NPU case. Canon-lane item, not an engineering one |
| U-38 | The OS / desktop-compositor / driver VRAM reserve on the anchor class | UNMEASURED and load-bearing. It is the unresolved term in §2.7.3's carve arithmetic: 8.00 − 7.00 = 1.00 GB nominal remainder, minus this. It decides whether even canon's 0.5–1B draft tier (~0.6–0.9 GB) fits on an 8 GB card, and it moves the ≥12 GB threshold. Read it off the running build on benchmark day; do not assume it. No practitioner figure is adopted here — that would be a §8.3 U-27 violation |
| U-39 | Whether a RUNTIME (on the player's machine) TTS lane is ever built, and what it costs | DOES NOT EXIST IN CANON TODAY, and this doctrine does not create one. VO is bake-time and PARTIAL (ROADMAP_POST_5090_TO_SHIP.md D-VO-POLICY); bulk/ambient is subtitle; the audio stack's local models are production tools; UE runtime audio is MetaSounds DSP, not inference. §2.7.2 permits local TTS as a *tierable enrichment* — so if it is ever built it is a second model consumer, and §2.7.3 binds it: same carve, never a second one, barred wherever the language model is barred |
| U-40 | Whether the HIGH-END class has a VRAM working-set ceiling at all | UNSTATED BY THIS DOCTRINE. §1.2's 7.0 GB ceiling is a recommended-class number; §7.3 raises the high-end texture pool from ~2.5 to ~5.0 GB, so the high-end working set is materially above 7.0 — §2.7.3 uses 9.50 GB [DERIVED] from that delta and says so. This is what leaves the 8B prestige rung [DECISION AT BENCHMARK DAY]: at 9.50 GB the 3–4B baseline clears with ~3 GB of margin and the 8B (~5.9–6.9 GB) does not clear at the top of its range |
| U-41 | Whether a CPU-side small-model lane is viable at minimum / 8 GB-recommended without touching §2.6's game-thread budget | UNMEASURED — and therefore NOT AUTHORISED (§2.7.3). It is the only route to generated barks on a machine the VRAM arithmetic excludes, short of an NPU. The questions are core affinity, thread priority against UE's task graph on the ruled 6-core/12-thread part, and the measured game-thread delta against the 13.5 ms target. INCONCLUSIVE is the honest status; the min-spec class ships authored-floor meanwhile, which costs nothing because canon pre-authorises it. The pc_minimum tooth added at criteria v1.2.0 does not reach this lane, by construction: it keys on vram_peak_mb, so a CPU-side lane declaring 0 passes it. Binding a CPU lane needs its own record field before any tooth can hold it — recorded here so the tooth is not read as coverage it does not have [game-lane cross-repo flag, TODO_CANON.md "Cross-repo flags raised by GSW2 (2026-07-28)"; corroborated in Tools/bench/bench_criteria.json kt5 pc_minimum._note, which states the same scope limit in the fixture itself] |
| U-43 | Whether any PC-class inference LATENCY bar is ruled at all | UNRULED, and deliberately left null rather than inherited. The pc_recommended / pc_high tiers added to bench_criteria.json at v1.2.0 (game systems wave 2, 2026-07-28) carry null for max_p99_total_ms and max_p99_ttft_ms instead of inheriting the console ship tier's 1200 ms / 400 ms — because those numbers were pre-registered against console shared memory and reading them as PC bars is the same over-claim §2.7.5's closing note warns against. So the PC tiers today enforce co-residency and VRAM, never latency. Closing this needs a ruled PC-class latency target measured on the §1.2 recommended anchor, not a number carried across from a different machine class. Minted here because the game lane cites this id and §8.2 is where an open uncertainty is registered [game-lane cross-repo flag, TODO_CANON.md "Cross-repo flags raised by GSW2 (2026-07-28)"] |
| U-44 | The realm-class budgets — what a generation_tier='realm' surface may spend above the region floor | UNMEASURED, reserved at the DR-2 freeze (reservation clause 6 — ASSET_DROPIN_CONTRACT.md §7 item 6). Realm surfaces inherit the §5.3/§5.4/§5.6 region budgets as their FLOOR until this closes; the realm CEILING — the reason the class exists — is a benchmark-day derivation on realm greybox content, never an assumed headroom. Blocks nothing pre-hardware; must close before the first realm art pass |
|---|
| # | Claim | Status |
|---|---|---|
| U-17 | "UE 5.8 requires Windows 11 24H2 + DirectX 12 Ultimate + Shader Model 6.7" | CONTRADICTED by Epic's own hardware-and-software-specifications page (Win10 22H2 minimum; SM6 + DX12 Agility SDK; RTX 2000 / RX 6000 / Arc A minimum). Do not adopt. Track it — if true it would eliminate a meaningful slice of our min-spec audience [W4 §10] |
| U-18 | "DLSS 4.5 = ~5× the compute of DLSS 4.0's transformer" | UNVERIFIED — secondary only; NVIDIA says only "more computationally intense" [W1 §8.3] |
| U-19 | DLSS 5 support on RTX 40 (partial) and non-support on RTX 30 | UNVERIFIED — secondary explainers only; NVIDIA primary says "GeForce RTX" and "a single RTX 50-series GPU at launch" [W4 §10] |
| U-20 | "First games with full NTC support by end of 2026" | UNVERIFIED — secondary only. Zero shipped games use NTC today [W4 §10] |
| U-21 | UE6 "Foliage Nanite 10×" / "Dynamic Nanite out of the box" | UNVERIFIED as numbers — secondary coverage of State of Unreal 2026, not traced to an Epic release note [W4 §10] |
| U-22 | All GTA 6 technique claims, and the Take-Two/Rockstar patent-derived claims | [SPECULATION] until 2026-11-19; the patent rows are additionally UNVERIFIED (no outlet reachable quoted patent numbers or claim text). No Humanity decision rests on them [W2 §A.3, §F.1] |
| U-23 | All PS6 / next-Xbox specs and dates; RTX 60 "Rubin" H2 2027 | RUMOR. Sony's CEO confirmed 2026-05-08 that no date or price is decided [W4 §7, §10] |
| U-24 | STALKER 2's traversal hitches attributed to aggressive asset streaming | UNVERIFIED, community-forum tier — listed as a pattern, not a fact [W2 §C.4] |
| U-25 | Whether an NvRTX branch exists for UE 5.8 | UNVERIFIED — NvRTX was at 5.7.4 as of 2026-05-27. Check before assuming the NVIDIA branch is an option for our engine version [W4 §10] |
| U-26 | Whether Epic exposes Cooperative Vectors / a neural-shading authoring path in mainline UE | UNVERIFIED GAP — no evidence either way as of 2026-07-27. Re-check at 5.8 hotfixes, NvRTX 5.8, and UE6 EA [W4 §2, §10] |
| U-27 | Practitioner-blog numbers generally (Nanite <4 ms on a 3060; masked ≈2× cost; 20–300 ms RT PSO compiles; ~1M PSO permutations from 20 landscape layers; the 75% VRAM texture rule; ~30-emitter ceiling) | Directionally useful, individually unverified. Use them to size SPIKES, never to size BUDGETS [W2 §F.8] |
SELF-CHECK ON THIS TABLE — run it whenever §4 or §5 is edited (critique F-3; the same check is
stated in "How to read" so it binds from either end):
No item listed in U-27 may appear as a rule in §4 or §5.
Rev 1 failed this twice, and both are corrected above: the **"~75% of minimum-spec VRAM" texture
rule had become part of requirement R-29** (now replaced by §7.3's per-class pools), and the
"~30-emitter ceiling" had become §5.4 D2, a DON'T row with a named stat Niagara check —
which *is* a gate, built on a number this very table forbids building gates on (now demoted to a
spike-investigation trigger and tagged [UNVERIFIED — NOT A GATE]). The failure mode is worth
naming because it is quiet: an unverified number acquires authority simply by being written in the
imperative next to verified ones.
| # | Decision | Note |
|---|---|---|
| U-28 | A named recommended-spec machine in the lane | FINDING W3-2 / §7.4. The whole budget is proxy-anchored until it exists |
| U-29 | Mesh Terrain (Experimental 5.8) vs the researched Gaea + heightmap chain | Overlaps RESEARCHED_STACK.md Chain 2 step 3. Separate evaluation; out of scope for the slice [W1 §12 item 12, W4 §8] |
| U-30 | Whether the ARC-Raiders / AC-Shadows hybrid baked × RT GI fallback is ever invoked | Only if §2.2 row 5 will not close on the anchor hardware, by director ruling with capture evidence (§3.3). It is the one authorized exception to the no-baked-lighting ban |
| U-31 | UE5 FSR plugin update from 4.0.3 to the 4.1.1 SDK | PENDING — AMD says "at a later date." Plan for either the plugin catching up or a direct SDK integration [W1 §12 item 9] |
| U-32 | Whether incremental GC becomes thread-safe in a 5.9 / UE6 | Re-check trigger. If Epic makes it thread-safe, one whole class of architectural constraint (R-23) relaxes [W2 §F.7] |
| U-36 | Does the "unrestricted-components-by-default" ruling extend from the shipped-ASSET path to runtime REDISTRIBUTABLES shipped in the binary? (DLSS/Streamline, FSR, XeSS, Reflex, any NvRTX branch) | DIRECTOR'S CALL — routed, not assumed. Josh's 2026-07-15 ruling is written against the asset path (docs/translation/T99_Translation_AssetGen.md; enforced at docs/BUILD_PLAN_END_TO_END.md:478) and this doctrine governs the binary path. §3.4 deliberately makes the answer cheap either way by keeping the all-vendor-stripped build permanently shippable and on-budget — but the scope question is genuinely Josh-shaped and neither answer is assumed here [critique F-10] |
scored evaluation against NTBC.
RT viable at 60 fps and would change the R-40 instance budgets.
→ re-evaluate the engine-line ruling.
and it carries the entire minimum-spec GI rung. Promotion to Production retires the §2.4 retreat
clause; a regression or withdrawal invokes it (min-spec falls back to Lumen High at reduced
quality and lower internal resolution). Either way the 30 fps GI row is re-derived, and this is
the trigger rev 1 omitted entirely [critique F-8].
or the conventional World Partition actor-streaming fallback becomes the plan (§3.1, R-05, U-15).
AI-1's frame-critical ban is unaffected (it is not a performance preference), but the lane is
forced onto a class §2.2 budgets and §2.7.4's contingency fires: a ms row is added and the
pre-named donor is internal resolution at that class, not a content row and not the reserve.
path in UE** → §2.7.3's NPU lane moves from a bonus to a stated tier. It is the only path that
puts the baseline model on an 8 GB recommended machine at zero VRAM and zero graphics ms, so
it would materially change the policy table. Present breadth is [UNVERIFIED] — the Steam
survey does not report NPUs.
bench_criteria.json is relocked with PC spec-class tiers for KT-5 (§2.7.5 fixtures 2–3) →AI-3 moves from arithmetic to measurement, and the "a KT-5 PASS is not a PC-class clearance"
caveat retires.
evidence: a frame-time failure revises the recommended-spec anchor and re-instantiates §2's table
against the new anchor; a VRAM failure revises the 7.0 GB ceiling (§1.4b). Rev 1's trigger
covered the frame only, which is how the ceiling escaped falsification for a whole draft.
---
Nothing below is applied by this draft; the canon repo was read-only for this pass and a gate chain
was running. Ranked by "cannot be retrofitted" first.
1. Tools/bench/bench_criteria.json — relock the KT-3 spec classes to the §6.1 mapping via
run_all_benchmarks.py --relock with a dated, attributed reason. **Highest priority: today the
high-end bar cannot fail.**
2. Config/DefaultEngine.ini + a new Config/DefaultScalability.ini — the §7.3 three-level
architecture plus the §4.8 four wrong-by-default settings. This is the build-it-right deliverable
in its most literal form, and it is the cheapest item on this list.
3. Tools/bench/kt3_wp_pcg_runtime.py + UHumanityBenchSamplerSubsystem — the §6.2 record
schema (coordinated with the Humanity.Bench.SamplerRecordContract test).
4. T1_Build_Pipeline_Contracts §1.7 — the enum stays; the frame-rate targets per class are
new canon from Josh's 2026-07-27 ruling and are written in beside the existing budget table, with
the illustrative-until-benchmark framing preserved.
5. docs/FACTORY_CONTRACT.md — §5's rows become consume rows; the widened first wave
(A1, A3, A4, A5, B1, B2, B4, C1, C2, C3, E4, E6, E7, F1, F6, F7, F8, F9) lands first as
deterministic zero-token gate teeth. Rev 1 listed only A3/A4/B1/B2/B4/C3; A5, C1 and C2 are
equally deterministic, A1's 2:1 vertex:tri threshold was hiding in the grounding column, and
§5.5/§5.6 had no checks at all until rev 2. **Rows marked CRITIC-TIER (A2, A6, B6-until-baselined,
D2, D10, E2, E3, E5, F3, F5, and F4's second half) do NOT land as gates** — they land in the
review rubric, and saying so is the point.
6. A hard-cross-cell-reference scan added to the region-gate belt (R-14) — deterministic, and it
catches a defect class that frame captures misattribute.
7. This document → docs/pipeline_review/tech_research/UE58_PERF_DOCTRINE.md in house style,
with a docs/DOC_MAP.md row. Nothing is done until it is in DOC_MAP.
8. Tools/bench/bench_criteria.json — a SECOND relock, for KT-5 (§2.7.5): add the
pc_minimum tier at max_vram_mb: 0 (fixture 2 — zero evaluator change, the existing
vram > max_vram_mb check does the work), then pc_recommended / pc_high with a
game_working_set_mb co-residency check (fixture 3, small evaluator change). Same
--relock discipline as item 1: dated and attributed, never a silent edit.
9. Tools/bench/kt5_runtime_inference.py + its fixtures — the AI-2 decision-parity tooth
(§2.7.5 fixture 1): a decision_trace field and kt5/kill_decision_divergence.json, on the same
"outranks perf" footing as the escaped-line tooth. **A director ruling with zero enforcement
today is the highest-priority gap this amendment opens.**
10. UHumanityBenchSamplerSubsystem + the §6.2 record — three fields for §2.7:
inference_enabled (bool, the AI-4b A/B), inference_scheduling_class and
inference_queue_priority (AI-1). Coordinated with the Humanity.Bench.SamplerRecordContract
test, exactly as item 3.
---
1. **60 fps means: 1080p internal, upscaled, frame generation OFF, on an RTX 5060 / 4060 Ti /
RX 7600 XT at High, under 7.0 GB VRAM — HELD, not averaged.**
2. **The frame is a 16.67 ms table with named rows, measurement handles, and ≥10% reserved
headroom. Lumen's row is 4.0 ms — Epic's own number. A feature with no budget row does not get
built.**
3. **Design GPU-bound on purpose; the game thread is capped at 80% and personas/crowds/raids live
in Mass fragments, never in per-actor Blueprints — and MODEL INFERENCE NEVER TOUCHES THE FRAME'S
CRITICAL PATH at any spec class, carries a VRAM carve rather than a ms row, and never changes a
gameplay DECISION between two machines (§2.7).**
4. **The lighting architecture is MegaLights (local, RT shadows) + VSM clipmaps (sun, quantized
movement) + contact shadows (foliage) — which sets the minimum spec at RT-capable and rewrites
the art bible to "unlimited lights, low per-pixel overlap."**
5. **Substrate stays on Blendable GBuffer. Lumen tiers are Lite / High / Epic, Epic's own ladder.
No baked lighting — it is the one decision that permanently closes every GI future.**
6. **Every WPO material carries a WPO Disable Distance; foliage wind goes through Nanite Skinning;
motion vectors are correct on everything that moves, from day one.**
7. **Stutter is architecture, not tuning: PSO precaching + bundled cache + -clearPSODriverCache
on every test + a CI gate that fails on new uncached PSOs — and the QA loop's playthroughs are
the PSO/SODB trace collector, which is our unfair advantage.**
8. **Every perf and image record declares its feature state — upscaler, frame gen, GI rung,
scalability level, texture pool, AND reflex_mode + measured latency; a record missing the
declaration is INCONCLUSIVE, not a pass. The 5090 produces throughput, never verdicts.**
9. **Neural rendering, NTC, DXR 2.0 and UE6 are compatibility surfaces, not dependencies, and every
vendor SDK (DLSS/Streamline, FSR, XeSS, Reflex) is a removable plugin behind a build flag: the
all-stripped build must be shippable, on-budget and correct-looking at every class — one pass
that is simultaneously the art-direction guard, the licensing guard and the R-49 latency
baseline. Parry windows are authored on that path, never on Reflex's.**
10. **The recommended-spec anchor AND the 7.0 GB ceiling are one hypothesis with one scheduled
falsification test — build one region, measure it on the anchor card on both axes, and revise
once with evidence rather than assume for three years. Where our own evidence rules were
applied loosely to non-Epic numbers, rev 2 withdrew the numbers rather than the rules.**