decisions/HARDWARE_DECISION.md
Josh's question: is a ~$30k RTX-6000-class workstation (or larger) needed for AAA quality,
versus the planned ~$15k 5090 build? Constraints he set: ~90M-word game; 2-4 Claude Max x20
subscriptions do the heavy lifting; everything else local, free of API costs and restrictions.
What the machine actually runs, and what each leg needs:
32GB 5090 is literally the first consumer card that clears it; TRELLIS.2/Hi3DGen run lighter.
Nothing in the shipped-asset path needs more than 32GB today. The only >32GB tools found
(Lyra 2.0 at 43GB; Cosmos-3 Nano unquantized BF16) are PREVIZ-ONLY lanes anyway — Lyra's
weights are non-commercial by license and Cosmos outputs video-class media, so neither sits
in the shipped path regardless of VRAM. Cosmos-3 Super (64B) is datacenter-class — no
workstation at any price runs it; that tier is cloud if ever wanted.
for bulk mechanical work (triage, embeddings, OCR assist, the 14B-class runtime REFERENCE).
On 32GB, 14B-32B quantized models run comfortably. The thing a 96GB card buys — local
70B+-class reasoning — is precisely the job Josh has assigned to the Max subscriptions. The
ship-side runtime target is 3-8B on PLAYER hardware; a bigger dev GPU does not change it.
Claude Max (the subscriptions ARE the heavy lift, by design). Local GPU touches text only for
embeddings (fastembed — runs on CPU today) and triage-class small models.
regions, cooks, lighting builds) are CPU cores + RAM + NVMe — the 9950X3D / 128GB / Gen5
spec is strong there, and a 5090 already exceeds Epic's heavy-dev GPU expectations (Lumen/
Nanite/virtual-texture work included). MetaHuman generation round-trips Epic's CLOUD per
character — no local GPU changes that throughput.
A · The planned ~$15k 5090 build (status quo) — clears every load-bearing tool in the
researched pipeline with headroom; matches the subscription-heavy architecture; the pipeline
was deliberately designed to it.
Strongest objection: model VRAM appetites grow — a must-have 2027 tool could want >32GB.
B · ~$30k with an RTX PRO 6000-class card (96GB) — buys: unquantized previz models, bigger
batch asset-gen, local 70B+ LLMs, insurance against future VRAM growth.
Against: none of the shipped-path tools need it; the 70B+ capability duplicates what the Max
subscriptions already do better; previz lanes are optional by definition; ~$15k of insurance
against a hypothetical, purchasable later at 2027 prices with 2027 options if the hypothesis
lands.
C · Larger still (multi-GPU / DGX-class) — only rational for local foundation-model
training/fine-tuning or Cosmos-Super-class inference. The project's architecture (deterministic-
first, canon-graph-guardrailed, small on-device runtime, Claude for judgment) explicitly does
not want that; most of the gen tools are single-GPU pipelines (a second card buys parallel
throughput of a bursty workload, not capability).
Option A — keep the ~$15k 5090 build — RATIFIED + BUILD FACTS CONFIRMED (Josh, 2026-07-15):
the system is a Puget Systems build with a 1700W PSU (exceeds the 1600W insurance spec),
thermal-tested and guaranteed for sustained heavy-AI workloads. That covers the two halves of
the decision cleanly: the 32GB capacity clears the workload math (the 29GB peak), and the Puget
validation is what makes the tight 29-of-32GB fit deliverable SUSTAINED (hours-to-days
asset-gen/terrain/index runs are exactly the throttling/transient-spike regime where unvalidated
builds silently lose 20-30% throughput). Slot 2 stays free — a future 96GB card or second 5090
is a drop-in on the same PSU if — and only if — a named trigger fires. Re-validate at the
already-scheduled 5090 benchmark day (the asset-gen gate) — if any winning model is VRAM-bound
at 32GB in practice, that is the evidence that flips this decision, at the moment it flips.
Upgrade triggers (the named conditions that would justify the 6000-class buy later):
1. The benchmark day shows the winning commercially-clean asset pipeline genuinely VRAM-bound.
2. A must-have SHIPPED-PATH tool (not previz) ships with a >32GB floor.
3. The runtime-layer work needs local fine-tuning at scale (currently out of architecture).
Where the saved ~$15k actually buys AAA quality (the honest answer to "what's needed for
AAA"): the 3rd/4th Max x20 subscription (the words + orchestration are the game); hero-voice
budget (ElevenLabs) + licensed SFX libraries; the one sanctioned paid gap-filler (Tripo
rigging); the legal reads (Epic EULA gen-AI clause + capture-rights); contingency for the
PBR-texture engineering spike; permissioned photogrammetry/reference captures for real cultural
sites. AAA quality in this project comes from the canon graph, the pipeline discipline, and
iteration count — not from VRAM the researched tools never ask for.
Puget's "Accelerators" line offers a second independent card (up to the RTX PRO 6000 Blackwell
Max-Q 96GB, +$13,010). The rulings:
pool + compute — jobs pin per-card per-process (every chain tool is single-GPU); the only
physical interaction is x8/x8 lane bifurcation (x8 Gen5 = x16 Gen4 — invisible for AI load,
~1-3% worst-case in rendering). The Max-Q (300W) is the multi-GPU-thermals variant; 1700W
covers 575+300 with room. VRAM never pools across cards.
DOWN and would hurt the 9950X3D's CPU-bound work. A true 256GB+ need would mean a
Threadripper-class platform — nothing researched asks for it.
9950X3D's integrated GPU carries Windows/DWM/browser VRAM), leaving the 5090 headless with
~its full 32GB; stages can also run sequentially to cut the peak. An OOM on a real run = the
literal trigger #1.
pipeline requirement; the fit is our arithmetic + community precedent; Puget guarantees the
SYSTEM under sustained load, not a model's fit; OUR benchmark day is the test — and the
ruled default (TRELLIS.2 geometry) runs lighter than Hunyuan regardless.