HARDWARE_DECISION.md

decisions/HARDWARE_DECISION.md

Hardware decision — the 5090 build vs a 6000-class workstation [2026-07-15]

Josh's question: is a ~$30k RTX-6000-class workstation (or larger) needed for AAA quality,

versus the planned ~$15k 5090 build? Constraints he set: ~90M-word game; 2-4 Claude Max x20

subscriptions do the heavy lifting; everything else local, free of API costs and restrictions.

The workload analysis (grounded in the nine tech_research briefs — per-tool VRAM cited there)

What the machine actually runs, and what each leg needs:

32GB 5090 is literally the first consumer card that clears it; TRELLIS.2/Hi3DGen run lighter.

Nothing in the shipped-asset path needs more than 32GB today. The only >32GB tools found

(Lyra 2.0 at 43GB; Cosmos-3 Nano unquantized BF16) are PREVIZ-ONLY lanes anyway — Lyra's

weights are non-commercial by license and Cosmos outputs video-class media, so neither sits

in the shipped path regardless of VRAM. Cosmos-3 Super (64B) is datacenter-class — no

workstation at any price runs it; that tier is cloud if ever wanted.

for bulk mechanical work (triage, embeddings, OCR assist, the 14B-class runtime REFERENCE).

On 32GB, 14B-32B quantized models run comfortably. The thing a 96GB card buys — local

70B+-class reasoning — is precisely the job Josh has assigned to the Max subscriptions. The

ship-side runtime target is 3-8B on PLAYER hardware; a bigger dev GPU does not change it.

Claude Max (the subscriptions ARE the heavy lift, by design). Local GPU touches text only for

embeddings (fastembed — runs on CPU today) and triage-class small models.

regions, cooks, lighting builds) are CPU cores + RAM + NVMe — the 9950X3D / 128GB / Gen5

spec is strong there, and a 5090 already exceeds Epic's heavy-dev GPU expectations (Lumen/

Nanite/virtual-texture work included). MetaHuman generation round-trips Epic's CLOUD per

character — no local GPU changes that throughput.

The options

A · The planned ~$15k 5090 build (status quo) — clears every load-bearing tool in the

researched pipeline with headroom; matches the subscription-heavy architecture; the pipeline

was deliberately designed to it.

Strongest objection: model VRAM appetites grow — a must-have 2027 tool could want >32GB.

B · ~$30k with an RTX PRO 6000-class card (96GB) — buys: unquantized previz models, bigger

batch asset-gen, local 70B+ LLMs, insurance against future VRAM growth.

Against: none of the shipped-path tools need it; the 70B+ capability duplicates what the Max

subscriptions already do better; previz lanes are optional by definition; ~$15k of insurance

against a hypothetical, purchasable later at 2027 prices with 2027 options if the hypothesis

lands.

C · Larger still (multi-GPU / DGX-class) — only rational for local foundation-model

training/fine-tuning or Cosmos-Super-class inference. The project's architecture (deterministic-

first, canon-graph-guardrailed, small on-device runtime, Claude for judgment) explicitly does

not want that; most of the gen tools are single-GPU pipelines (a second card buys parallel

throughput of a bursty workload, not capability).

RECOMMENDATION (applied)

Option A — keep the ~$15k 5090 build — RATIFIED + BUILD FACTS CONFIRMED (Josh, 2026-07-15):

the system is a Puget Systems build with a 1700W PSU (exceeds the 1600W insurance spec),

thermal-tested and guaranteed for sustained heavy-AI workloads. That covers the two halves of

the decision cleanly: the 32GB capacity clears the workload math (the 29GB peak), and the Puget

validation is what makes the tight 29-of-32GB fit deliverable SUSTAINED (hours-to-days

asset-gen/terrain/index runs are exactly the throttling/transient-spike regime where unvalidated

builds silently lose 20-30% throughput). Slot 2 stays free — a future 96GB card or second 5090

is a drop-in on the same PSU if — and only if — a named trigger fires. Re-validate at the

already-scheduled 5090 benchmark day (the asset-gen gate) — if any winning model is VRAM-bound

at 32GB in practice, that is the evidence that flips this decision, at the moment it flips.

Upgrade triggers (the named conditions that would justify the 6000-class buy later):

1. The benchmark day shows the winning commercially-clean asset pipeline genuinely VRAM-bound.

2. A must-have SHIPPED-PATH tool (not previz) ships with a >32GB floor.

3. The runtime-layer work needs local fine-tuning at scale (currently out of architecture).

Where the saved ~$15k actually buys AAA quality (the honest answer to "what's needed for

AAA"): the 3rd/4th Max x20 subscription (the words + orchestration are the game); hero-voice

budget (ElevenLabs) + licensed SFX libraries; the one sanctioned paid gap-filler (Tripo

rigging); the legal reads (Epic EULA gen-AI clause + capture-rights); contingency for the

PBR-texture engineering spike; permissioned photogrammetry/reference captures for real cultural

sites. AAA quality in this project comes from the canon graph, the pipeline discipline, and

iteration count — not from VRAM the researched tools never ask for.

Addendum — the accelerator-slot Q&A (Josh, 2026-07-15; the Puget configurator open)

Puget's "Accelerators" line offers a second independent card (up to the RTX PRO 6000 Blackwell

Max-Q 96GB, +$13,010). The rulings:

pool + compute — jobs pin per-card per-process (every chain tool is single-GPU); the only

physical interaction is x8/x8 lane bifurcation (x8 Gen5 = x16 Gen4 — invisible for AI load,

~1-3% worst-case in rendering). The Max-Q (300W) is the multi-GPU-thermals variant; 1700W

covers 575+300 with room. VRAM never pools across cards.

DOWN and would hurt the 9950X3D's CPU-bound work. A true 256GB+ need would mean a

Threadripper-class platform — nothing researched asks for it.

9950X3D's integrated GPU carries Windows/DWM/browser VRAM), leaving the 5090 headless with

~its full 32GB; stages can also run sequentially to cut the peak. An OOM on a real run = the

literal trigger #1.

pipeline requirement; the fit is our arithmetic + community precedent; Puget guarantees the

SYSTEM under sustained load, not a model's fit; OUR benchmark day is the test — and the

ruled default (TRELLIS.2 geometry) runs lighter than Hunyuan regardless.

Generated by harness/site/structure_site.py — the URL path is the repo path. review root