pipelines/Q3_2026_MODELS_REFRESH.md
Status: RESEARCH BRIEF — a refresh + gap-fill pass over the tech_research corpus, not canon, not a
build order. Research date 2026-07-15 (same calendar day as LOCAL_3D_ASSET_GEN.md and
NEOSTACK_AI.md — this is a same-day deepen-and-widen pass, not a months-later refresh; where it
finds nothing changed, it says so explicitly rather than padding). Training-data cutoff for the
assistant that wrote this is ~January 2026; every load-bearing claim below was checked against the
live web on 2026-07-15 and is tagged VERIFIED (fetched from a primary source — official repo/
license file/model card/vendor pricing-terms page — and quotable) or INFERRED (secondary/
aggregator source, community report, or synthesis — not independently cross-checked against a
primary document). Where sources conflicted, the conflict is stated rather than silently resolved.
Two lanes (§3 local LLMs, §4 voice/audio) were researched by parallel sub-agents under this same
tagging discipline, then spot-checked directly against primary sources by the lead researcher
(Llama 4/Gemma 4/Phi-4/DeepSeek license claims and the NVIDIA ACE Game Agent SDK were independently
re-fetched and confirmed) — treat their VERIFIED tags with the same weight as the rest of this doc.
Scope: five deep-dive targets set by the mission brief — (1) refresh LOCAL_3D_ASSET_GEN.md against
anything newer, (2) identify what NeoStack's own bundled cloud generation actually runs, (3) the
current local-LLM landscape for the 5090 (runtime layer + build-agent tiers), (4) a light pass on
voice/music/SFX vendors, (5) credible 2026 H2 announcements. A dedicated §DELTAS section closes the
loop against every prior conclusion in LOCAL_3D_ASSET_GEN.md, NEOSTACK_AI.md,
RUNTIME_GENERATIVE_LAYER.md, RUNTIME_GUARDRAIL_ARCHITECTURE.md, STACK_FACTS_QUESTIONS.md, and
RESEARCHED_STACK.md — confirming or superseding each.
---
**Nothing in the existing tech_research corpus was WRONG when written earlier today; several things
are already dated, and one finding (§3.5) is big enough to change how the runtime-generative-layer
design doc should be evaluated.**
1. 3D asset gen (§1): the core §0 verdict of LOCAL_3D_ASSET_GEN.md — split TRELLIS.2
(geometry) / Hunyuan3D-2.1 (finished PBR, territory-capped) / Hi3DGen (sharp-geometry fallback) —
is RECONFIRMED, not superseded. TRELLIS.2's nvdiffrast/nvdiffrec commercial trap is still
unresolved (issue #22, no maintainer response). Hunyuan3D 2.5/3.x are still not open-sourced. The
"commercially-clean-PBR gap" the sibling brief flagged is STILL OPEN — two new SIGGRAPH-2026
entrants (Pixal3D, Ubisoft CHORD) both looked at first glance like they might close it and **both
turn out not to**: Pixal3D's own code is clean MIT but its install chain runs through TRELLIS.2
(same nvdiffrast dependency, inherited not escaped); CHORD ships real open weights but under a
Research-Only Copyleft license. New, real, and worth benchmarking: Pixal3D (near-
reconstruction-fidelity pixel-aligned texturing) and Direct3D-S2 (the sparse-attention
geometry model Pixal3D is built on) both belong in the §5.2 benchmark set alongside the existing
three. SIGGRAPH 2026 is NOT "next month" — it is 2026-07-19 to 23 in Los Angeles, i.e., this
week — corrected below.
2. NeoStack (§2): the Discord-reported "3D asset generation" and "voice generation" bundled into
the lifetime tier are identified concretely for the first time: 3D gen is a cloud passthrough
to Tripo and Meshy (both closed SaaS, both already known to this project as reference-tier,
not local-pipeline candidates); voice gen is a cloud passthrough to ElevenLabs — the exact
same vendor already locked into this project's own audio stack. **Neither duplicates the local
TRELLIS-class pipeline** — they're a convenience UI over already-evaluated closed vendors, not a
new model choice. The "4 new products / own IDE" Discord datapoint most plausibly resolves to
Betide Studio's wider 9-product integration-kit catalog (Steam/EOS/Edgegap/PlayFab/Crossplay/
GameCenter/PlayServices/Matchmaking + Agent Integration Kit) rather than a genuinely new product
category — no public evidence of a standalone "NeoStack IDE" was found; the IDE reference most
likely means the already-known NeoStack Cloud hosted-IDE feature. A genuine new flag: the
Discord quote Josh has on record naming UE 5.8 support verbatim names "SIK" and "EIK"
(Steam/EOS Integration Kit) — not "AIK" (Agent Integration Kit = NeoStack). Every public source
checked (official docs, changelog, Fab listings) still shows NeoStack itself capped at UE 5.5-5.7.
This doesn't overturn the "keep NeoStack" ruling — it sharpens exactly why the standing "verify at
first use" pre-flight check matters.
3. Local LLMs (§3): the runtime-layer doc's named candidates (Qwen3-3B, Phi-4-mini, Qwen3-8B,
Llama-3.2-1B) are still real and still fine choices, but no longer the sharpest ones available,
and one licensing fact has flipped in the project's favor: **Gemma moved to Apache 2.0 with Gemma
4 (2026-03-31)** — the historical Gemma restricted-license problem is gone for anyone using
Gemma 4+. Blackwell ships a genuinely better-than-Q4_K_M quantization path (NVFP4/MXFP4). Most
importantly: **NVIDIA shipped an open-source (Apache 2.0), on-device, UE5-plugin-integrated
"ACE Game Agent SDK"** (v0.5.0, 2026-06-15) that is architecturally close to what
RUNTIME_GENERATIVE_LAYER.md designed independently — small local SLMs (Qwen3.5-4B / Nemotron-3-
Nano-4B), on-device TTS/ASR, an agent/chat/RAG API split, already adopted in three shipped/beta
titles (PUBG Battlegrounds, Total War Pharaoh, Mecha BREAK). This deserves a dedicated evaluation
pass, not just a footnote — see §3.5.
4. Voice/audio (§4): the named stack (ElevenLabs, AIVA, Stable Audio Open, licensed libraries)
all still stand — no vendor needs replacing. Two real updates: Stable Audio Open 1.0 has a
successor (Stable Audio 3.0, May 2026, adds an open-weight dedicated SFX model) and the open-
TTS field now has genuinely commercial-clean local options (Kokoro, Apache 2.0, no cloning —
and, notably, Chatterbox — MIT — is the same TTS model NVIDIA's ACE SDK bundles) worth
evaluating for bulk ambient NPC lines to cut ElevenLabs per-call spend. Suno/Udio remain usable but
carry residual legal risk (no indemnification, US fair-use ruling still pending, likely into 2027)
that AIVA's full-copyright-assignment license does not.
5. Upcoming (§5): SIGGRAPH 2026 is this week, not next month, and is already producing usable
artifacts (Pixal3D, SimArt, SATO all ship with GitHub repos, though SATO's code is "being prepared
for public release" — not yet usable). The single most consequential near-term item for this
project's own architecture is the NVIDIA ACE Game Agent SDK (§3.5) — it should be evaluated before
the runtime-generative-layer build starts, not discovered after.
---
LOCAL_3D_ASSET_GEN.mdGitHub issue #22 (the nvdiffrast non-commercial-license question) — VERIFIED, re-fetched directly —
remains open, opened by a user on 2025-12-18, no maintainer response of any kind as of today.
Microsoft has not patched the dependency, has not responded to the issue, and no newer TRELLIS
version exists to check instead (VERIFIED — github.com/microsoft/TRELLIS.2/issues/22).
Tencent-Hunyuan GitHub org listing: it shows only Hunyuan3D-2.1 (3.7k stars) as a public 3D-mesh repo — no
2.5, no 3.0, no 3.1 repo exists (VERIFIED — github.com/Tencent-Hunyuan).
This directly reconfirms LOCAL_3D_ASSET_GEN.md §1.2's finding; **"Hunyuan3D-2.1 is the version to
actually use, nothing past it is open" still holds exactly as written.**
March 2025 paper/repo the sibling brief already catalogued (VERIFIED — repo/release pages checked,
no new tags).
Three genuinely new items surfaced, all tagged [SIGGRAPH 2026] on their own repos, none of which
were findable when the sibling brief was researched (they are today's/this-week's news given the
SIGGRAPH date correction in §1.3):
**Pixal3D (TencentARC) — MIT on its own code, but inherits TRELLIS.2's nvdiffrast dependency; do NOT
treat as a clean escape from the PBR-licensing trap.**
Pixal3D: Pixel-Aligned 3D Generation from Images." Weights on HuggingFace
(huggingface.co/TencentARC/Pixal3D). VERIFIED.
correspondence) rather than loosely injecting image features via attention — claimed near-
reconstruction-level fidelity with detailed geometry and PBR textures (VERIFIED, repo/paper
description, arxiv.org/html/2605.10922v1).
LICENSE file is a standard,unmodified MIT license (VERIFIED —
github.com/TencentARC/Pixal3D/blob/master/LICENSE).
Its NOTICE file lists only Apache-2.0 (dinov2) and MIT (TRELLIS.2, Direct3D-S2, MoGe) dependencies
— no NVIDIA non-commercial component is disclosed there (VERIFIED, direct fetch). But the
actual install instructions are: "Step 1: Follow TRELLIS.2 Installation" — i.e., Pixal3D
requires installing TRELLIS.2 first as a hard prerequisite (VERIFIED, direct fetch of the raw
README). A separate web search independently found that Pixal3D's own GPU-extension-wheel
requirements include nvdiffrast and nvdiffrec_render (INFERRED — search-aggregated, not
independently re-confirmed against a Pixal3D-specific install script beyond the TRELLIS.2-delegation
step), which is fully consistent with the TRELLIS.2-first install chain. **Conclusion: Pixal3D's own
code is clean MIT, but its practical dependency chain very likely reruns TRELLIS.2's PBR-bake path
(and therefore nvdiffrast/nvdiffrec) to produce its textured output — the same commercial-use
caveat flagged for TRELLIS.2 in LOCAL_3D_ASSET_GEN.md §1.1 should be assumed to extend to
Pixal3D's textured output until directly re-verified at benchmark time, not treated as bypassed.**
This is a nuanced, partially-INFERRED finding — flag it exactly this way, don't round it to either
"clean" or "blocked."
recommended, low-VRAM mode targets 24GB cards" while a YouTube video is titled "Pixal3D on 6GB VRAM
— Better Than Trellis 2!" (both INFERRED, unreconciled — treat 6GB as an aggressive/cut-down claim
and 16-24GB as the more conservative real-world figure pending direct benchmarking). Generation time
claimed at 3-5 minutes per asset in one ComfyUI-integration repo description (INFERRED,
github.com/dreamrec/ComfyUI-Pixal3D).
entry: a sparse-volume 3D generation framework using Spatial Sparse Attention (SSA), enabling
1024³-resolution training on 8 GPUs where prior volumetric approaches needed 32+ (VERIFIED license
and claim — github.com/DreamTechAI/Direct3D-S2,
cross-confirmed on HuggingFace). Not independently benchmarked here; worth a bench-day look as a
possible TRELLIS.2 alternative for the geometry stage specifically, since it does NOT carry
TRELLIS.2's nvdiffrast dependency chain (it's Pixal3D that adds that, by choosing to build on
TRELLIS.2 rather than on Direct3D-S2 alone for its render step).
**Ubisoft La Forge CHORD — resolves the sibling brief's open question, but the answer is "still not
commercially usable."**
LOCAL_3D_ASSET_GEN.md §3.6 flagged as "debuted at SIGGRAPH Asia 2025... maturity/completeness of
the actual released code not independently verified" — that flag is now resolved: it shipped, with
real code, real weights, and a ComfyUI integration (VERIFIED — released 2025-12-09,
github.com/ubisoft/ubisoft-laforge-chord,
huggingface.co/Ubisoft/ubisoft-laforge-chord,
ComfyUI node at github.com/ubisoft/ComfyUI-Chord).
Normal, Height, Roughness, Metalness — via chained rendering-decomposition + single-step diffusion
(VERIFIED, repo/paper description).
— explicitly not permitted for commercial use (VERIFIED, repo license page). **This confirms,
rather than closes, the sibling brief's "commercially-clean-PBR gap" finding**: the field's most
credible new open PBR-material-estimation release in the last several months is real, working, and
explicitly non-commercial.
ByteDance-Seed SimArt — adjacent to, not a solve for, the game-readiness gap.
— Apache 2.0, real weights on HuggingFace (VERIFIED —
github.com/ByteDance-Seed/SimArt). Does part-level mesh
segmentation + kinematic/joint prediction + URDF (robotics-simulator format) generation via a
sparse 3D VQ-VAE (VERIFIED, repo description).
LOCAL_3D_ASSET_GEN.md §2.1named as one of the unsolved stages in the Hunyuan3D Studio research pipeline (which itself never
released code). SimArt is real, open, and does part decomposition + articulation prediction — but
its native output (URDF) targets robotics physics simulation, not UE skeletal-mesh rigging, and
the repo's own documentation frames coordinate-system alignment as a preprocessing concern rather
than addressing retopology or game-engine rig authoring directly (VERIFIED via direct fetch). **Net:
a genuinely promising building block to prototype against the game-readiness gap, not a drop-in
fix** — the §2 conclusion of LOCAL_3D_ASSET_GEN.md ("no mature open AI system automates
retopology + UVs + LODs + collision end-to-end today; budget Blender-headless scripting + human QC")
still stands as the operative default, with SimArt now flagged as worth a bench-day evaluation
specifically for the part-decomposition/rigging-prep sub-step.
artist-style mesh generation with native UV segmentation built into the same autoregressive
token stream (VERIFIED, arxiv.org/abs/2604.09132,
github.com/Xrvitd/SATO) — this is close to exactly the retopology-
plus-UV-unwrap gap. Not yet actionable: the repo's own status is "the codebase is being prepared
for public release" (VERIFIED, direct fetch) — a paper-with-a-repo-stub, same non-actionable status
the sibling brief gave Hunyuan3D Studio's PolyGen/SeamGPT. Worth tracking; re-check at benchmark-gate
time since a real release could land any time given SIGGRAPH is this week.
SIGGRAPH 2026 is 2026-07-19 to 2026-07-23 in Los Angeles (exhibition July 21-23) — VERIFIED,
directly fetched from the official conference site
(s2026.siggraph.org). **This is not "next month" relative to this
brief's research date (2026-07-15) — it is this week, essentially concurrent with this research
pass.** The mission brief's framing should be corrected going forward. Practical consequence: this
is not a "what's teased" forward-look question so much as a "what's already dropping" question — the
three items in §1.2 (Pixal3D, CHORD-adjacent lineage, SimArt) and SATO in §1.2 are evidence this wave
is already live, not upcoming. Technical Papers program stats: 1,120+ submissions this cycle
(VERIFIED, s2026.siggraph.org),
covering generative AI/ML for visual computing among its named tracks; the live conference-schedule
tool (s2026.conference-schedule.org) did not return a
populated session list at fetch time (client-side filtering, not yet browsable via a plain fetch) —
re-check that tool directly during/after the conference week for the full accepted-papers list rather
than relying on search-engine coverage, which is necessarily partial and skewed toward whatever
already has a public GitHub repo.
LOCAL_3D_ASSET_GEN.md §5.2's benchmark protocol (re-check licenses, run a fixed 2-per-class test set
through TRELLIS.2/Hunyuan3D-2.1/Hi3DGen/TripoSG/SF3D-SPAR3D) stands as written — add Pixal3D
(as a geometry+texture challenger to benchmark specifically for whether its output actually routes
through nvdiffrec at runtime — instrument the process, don't just trust the NOTICE file) and
Direct3D-S2 (as a possible TRELLIS.2-alternative geometry stage that doesn't carry the nvdiffrast
baggage) to the roster. Keep CHORD off the commercial-asset roster entirely (research-only license) —
it remains useful only as an internal-reference/previz tool, same treatment as the territory-capped
Hunyuan3D under the existing JOSH-RULED posture (commercially-unrestricted components are the
shipped-asset default; anything else is previz-only). Flag SimArt and SATO for a later, separate
evaluation specifically against the game-readiness/retopology gap once the workstation lands, rather
than folding them into the mesh-generation benchmark itself — they solve a different sub-problem.
---
This section answers the mission's Q1-addendum data point ("the lifetime purchase bundles weekly
usage incl. 4-5 agentic models, 3D ASSET GENERATION, and VOICE GENERATION") that NEOSTACK_AI.md
flagged but did not resolve to specific vendors.
Directly confirmed via NeoStack's own official changelog (VERIFIED, direct fetch of
A separate search independently corroborates a "Studio" tab inside the plugin that surfaces
"multiple AI generation providers including Meshy, Tripo, and ElevenLabs for 3D generation, rigging,
animation, and text-to-speech" (INFERRED — search-aggregated summary of the docs, not a direct fetch
of the Studio-tab page itself, which was not independently reached). **This is a router/aggregator UI
over closed third-party SaaS providers, not a NeoStack-proprietary model** — the same Meshy and Tripo
this project's own LOCAL_3D_ASSET_GEN.md §1.8 already evaluated and deliberately positioned as
"reference/quality bar, not a candidate" for the local, open-source pipeline (Meshy is the incumbent
the whole local-3D-gen research effort exists to replace; Tripo 3.0 is closed/API-only).
Same changelog, same version: **v1.0.49 (2026-03-24): "ElevenLabs (new) — text-to-speech and sound
effects."** VERIFIED, direct fetch. This is not a new or different voice vendor — it is the exact
same ElevenLabs already locked into this project's audio stack per STACK_FACTS_QUESTIONS.md Q4 and
reconfirmed current in §4.1 below, just reachable through NeoStack's in-editor chat UI (and billed
against NeoStack Cloud's usage allowance) rather than a direct API call.
Complement, not duplicate — and only partially even that. NeoStack's 3D-gen surface gives
in-editor, conversational access to two closed SaaS generators this project already decided NOT to
depend on for the shipped-asset pipeline (§17-ruled commercial-clean-by-default posture,
STACK_FACTS_QUESTIONS.md Q3 Josh ruling). It doesn't touch TRELLIS.2/Hunyuan3D-2.1/Hi3DGen at all —
different vendors entirely, no overlap in model, license, or output path. Where it's genuinely useful
despite that: Tripo's rigging/animation capability is a real gap-filler. LOCAL_3D_ASSET_GEN.md
§3.2 already flagged that none of the open local mesh-generators produce rigs, and separately named
"Tripo AI's 'universal rig'... explicitly advertised as working across humanoid-to-creature ranges" as
a worth-knowing paid point-solution for the 132-creature roster. NeoStack's Studio tab is simply a
more convenient way to reach that same Tripo capability from inside the editor, at whatever
credit-cost NeoStack Cloud's allowance provides — an operational convenience on an already-identified
tool, not a new research finding about the tool itself.
had visibility into (VERIFIED, direct fetch of
betide.studio/plugins): EOS Integration Kit, Steam Integration Kit,
Edgegap Integration Kit, PlayFab Integration Kit (deprecated), Ultimate Crossplay Integration Kit,
Game Center Integration Kit, Play Services Integration Kit, Matchmaking Integration Kit, and Agent
Integration Kit (NeoStack). **This most plausibly explains the "expanded product line, 4 new
products" Discord datapoint** — several of these (Edgegap/Crossplay/GameCenter/PlayServices/
Matchmaking) read as a newer wave layered onto the original EOS/Steam/AIK trio — rather than
indicating Betide is diversifying outside Unreal platform-integration tooling. **No standalone
"IDE" product was found anywhere in this catalog or in NeoStack's own public docs.** The most likely
resolution: the "own IDE" reference is the already-documented NeoStack Cloud hosted-IDE feature
(NEOSTACK_AI.md §6 — the $20/mo Cloud tier bundles "a hosted IDE, asset-aware version control,
CI/CD"), not a new fourth product. This is INFERRED — a reasonable reading of public evidence, not a
direct confirmation from Betide Studio; if Josh's Discord screenshot names something more specific,
that should override this inference.
(Claude Code, Gemini CLI, Codex, Cursor, GitHub Copilot = 5) from NEOSTACK_AI.md §3, not a new
finding — no marketing copy using that exact phrase was found (INFERRED).
changelog page, directly fetched twice, tops out at v1.0.50 (2026-03-24) with no entries between
then and today. One unrelated secondary aggregator's snippet claims "v2.0.45 as of June 6" exists
(INFERRED, uncorroborated by any other source, possibly confusing NeoStack with one of Betide's
other 8 products or simply wrong) — this conflict is presented, not resolved. Do not assume V2
has shipped for NeoStack specifically without checking aik.betide.studio directly at build time;
equally, do not assume it hasn't — the changelog fetch tool may be serving cached/incomplete content
for a fast-moving product page.
STACK_FACTS_QUESTIONS.md Q1 records Josh's own Discord screenshot, quoted verbatim: *"SIK currently
have 5.8 on FAB, EIK ... soon also will be 5.8. We focus updating marketplace version first."*
**Read literally, this quote names Steam Integration Kit (SIK) and EOS Integration Kit (EIK) — not
Agent Integration Kit (AIK = NeoStack).** Every public source reachable in this research pass
(official docs at aik.betide.studio, the changelog, Fab-listing search snippets, `betide.studio/
neostack`) still states NeoStack/AIK support caps at UE 5.5, 5.6, 5.7 — none show a 5.8 tag
(VERIFIED across four independent fetches today). This does not contradict Josh's screenshot — it's
entirely possible SIK/EIK reached 5.8 first and AIK follows "soon," exactly as the quote's own phrasing
suggests with "we focus updating marketplace version first." This sharpens, rather than reverses,
the existing "verify NeoStack-on-5.8 directly before the Phase 6 vertical slice" action item already
standing in NEOSTACK_AI.md §0 and RESEARCHED_STACK.md's ENGINEERING-SPIKES list — it is now a
clearly evidenced check, not a generic caution.
No change to the Q1 "keep NeoStack" ruling — nothing here is disqualifying. Three concrete refinements
for the correction pass: (1) when the Translation Doc specifies NeoStack's generative capabilities,
name Meshy/Tripo (3D) and ElevenLabs (voice) explicitly as the underlying vendors rather than
treating "NeoStack generation" as its own model class — this matters because it means NeoStack's
generative output inherits Meshy's and Tripo's own commercial terms, not a NeoStack-specific license;
(2) the UE-5.8-pending status is now evidenced specifically for AIK, not just generically flagged —
keep the "verify at first use" gate but don't treat Josh's Discord screenshot as having already
cleared it for NeoStack itself; (3) if Tripo's rigging capability is wanted for the 132-creature
roster, reaching it via NeoStack's existing Cloud allowance is a reasonable default rather than a
separate direct Tripo subscription — but confirm the credit-cost math once the "10x the free tier of
NeoStack Cloud" allowance (VERIFIED phrase, found via search, exact numeric limits not located in this
pass) is understood in practice.
---
Researched via a dedicated parallel sub-lane against the same VERIFIED/INFERRED discipline as the
rest of this brief, then spot-checked directly (Llama 4, Gemma 4, Phi-4, DeepSeek license claims and
the NVIDIA ACE Game Agent SDK were independently re-fetched by the lead researcher and confirmed
consistent). This directly targets RUNTIME_GENERATIVE_LAYER.md's named tiers (draft 0.5-1B, min-spec
3-4B, prestige 7-8B, 14B = the ruled reference/ship tier, 28-32B optional prestige) and
RUNTIME_GUARDRAIL_ARCHITECTURE.md's matching model+VRAM line.
| Family | Current open line | Sizes relevant to Josh's tiers | License |
|---|---|---|---|
| Meta Llama | Llama 4 (Scout/Maverick); Behemoth shelved, never shipped | No small dense model — Scout is ~109B MoE/17B-active; smallest usable local size is still the older Llama-3.2-1B/3B | Llama Community License — 700M-MAU clause, attribution required, EU vision carve-out (VERIFIED, license text quoted below) |
| Alibaba Qwen | Qwen3 (base) + Qwen3.5 (Feb 2026) + Qwen3.6 (Apr 2026) | Qwen3: 0.6B/1.7B/4B/8B/14B/32B dense + MoE variants; Qwen3.6 adds a 27B dense. Qwen3.7 (May-Jun 2026) is API-only, no open weights — the flagship is closing even as the small/dense line stays open | Apache 2.0 across the dense/open line (VERIFIED, LICENSE files on HF) |
| Google Gemma | Gemma 4 (2026-03-31) supersedes Gemma 3/3n | Gemma 4: E2B (~2.3B), E4B (~4.5B), 26B-A4B MoE, 31B dense, +12B "Unified" multimodal (Jun 2026) | Apache 2.0 — a genuine license flip from Gemma 1-3's restrictive custom terms (VERIFIED, quoted below) |
| Microsoft Phi | Phi-4 family, now 7 variants | Phi-4-mini (3.8B), Phi-4 (14B), Phi-4-reasoning/-plus (14B), Phi-4-mini-reasoning (3.8B), Phi-4-reasoning-vision (~15B, Mar 2026) | MIT across all variants (VERIFIED for Phi-4-mini's HF frontmatter) |
| Mistral | Ministral 3 (Dec 2025) + Mistral Small 4 (Mar 2026) + Magistral Small | Ministral 3: dense 3B/8B/14B (base+instruct+reasoning) | Apache 2.0 for Ministral 3 (VERIFIED HF frontmatter) — Mistral is "progressively moving away" from its older non-commercial Research License; check each model card, not uniform |
| DeepSeek | R1-Distill line; V4 Pro/Flash (Apr 2026, too large for local) | Distills: 1.5B/7B/14B/32B (Qwen-based) + 8B/70B (Llama-based) | Qwen-based distills inherit Apache 2.0; Llama-based distills inherit the Llama license (MAU clause) — split by base model, not uniform |
| NVIDIA Nemotron (not in the doc's original list — a genuine addition) | Nemotron 3 Nano (Dec 2025) — hybrid Mamba-Transformer, 1M context, 4B variant used in NVIDIA's own game SDK (§3.5) | 4B / 9B / 12B | NVIDIA Open Model License — permissive/commercial, derivatives allowed, but a bespoke license, not Apache/MIT — read before shipping |
| Out of local range, context only | GLM-5.2 (Zhipu, MIT, rated a top open-weight model on several boards), Kimi K2.7, MiniMax M3 | 40B+ active / 700B+ total | Large MoE — not a 32GB-card candidate regardless of license |
License texts, directly quoted (VERIFIED):
active users of the products or services made available by or for Licensee… is greater than 700
million monthly active users in the preceding calendar month, you must request a license from
Meta."* Plus a mandatory "Built with Llama" attribution and an EU-domiciled-licensee carve-out
excluding multimodal/vision use.
permissive Apache 2.0 license."* The HuggingFace checkpoint itself shows license: apache-2.0 with
no gated agreement — a genuine, structural change from Gemma 1-3's custom Gemma Terms of Use (which
carries a Prohibited Use Policy, a flow-down obligation to downstream users, and a unilateral
Google termination right, and which Gemma 3/3n still use).
Two different legal questions: (a) redistributing model weights to end users and running them
offline inside a shipped commercial game, vs (b) using a model purely as the developer's own
build-agent tooling, never redistributed. Nearly everything here clears (b) trivially; the prize is
what clears (a) cleanly.
Gemma 4 (Apache 2.0 — newly clean, was NOT true of Gemma 1-3), Phi-4/Phi-4-mini (MIT), Ministral
3 (Apache 2.0), DeepSeek's Qwen-based R1-distills (Apache 2.0 inherited). **This set alone covers
every tier RUNTIME_GENERATIVE_LAYER.md needs** (0.6B through 32B) with zero licensing exposure at
ship time.
plausibly fine for embedding given NVIDIA ships it inside their own game SDK, but confirm the
redistribution clause directly); Gemma 3/3n (the old custom Gemma Terms of Use — workable but
strictly worse than just using Gemma 4 instead).
DeepSeek distill (8B/70B) — a solo dev is nowhere near the 700M-MAU threshold, so shipping is legal,
but it requires "Built with Llama" attribution, redistributing license text with the weights,
Acceptable Use Policy adherence, and (if EU-domiciled) forgoing multimodal/vision use. **Given the
Apache/MIT set above covers every needed tier with none of this overhead, there is no longer a
reason to reach for Llama for the shipped-embedded tiers** — keep it as a dev-side/build-agent-only
option if used at all.
RUNTIME_GUARDRAIL_ARCHITECTURE.md's own estimate,fits comfortably in 32GB alongside a AAA renderer (INFERRED, consistent across benchmark sources).
settings routinely wants 12-20GB+ of its own VRAM — meaning the 28-32B "prestige" tier cannot
reliably co-reside with a real renderer on a single 32GB card today.** This is a genuine, concrete
refinement to the existing "14B = reference, 28-32B = optional prestige if 2030 hardware allows"
ruling: on THIS hardware (the 5090, now), 14B is the practical in-game co-resident ceiling; 32B
is realistically a dev-side/offline-evaluation tier or a 2030-hardware ship tier exactly as already
ruled, not something to expect working live next to the renderer on the 5090 itself. This
*reinforces* the existing ruling's own hedge ("NEVER a content dependency") rather than contradicting
it — it just makes the "why" concrete with real numbers.
standard MXFP4).** NVIDIA's 4-bit float format with two-level micro-scaling lands roughly 25%
smaller than Q4_K_M at within ~1% of FP8 accuracy on reasoning/knowledge benchmarks (INFERRED,
NVIDIA's own technical blog + secondary corroboration) — e.g. a 27B model at ~14GB instead of the
Q4_K_M-equivalent ~18-20GB. The catch: NVFP4 is Blackwell-exclusive (5090/Blackwell-generation
cards only), so it can only ever be a Blackwell-tier optimization layered on top of a portable
GGUF Q4_K_M baseline for players on older hardware — it doesn't change the shipped min-spec quant
choice, but it meaningfully improves what's achievable on the dev's own 5090 specifically (more
VRAM headroom for renderer co-residency at the 14B-and-above tiers).
with CUDA 12.8, not CUDA 13.1** — 13.1 is reported to segfault llama.cpp's Blackwell MMQ kernel,
silently falling back to slower cuBLAS (VERIFIED as a real, documented build issue). NVFP4/MXFP4
kernel support is actively merging for native Blackwell dispatch.
community workarounds; heavier/server-oriented, more than a single in-game agent needs.
consumer-Blackwell kernel support lagged datacenter Blackwell into early 2026 and needs custom
builds — a heavier toolchain than a solo dev needs versus llama.cpp.
5090 — the natural day-to-day dev-tooling choice.
successor to the old ChatRTX concept), with an NVIDIA-published (so vendor-sourced, treat with
appropriate skepticism) claim of ~7.3x faster than Ollama in one comparison.
llama.cpp + GGUF Q4_K_M reference-runtime pick — pin CUDA 12.8,and treat NVFP4/MXFP4 as an additive Blackwell-tier optimization once it's fully merged, not a
replacement recommendation.
This is the single most consequential finding in this lane and deserves prominent treatment rather
than a passing mention. Independently re-fetched and confirmed (not solely relying on the research
sub-lane):
released 2026-06-15, beta unveiled at Unreal Fest 2026 (2026-06-16) with NVIDIA researchers
presenting at SIGGRAPH 2026 (2026-07-20 — this week). Repo:
github.com/NVIDIA/game-agent-sdk — VERIFIED to contain
real C99-ABI/C++ code (85.1% C++/7.8% C), a working build system, three functional sample
executables, and ~3GB of bundled ML models — not a stub.
shipping-embedded bar cleanly, same tier as the best options in §3.2.
RUNTIME_GENERATIVE_LAYER.md independently designed: anAgent/Chat/RAG API split (stateful autonomous-reasoning agent vs. stateless direct-control chat vs.
semantic/lexical knowledge retrieval), fully on-device ("Unlike cloud-based services... local,
RTX-optimized workflows"), with reference models sized for exactly the small-local tier this project
already anchors on: Qwen3.5-4B (local GGUF, the decision-making SLM), Nemotron-3-Nano-4B
(alternative SLM), Chatterbox-Turbo-350M (on-device TTS — the same Chatterbox family the voice-
gen sub-lane independently flagged in §4.2, cross-confirming it as a real, current, commercially
viable option), and NeMo-Conformer-CTC-120M (on-device ASR). Targets an ~8GB VRAM budget,
explicitly runs down to an RTX 3060-class card, not just the 5090.
character dialogue, voice commands, and contextual action responses — distributed via
developer.nvidia.com/ace-for-games rather than inside
the core game-agent-sdk GitHub repo itself (the core repo is engine-agnostic C/C++; the UE5 plugin
layer sits on top of it — this distinction matters if evaluating the SDK, since the two are separate
downloads).
teammate, open beta), Total War: PHARAOH (experimental in-game advisor using RAG over "1,200+
interlinked game data tables" — a scale point directly relevant to how this project's own canon
graph/registries could feed an equivalent RAG layer), and Mecha BREAK (shipped).
RUNTIME_GENERATIVE_LAYER.md specifically: the doc's own thesis (§0,"deterministic sim + a bounded strategy-token LLM layer, validated, fallback-guarded, on-device,
small models") was designed bespoke, independent of this SDK. NVIDIA has now shipped a real,
Apache-2.0, on-device, UE5-native framework built on nearly the same size-class models
(Qwen3.5-4B / Nemotron-3-Nano-4B sit almost exactly on the doc's "min-spec 3-4B" tier) with a
working Agent/Chat/RAG API split and actual shipped-game validation. **This does not mean the
bespoke architecture should be abandoned** — the doc's deterministic-first guardrail spine (F1-F10,
the canon-graph vocabulary mask, the strategy-token-never-writes-to-sim rule) is considerably more
rigorous than what a general-purpose game-agent SDK provides out of the box, and the project's canon-
safety requirements (§17, do-not-invent, the whitelist-per-character fact model) are bespoke by
necessity. But treating ACE purely as market trivia would be a missed opportunity: it is worth a
direct evaluation pass — either as a reference architecture to validate design choices against, or
as a foundation layer to build the bespoke guardrail spine *on top of* rather than writing the
on-device inference/API/UE5-plugin substrate from scratch. **Recommendation: add an explicit
evaluation task ahead of the runtime-generative-layer build** (this is a scoping/architecture
decision with real engineering-time stakes, not a pure research question — flag it for Josh's
ruling per the standing decision protocol rather than silently deciding build-from-scratch vs
build-on-ACE here).
RUNTIME_GENERATIVE_LAYER.md / RUNTIME_GUARDRAIL_ARCHITECTURE.mdRefreshed candidate list per tier (all Apache 2.0 or MIT unless noted — no MAU/field-of-use exposure
at ship):
| Tier | Refreshed candidates | Change from the doc's current picks |
|---|---|---|
| Draft/ambient (0.5-1B) | Qwen3-0.6B, Qwen3-1.7B, Llama-3.2-1B | Qwen3-0.6B/1.7B are current-generation replacements for the doc's "Qwen3-0.8B-class" placeholder |
| Min-spec (3-4B) | Qwen3-4B, Gemma-4 E4B, Phi-4-mini (3.8B), Ministral-3-3B, Nemotron-3-Nano-4B | Gemma is now addable (license flip); Nemotron-3-Nano-4B is a new NVIDIA-game-SDK-validated option |
| Prestige PC (7-8B) | Qwen3-8B, Ministral-3-8B | Largely unchanged, Ministral 3 added as a second clean option |
| 14B reference/ship | Qwen3-14B, Phi-4 (14B), Ministral-3-14B, Gemma-4 12B-class | Same tier, more (and more current) candidates — no change to the 14B ruling itself |
| 28-32B prestige | Qwen3-32B, Qwen3.6-27B, DeepSeek-R1-Distill-Qwen-32B | Confirmed as dev-side/2030-tier only given the §3.3 VRAM-co-residency finding — this is a sharpening of the existing ruling, not a reversal |
No change to the 14B-is-the-reference ruling itself — this refresh adds current, cleanly-licensed
candidates at every tier and confirms the ruling's own "never a content dependency" hedge on 28-32B
with concrete VRAM-contention numbers, but does not disturb the anchor decision.
---
Researched via a dedicated parallel sub-lane against the same discipline; reproduced here with light
reformatting to match this document's structure. **Bottom line: all four named vendors (ElevenLabs,
AIVA, Stable Audio Open, licensed libraries) still stand — no decision-forcing change.**
10,000 credits; Starter $6/30,000; Creator $11/121,000; Pro $99/600,000; Scale $299/1,800,000;
Business $990/6,000,000; Enterprise custom. (Some SEO aggregators still cite stale $5/$22 figures —
the live pricing page supersedes them.)
paid users retain full output rights, no volume cap, no per-line restriction — NPC-scale
many-thousands-of-lines use is permitted. The existing clone-authorization-gate design remains
correctly aligned to ElevenLabs' own consent requirement.
seamless looping, explicitly pitched for games — an upgrade (SFX V2, Sept 2025) over what the
existing stack docs assumed.
(April 2026) unifies voice+music+SFX under one vendor — noted as an FYI, not a reason to displace
AIVA's composer-in-the-loop workflow.
(fixed preset voices only), ~2-3GB VRAM or CPU-runnable, faster than real-time. **The safest
bulk-ambient-NPC option**: zero cloning-consent exposure because it cannot clone at all.
23+ languages, built-in "PerTh" neural watermarking. **This is the same model family NVIDIA's ACE
Game Agent SDK bundles as its on-device TTS** (§3.5, "Chatterbox-Turbo-350M") — independent
cross-confirmation that it's a real, current, production-viable choice. Consent flag: trivial-
to-use zero-shot cloning with no platform-side gatekeeping (unlike ElevenLabs) means adopting
Chatterbox for any real-person-adjacent voice work shifts the entire consent burden onto this
project's own clone-authorization gate — the watermarking is a partial mitigation, not a substitute.
(CC-BY-NC-SA-4.0), XTTS v2/Coqui (CPML) — all non-commercial, a real trap since these are
otherwise well-known names.
ambient/crowd lines generated locally on the 5090 to cut per-call cost — additive, not a
replacement.
tier genuinely assigns full copyright, suitable for a shipped game soundtrack; no material change.
(VERIFIED), but no indemnification against infringement claims, and the underlying label
litigation is partially, not fully, settled — Warner and UMG have settled with one or both
platforms, but Sony has not settled with either, and the core US fair-use question now looks
likely to slip into 2027 (INFERRED, well-corroborated). Usable for scratch/reference; AIVA's
copyright-assignment license remains the lower-risk choice for final shipped assets.
data specifically to sidestep the Suno/Udio-style infringement question; relevant primarily via its
open SFX model (§4.4).
$1M annual revenue with output ownership (VERIFIED) — a solo dev is cleanly inside this.
the direct successor to evaluate for the SFX chain; one item to verify directly before relying on
it: whether the weights (not just the inference code, which is MIT) carry the same $1M-revenue
Community License threshold or something else — a genuine unresolved conflict between the repo's
MIT code license and Stability's typical weights-licensing pattern, flagged rather than resolved.
No vendor replacement needed. Two concrete additive actions: (1) upgrade the SFX-chain reference from
Stable Audio Open 1.0 to Stable Audio 3.0's open SFX model (after confirming its weights license
directly); (2) log Kokoro-82M as an evaluated option for bulk NPC ambient dialogue to reduce
ElevenLabs per-call spend at scale, with Chatterbox as the higher-quality/cloning-capable alternative
whose consent risk must be owned entirely by this project's own gate design.
---
Consolidated across all four lanes rather than repeated per-section:
1. NVIDIA ACE Game Agent SDK (§3.5) — not really "upcoming," already shipped (v0.5.0,
2026-06-15) and in beta at Unreal Fest 2026, with NVIDIA presenting on it at SIGGRAPH 2026 this
week. The single highest-priority item to evaluate before the runtime-generative-layer build
starts.
2. SIGGRAPH 2026 itself (§1.3) — this week (July 19-23), not next month. Expect more
papers-with-code to land through and just after the conference; the three found in this pass
(Pixal3D, CHORD-lineage, SimArt, plus the not-yet-released SATO) are very likely not the last —
worth a direct re-check of s2026.conference-schedule.org and the kesen.realtimerendering.com
SIGGRAPH papers tracker after the conference closes, since search-engine coverage during the
conference week itself is necessarily partial.
3. Qwen's flagship going closed (Qwen3.7, May-June 2026, API-only) while the open dense/small line
(Qwen3/Qwen3.6) stays Apache 2.0 — a pattern worth watching across other vendors; if it continues,
the open small-model tier this project depends on could face future pressure even if today's
picks are safe.
4. Suno/Udio litigation — Munich's GEMA v. Suno verdict was due 2026-07-31 (right after this
research date) and a US fair-use summary-judgment hearing was underway in July 2026, with a
definitive US ruling now expected to slip into 2027. Worth a calendar re-check regardless of
whether this project uses Suno/Udio directly, since a ruling either direction will reshape the
entire AI-music-licensing landscape AIVA and Stable Audio currently sit inside.
5. A rumored higher-end Blackwell card ("RTX 5090 Ti" / "Titan Blackwell") teased for Q3 2026
(INFERRED, rumor-tier sourcing, not a product announcement) — not actionable, but worth knowing the
5090 may not stay the top consumer card for the full dev cycle.
6. NVFP4/MXFP4 kernel support actively merging into llama.cpp (§3.4) — re-check at benchmark-gate
time; if fully landed by then, it changes the practical VRAM math for the 5090 in this project's
favor.
---
Against LOCAL_3D_ASSET_GEN.md:
still no maintainer response.
directly against the live GitHub org listing.
ArmorLab" and the LaForge "Generative Base Material" maturity flag — SUPERSEDED/RESOLVED: it
shipped as CHORD (Dec 2025), real code+weights exist now, but it's Research-Only Copyleft — so the
underlying conclusion (no free+open+commercial PBR-material option beats the existing Hunyuan3D+
ArmorLab combo) is actually REINFORCED, just with a concrete, named, verified example rather
than an open question.
the operative default**, with two new promising-but-not-yet-actionable candidates added to watch
(SimArt — real but robotics-URDF-targeted; SATO — code not yet released).
benchmark roster.
Against NEOSTACK_AI.md:
unverified NeoStack categories" — CONFIRMED, no change to the core recommendation.
generation actually run") — RESOLVED: Meshy+Tripo (3D), ElevenLabs (voice), per §2.1-2.2.
quote on record literally names SIK/EIK, not AIK; every public AIK-specific source still shows
5.5-5.7 only. Strengthens rather than removes the existing "verify at first use" gate.
Against STACK_FACTS_QUESTIONS.md:
RE-READ**: the quoted Discord text names SIK/EIK specifically; no independent public source confirms
AIK/NeoStack itself is on 5.8 yet. This is presented as evidence for Josh's own re-check, not as an
autonomous override of a Josh-ruled item — per the standing decision protocol, a technical
verification finding surfaces here with full evidence; the ruling itself is Josh's to revisit or
reaffirm.
voice generation" — RESOLVED to specific vendors (§2.1-2.2), "4-5 agentic models" and "4 new
products/IDE" INFERRED-resolved to already-known quantities (§2.4) pending Josh confirming
against his own Discord screenshots if a sharper source is available.
as sharpest-geometry fallback... avoid Sparc3D... the JOSH-RULED commercially-unrestricted-by-default
posture" — CONFIRMED, no change; NeoStack's own bundled 3D-gen (Meshy/Tripo, both closed SaaS)
now explicitly folded into the "previz/reference only" tier of that same ruling, consistent with how
Meshy/Tripo were already treated pre-NeoStack.
CONFIRMED, per §4 — no vendor swap needed; two additive upgrades logged (Stable Audio 3.0's SFX
model, Kokoro for bulk NPC lines) for whenever that first audio build happens.
Against RUNTIME_GENERATIVE_LAYER.md / RUNTIME_GUARDRAIL_ARCHITECTURE.md:
"never a content dependency" prestige-if-2030-hardware-allows tier — CONFIRMED, no change to the
ruled anchor.
— PARTIALLY SUPERSEDED: still real, valid, currently-shippable options, but no longer the
sharpest available at each tier — refreshed roster in §3.6 (adds Gemma-4, Ministral-3, Qwen3.6-27B,
Nemotron-3-Nano-4B as newly-relevant or newly-licensed-clean options).
(CUDA 12.8, not 13.1) and an additive Blackwell-tier optimization path (NVFP4/MXFP4) not previously
known.
numbers now show explicitly why 32B can't co-reside with a demanding renderer TODAY, giving the
existing "never a content dependency" hedge a measured basis rather than a projected one.
shipped, Apache-2.0, on-device, UE5-integrated architecture close to this doc's own independent
design. This is a genuine addition requiring a scoping decision (evaluate-as-reference vs
build-on-top vs bespoke-regardless) that this brief flags for Josh's ruling rather than resolving
unilaterally, given it has real engineering-time stakes for how the runtime-generative-layer build
actually starts.
Against RESEARCHED_STACK.md's residual-set triage:
the right posture**, roster expanded per §1.4.
evidence for why** (§2.5).
evaluation, the Stable-Audio-3.0-weights-license check) should be added to it.
---
3daistudio.com state-of-AI-3D-2026 · pixazo.ai open-source-3D leaderboard · trellis-2.com/trellis2.app/
trellis2.org (vendor-adjacent) · hunyuan3d.cc (already flagged unreliable by the sibling brief) ·
various 2026 SEO/aggregator roundups on "best open source 3D model" and "state of 3D generation 2026" ·
YouTube titles (Pixal3D VRAM claims) · s2026.conference-schedule.org (tool did not return populated
session data at fetch time) · kesen.realtimerendering.com SIGGRAPH papers tracker (referenced, not
directly fetched) · aggregator search snippets for NeoStack v2.0.45 (uncorroborated, flagged as
conflicting) · various RTX 5090 local-AI benchmark blogs.
own report; headline primary sources independently re-verified by the lead researcher: