Q3_2026_MODELS_REFRESH.md

pipelines/Q3_2026_MODELS_REFRESH.md

Q3-2026 Models Refresh — Generative Pipeline Sweep

Status: RESEARCH BRIEF — a refresh + gap-fill pass over the tech_research corpus, not canon, not a

build order. Research date 2026-07-15 (same calendar day as LOCAL_3D_ASSET_GEN.md and

NEOSTACK_AI.md — this is a same-day deepen-and-widen pass, not a months-later refresh; where it

finds nothing changed, it says so explicitly rather than padding). Training-data cutoff for the

assistant that wrote this is ~January 2026; every load-bearing claim below was checked against the

live web on 2026-07-15 and is tagged VERIFIED (fetched from a primary source — official repo/

license file/model card/vendor pricing-terms page — and quotable) or INFERRED (secondary/

aggregator source, community report, or synthesis — not independently cross-checked against a

primary document). Where sources conflicted, the conflict is stated rather than silently resolved.

Two lanes (§3 local LLMs, §4 voice/audio) were researched by parallel sub-agents under this same

tagging discipline, then spot-checked directly against primary sources by the lead researcher

(Llama 4/Gemma 4/Phi-4/DeepSeek license claims and the NVIDIA ACE Game Agent SDK were independently

re-fetched and confirmed) — treat their VERIFIED tags with the same weight as the rest of this doc.

Scope: five deep-dive targets set by the mission brief — (1) refresh LOCAL_3D_ASSET_GEN.md against

anything newer, (2) identify what NeoStack's own bundled cloud generation actually runs, (3) the

current local-LLM landscape for the 5090 (runtime layer + build-agent tiers), (4) a light pass on

voice/music/SFX vendors, (5) credible 2026 H2 announcements. A dedicated §DELTAS section closes the

loop against every prior conclusion in LOCAL_3D_ASSET_GEN.md, NEOSTACK_AI.md,

RUNTIME_GENERATIVE_LAYER.md, RUNTIME_GUARDRAIL_ARCHITECTURE.md, STACK_FACTS_QUESTIONS.md, and

RESEARCHED_STACK.md — confirming or superseding each.

---

0. VERDICT — read this first

**Nothing in the existing tech_research corpus was WRONG when written earlier today; several things

are already dated, and one finding (§3.5) is big enough to change how the runtime-generative-layer

design doc should be evaluated.**

1. 3D asset gen (§1): the core §0 verdict of LOCAL_3D_ASSET_GEN.md — split TRELLIS.2

(geometry) / Hunyuan3D-2.1 (finished PBR, territory-capped) / Hi3DGen (sharp-geometry fallback) —

is RECONFIRMED, not superseded. TRELLIS.2's nvdiffrast/nvdiffrec commercial trap is still

unresolved (issue #22, no maintainer response). Hunyuan3D 2.5/3.x are still not open-sourced. The

"commercially-clean-PBR gap" the sibling brief flagged is STILL OPEN — two new SIGGRAPH-2026

entrants (Pixal3D, Ubisoft CHORD) both looked at first glance like they might close it and **both

turn out not to**: Pixal3D's own code is clean MIT but its install chain runs through TRELLIS.2

(same nvdiffrast dependency, inherited not escaped); CHORD ships real open weights but under a

Research-Only Copyleft license. New, real, and worth benchmarking: Pixal3D (near-

reconstruction-fidelity pixel-aligned texturing) and Direct3D-S2 (the sparse-attention

geometry model Pixal3D is built on) both belong in the §5.2 benchmark set alongside the existing

three. SIGGRAPH 2026 is NOT "next month" — it is 2026-07-19 to 23 in Los Angeles, i.e., this

week — corrected below.

2. NeoStack (§2): the Discord-reported "3D asset generation" and "voice generation" bundled into

the lifetime tier are identified concretely for the first time: 3D gen is a cloud passthrough

to Tripo and Meshy (both closed SaaS, both already known to this project as reference-tier,

not local-pipeline candidates); voice gen is a cloud passthrough to ElevenLabs — the exact

same vendor already locked into this project's own audio stack. **Neither duplicates the local

TRELLIS-class pipeline** — they're a convenience UI over already-evaluated closed vendors, not a

new model choice. The "4 new products / own IDE" Discord datapoint most plausibly resolves to

Betide Studio's wider 9-product integration-kit catalog (Steam/EOS/Edgegap/PlayFab/Crossplay/

GameCenter/PlayServices/Matchmaking + Agent Integration Kit) rather than a genuinely new product

category — no public evidence of a standalone "NeoStack IDE" was found; the IDE reference most

likely means the already-known NeoStack Cloud hosted-IDE feature. A genuine new flag: the

Discord quote Josh has on record naming UE 5.8 support verbatim names "SIK" and "EIK"

(Steam/EOS Integration Kit) — not "AIK" (Agent Integration Kit = NeoStack). Every public source

checked (official docs, changelog, Fab listings) still shows NeoStack itself capped at UE 5.5-5.7.

This doesn't overturn the "keep NeoStack" ruling — it sharpens exactly why the standing "verify at

first use" pre-flight check matters.

3. Local LLMs (§3): the runtime-layer doc's named candidates (Qwen3-3B, Phi-4-mini, Qwen3-8B,

Llama-3.2-1B) are still real and still fine choices, but no longer the sharpest ones available,

and one licensing fact has flipped in the project's favor: **Gemma moved to Apache 2.0 with Gemma

4 (2026-03-31)** — the historical Gemma restricted-license problem is gone for anyone using

Gemma 4+. Blackwell ships a genuinely better-than-Q4_K_M quantization path (NVFP4/MXFP4). Most

importantly: **NVIDIA shipped an open-source (Apache 2.0), on-device, UE5-plugin-integrated

"ACE Game Agent SDK"** (v0.5.0, 2026-06-15) that is architecturally close to what

RUNTIME_GENERATIVE_LAYER.md designed independently — small local SLMs (Qwen3.5-4B / Nemotron-3-

Nano-4B), on-device TTS/ASR, an agent/chat/RAG API split, already adopted in three shipped/beta

titles (PUBG Battlegrounds, Total War Pharaoh, Mecha BREAK). This deserves a dedicated evaluation

pass, not just a footnote — see §3.5.

4. Voice/audio (§4): the named stack (ElevenLabs, AIVA, Stable Audio Open, licensed libraries)

all still stand — no vendor needs replacing. Two real updates: Stable Audio Open 1.0 has a

successor (Stable Audio 3.0, May 2026, adds an open-weight dedicated SFX model) and the open-

TTS field now has genuinely commercial-clean local options (Kokoro, Apache 2.0, no cloning —

and, notably, Chatterbox — MIT — is the same TTS model NVIDIA's ACE SDK bundles) worth

evaluating for bulk ambient NPC lines to cut ElevenLabs per-call spend. Suno/Udio remain usable but

carry residual legal risk (no indemnification, US fair-use ruling still pending, likely into 2027)

that AIVA's full-copyright-assignment license does not.

5. Upcoming (§5): SIGGRAPH 2026 is this week, not next month, and is already producing usable

artifacts (Pixal3D, SimArt, SATO all ship with GitHub repos, though SATO's code is "being prepared

for public release" — not yet usable). The single most consequential near-term item for this

project's own architecture is the NVIDIA ACE Game Agent SDK (§3.5) — it should be evaluated before

the runtime-generative-layer build starts, not discovered after.

---

1. 3D asset generation — refresh of LOCAL_3D_ASSET_GEN.md

1.1 What's unchanged — reconfirmed, not superseded

GitHub issue #22 (the nvdiffrast non-commercial-license question) — VERIFIED, re-fetched directly —

remains open, opened by a user on 2025-12-18, no maintainer response of any kind as of today.

Microsoft has not patched the dependency, has not responded to the issue, and no newer TRELLIS

version exists to check instead (VERIFIED — github.com/microsoft/TRELLIS.2/issues/22).

GitHub org listing: it shows only Hunyuan3D-2.1 (3.7k stars) as a public 3D-mesh repo — no

2.5, no 3.0, no 3.1 repo exists (VERIFIED — github.com/Tencent-Hunyuan).

This directly reconfirms LOCAL_3D_ASSET_GEN.md §1.2's finding; **"Hunyuan3D-2.1 is the version to

actually use, nothing past it is open" still holds exactly as written.**

March 2025 paper/repo the sibling brief already catalogued (VERIFIED — repo/release pages checked,

no new tags).

1.2 What's new since that brief — the SIGGRAPH-2026 wave

Three genuinely new items surfaced, all tagged [SIGGRAPH 2026] on their own repos, none of which

were findable when the sibling brief was researched (they are today's/this-week's news given the

SIGGRAPH date correction in §1.3):

**Pixal3D (TencentARC) — MIT on its own code, but inherits TRELLIS.2's nvdiffrast dependency; do NOT

treat as a clean escape from the PBR-licensing trap.**

Pixal3D: Pixel-Aligned 3D Generation from Images." Weights on HuggingFace

(huggingface.co/TencentARC/Pixal3D). VERIFIED.

correspondence) rather than loosely injecting image features via attention — claimed near-

reconstruction-level fidelity with detailed geometry and PBR textures (VERIFIED, repo/paper

description, arxiv.org/html/2605.10922v1).

unmodified MIT license (VERIFIED —

github.com/TencentARC/Pixal3D/blob/master/LICENSE).

Its NOTICE file lists only Apache-2.0 (dinov2) and MIT (TRELLIS.2, Direct3D-S2, MoGe) dependencies

— no NVIDIA non-commercial component is disclosed there (VERIFIED, direct fetch). But the

actual install instructions are: "Step 1: Follow TRELLIS.2 Installation" — i.e., Pixal3D

requires installing TRELLIS.2 first as a hard prerequisite (VERIFIED, direct fetch of the raw

README). A separate web search independently found that Pixal3D's own GPU-extension-wheel

requirements include nvdiffrast and nvdiffrec_render (INFERRED — search-aggregated, not

independently re-confirmed against a Pixal3D-specific install script beyond the TRELLIS.2-delegation

step), which is fully consistent with the TRELLIS.2-first install chain. **Conclusion: Pixal3D's own

code is clean MIT, but its practical dependency chain very likely reruns TRELLIS.2's PBR-bake path

(and therefore nvdiffrast/nvdiffrec) to produce its textured output — the same commercial-use

caveat flagged for TRELLIS.2 in LOCAL_3D_ASSET_GEN.md §1.1 should be assumed to extend to

Pixal3D's textured output until directly re-verified at benchmark time, not treated as bypassed.**

This is a nuanced, partially-INFERRED finding — flag it exactly this way, don't round it to either

"clean" or "blocked."

recommended, low-VRAM mode targets 24GB cards" while a YouTube video is titled "Pixal3D on 6GB VRAM

— Better Than Trellis 2!" (both INFERRED, unreconciled — treat 6GB as an aggressive/cut-down claim

and 16-24GB as the more conservative real-world figure pending direct benchmarking). Generation time

claimed at 3-5 minutes per asset in one ComfyUI-integration repo description (INFERRED,

github.com/dreamrec/ComfyUI-Pixal3D).

entry: a sparse-volume 3D generation framework using Spatial Sparse Attention (SSA), enabling

1024³-resolution training on 8 GPUs where prior volumetric approaches needed 32+ (VERIFIED license

and claim — github.com/DreamTechAI/Direct3D-S2,

cross-confirmed on HuggingFace). Not independently benchmarked here; worth a bench-day look as a

possible TRELLIS.2 alternative for the geometry stage specifically, since it does NOT carry

TRELLIS.2's nvdiffrast dependency chain (it's Pixal3D that adds that, by choosing to build on

TRELLIS.2 rather than on Direct3D-S2 alone for its render step).

**Ubisoft La Forge CHORD — resolves the sibling brief's open question, but the answer is "still not

commercially usable."**

LOCAL_3D_ASSET_GEN.md §3.6 flagged as "debuted at SIGGRAPH Asia 2025... maturity/completeness of

the actual released code not independently verified" — that flag is now resolved: it shipped, with

real code, real weights, and a ComfyUI integration (VERIFIED — released 2025-12-09,

github.com/ubisoft/ubisoft-laforge-chord,

huggingface.co/Ubisoft/ubisoft-laforge-chord,

ComfyUI node at github.com/ubisoft/ComfyUI-Chord).

Normal, Height, Roughness, Metalness — via chained rendering-decomposition + single-step diffusion

(VERIFIED, repo/paper description).

— explicitly not permitted for commercial use (VERIFIED, repo license page). **This confirms,

rather than closes, the sibling brief's "commercially-clean-PBR gap" finding**: the field's most

credible new open PBR-material-estimation release in the last several months is real, working, and

explicitly non-commercial.

ByteDance-Seed SimArt — adjacent to, not a solve for, the game-readiness gap.

— Apache 2.0, real weights on HuggingFace (VERIFIED —

github.com/ByteDance-Seed/SimArt). Does part-level mesh

segmentation + kinematic/joint prediction + URDF (robotics-simulator format) generation via a

sparse 3D VQ-VAE (VERIFIED, repo description).

named as one of the unsolved stages in the Hunyuan3D Studio research pipeline (which itself never

released code). SimArt is real, open, and does part decomposition + articulation prediction — but

its native output (URDF) targets robotics physics simulation, not UE skeletal-mesh rigging, and

the repo's own documentation frames coordinate-system alignment as a preprocessing concern rather

than addressing retopology or game-engine rig authoring directly (VERIFIED via direct fetch). **Net:

a genuinely promising building block to prototype against the game-readiness gap, not a drop-in

fix** — the §2 conclusion of LOCAL_3D_ASSET_GEN.md ("no mature open AI system automates

retopology + UVs + LODs + collision end-to-end today; budget Blender-headless scripting + human QC")

still stands as the operative default, with SimArt now flagged as worth a bench-day evaluation

specifically for the part-decomposition/rigging-prep sub-step.

artist-style mesh generation with native UV segmentation built into the same autoregressive

token stream (VERIFIED, arxiv.org/abs/2604.09132,

github.com/Xrvitd/SATO) — this is close to exactly the retopology-

plus-UV-unwrap gap. Not yet actionable: the repo's own status is "the codebase is being prepared

for public release" (VERIFIED, direct fetch) — a paper-with-a-repo-stub, same non-actionable status

the sibling brief gave Hunyuan3D Studio's PolyGen/SeamGPT. Worth tracking; re-check at benchmark-gate

time since a real release could land any time given SIGGRAPH is this week.

1.3 SIGGRAPH 2026 — a factual date correction

SIGGRAPH 2026 is 2026-07-19 to 2026-07-23 in Los Angeles (exhibition July 21-23) — VERIFIED,

directly fetched from the official conference site

(s2026.siggraph.org). **This is not "next month" relative to this

brief's research date (2026-07-15) — it is this week, essentially concurrent with this research

pass.** The mission brief's framing should be corrected going forward. Practical consequence: this

is not a "what's teased" forward-look question so much as a "what's already dropping" question — the

three items in §1.2 (Pixal3D, CHORD-adjacent lineage, SimArt) and SATO in §1.2 are evidence this wave

is already live, not upcoming. Technical Papers program stats: 1,120+ submissions this cycle

(VERIFIED, s2026.siggraph.org),

covering generative AI/ML for visual computing among its named tracks; the live conference-schedule

tool (s2026.conference-schedule.org) did not return a

populated session list at fetch time (client-side filtering, not yet browsable via a plain fetch) —

re-check that tool directly during/after the conference week for the full accepted-papers list rather

than relying on search-engine coverage, which is necessarily partial and skewed toward whatever

already has a public GitHub repo.

1.4 Updated recommendation for §5's benchmark gate

LOCAL_3D_ASSET_GEN.md §5.2's benchmark protocol (re-check licenses, run a fixed 2-per-class test set

through TRELLIS.2/Hunyuan3D-2.1/Hi3DGen/TripoSG/SF3D-SPAR3D) stands as written — add Pixal3D

(as a geometry+texture challenger to benchmark specifically for whether its output actually routes

through nvdiffrec at runtime — instrument the process, don't just trust the NOTICE file) and

Direct3D-S2 (as a possible TRELLIS.2-alternative geometry stage that doesn't carry the nvdiffrast

baggage) to the roster. Keep CHORD off the commercial-asset roster entirely (research-only license) —

it remains useful only as an internal-reference/previz tool, same treatment as the territory-capped

Hunyuan3D under the existing JOSH-RULED posture (commercially-unrestricted components are the

shipped-asset default; anything else is previz-only). Flag SimArt and SATO for a later, separate

evaluation specifically against the game-readiness/retopology gap once the workstation lands, rather

than folding them into the mesh-generation benchmark itself — they solve a different sub-problem.

---

2. NeoStack's bundled generation — the Discord/cloud-lane deep dive

This section answers the mission's Q1-addendum data point ("the lifetime purchase bundles weekly

usage incl. 4-5 agentic models, 3D ASSET GENERATION, and VOICE GENERATION") that NEOSTACK_AI.md

flagged but did not resolve to specific vendors.

2.1 The cloud 3D-gen lane — identified: Meshy + Tripo passthrough, not a NeoStack-owned model

Directly confirmed via NeoStack's own official changelog (VERIFIED, direct fetch of

aik.betide.studio/changelog):

A separate search independently corroborates a "Studio" tab inside the plugin that surfaces

"multiple AI generation providers including Meshy, Tripo, and ElevenLabs for 3D generation, rigging,

animation, and text-to-speech" (INFERRED — search-aggregated summary of the docs, not a direct fetch

of the Studio-tab page itself, which was not independently reached). **This is a router/aggregator UI

over closed third-party SaaS providers, not a NeoStack-proprietary model** — the same Meshy and Tripo

this project's own LOCAL_3D_ASSET_GEN.md §1.8 already evaluated and deliberately positioned as

"reference/quality bar, not a candidate" for the local, open-source pipeline (Meshy is the incumbent

the whole local-3D-gen research effort exists to replace; Tripo 3.0 is closed/API-only).

2.2 The cloud voice-gen lane — identified: ElevenLabs passthrough, same vendor already in the stack

Same changelog, same version: **v1.0.49 (2026-03-24): "ElevenLabs (new) — text-to-speech and sound

effects."** VERIFIED, direct fetch. This is not a new or different voice vendor — it is the exact

same ElevenLabs already locked into this project's audio stack per STACK_FACTS_QUESTIONS.md Q4 and

reconfirmed current in §4.1 below, just reachable through NeoStack's in-editor chat UI (and billed

against NeoStack Cloud's usage allowance) rather than a direct API call.

2.3 Does NeoStack's gen lane duplicate or complement the local TRELLIS-class pipeline

Complement, not duplicate — and only partially even that. NeoStack's 3D-gen surface gives

in-editor, conversational access to two closed SaaS generators this project already decided NOT to

depend on for the shipped-asset pipeline (§17-ruled commercial-clean-by-default posture,

STACK_FACTS_QUESTIONS.md Q3 Josh ruling). It doesn't touch TRELLIS.2/Hunyuan3D-2.1/Hi3DGen at all —

different vendors entirely, no overlap in model, license, or output path. Where it's genuinely useful

despite that: Tripo's rigging/animation capability is a real gap-filler. LOCAL_3D_ASSET_GEN.md

§3.2 already flagged that none of the open local mesh-generators produce rigs, and separately named

"Tripo AI's 'universal rig'... explicitly advertised as working across humanoid-to-creature ranges" as

a worth-knowing paid point-solution for the 132-creature roster. NeoStack's Studio tab is simply a

more convenient way to reach that same Tripo capability from inside the editor, at whatever

credit-cost NeoStack Cloud's allowance provides — an operational convenience on an already-identified

tool, not a new research finding about the tool itself.

2.4 "4 new products" / IDE / V2 / V3 — resolved as far as public sources allow

had visibility into (VERIFIED, direct fetch of

betide.studio/plugins): EOS Integration Kit, Steam Integration Kit,

Edgegap Integration Kit, PlayFab Integration Kit (deprecated), Ultimate Crossplay Integration Kit,

Game Center Integration Kit, Play Services Integration Kit, Matchmaking Integration Kit, and Agent

Integration Kit (NeoStack). **This most plausibly explains the "expanded product line, 4 new

products" Discord datapoint** — several of these (Edgegap/Crossplay/GameCenter/PlayServices/

Matchmaking) read as a newer wave layered onto the original EOS/Steam/AIK trio — rather than

indicating Betide is diversifying outside Unreal platform-integration tooling. **No standalone

"IDE" product was found anywhere in this catalog or in NeoStack's own public docs.** The most likely

resolution: the "own IDE" reference is the already-documented NeoStack Cloud hosted-IDE feature

(NEOSTACK_AI.md §6 — the $20/mo Cloud tier bundles "a hosted IDE, asset-aware version control,

CI/CD"), not a new fourth product. This is INFERRED — a reasonable reading of public evidence, not a

direct confirmation from Betide Studio; if Josh's Discord screenshot names something more specific,

that should override this inference.

(Claude Code, Gemini CLI, Codex, Cursor, GitHub Copilot = 5) from NEOSTACK_AI.md §3, not a new

finding — no marketing copy using that exact phrase was found (INFERRED).

changelog page, directly fetched twice, tops out at v1.0.50 (2026-03-24) with no entries between

then and today. One unrelated secondary aggregator's snippet claims "v2.0.45 as of June 6" exists

(INFERRED, uncorroborated by any other source, possibly confusing NeoStack with one of Betide's

other 8 products or simply wrong) — this conflict is presented, not resolved. Do not assume V2

has shipped for NeoStack specifically without checking aik.betide.studio directly at build time;

equally, do not assume it hasn't — the changelog fetch tool may be serving cached/incomplete content

for a fast-moving product page.

2.5 UE 5.8 support — the SIK/EIK vs AIK evidence nuance

STACK_FACTS_QUESTIONS.md Q1 records Josh's own Discord screenshot, quoted verbatim: *"SIK currently

have 5.8 on FAB, EIK ... soon also will be 5.8. We focus updating marketplace version first."*

**Read literally, this quote names Steam Integration Kit (SIK) and EOS Integration Kit (EIK) — not

Agent Integration Kit (AIK = NeoStack).** Every public source reachable in this research pass

(official docs at aik.betide.studio, the changelog, Fab-listing search snippets, `betide.studio/

neostack`) still states NeoStack/AIK support caps at UE 5.5, 5.6, 5.7 — none show a 5.8 tag

(VERIFIED across four independent fetches today). This does not contradict Josh's screenshot — it's

entirely possible SIK/EIK reached 5.8 first and AIK follows "soon," exactly as the quote's own phrasing

suggests with "we focus updating marketplace version first." This sharpens, rather than reverses,

the existing "verify NeoStack-on-5.8 directly before the Phase 6 vertical slice" action item already

standing in NEOSTACK_AI.md §0 and RESEARCHED_STACK.md's ENGINEERING-SPIKES list — it is now a

clearly evidenced check, not a generic caution.

2.6 Updated recommendation

No change to the Q1 "keep NeoStack" ruling — nothing here is disqualifying. Three concrete refinements

for the correction pass: (1) when the Translation Doc specifies NeoStack's generative capabilities,

name Meshy/Tripo (3D) and ElevenLabs (voice) explicitly as the underlying vendors rather than

treating "NeoStack generation" as its own model class — this matters because it means NeoStack's

generative output inherits Meshy's and Tripo's own commercial terms, not a NeoStack-specific license;

(2) the UE-5.8-pending status is now evidenced specifically for AIK, not just generically flagged —

keep the "verify at first use" gate but don't treat Josh's Discord screenshot as having already

cleared it for NeoStack itself; (3) if Tripo's rigging capability is wanted for the 132-creature

roster, reaching it via NeoStack's existing Cloud allowance is a reasonable default rather than a

separate direct Tripo subscription — but confirm the credit-cost math once the "10x the free tier of

NeoStack Cloud" allowance (VERIFIED phrase, found via search, exact numeric limits not located in this

pass) is understood in practice.

---

3. Local LLMs for the 5090 — runtime layer + build-agent tiers

Researched via a dedicated parallel sub-lane against the same VERIFIED/INFERRED discipline as the

rest of this brief, then spot-checked directly (Llama 4, Gemma 4, Phi-4, DeepSeek license claims and

the NVIDIA ACE Game Agent SDK were independently re-fetched by the lead researcher and confirmed

consistent). This directly targets RUNTIME_GENERATIVE_LAYER.md's named tiers (draft 0.5-1B, min-spec

3-4B, prestige 7-8B, 14B = the ruled reference/ship tier, 28-32B optional prestige) and

RUNTIME_GUARDRAIL_ARCHITECTURE.md's matching model+VRAM line.

3.1 Current SOTA by family, Q3 2026

FamilyCurrent open lineSizes relevant to Josh's tiersLicense
Meta LlamaLlama 4 (Scout/Maverick); Behemoth shelved, never shippedNo small dense model — Scout is ~109B MoE/17B-active; smallest usable local size is still the older Llama-3.2-1B/3BLlama Community License — 700M-MAU clause, attribution required, EU vision carve-out (VERIFIED, license text quoted below)
Alibaba QwenQwen3 (base) + Qwen3.5 (Feb 2026) + Qwen3.6 (Apr 2026)Qwen3: 0.6B/1.7B/4B/8B/14B/32B dense + MoE variants; Qwen3.6 adds a 27B dense. Qwen3.7 (May-Jun 2026) is API-only, no open weights — the flagship is closing even as the small/dense line stays openApache 2.0 across the dense/open line (VERIFIED, LICENSE files on HF)
Google GemmaGemma 4 (2026-03-31) supersedes Gemma 3/3nGemma 4: E2B (~2.3B), E4B (~4.5B), 26B-A4B MoE, 31B dense, +12B "Unified" multimodal (Jun 2026)Apache 2.0 — a genuine license flip from Gemma 1-3's restrictive custom terms (VERIFIED, quoted below)
Microsoft PhiPhi-4 family, now 7 variantsPhi-4-mini (3.8B), Phi-4 (14B), Phi-4-reasoning/-plus (14B), Phi-4-mini-reasoning (3.8B), Phi-4-reasoning-vision (~15B, Mar 2026)MIT across all variants (VERIFIED for Phi-4-mini's HF frontmatter)
MistralMinistral 3 (Dec 2025) + Mistral Small 4 (Mar 2026) + Magistral SmallMinistral 3: dense 3B/8B/14B (base+instruct+reasoning)Apache 2.0 for Ministral 3 (VERIFIED HF frontmatter) — Mistral is "progressively moving away" from its older non-commercial Research License; check each model card, not uniform
DeepSeekR1-Distill line; V4 Pro/Flash (Apr 2026, too large for local)Distills: 1.5B/7B/14B/32B (Qwen-based) + 8B/70B (Llama-based)Qwen-based distills inherit Apache 2.0; Llama-based distills inherit the Llama license (MAU clause) — split by base model, not uniform
NVIDIA Nemotron (not in the doc's original list — a genuine addition)Nemotron 3 Nano (Dec 2025) — hybrid Mamba-Transformer, 1M context, 4B variant used in NVIDIA's own game SDK (§3.5)4B / 9B / 12BNVIDIA Open Model License — permissive/commercial, derivatives allowed, but a bespoke license, not Apache/MIT — read before shipping
Out of local range, context onlyGLM-5.2 (Zhipu, MIT, rated a top open-weight model on several boards), Kimi K2.7, MiniMax M340B+ active / 700B+ totalLarge MoE — not a 32GB-card candidate regardless of license

License texts, directly quoted (VERIFIED):

active users of the products or services made available by or for Licensee… is greater than 700

million monthly active users in the preceding calendar month, you must request a license from

Meta."* Plus a mandatory "Built with Llama" attribution and an EU-domiciled-licensee carve-out

excluding multimodal/vision use.

permissive Apache 2.0 license."* The HuggingFace checkpoint itself shows license: apache-2.0 with

no gated agreement — a genuine, structural change from Gemma 1-3's custom Gemma Terms of Use (which

carries a Prohibited Use Policy, a flow-down obligation to downstream users, and a unilateral

Google termination right, and which Gemma 3/3n still use).

3.2 Shipping-embedded weights vs dev-side-tooling-only — the license split that matters for this project

Two different legal questions: (a) redistributing model weights to end users and running them

offline inside a shipped commercial game, vs (b) using a model purely as the developer's own

build-agent tooling, never redistributed. Nearly everything here clears (b) trivially; the prize is

what clears (a) cleanly.

Gemma 4 (Apache 2.0 — newly clean, was NOT true of Gemma 1-3), Phi-4/Phi-4-mini (MIT), Ministral

3 (Apache 2.0), DeepSeek's Qwen-based R1-distills (Apache 2.0 inherited). **This set alone covers

every tier RUNTIME_GENERATIVE_LAYER.md needs** (0.6B through 32B) with zero licensing exposure at

ship time.

plausibly fine for embedding given NVIDIA ships it inside their own game SDK, but confirm the

redistribution clause directly); Gemma 3/3n (the old custom Gemma Terms of Use — workable but

strictly worse than just using Gemma 4 instead).

DeepSeek distill (8B/70B) — a solo dev is nowhere near the 700M-MAU threshold, so shipping is legal,

but it requires "Built with Llama" attribution, redistributing license text with the weights,

Acceptable Use Policy adherence, and (if EU-domiciled) forgoing multimodal/vision use. **Given the

Apache/MIT set above covers every needed tier with none of this overhead, there is no longer a

reason to reach for Llama for the shipped-embedded tiers** — keep it as a dev-side/build-agent-only

option if used at all.

3.3 VRAM and quantization on a 32GB RTX 5090

fits comfortably in 32GB alongside a AAA renderer (INFERRED, consistent across benchmark sources).

settings routinely wants 12-20GB+ of its own VRAM — meaning the 28-32B "prestige" tier cannot

reliably co-reside with a real renderer on a single 32GB card today.** This is a genuine, concrete

refinement to the existing "14B = reference, 28-32B = optional prestige if 2030 hardware allows"

ruling: on THIS hardware (the 5090, now), 14B is the practical in-game co-resident ceiling; 32B

is realistically a dev-side/offline-evaluation tier or a 2030-hardware ship tier exactly as already

ruled, not something to expect working live next to the renderer on the 5090 itself. This

*reinforces* the existing ruling's own hedge ("NEVER a content dependency") rather than contradicting

it — it just makes the "why" concrete with real numbers.

standard MXFP4).** NVIDIA's 4-bit float format with two-level micro-scaling lands roughly 25%

smaller than Q4_K_M at within ~1% of FP8 accuracy on reasoning/knowledge benchmarks (INFERRED,

NVIDIA's own technical blog + secondary corroboration) — e.g. a 27B model at ~14GB instead of the

Q4_K_M-equivalent ~18-20GB. The catch: NVFP4 is Blackwell-exclusive (5090/Blackwell-generation

cards only), so it can only ever be a Blackwell-tier optimization layered on top of a portable

GGUF Q4_K_M baseline for players on older hardware — it doesn't change the shipped min-spec quant

choice, but it meaningfully improves what's achievable on the dev's own 5090 specifically (more

VRAM headroom for renderer co-residency at the 14B-and-above tiers).

3.4 Inference runtime maturity on Blackwell/5090

with CUDA 12.8, not CUDA 13.1** — 13.1 is reported to segfault llama.cpp's Blackwell MMQ kernel,

silently falling back to slower cuBLAS (VERIFIED as a real, documented build issue). NVFP4/MXFP4

kernel support is actively merging for native Blackwell dispatch.

community workarounds; heavier/server-oriented, more than a single in-game agent needs.

consumer-Blackwell kernel support lagged datacenter Blackwell into early 2026 and needs custom

builds — a heavier toolchain than a solo dev needs versus llama.cpp.

5090 — the natural day-to-day dev-tooling choice.

successor to the old ChatRTX concept), with an NVIDIA-published (so vendor-sourced, treat with

appropriate skepticism) claim of ~7.3x faster than Ollama in one comparison.

and treat NVFP4/MXFP4 as an additive Blackwell-tier optimization once it's fully merged, not a

replacement recommendation.

3.5 The NVIDIA ACE Game Agent SDK — verified directly, and it matters

This is the single most consequential finding in this lane and deserves prominent treatment rather

than a passing mention. Independently re-fetched and confirmed (not solely relying on the research

sub-lane):

released 2026-06-15, beta unveiled at Unreal Fest 2026 (2026-06-16) with NVIDIA researchers

presenting at SIGGRAPH 2026 (2026-07-20 — this week). Repo:

github.com/NVIDIA/game-agent-sdk — VERIFIED to contain

real C99-ABI/C++ code (85.1% C++/7.8% C), a working build system, three functional sample

executables, and ~3GB of bundled ML models — not a stub.

shipping-embedded bar cleanly, same tier as the best options in §3.2.

Agent/Chat/RAG API split (stateful autonomous-reasoning agent vs. stateless direct-control chat vs.

semantic/lexical knowledge retrieval), fully on-device ("Unlike cloud-based services... local,

RTX-optimized workflows"), with reference models sized for exactly the small-local tier this project

already anchors on: Qwen3.5-4B (local GGUF, the decision-making SLM), Nemotron-3-Nano-4B

(alternative SLM), Chatterbox-Turbo-350M (on-device TTS — the same Chatterbox family the voice-

gen sub-lane independently flagged in §4.2, cross-confirming it as a real, current, commercially

viable option), and NeMo-Conformer-CTC-120M (on-device ASR). Targets an ~8GB VRAM budget,

explicitly runs down to an RTX 3060-class card, not just the 5090.

character dialogue, voice commands, and contextual action responses — distributed via

developer.nvidia.com/ace-for-games rather than inside

the core game-agent-sdk GitHub repo itself (the core repo is engine-agnostic C/C++; the UE5 plugin

layer sits on top of it — this distinction matters if evaluating the SDK, since the two are separate

downloads).

teammate, open beta), Total War: PHARAOH (experimental in-game advisor using RAG over "1,200+

interlinked game data tables" — a scale point directly relevant to how this project's own canon

graph/registries could feed an equivalent RAG layer), and Mecha BREAK (shipped).

"deterministic sim + a bounded strategy-token LLM layer, validated, fallback-guarded, on-device,

small models") was designed bespoke, independent of this SDK. NVIDIA has now shipped a real,

Apache-2.0, on-device, UE5-native framework built on nearly the same size-class models

(Qwen3.5-4B / Nemotron-3-Nano-4B sit almost exactly on the doc's "min-spec 3-4B" tier) with a

working Agent/Chat/RAG API split and actual shipped-game validation. **This does not mean the

bespoke architecture should be abandoned** — the doc's deterministic-first guardrail spine (F1-F10,

the canon-graph vocabulary mask, the strategy-token-never-writes-to-sim rule) is considerably more

rigorous than what a general-purpose game-agent SDK provides out of the box, and the project's canon-

safety requirements (§17, do-not-invent, the whitelist-per-character fact model) are bespoke by

necessity. But treating ACE purely as market trivia would be a missed opportunity: it is worth a

direct evaluation pass — either as a reference architecture to validate design choices against, or

as a foundation layer to build the bespoke guardrail spine *on top of* rather than writing the

on-device inference/API/UE5-plugin substrate from scratch. **Recommendation: add an explicit

evaluation task ahead of the runtime-generative-layer build** (this is a scoping/architecture

decision with real engineering-time stakes, not a pure research question — flag it for Josh's

ruling per the standing decision protocol rather than silently deciding build-from-scratch vs

build-on-ACE here).

3.6 Updated tier-ladder for RUNTIME_GENERATIVE_LAYER.md / RUNTIME_GUARDRAIL_ARCHITECTURE.md

Refreshed candidate list per tier (all Apache 2.0 or MIT unless noted — no MAU/field-of-use exposure

at ship):

TierRefreshed candidatesChange from the doc's current picks
Draft/ambient (0.5-1B)Qwen3-0.6B, Qwen3-1.7B, Llama-3.2-1BQwen3-0.6B/1.7B are current-generation replacements for the doc's "Qwen3-0.8B-class" placeholder
Min-spec (3-4B)Qwen3-4B, Gemma-4 E4B, Phi-4-mini (3.8B), Ministral-3-3B, Nemotron-3-Nano-4BGemma is now addable (license flip); Nemotron-3-Nano-4B is a new NVIDIA-game-SDK-validated option
Prestige PC (7-8B)Qwen3-8B, Ministral-3-8BLargely unchanged, Ministral 3 added as a second clean option
14B reference/shipQwen3-14B, Phi-4 (14B), Ministral-3-14B, Gemma-4 12B-classSame tier, more (and more current) candidates — no change to the 14B ruling itself
28-32B prestigeQwen3-32B, Qwen3.6-27B, DeepSeek-R1-Distill-Qwen-32BConfirmed as dev-side/2030-tier only given the §3.3 VRAM-co-residency finding — this is a sharpening of the existing ruling, not a reversal

No change to the 14B-is-the-reference ruling itself — this refresh adds current, cleanly-licensed

candidates at every tier and confirms the ruling's own "never a content dependency" hedge on 28-32B

with concrete VRAM-contention numbers, but does not disturb the anchor decision.

---

4. Voice/audio refresh — light pass

Researched via a dedicated parallel sub-lane against the same discipline; reproduced here with light

reformatting to match this document's structure. **Bottom line: all four named vendors (ElevenLabs,

AIVA, Stable Audio Open, licensed libraries) still stand — no decision-forcing change.**

4.1 ElevenLabs — still the right pick for character/NPC voice

10,000 credits; Starter $6/30,000; Creator $11/121,000; Pro $99/600,000; Scale $299/1,800,000;

Business $990/6,000,000; Enterprise custom. (Some SEO aggregators still cite stale $5/$22 figures —

the live pricing page supersedes them.)

paid users retain full output rights, no volume cap, no per-line restriction — NPC-scale

many-thousands-of-lines use is permitted. The existing clone-authorization-gate design remains

correctly aligned to ElevenLabs' own consent requirement.

seamless looping, explicitly pitched for games — an upgrade (SFX V2, Sept 2025) over what the

existing stack docs assumed.

(April 2026) unifies voice+music+SFX under one vendor — noted as an FYI, not a reason to displace

AIVA's composer-in-the-loop workflow.

4.2 Open/local TTS — a genuinely new option for bulk NPC dialogue

(fixed preset voices only), ~2-3GB VRAM or CPU-runnable, faster than real-time. **The safest

bulk-ambient-NPC option**: zero cloning-consent exposure because it cannot clone at all.

23+ languages, built-in "PerTh" neural watermarking. **This is the same model family NVIDIA's ACE

Game Agent SDK bundles as its on-device TTS** (§3.5, "Chatterbox-Turbo-350M") — independent

cross-confirmation that it's a real, current, production-viable choice. Consent flag: trivial-

to-use zero-shot cloning with no platform-side gatekeeping (unlike ElevenLabs) means adopting

Chatterbox for any real-person-adjacent voice work shifts the entire consent burden onto this

project's own clone-authorization gate — the watermarking is a partial mitigation, not a substitute.

(CC-BY-NC-SA-4.0), XTTS v2/Coqui (CPML) — all non-commercial, a real trap since these are

otherwise well-known names.

ambient/crowd lines generated locally on the 5090 to cut per-call cost — additive, not a

replacement.

4.3 Music generation

tier genuinely assigns full copyright, suitable for a shipped game soundtrack; no material change.

(VERIFIED), but no indemnification against infringement claims, and the underlying label

litigation is partially, not fully, settled — Warner and UMG have settled with one or both

platforms, but Sony has not settled with either, and the core US fair-use question now looks

likely to slip into 2027 (INFERRED, well-corroborated). Usable for scratch/reference; AIVA's

copyright-assignment license remains the lower-risk choice for final shipped assets.

data specifically to sidestep the Suno/Udio-style infringement question; relevant primarily via its

open SFX model (§4.4).

4.4 SFX generation

$1M annual revenue with output ownership (VERIFIED) — a solo dev is cleanly inside this.

the direct successor to evaluate for the SFX chain; one item to verify directly before relying on

it: whether the weights (not just the inference code, which is MIT) carry the same $1M-revenue

Community License threshold or something else — a genuine unresolved conflict between the repo's

MIT code license and Stability's typical weights-licensing pattern, flagged rather than resolved.

4.5 Updated recommendation

No vendor replacement needed. Two concrete additive actions: (1) upgrade the SFX-chain reference from

Stable Audio Open 1.0 to Stable Audio 3.0's open SFX model (after confirming its weights license

directly); (2) log Kokoro-82M as an evaluated option for bulk NPC ambient dialogue to reduce

ElevenLabs per-call spend at scale, with Chatterbox as the higher-quality/cloning-capable alternative

whose consent risk must be owned entirely by this project's own gate design.

---

5. Upcoming — credible 2026 H2 developments worth designing around

Consolidated across all four lanes rather than repeated per-section:

1. NVIDIA ACE Game Agent SDK (§3.5) — not really "upcoming," already shipped (v0.5.0,

2026-06-15) and in beta at Unreal Fest 2026, with NVIDIA presenting on it at SIGGRAPH 2026 this

week. The single highest-priority item to evaluate before the runtime-generative-layer build

starts.

2. SIGGRAPH 2026 itself (§1.3) — this week (July 19-23), not next month. Expect more

papers-with-code to land through and just after the conference; the three found in this pass

(Pixal3D, CHORD-lineage, SimArt, plus the not-yet-released SATO) are very likely not the last —

worth a direct re-check of s2026.conference-schedule.org and the kesen.realtimerendering.com

SIGGRAPH papers tracker after the conference closes, since search-engine coverage during the

conference week itself is necessarily partial.

3. Qwen's flagship going closed (Qwen3.7, May-June 2026, API-only) while the open dense/small line

(Qwen3/Qwen3.6) stays Apache 2.0 — a pattern worth watching across other vendors; if it continues,

the open small-model tier this project depends on could face future pressure even if today's

picks are safe.

4. Suno/Udio litigation — Munich's GEMA v. Suno verdict was due 2026-07-31 (right after this

research date) and a US fair-use summary-judgment hearing was underway in July 2026, with a

definitive US ruling now expected to slip into 2027. Worth a calendar re-check regardless of

whether this project uses Suno/Udio directly, since a ruling either direction will reshape the

entire AI-music-licensing landscape AIVA and Stable Audio currently sit inside.

5. A rumored higher-end Blackwell card ("RTX 5090 Ti" / "Titan Blackwell") teased for Q3 2026

(INFERRED, rumor-tier sourcing, not a product announcement) — not actionable, but worth knowing the

5090 may not stay the top consumer card for the full dev cycle.

6. NVFP4/MXFP4 kernel support actively merging into llama.cpp (§3.4) — re-check at benchmark-gate

time; if fully landed by then, it changes the practical VRAM math for the 5090 in this project's

favor.

---

DELTAS — confirm or supersede, conclusion by conclusion

Against LOCAL_3D_ASSET_GEN.md:

still no maintainer response.

directly against the live GitHub org listing.

ArmorLab" and the LaForge "Generative Base Material" maturity flag — SUPERSEDED/RESOLVED: it

shipped as CHORD (Dec 2025), real code+weights exist now, but it's Research-Only Copyleft — so the

underlying conclusion (no free+open+commercial PBR-material option beats the existing Hunyuan3D+

ArmorLab combo) is actually REINFORCED, just with a concrete, named, verified example rather

than an open question.

the operative default**, with two new promising-but-not-yet-actionable candidates added to watch

(SimArt — real but robotics-URDF-targeted; SATO — code not yet released).

benchmark roster.

Against NEOSTACK_AI.md:

unverified NeoStack categories" — CONFIRMED, no change to the core recommendation.

generation actually run") — RESOLVED: Meshy+Tripo (3D), ElevenLabs (voice), per §2.1-2.2.

quote on record literally names SIK/EIK, not AIK; every public AIK-specific source still shows

5.5-5.7 only. Strengthens rather than removes the existing "verify at first use" gate.

Against STACK_FACTS_QUESTIONS.md:

RE-READ**: the quoted Discord text names SIK/EIK specifically; no independent public source confirms

AIK/NeoStack itself is on 5.8 yet. This is presented as evidence for Josh's own re-check, not as an

autonomous override of a Josh-ruled item — per the standing decision protocol, a technical

verification finding surfaces here with full evidence; the ruling itself is Josh's to revisit or

reaffirm.

voice generation" — RESOLVED to specific vendors (§2.1-2.2), "4-5 agentic models" and "4 new

products/IDE" INFERRED-resolved to already-known quantities (§2.4) pending Josh confirming

against his own Discord screenshots if a sharper source is available.

as sharpest-geometry fallback... avoid Sparc3D... the JOSH-RULED commercially-unrestricted-by-default

posture" — CONFIRMED, no change; NeoStack's own bundled 3D-gen (Meshy/Tripo, both closed SaaS)

now explicitly folded into the "previz/reference only" tier of that same ruling, consistent with how

Meshy/Tripo were already treated pre-NeoStack.

CONFIRMED, per §4 — no vendor swap needed; two additive upgrades logged (Stable Audio 3.0's SFX

model, Kokoro for bulk NPC lines) for whenever that first audio build happens.

Against RUNTIME_GENERATIVE_LAYER.md / RUNTIME_GUARDRAIL_ARCHITECTURE.md:

"never a content dependency" prestige-if-2030-hardware-allows tier — CONFIRMED, no change to the

ruled anchor.

PARTIALLY SUPERSEDED: still real, valid, currently-shippable options, but no longer the

sharpest available at each tier — refreshed roster in §3.6 (adds Gemma-4, Ministral-3, Qwen3.6-27B,

Nemotron-3-Nano-4B as newly-relevant or newly-licensed-clean options).

(CUDA 12.8, not 13.1) and an additive Blackwell-tier optimization path (NVFP4/MXFP4) not previously

known.

numbers now show explicitly why 32B can't co-reside with a demanding renderer TODAY, giving the

existing "never a content dependency" hedge a measured basis rather than a projected one.

shipped, Apache-2.0, on-device, UE5-integrated architecture close to this doc's own independent

design. This is a genuine addition requiring a scoping decision (evaluate-as-reference vs

build-on-top vs bespoke-regardless) that this brief flags for Josh's ruling rather than resolving

unilaterally, given it has real engineering-time stakes for how the runtime-generative-layer build

actually starts.

Against RESEARCHED_STACK.md's residual-set triage:

the right posture**, roster expanded per §1.4.

evidence for why** (§2.5).

evaluation, the Stable-Audio-3.0-weights-license check) should be added to it.

---

Appendix: primary sources fetched directly (VERIFIED tier, this lead researcher only)

Appendix: secondary/aggregator sources used (INFERRED tier, this lead researcher's own searches)

3daistudio.com state-of-AI-3D-2026 · pixazo.ai open-source-3D leaderboard · trellis-2.com/trellis2.app/

trellis2.org (vendor-adjacent) · hunyuan3d.cc (already flagged unreliable by the sibling brief) ·

various 2026 SEO/aggregator roundups on "best open source 3D model" and "state of 3D generation 2026" ·

YouTube titles (Pixal3D VRAM claims) · s2026.conference-schedule.org (tool did not return populated

session data at fetch time) · kesen.realtimerendering.com SIGGRAPH papers tracker (referenced, not

directly fetched) · aggregator search snippets for NeoStack v2.0.45 (uncorroborated, flagged as

conflicting) · various RTX 5090 local-AI benchmark blogs.

Appendix: sub-lane research (Local LLMs §3, Voice/Audio §4) — full source lists in each sub-agent's

own report; headline primary sources independently re-verified by the lead researcher:

Generated by harness/site/structure_site.py — the URL path is the repo path. review root