pipelines/RC4_EXTERNAL_LANDSCAPE_2026-08-08.md
Lane: RC4 of the WHY-SO-FAR-OFF post-mortem. READ-ONLY; nothing applied.
Research date 2026-08-08. Every claim below is tagged VERIFIED (fetched live, URL inline),
REPORTED (secondary source, named), or UNVERIFIED (could not confirm at source — say so).
---
DERIVED FROM:
- docs/HANDOFF_2026-08-05_ACCOUNT_SWITCH.md §8.14 (L2811-2842) — the charter, Josh's
verbatim grade, and RC4's exact wording: "EVALUATE NVIDIA-CLASS 3D (Edify 3D via
partner access; Cosmos is scene/video not character) + any Aug-2026 leader on
CHARACTER LIKENESS specifically - the bake-off scored symmetry/topology, never
likeness-to-exemplar." (L2833-2835)
- docs/pipeline_review/tech_research/PIPE_CHARACTER_MODELS_2026-08-06.md — read IN FULL.
This lane's job is to extend/contest it, not repeat it. Load-bearing lines used:
§3.1 "No open-weight general 3D generator's own gallery shows a photoreal human
face... That silence is the finding, not an absence of evidence."
§4 "Edify 3D - quad meshes with PBR, architecturally right, weights never released
and the hosted NIM preview retired."
§9 the Hunyuan territorial disqualification + the recorded lesson that the FIRST
licence read was incomplete.
§11 "Nothing in §2's bench is enrolled in a blind join" / Pixal3D+TRELLIS.2 never run.
- docs/HANDOFF §8.8 (L431-433, L456) — the char-survey lane landing: "TripoSG holds,
Hunyuan licence-killed".
CONTRACT ITEM 2 NOTE (required by the standing subagent contract):
harness/route.py carries NO work-kind for an external-model post-mortem — this lane
produces no canon and realizes no canon node. I therefore ran no route.py read-set.
What I read instead is named above. No canon node was substituted by a proposal.
NOT DERIVED (authored judgment, and why it had no canon home):
- The entire ranking in §7. Canon specifies the character, never the generator vendor.
This is an engineering read of the public record as of 2026-08-08.
- The claim in §1 that PIPE_CHARACTER_MODELS §3.1 searched the wrong family. That is my
finding against a proposal-tier doc, which the contract permits (proposals are
subordinate); it is grounded in named papers, not asserted.
---
1. §3.1's class finding was derived from the wrong search. The prior survey looked at
*general* image-to-3D galleries and concluded the field does not do characters. A
character-specific sub-field exists, with its own benchmark protocol that scores
facial likeness to the reference directly. We benched the wrong family. (§1)
2. Seed3D 1.0 (ByteDance) is a total omission from the prior survey and is the single
highest-expected-value open candidate — 1.5B params beating Hunyuan3D-2.1's 3B on
geometry, full PBR, and the vendor's own character/face example. Its licence is
UNVERIFIED and that is a blocking read, not a footnote. (§2)
3. **NVIDIA is confirmed closed for this use, and Josh's "open world generation one" is
the wrong tool class** — Cosmos 3 generates video of scenes, not character meshes;
Edify's NIM preview was retired 2025-06-06 and the only live path is partner-mediated
through Shutterstock. The prior survey's §4 holds and hardens. (§3)
4. **The one independent likeness-specific head-to-head in the public record ranks closed
SaaS above everything we run** — and flags TRELLIS 2 for *broken hair strands*, which
independently corroborates our own hair finding. (§4)
The honest counterweight, stated up front: not one paper or gallery in this entire
sweep shows a photoreal human face. DreamCharacter-1's gallery is stylized. So the
prior survey's *deeper* point survives even though its stated reason is wrong: a photoreal
child face is unevidenced across the whole field. Nothing below promises to close that.
---
§3.1 concluded from silence across TRELLIS.2 / Pixal3D / Step1X-3D / LATO.2 / AniGen /
Hunyuan3D-2.1 — all general object generators. Searching on *character likeness*
instead of *image-to-3D* returns a different literature that evaluates exactly our axis.
| Work | What it is | Likeness evidence | Weights / licence |
|---|---|---|---|
| DreamCharacter-1 (ByteDance Intelligent Creation, arXiv 2607.07817, 2026-07-08) | post-adaptation on a 3D foundation model: geometry post-training via preference optimization, texture post-training for occluded regions, inference acceleration | Beats every open baseline on ULIP/Uni3D — see table below. Human study on PANIC3D, 60 test images | No code/weights released. Paper CC BY 4.0. Gallery is stylized, not photoreal |
| StdGEN (CVPR 2025, hyz317/StdGEN) | semantic-decomposed character: separates body, clothes, hair as distinct components in ~3 min, via a Semantic-aware Large Reconstruction Model | This is literally the thing §3.1 said no generator does | DISQUALIFIED — checkpoints are "for research purposes only" (explicit triggered restriction) |
| CharacterGen (SIGGRAPH'24) | multi-view pose canonicalization → 3D character | the original of this line; StdGEN borrows its code | weakest scores of the group |
| IDOL (arXiv 2412.14963) | "Instant Photorealistic 3D Human Creation from a Single Image"; feed-forward, 100K multi-view subject dataset; avatars directly animatable and editable | photoreal *humans* is its explicit claim | MIT (REPORTED) — commercially permissive |
| Seed3D 1.0 (ByteDance) | the foundation model DreamCharacter-1 post-trains on | §2 | §2 |
DreamCharacter-1 Table 1 (geometry, ULIP / Uni3D) — VERIFIED from arxiv.org/html/2607.07817:
| Model | ULIP | Uni3D |
|---|---|---|
| CharacterGen | 0.5516 | 0.6486 |
| StdGEN | 0.6697 | 0.7278 |
| TRELLIS-1.0 | 0.8016 | 0.7953 |
| Hunyuan3D-2.0 | 0.8025 | 0.7968 |
| Hunyuan3D-2.1 | 0.8273 | 0.8195 |
| Pixal3D | 0.8276 | 0.8247 |
| TRELLIS-2.0 | 0.8329 | 0.8261 |
| DreamCharacter-1 | 0.8532 | 0.8497 |
Texture (Table 2): DreamCharacter-1 leads SSIM 0.9349 / LPIPS 0.0686 / FID 0.0319 /
CLIP-Sim 0.9576. No commercial baselines (Tripo/Rodin/Meshy) were compared.
Two things this table buys us even though the model is unavailable:
above Hunyuan3D-2.1. The prior survey listed TRELLIS.2 as an "untested upgrade lead"
(§11); this is published evidence that the lead is real *on characters specifically*.
It also directly supports RC3's charted first test.
character table Pixal3D sits *below* TRELLIS-2.0, inverting the §3.2 ranking.
The benchmark protocol is the reusable asset (RC5's answer). A character benchmark in
this line evaluates 100 character reference images on dimensions general benchmarks do not
score — **facial likeness, anatomical hand correctness, and full-body fidelity to the
reference — using VLM- and CLIP-based protocols, and for the face it selects a
VLM-verified camera distance that keeps the full face visible, then evaluates frontal and
profile views** (REPORTED, via search synthesis; the exact source paper was not pinned —
see §8 OWED). That camera rule is precisely the thing our turntable lacked. **RC5 should
copy this protocol rather than invent a likeness lens from scratch.**
---
Absent from PIPE_CHARACTER_MODELS entirely. It is the backbone of the current SOTA
character paper.
Seed3D 1.0 (ByteDance Seed) — VERIFIED from seed.bytedance.com blog:
metalness, and roughness"**.
detail preservation. Texture/material claimed SOTA. Human eval, 14 evaluators,
"consistently higher ratings across all dimensions."
character generation example (facial features and fabric textures) where Seed3D 1.0
"accurately reconstructs facial features" while baselines "tend to lose reference
consistency." That is the one vendor face-likeness claim I found in the open-weight tier.
ByteDance-Seed's house licence is Apache-2.0 across Seed-OSS, SeedVR2 and Seed2.0. But
I could not confirm at file level — raw.githubusercontent.com/Seed3D/Seed3D/main/LICENSE
404s (branch name likely differs), the HF model card 401s, and the GitHub HTML render
came back incomplete. This is exactly the §9 Hunyuan trap — that lane read §6(d),
concluded clean, queued a 14 GB download, and found §5(c) afterwards. Do not spend a GPU
hour on Seed3D before the LICENSE file is in hand and pinned to docs/licence_records/.
Seed3D 2.0 — released 2026-04-23, API-only on Volcano Engine. Coarse-to-fine
geometry, unified PBR material model, MoE texture detail, VLM material priors,
part-level generation, articulated assets, scene composition. Blind human eval texture
preference >69%. Part-level + articulated is architecturally the answer to §3.1's
"one fused implicit surface" complaint — but there are no weights, so it is an API lane
only, and no character/face evidence was found for 2.0 specifically.
---
Cosmos is confirmed scene/video, not character. Cosmos 3 launched 2026-05-31 (GTC
Taipei keynote at Computex): an open physical-AI omnimodel that generates text, images,
video and ambient sound plus robot action signals; 16B Nano and 65B Super. It is trained on
physical scenes — factories, warehouses, roads, manipulation workspaces — to produce
physically plausible video. It emits no character mesh. For character *motion* NVIDIA
points at ARDY (SIGGRAPH 2026, text→3D human/humanoid motion, real time), which the
prior survey already has under Apache code + Open Model License weights.
→ **Josh's "nvidias open world generation one applied to this" names a real and impressive
model that is the wrong tool class for a character asset.** That should be said plainly and
without hedging in the post-mortem; it is not a dodge, it is the answer.
Edify 3D — the partner access path, precisely. VERIFIED: NVIDIA's own editor's note,
dated 2025-06-06, reads that Edify is **no longer available as an NVIDIA NIM
microservice preview**, directing users to build.nvidia.com for other visual AI models.
No weights were ever released. The surviving path is partner-mediated, not NVIDIA-direct:
exclusively on Shutterstock content including >500K ethically-sourced 3D models —
the "commercially safe / ethical" positioning is the product.
Read for us: the prior survey's §4 verdict stands unchanged — *NVIDIA has no local,
commercially-licensed, weights-downloadable character mesh generator.* The Edify path is a
third-party stock-media API with no local run, no character-likeness evidence, and a
training corpus of stock 3D objects. It is not a likeness route. RC4 closes NVIDIA.
---
Vitalify Asia, July 2026 — seven image-to-3D tools, the same portrait image,
default settings, no post-processing. This is the only public head-to-head I found that
judges *facial similarity to a reference portrait* rather than object IoU.
| Tool | Facial similarity | Texture | Hair | Clothing |
|---|---|---|---|---|
| Prism 3.1 | best — "facial structure closely matched the original image" | sharp, detailed | "remarkable detail", slight strand distortion | best preserved |
| Tripo P1 | "impressive accuracy" reconstructing facial features | sharp, detailed | — | — |
| Meshy 6 | — | — | volume preserved, no individual strands | — |
| Hunyuan3D | — | clean/consistent but lacked depth | — | — |
| Trellis 2 | — | — | "broken hair strands" | — |
| Rodin 2.5 | — | — | — | significant loss of clothing information |
| Forge | worst — "distorted facial features" | heavy noise, artifacts | — | — |
Overall: Prism 3.1 for realism; Tripo P1 best speed/quality balance, called
"especially suitable for game development workflows."
How much weight this carries: n=1 per tool, sighted, single portrait, published by a
dev-services company — by our own §2.1 standard this is a LOOK, not a scored arm, and
it is enough to justify a probe, not to ratify a swap. Two things still make it the most
useful external datum in this report:
1. It is the only source that scores the axis Josh graded us on.
2. Its TRELLIS 2 hair finding — *broken strands* — was reached independently of us and
agrees with our own measured hair result (§2 of the prior survey found Hunyuan
resolves hair as flat sheets while TripoSG resolves strands). Two independent observers
landing on hair as the failure axis is worth more than either alone.
Note on identity: "Prism 3.1" is a Tripo model surfaced under 3D AI Studio's naming
(35 credits/generation there). Treat Prism 3.1 and Tripo P1 as the same vendor family.
---
Standing Josh ruling keeps these reference-tier. Recorded because RC4 was asked to rank
them and because one of them is architecturally the closest fit to our actual problem.
Rodin Gen-2.5 (Hyper3D) — the best *architectural* match in the entire sweep:
model in the prior survey is single-image conditioned. This alone is a structural
advantage on likeness.
24k [CHAR_0001_CANON_SHEET.md L192-198]. A quad output inside that budget removes
the entire retopology stage that §3.1 named as the class blocker — the "~2 M-triangle
marching-cubes shell with no edge loops" problem simply does not occur.
symmetry forcing, micro detail control, HD texture enhancement. Exports GLB/FBX/OBJ/USDZ.
Tripo P1 / Prism 3.1 — native 3D diffusion, engine-ready, clean low-poly topology,
full PBR set (albedo, normal, metallic, roughness). Licence: **Free plan grants NO
commercial use**; paid plans (Pro $19.90/mo, 3000 credits) grant broad rights to use,
distribute and derive revenue from outputs. Explicit, readable, and adequate for us.
CSM (Common Sense Machines) — strong on organic shapes and characters; all paid plans
include commercial rights; API access. Acquired by Google, January 2026 (REPORTED) —
worth noting beside the prior survey's other Google datum, GNM Head published 2026-07-28.
Sparc3D → Hitem3D — commercialized by Math Magic; Sparcubes/Sparconv-VAE at 1024³,
Hitem3D "1536p Pro". Possible correction to the prior survey, which recorded that
Sparc3D's weights were never released and the demo was pulled: a source states the GitHub
repo (lizhihao6/Sparc3D) ships pretrained weights, and a HF Space exists. **Low confidence,
UNVERIFIED** — flagged for a five-minute check, not asserted.
Hunyuan — corroborated, no change. 2.5 / 3.0 / 3.1 are hosted-only and were never
open-sourced; 2.1 (June 2025) remains the last open release. The prior survey's §9
territorial disqualification of 2.1 is therefore still the whole story, and there is no
newer open Hunyuan to reconsider.
---
| Candidate | Local on 5090? | Basis |
|---|---|---|
| Seed3D 1.0 | Very likely — best fit in the sweep. 1.5B params, i.e. half Hunyuan3D-2.1's 3B, and 2.1 measured 7.63 GiB on this box. Expect well under 16 GiB | prior survey §2 measured VRAM; Seed3D param count VERIFIED |
| TRELLIS.2 | Yes — 24 GB stated, fits | prior survey §3.2 |
| IDOL | Likely; feed-forward single-image, MIT | REPORTED |
| StdGEN | Yes technically — but licence-blocked (research-only checkpoints) | VERIFIED |
| Seed3D 2.0 | No — API-only, Volcano Engine | VERIFIED |
| DreamCharacter-1 | No — no weights | VERIFIED |
| Edify 3D | No — no weights, NIM retired | VERIFIED |
| Rodin 2.5 / Tripo P1 / Prism 3.1 / CSM / Hitem3D | No — closed SaaS | VERIFIED |
Standing hazard that applies to every open candidate here: sm_120 is untested for all
of them, and the prior survey §5 recorded gsplat PR #1034 (5090 support) still open with an
sm_120 illegal-memory-access issue filed 2026-08-04. Budget for a build fight.
---
1. Seed3D 1.0 — highest expected value, GATED ON A LICENCE READ. Only open-weight
candidate with a vendor face-likeness claim; full PBR; smallest model; and it is the
backbone the SOTA character paper chose to build on. If the licence is Apache-2.0 as
the secondary sources say, this is the model to bench against TripoSG immediately.
2. Rodin Gen-2.5 — best architectural fit, closed. Multi-view input matches our plate
stack; quad topology at 4K–50K lands inside the 24k LOD0 budget and deletes the
retopology stage that §3.1 identified as the class blocker; texture delighting is a
stage we already run. Buy one generation and look.
3. **Tripo P1 / Prism 3.1 — the measured likeness leaders, closed, licence-clean on a paid
tier.** Cheapest possible empirical answer to "is our chain the problem, or is 3D just
not there yet."
4. **TRELLIS.2 full PBR — already charted as RC3's first test, and now externally
corroborated** as the strongest *open* model on the character axis (0.8329/0.8261,
above Pixal3D and Hunyuan3D-2.1). Run it. Expect hair to be its weak axis.
5. IDOL (MIT) — reference tier for the photoreal-human/body half; animatable output.
6. DreamCharacter-1 / StdGEN / Seed3D 2.0 — architecture proof-points only. They show
the field's answer is post-training + semantic decomposition + part-level generation,
none of which we can download today.
DISQUALIFIED, with the clause:
never opened.
---
most load-bearing unverified claim in this report. Zero GPU cost. Do it first.
VLM-camera-distance protocol described in §1 came through search synthesis; I could not
attribute it to a specific arXiv ID with confidence (the name collides with an unrelated
LLM character-customization benchmark and a 4D animation benchmark). RC5 should pin it
before copying the protocol.
on characters. Any comparison we make has to be run by us.
§2.1 an A/B is only an A/B when both arms are re-rendered by the instrument doing the
comparing. Nothing here is a measured result.
---
Chasing a better generator may be treating the wrong root cause. RC1 and RC2 say the
chain degraded a good sculpt and grafted a bare cage head onto it; RC3 says the native
texture was never even shown. If any of those hold, then a new generator changes nothing —
we would feed a better mesh into the same quality-destroying assembly and grade the same
turntable again. **The cheapest decisive experiment in this whole post-mortem is still
RC3's** (render TRELLIS.2's native PBR straight, beside the assembled chain), because it
costs no licence read, no purchase and no new dependency, and it discriminates between
"our generator is weak" and "our assembly throws quality away." RC4's candidates only
become the answer *after* RC3 shows the native output is also short.
The counter-argument for doing both in parallel: the two §5 probes (Rodin, Tripo) cost
roughly $20–40 total, need no GPU, no install and no sm_120 fight, and would put a
likeness-leading external result beside our own plate within an afternoon. That is a
genuinely cheap upper bound on what the field can do with our exemplars — and if the paid
tools *also* miss the plates, that reframes the entire post-mortem away from tooling and
toward the head/hair/skin stages, which is a far more valuable finding than a model swap.