Rodin Image-to-3D Generator
Turn one photo - or up to five - into a textured, production-ready 3D model. Here's exactly how Rodin's image-to-3D pipeline works: what it needs from your source image, how Fuse and Concat modes differ, which settings actually change the output, and how to get a cleaner result on the first try.
From photo to 3D model, in one pass
Rodin's image-to-3D generator reads a photo's depth, silhouette, structure and surface detail, then builds matching geometry and a matching texture in the same generation. Hyper3D's own framing is simple: "one image is enough for Rodin to generate most objects." Extra reference views are optional, not required - they exist purely to give the model more visual information when a single shot leaves something ambiguous (the back of an object, a hidden side, fine surface detail).
Speed is close to instant at default settings: base geometry typically returns in about 4 seconds, with a fully textured model ready around 5 seconds in. Every model ships with PBR texture maps out of the box, so there's no separate texturing pass before export. This page covers the mechanics in depth - the companion Image to 3D product page has the short version, and the Gen-2.5 guide covers the underlying model powering all of this.
Three steps, start to export
- Upload your reference image(s). JPG, PNG and WebP are all accepted. One photo works; add up to four more angles if you want extra guidance.
- Generate. Pick an effort tier, and - if you uploaded more than one image - a condition mode (Concat or Fuse, covered below). Rodin analyzes depth, silhouette, structure and surface detail and produces geometry and PBR textures together.
- Preview, refine and export. Review the result, use Partial Edit or a redo if something's off, then export to whichever format your pipeline needs.
Fuse and Concat: two different jobs
When you upload more than one image, Rodin needs to know what relationship those images have to each other. That's what the condition mode controls:
Concat mode (default)
Treats your images as multiple views of the same object - front, side, back, whatever angles you have - and combines them into one coherent model. This is the mode to use for the common case: you have a few photos of one thing and want the most accurate possible reconstruction of it.
Fuse mode
Extracts and blends features from different objects across your images rather than treating them as one subject. Useful when you want to combine design elements from separate references into a single new asset - for example, pulling a body shape from one reference and a surface pattern from another.
Full guidance on framing, image count and when multi-image actually helps is on the dedicated multi-image reference generation page.
What actually improves the output
Rodin doesn't require a clean studio photo - resolution, framing, background and lighting are optional capture choices, not submission requirements. That said, a few habits reliably produce better geometry:
Front-facing, centered subject
Works best as your primary image, especially for a single-image generation.
Angles for complex objects
Multiple views from different angles produce dramatically better geometry than one shot alone.
Clutter in the background
A neutral background and clean, even lighting make the subject easier to isolate.
If you don't have a photo of what you're trying to model, generating a clean reference image with an AI image tool first - then feeding that into Rodin - is a documented workaround. See the prompt-writing guide for how to describe an object so a generated reference comes out usable.
Match the tier to the asset
| Tier | What it's for |
|---|---|
| Extreme-Low | Smallest, fastest output - rough iteration and concept checks |
| Low | Light meshes for background and secondary assets |
| Medium | Default, balanced tier for most production work |
| High | Denser meshes with crisper geometric detail, for close-up assets |
| Extreme-High | Maximum detail, including micro-detail only available at this tier |
A common workflow: iterate on Extreme-Low or Low to lock in the right input and framing, then promote just the final version to a higher tier for the deliverable - see the full Gen-2.5 effort tier breakdown for exact timing per tier.
Topology, texture modes, and character options
Mesh topology
Two topology choices cover most needs: Quad topology is cleaner for sculpting, rigging and downstream animation; Triangle topology produces smaller files and renders faster in real time. Face-count range runs from roughly a few thousand up into the hundreds of thousands depending on tier, with a RAW ceiling past 10 million polygons at maximum effort - see the low-poly guide for taking a dense result down to a game-ready count with Smart Low-Poly.
Texture modes
Every model generates with PBR maps (albedo, roughness, metalness) by default; a Shaded mode and a combined "All" mode are also available, or textures can be skipped entirely. Two extra options matter for production: hdTexture for enhanced post-processed detail, and textureDelight, which removes baked-in lighting from the source photo so the asset can be relit cleanly in your own scene instead of carrying the original photo's shadows with it.
Character-specific option
For humanlike subjects, a T/A-pose option outputs the character in a standard rigging pose automatically, rather than in whatever pose the reference photo happened to show - useful when the model is headed straight into an animation pipeline.
Faithful vs creative geometry
Switch to creative mode when the reference image is meant as a starting point rather than an exact target - it gives the model more room to deviate from the photo. Keep the default, more literal mode when you need the output to match the reference as closely as possible.
ControlNet for image-to-3D
The same three spatial guidance tools available across Rodin apply to image-driven generation: a Bounding Box ControlNet to constrain overall scale and silhouette, a Voxel ControlNet for coarse volumetric structure, and a Point Cloud ControlNet for guiding generation from existing 3D scan data alongside your photo. These matter most when a single image leaves proportions ambiguous and you already know the rough shape you're targeting. Full detail on all three is in the Gen-2.5 complete guide.
Formats supported out of the image-to-3D pipeline
| Format | Typical use |
|---|---|
| GLB / glTF | Web, real-time engines, default export |
| FBX | Game engines, DCC round-tripping, rigged characters |
| OBJ | Universal static-mesh interchange |
| USDZ | AR/VR, Apple ecosystem previews |
| STL / 3MF | 3D printing |
| DXF | CAD-adjacent 2D/3D interchange |
Full details on what each format does and doesn't carry (rigs, materials, UVs) are on the export formats page.
Which one should you use?
| Situation | Better fit |
|---|---|
| You have a real photo or reference of the object | Image-to-3D |
| You're starting from a concept with no existing reference | Text-to-3D |
| You need the output to match a specific real-world item closely | Image-to-3D |
| You want fast concept exploration across many variations | Text-to-3D, then refine with image-to-3D |
| You have design references from multiple sources to combine | Image-to-3D, Fuse mode |
See the dedicated Text to 3D guide for the prompt-driven side of Rodin, and the prompt-writing guide for getting good results from either path.
What people are turning into 3D models
Product photography
Turn a catalog photo into a 3D asset for AR or e-commerce
Character reference art
Concept sketches or portraits into riggable T/A-pose characters
Physical objects
Scan-free 3D printing of real props, parts and collectibles
Game & VFX props
Reference photos into engine-ready assets with Smart Low-Poly
See a full production pipeline example in the game development workflow guide.
Image-to-3D walkthroughs
Every Rodin guide on rodin3ds.com
Image-to-3D & related features
Getting started
Image-to-3D: frequently asked questions
How many images does Rodin need to generate a 3D model?+
Just one. Rodin's Image-to-3D generator can produce most objects from a single photo. Adding up to 5 reference images from different angles is optional but gives the model noticeably better geometry and detail, especially on complex or asymmetric objects.
What's the difference between Fuse and Concat mode?+
Concat mode (the default) treats your uploaded images as multiple views of the same object and combines them into one coherent model. Fuse mode instead extracts and blends features from different objects across the images, useful for combining design elements from separate references into a single new asset.
How fast is Rodin's image-to-3D generation?+
Base geometry typically returns in around 4 seconds, with a fully textured model ready in about 5 seconds at default settings. Actual time varies by the effort tier selected, from Extreme-Low for quick iteration up to Extreme-High for maximum detail.
What image should I use for the best result?+
A single image works, but a front-facing, centered subject with a neutral background and even lighting gives the most reliable results. For anything with detail on multiple sides, 3 to 5 angled reference views produce noticeably better geometry than one shot alone.
Does the generated model come with textures?+
Yes. Every model includes PBR texture maps (albedo, roughness, metalness) by default, generated directly from the reference image's colors and materials, so there's no separate texturing step required before export.