Type a description, get a textured 3D model in seconds - no reference photo, no sketch, no 3D experience required. Here's how it works and how to prompt it well.
In the real Tripo AI generator, this becomes an editable, textured 3D mesh you can refine and export.
Generate the real thing →Text to 3D is the input mode where you describe an object in plain language and Tripo generates a textured 3D mesh from that description alone - no photo, no drawing, nothing but words. It's the lowest-friction way into AI 3D creation: if you can write a sentence, you can generate a model. Under the hood, a 3D foundation model interprets your prompt the way an image generator interprets a caption, inferring shape, proportions, style, and materials simultaneously, then outputs a mesh with PBR textures already applied.
It's the natural starting point for anyone without a reference image - concept exploration, game props imagined from scratch, characters that exist only in your head - and it's also the fastest way to iterate, since there's no photo to stage or reshoot between attempts. The trade-off, covered in full below, is that text alone gives the model less to work with than a photo does, which is exactly why prompting technique matters more here than in any other input mode.
1. You write a prompt. A short natural-language description - object, style, materials, details. 2. The model interprets intent. Tripo's foundation model (v3.1/H3.1 by default) parses the text into an implied 3D shape and material set, drawing on patterns learned from vast 3D and image data. 3. Geometry and texture generate together. Unlike a two-stage pipeline, Tripo produces base mesh and PBR texture in the same pass, which is part of why results return in seconds. 4. You get an editable result. The output drops into the same refinement tools as any other input mode - segmentation, Magic Brush texturing, auto-rigging, and export - so text-to-3D isn't a lesser starting point, just a different one.
Write it like a brief to a junior artist, in this order, and skip nothing:
One object per prompt. The generator makes a single coherent asset per run - "a knight and his horse and a castle" forces it to compromise on all three. Compose scenes afterward in your engine or DCC tool, one clean asset at a time.
Generic, ambiguous - the model must guess style, era, and materials from nothing.
Adds era and category, but still leaves shape, condition, and finish undefined.
Names the object, style, materials, and wear - every clause gives the generator something concrete to render.
Mix and match across these four categories - you rarely need more than one or two words from each.
"a medieval longsword with ornate gilded crossguard, worn leather grip, weathered steel blade"
"a cute chibi robot companion, rounded plastic shell, glowing blue chest panel, stubby arms"
"a cracked ceramic vase with blue floral relief pattern, rustic terracotta base"
"a low-poly stylized pine tree, flat-shaded foliage, game-ready for a forest scene"
"a small stone watchtower with a conical wooden roof, moss on the lower stones"
"a small horned dragon whelp, leathery wings folded, glossy scales, curious pose"
They're not competitors so much as answers to different starting points. Use text when you have an idea but no reference - pure concept work, fantasy creatures, stylized props that don't exist yet. Use an image when you have one - a photo of a real object almost always reconstructs more accurately than describing it in words, because the model has actual geometry to work from instead of an inference. If you have both a rough idea and a loose reference, sketch-to-3D or multi-view can split the difference.
| Situation | Better mode | Why |
|---|---|---|
| Pure concept, nothing exists yet | Text to 3D | No reference to work from - this is the only option |
| You have a real object to digitize | Image to 3D | A photo carries real geometry; description carries only inference |
| Fast iteration on many directions | Text to 3D | No staging or reshooting between attempts |
| Precise proportions matter | Image (ideally multi-view) | Text can't specify exact dimensions reliably |
| Stylized / fantastical subject | Text to 3D | Nothing to photograph - describe the style directly |
Precise proportions. "A 15cm tall figure with a 3:1 leg-to-torso ratio" doesn't reliably translate - text gives the model style and character, not exact measurements. For precise dimensions, model in CAD or start from an image and scale explicitly. Multi-object scenes. As above, one asset per prompt is the reliable unit; scene composition happens afterward. Very specific real-world objects. Describing "my grandmother's 1962 Singer sewing machine" won't reconstruct that exact machine - text conjures a plausible interpretation of the category, not a specific real instance. For that, image-to-3D with an actual photo is the right tool. Organic complexity. Faces, hands, and complex creature anatomy remain the frontier for every AI 3D generator in 2026, text-driven or not - expect to iterate more on these subjects regardless of input mode.
A standard text-to-3D generation costs the same as any other base generation - roughly 25 credits in Studio, which is where the "~200 free credits ≈ 8 models" and "~3,000 Professional credits ≈ 120 models" math comes from. Because text prompts cost nothing to "reshoot" the way a photo does, it's the cheapest input mode to iterate on: refine the wording and regenerate rather than spending on a heavier pass when the shape isn't right yet. Full plan breakdown and a plan-picker calculator live in our pricing guide.
The free tier covers dozens of text-to-3D generations - enough to find your prompting style in one sitting.
Generate free →