How Tripo Is Tackling Clean Topology for its 3D Asset Pipeline
Tripo Co-Founder and Chief Scientist Dr. Yanpei Cao discusses Smart Mesh, generating usable topology in seconds, giving artists more control over generated assets, Gaussian Splats, production-ready 3D, and the shift from individual asset generation toward complete interactive worlds.
Generating a 3D model from an image has become dramatically easier over the past few years, but generating an asset is not the same thing as producing something a professional artist can actually use and iterate upon to achieve their artistic vision. AI-generated models can arrive with messy topology, difficult-to-edit materials, inconsistent geometry, and other problems that turn an impressive-looking result into considerable cleanup work before it can enter a production pipeline.
For Tripo, solving that gap between generation and usability has become one of the company's primary goals. Co-Founder and Chief Scientist Dr. Yanpei Cao comes to that problem from a background in geometry processing, reconstruction, and 3D digitization. We spoke with Dr. Yanpei Cao about how that technology is evolving from simple model generation toward production-ready assets.
Can you introduce yourself and tell us a little bit about the company, your career, and how it all built up into Tripo?
Dr. Yanpei Cao, Chief Scientist and Co-Founder at Vast (Tripo AI): I got my bachelor's and PhD degrees from Tsinghua University, and at Tsinghua I did geometry processing, geometry understanding, and reconstruction. After graduation, I co-founded a company called Owlii, and later it got acquired by Kuaishou.
Back then, we were doing 3D digitization, like capturing and reconstructing dynamic environments. That experience taught me how to transfer a research project into a system that our users can use. But during the process, I became convinced of the bottleneck in 3D creation. It's slow, it's fragmented. Only a small fraction of highly expert artists can do that. I wanted to change that, and then generative AI emerged, and some of the founders and I decided to use generative AI to build a company, build a platform, to lower the barrier of 3D creation, let more users, more developers generate the 3D assets they need.
How would you describe Tripo to someone unfamiliar? Is it a tool for 3D artists? Is it a generative tool? Who is the customer? And what are the main features, or the main pain points that you're addressing?
Dr. Yanpei Cao: Currently, we see Tripo as both an AI native system and a series of AI 3D foundation models. We can provide tools and platforms to creators from the game industry, 3D printing industry, or even robotics, to any industry or vertical that needs 3D assets or 3D environments; they can use Tripo to create the kind of content they need.
We don't just stop at generating assets; we want to make them useful and stay functional, meaning that after generating 3D assets from a single prompt image or multiple images, the platform and our models also do the rest, like separating the 3D models into meaningful parts, doing the texturing, rigging, and even animation.
There are a number of companies that allow you to generate a 3D asset from an image. From my perspective, one of the biggest pain points is making sure that it's not just a mesh, but it actually has a rig and can be animated. It has materials that can be edited. It has the ability to add shaders and other elements. Can you tell us about the way that you can add creative control into what you're doing, and how you can work with it? Are most of the tools inside the software that you're building, or can it be connected with other software? How does Tripo live in the pipeline?
Dr. Yanpei Cao: I think I'll use one or two examples. Currently, we are developing our models into two pathways. The first is high-fidelity, high-quality models, where our AI model outputs highly detailed 3D models, up to a few million polygons. That was the more traditional pathway, but this year we introduced Tripo Smart Mesh, which is backed by our research at SIGGRAPH, called Nexus. So by using Smart Mesh or Nexus, we created a new paradigm for artists to create usable, game-ready assets from single images, which means the assets won't be a messy polygon soup; they'll have clean edge flows.
This will make them easier to animate, or they can control the polygon budget easily and make the model very fast. Tripo Smart Mesh can output game-ready assets in under 5 seconds, end-to-end. That changes the process a lot, because now game developers can send 20, or even 50, prompts or directions to the model and wait just a few seconds, then collect the results and decide which direction to go. That acceleration of the iteration process is a huge thing, and also the result is having usable topologies with bridges to send directly to software like Blender, Maya, or Unity.
By doing that, we are not building something around artists. We are building tools for the artists that cut out the cleanup as much as possible, and we do the initial modeling and let the artist decide the art direction.
When someone is working with Tripo, what is the input? Is the input just a text prompt, or can you add reference images? Can you add your concept art?
Dr. Yanpei Cao: We want to make the whole process as controllable as possible. You can use text, one single reference image, or even multiple images like an entire character sheet of multiple references, or even some zoomed-in angles of the character. We are developing more, like bounding box control, aspect ratio control, and even 3D to 3D if you have an old asset that you want to remaster.
One of the biggest challenges I think with AI-generated assets, especially 3D, is that you lose definition in a lot of details in materials, but also in geometry; a lot of definition is kind of mushy. I'm wondering, how do you address this problem? The materials that you're sharing are always like very crisp, very high-level, so I'm wondering: are these out-of-the-box results? Or do you do some extra work with them and clean them up? How are you solving this problem to make sure the geometry is crisp?
Dr. Yanpei Cao: We do it both ways. First, the purely automatic way we are pushing our models to build in-house tech from the ground up. The AI model can now compress very high-resolution geometry into a compact asset around 2,000 by 2,000 by 2,000 voxels, compressed into a latent state to generate the high-resolution geometry.
We are also building tools around artists that allow them to iterate the outputs together with the AI models, like we developed Magic Brush, which allows users to quickly tweak the details of geometry and texture materials as a collaboration between AI models and artistic intent.
You’ve worked a lot with Gaussian Splatting in the past. Are you using that background and that research in the stuff that you're generating now? You’re generating characters and props with Tripo now, but what about generating more complex entire environments, complete scenes, and so on? What's your take on that, and what are the problems with generating that type of content?
Dr. Yanpei Cao: The advances in research on things like NeRFs and Gaussian Splats have helped us a lot, and during the process we're still developing the ability for models to directly output 3D Gaussian Splats as a final format, because it's lighter and easier to render. Even for physical objects, we are collaborating with some manufacturing labs that can print out the 3D Gaussian Splats so that you can print out the Gaussian Splats. That's crazy! I mean, you can see the colors and geometry at the same time.
That's one thing we are developing, and the important thing is that we are open sourcing a lot of our progress, like TripoSplat, which we open sourced under the MIT license, and we made it as light as possible, so once it was out, it had day zero support by Comfy, and a lot of users are already using TripoSplat, our open source project, to create their assets as 3D Gaussians. Another thing is, of course, the environment we are building. I think two things: first, we are collaborating with companies like World Labs, through hackathons and other events, so users can see how to integrate polygon meshes with Gaussian Splats, the thing World Labs' Marble is building, and integrate it into a harmonious, unified experience.
And another thing is we are also iterating our world models to expand the ability to build environments. We're collaborating with some robotics companies and simulation companies to expand our ability to output simulation-ready assets and simulation-ready environments so that robots can directly manipulate things, like putting things down or picking things up, in our generative environments.
What would you say are the biggest challenges that AI is currently facing in terms of R&D? Is it geometry, materials, definitions, or cost? What are the biggest challenges, and how do you see this developing in the future? Only recently have we started seeing prompts to 2D to 3D pipelines, and a lot of the output that you're getting looks very good, but there are of course issues with control, hallucination, and so on.
Dr. Yanpei Cao: I think one thing for us is still the usefulness, or the production readiness, of the generative assets. Like we see geometry and topology, or wireframes, as two different things. Geometry is more continuous, and topology, or wireframes, are more discrete, like what connects to what. The topology or wireframes are what most artists are looking at right now. Most of the time, with Tripo Smart Mesh, we've already made a huge step forward, but we are putting more development resources into iterating in that direction. We want the models to output, or understand, the topology thinking behind the artists. That's one thing we are trying to build.
Model-wise, we do need to add more controls, allowing users to input arbitrary views of the reference images, and we're trying to figure out how to integrate the agentic side of the creative process. Currently, our APIs are a bit scattered, functionality-wise, as endpoints. But we are building agentic flows that allow users to input the final art direction and let the agents do as much as possible to create a workflow dynamically, on the fly, using AI, so the artists can focus even more on the art direction, not on the heavy lifting or chores.
You talked about the output. What is the time it takes to generate a complete, usable asset? For example, if you sit down to start working and you create a prompt and you have some reference images, what is the time to asset generation for a character that can be used in a game? Like, a full-blown 3D character?
Dr. Yanpei Cao: We have two types of models. We have the high-definition or H series. For this one, the time cost from your input prompt into a fully texturized character is probably still under one minute, and that's a high-poly model. For Smart Mesh, it's even faster. You can get the final output in under 10 seconds. It's a character, not that detailed, but with controllable polygons and usable edge flows in under 10 seconds. It would be around 20,000 polygons.
How are you charging for this? Is it a license model? Is it the token model?
Dr. Yanpei Cao: It's a token model. You spend some tokens to get the assets.
For my last question, what do you expect to see in the industry in the next few years from NVIDIA, for example, with what they’re doing, or what all these other companies are doing? What are you most excited about?
Dr. Yanpei Cao: For future challenges, there are a few main things. First is how the industry, more broadly, goes from single-asset generation into environments, into interaction and dynamics. That needs two things. First, high-quality asset generation, and second, we need to integrate code or logic to structure the whole environment, or to give the interaction more logic, and the third thing is world models. We need a model to predict what's happening in the next few minutes. So that's three things together, and that's what we are hoping the whole industry is developing toward. So by achieving that, I think the creative process can be revolutionized. Furthermore, artists can really focus on the creative direction, rather than spending time creating assets or coding and scripting the logic.