[Ep. 097]

Blog Image

Teaching Robots by Watching: Jagdeep Singh on Rhoda AI and the End of the Teleoperation Arms Race

Duration:

40min

Listen on:

Jagdeep Singh has built and sold companies across three deep tech frontiers. Optical networking at Infinera. Solid state batteries at QuantumScape. Now robotic intelligence at Rhoda AI, where he is founder and CEO.

Rhoda AI is building a foundation model for physical AI, robots that learn how the world works by watching. Most of the field trains robots on tens of thousands of hours of teleoperation, humans puppeteering machines through every task. Rhoda flips that. It pre-trains on hundreds of millions of internet videos to learn the physics of how things move, then fine-tunes each real task with just 10 to 20 hours of robot data. The wager is simple. Internet video already holds the scale, the diversity, and the physics a model needs to generalize.

That wager pays off in the dirty, dusty, dangerous work most robots still cannot touch. Factories, warehouses, supply chains. Fewer than 10% of American factories run robotic solutions today, and plenty that want them cannot deploy, because a specialized robot demands too much surrounding infrastructure and too much patience. A generalist you can wheel in and train in a day changes that math. Here is our conversation with Jagdeep, lightly edited for clarity.

The conversation

Jay: You've built and scaled companies across three deep tech frontiers. Optical networking with Infinera, solid state batteries with QuantumScape, and now robotic intelligence at Rhoda AI. Each time, you made a bet against what the market was saying about the fundamental physics. So let's start there. What did you see with Rhoda that made you feel this company needed to exist?

Jagdeep: It starts with an observation that AI is changing everything. Over the next 10 years, I think the society we live in is going to be fundamentally different because of AI. So of course you want to be in AI. That said, the impact of AI so far has been limited to what I call the symbolic domain. You're working on bits and symbols, not atoms. Language models make a huge impact. They write text, they write code. We've seen image generation, video generation, deepfakes. These are amazing developments. But the ability of AI models to manipulate atoms is effectively zero right now. The number of deployed intelligent robots capable of manipulation in the real world is close to zero. You have robots doing locomotion, inspection, monitoring. People run models on quadrupeds for those applications. But for applications that involve hands and arms, manipulation, AI has not made an impact. The question we asked ourselves is why. If you could have AI manipulate objects, you could change the way the global economy works. Almost all production today requires human beings. The automation that exists, and has for decades, is largely point to point. If you could put AI in the physical world, it would transform industries and transform the way the global economy works. The problem is today's models don't work well in the real world. You've probably seen lots of robots doing cool things in videos.

Jay: Sure. Boston Dynamics, all of those viral videos.

Jagdeep: They do backflips, they fold a T-shirt, they make coffee. But all of those demonstrations are in a lab setting. In a lab you can stage the scenario you're testing to match very closely the data set the model was trained on.

Jay: You give the robot the answer to the test.

Jagdeep: Pretty much. You've pre-trained on exactly what you're going to test it on. The real world doesn't work that way. The distribution of data is much broader. There's a much longer tail of rare events, corner cases that lead to failure. You can't help that the distribution is different. What you need is a model capable of dealing with that distribution. That was the problem we saw, and it was big enough that we thought there was an opportunity to really change the world.

Jay: Your Gen zero model, your prototype, is built on partner hardware. But longer term the vision is to bring your next model onto your own design. How do you phase that? Should I think about it like the Android model for phones, something that could eventually run on any machine? Or is the long-term goal to be fully vertically integrated?

Jagdeep: We want to build the best solutions for our customers. Those solutions require a full stack. They require an AI model and they require hardware. Neither exists in the form we want, so we're doing both. That said, the critical unsolved piece is the AI model. We're an AI-first company. The model needs to be robust enough to work in the real world, and that's an unsolved problem. We believe we're addressing it. Until now we really haven't seen models capable of crossing the chasm between the lab and the real world. Our model is in good shape in terms of maturity, and there's no reason to wait until we have our own hardware to make it available to customers. So we run the model on off-the-shelf hardware. The robots you see here are off-the-shelf, commercially available robots. But with our model they can do many more useful things than they could without it. With our own hardware, the models can do even more, because we also can't find good enough hardware to do everything we have in mind. The model is the key value creation of the company. It can work on any hardware platform. We're hardware agnostic. We've already run it on multiple hardware types, and that lets it perform a wide range of industrial and commercial tasks. With our own hardware, we simply increase the sophistication of the tasks. Over time the model gets more powerful, and together with the hardware it does more and more interesting tasks.

Jay: Let's talk about capital intensity. Building a full stack solution takes an enormous amount of capital, and you're one of many players in the category. Just to get the numbers right: Physical Intelligence has raised 1.1 billion. Skild 1.5. Figure 2.3, maybe more now. You recently announced a raise of 450 million. With frontier language models, training was a capital expenditure game. People put in more money and the models got better. In physical AI, is it the same arms race, or should we think about capital differently?

Jagdeep: First, physical AI models are not nearly as big as language models. So you don't need the same level of compute to train them. That said, the companies you mentioned are all doing what are known as VLAs, the Vision Language Action paradigm. In the VLA paradigm you start with a language model and post-train it on teleoperation data, collected by teleoperating or puppeteering robots. You run the robot through certain trajectories doing tasks like making coffee or folding T-shirts, and when you collect enough trajectories you train the model on them. This approach has a couple of problems. One, to your point about capital intensity, you need an enormous amount of capital to collect teleoperation data. That's where a lot of the capital these companies raise is going. And in our view, a lifetime spent teleoperating robots is not going to be enough. It's a drop in the bucket compared to what you need to truly generalize, to learn a strong prior on physics. So there's a double problem. It's expensive to collect teleoperation data, and even more severely, you can never collect the scale or diversity of data required to truly generalize. You're spending a lot of money on what we believe is a dead-end path. We're doing something different. We asked ourselves, what source of data has the scale, the diversity, and the ability to learn physics, that's already out there and we don't have to generate ourselves? There's only one answer. Internet video. By some estimates, something like 80% of all data on the internet is video. These videos cover virtually every conceivable scenario, from human beings doing crazy things to scenes with no humans present. The canonical example we use is waves crashing on a beach. Even scenes with no humans have something to teach the model about the physics of how the world works. So we train our model using what we call a data pyramid. The bottom of the pyramid is very general data, of which there's a lot, maybe not even a human in the scene. As you go up the pyramid you get more and more task-specific data, culminating in a video of an actual human doing the actual task you want. And a bit counterintuitively, the models that perform best in our experience are the ones with less curated data, because less curated data forces the model to learn more about how the world works. Once you have that prior, you can apply it to a specific task. That's what leads to the high data efficiency we announced in our launch a couple of weeks ago. We can train models to perform very sophisticated tasks. For example, an automotive decanting task. These boxes are full of ball bearings. As we showed in our videos, the robot can lift the box, a 10 kilogram, 22 pound box, pull open the little pull tab, which is a dexterous action, pop open the lid, open the plastic bag inside, pull out the desiccant paper, put it in the trash, and dump the contents. It does all that on the order of 10 to 20 hours of robot data. In our view that's a high watermark, maybe a low watermark, for how much data is required. And all of it is enabled because the model has already been trained on hundreds of millions of videos. It's learned a really strong prior on how things move. Without that you'd have to collect tens of thousands of hours of data, and even then, as we see with VLA models, you can only do the task if you recreate exactly the data set it was trained on.

Jay: I want to come back to the 10 hours in a second, because that's a very interesting claim, especially compared to what some VLA companies need.

Jagdeep: It's not a claim, it's a fact. That's what I mean.

Jay: The fact that you've been able to do it with 10 hours is something I want to spend more time on. But first, you started talking about teleoperation as one of the ways these models are trained. It's not the only way. As I understand it, there's teleoperation, synthetic data, world models, and then the method you're talking about. Lay out the landscape. Where do these play, and where do you fit?

Jagdeep: Great question. There are a few main approaches to collecting data. We've talked a lot about teleoperation, where a human puppeteers the robot to perform tasks. Think of it as remote control. As the robot performs the tasks you collect that data and assemble it into a large data set you post-train the model on. The downsides of teleoperation: one, it's expensive to collect. More fundamentally, the size and diversity of those data sets is always limited, because you're only collecting data you intentionally collected. Almost by definition you're not seeing the things that are unintentional.

Jay: You'd never test the edge cases. You're only testing exactly what you see.

Jagdeep: Any edge case you test is, by definition, the specific edge case you're testing for, not all the other infinite ways things can go wrong. That's a very serious problem with teleoperation, and it's why we think teleoperation alone is not going to lead to true generalization. The second approach is simulation. People say, what if you didn't have to use real robots? What if you ran them in simulation and used the power of computation to get more data? You probably could get more data that way, because you're not dealing with physical atoms. But you have the exact same diversity problem. You're collecting data you intentionally collect. And diversity turns out to be an even bigger issue for generalization than pure scale alone. That's ignoring the sim-to-real gap, the idea that even with all our years of experience in simulation, you really can't capture the full richness of the real world. But ignore that for a second. Diversity alone is fatal. Another approach is synthetic data, where you use a video model to generate scenes. The problem is it's synthetic. Any mistakes the model makes relative to how the real world's physics work end up in the model that learns from it. So all of these approaches have serious drawbacks. The only approach that meets the three requirements, internet scale, internet diversity, and the ability to learn real physics, is internet video. We keep falling back to that. One analogy: this approach may be radical in the robot space, but it's not radical at all in the AI space. Every generalist AI model you use every day, language models, image models, video models, is trained the same way. The pre-training phase involves tens of trillions of tokens from that domain, basically an entire internet's worth of data. Then a small amount of post-training on a data set aligned closely to what you want the model to produce. That's fine-tuning. The combination of massive internet-scale pre-training and a small amount of fine-tuning is what's produced generality everywhere else. There's no reason to think the robotic space is any different. And if you need an internet's worth of data, there's only one source that exists. Internet video.

Jay: How do you think about defensibility, given that this is publicly available? On one hand I think about where value accrues. Nvidia is a good example. They made all the investment in the picks and shovels, and then value accrued to the inference layer and the application layer. Is the same thing going to happen in physical AI? Are there going to be moats in what you're doing?

Jagdeep: There are two things I think are moats. It comes down to what's hard about this problem, what's hard for people to replicate. One: yes, the data itself is accessible, that's how we access it. But to actually harness and ingest that data you need a set of techniques not typically used in robot learning. These come from computer vision and generative AI. They require that you understand physics to a high level of accuracy and can reproduce it in real time. We need to make video predictions of what's going to happen that are physically accurate, that don't involve hallucinations, and that happen in real time, in a few hundred milliseconds. We're not aware of another video model that can make video predictions that are both that accurate and real time. Ordinarily those two things run at cross purposes. To make a model more accurate you typically make it bigger, and if you make it bigger it runs slower, not faster. How do you make a model both more accurate and faster? That's the central challenge our team figured out, because of their backgrounds in computer vision and generative modeling. That's just the first layer of the moat. Once you're running autonomously in real factories, you start to see all the corner cases, all the edge scenarios. As you deal with those using a technique called Dagger, you're getting data on the long tail of the distribution. That data is proprietary. It's not available on the internet. Only a robot that's already autonomous and operating can collect it. Using that data makes the model more robust, which lets us deploy it more and collect more data. It kicks off a classic data flywheel. That's a much harder moat to replicate. At some point it becomes, in our view, a winner-take-all scenario. The model that gets good enough to see more of these corner cases is the only one people want to deploy, and the model people deploy is the one that gets more data to make it even more robust.

Jay: So another way you're unlocking deployment: by using the Rhoda model you get to a point where someone would want to live-deploy it, and the more live deployments you have, the more you can strengthen the model.

Jagdeep: You said it better than I did. I see now why you're a storyteller.

Jay: A bit of what I do for a living. So let's talk about the 10-hour data point, which is so interesting given how much of the industry focuses on teleoperation and the number of hours. You'll see claims from the industry: with 10,000 hours of training we did this, with 5,000 hours we did that. And you just told me it takes 10, maybe 20 hours to confidently do this task. Help me square how you get to that low number.

Jagdeep: I'm really glad you brought this up, because it's one of the most interesting things about the Rhoda story. If you look at the evolution of VLAs over the past few years, you'll see claims about how much data they were trained on. Some people claim 70,000 hours of robot data. The next says 270,000 hours of UMI-collected data. Some say 500,000 hours. Then you come to Rhoda, which is later in time, and we say we train on 10 to 20 hours of data. It's almost like, wait a minute, why are you not in this game of trying to increase the number of robot hours? The answer is that game just doesn't lead to generality, no matter how many hundreds of thousands of hours you collect. It pales in comparison to the amount of data and diversity you need to understand how the world works. So we changed the rules of the game. Our approach was to learn fundamental physics from an internet-scale dataset that already exists, in video. If we succeed at learning a good prior, then in principle we should be able to learn specific tasks with a very small amount of data. To be candid, we surprised ourselves by how little data we needed. We've now run multiple POCs in real factory settings. This particular setup was run in a live factory, meaning they're processing real material flow into real cars. By definition you're not controlling it. It's not staged for a lab. You're dealing with the boxes and straps and parts as they come through the material flow pipeline. The fact that we could do that with so little data collected was really gratifying.

Jay: The error correction piece you mentioned earlier, I'd love to understand better. You said it runs a prediction in real time of what should be happening and then, every hundred milliseconds or so, course corrects. Is that different from how typical VLA approaches work, or is it unique to Rhoda?

Jagdeep: Every VLA approach has to predict what the robot should do next. They're predicting actions, and those actions are not based on a video-based prediction. They're trying to learn how to predict what should happen next based on a limited number of hours of robot data. Remember, with a VLA the pre-training is entirely language and image based. There's no knowledge of how things move in the pre-training. All the motion understanding a VLA has comes from that teleop dataset, and there just isn't enough data there for the model to really learn how things work.

Jay: The variance in those predictions could be anywhere, because you don't have the right amount of data.

Jagdeep: Exactly. You can make a prediction if the scenario matches what you've seen before. But as things diverge and you haven't seen it, once the model goes out of distribution, the output is random. There's nowhere for it to come back into distribution. The beauty of pre-training on hundreds of millions of videos is that the model has seen so much it always has a fallback. I'll give you one example that, for us, was the moment we felt this thing is going to happen for real. We were in the factory at this automotive customer, showing one of their top executives a task. The robot was picking up boxes and boxing them. The executive was a bit skeptical. He picked up a glove lying nearby and threw it in there. We had never trained the model on picking up a glove out of the box. It saw the glove, reached in, and picked it up. Now the question is, which trash can did it put it in? There's a can for cardboard recyclables, one for plastic bag recyclables, and one for regular trash. It put the glove in the regular trash, because it wasn't a cardboard box and it wasn't a plastic bag. The executive was unconvinced. He picked up the glove again, put it into a plastic bag, and threw that into the container.

Jay: Now you're really trying to break it.

Jagdeep: Right. What the model did was pick up the plastic bag, grab it from the ends, shake the glove out through the plastic bag into plastic recycling, then pick up the glove and put that in the trash. At that point I thought, we've never trained on a task this far out of distribution. How did it figure out what to do? That's where we thought, okay, this pre-training is definitely having an effect.

Jay: So the generalization piece. This is the tension I hear from a lot of the robotics industry. You're talking about building for generalization, but so far we've really talked about very specific tasks. How do you marry the story you tell investors, which is that this is a path to generalization, with what the customer wants, which is very specific tasks?

Jagdeep: There really isn't any dichotomy. Think of human beings. This task is currently done by human beings. Those people are not specialists. They have the same kind of brain you and I have, the same kind of limbs and embodiment. They're just doing a very specific task. That's how we think of our model. The model itself is a generalist model running on generalist hardware. We're not customizing it for any one task. The specialization comes from the post-training. The same way you hire a worker and train them on how your specific company does a process, you hire this generalist model and train it on how your company does that process. Over time, unlike a worker, the model can learn more and more different tasks, to the point where it can operate zero-shot, meaning without any training at all.

Jay: Let's talk about scaling. The POC is very impressive. You've got a few deployments collecting real-world data. But even within this specific use case, how do you go from 5 to 1,000 to 10,000, so you can collect data at the scale you need to solidify your position?

Jagdeep: There are a few key elements of scaling. The first is demand. On the demand side, this is about as scalable a use case as you can get, because virtually every business has some variant of this task. Every business receives incoming materials in boxes or containers. Those materials have to be decanted, taken out of the boxes, processed, and shipped out again. The second thing you need is hardware capable of doing the task. Right now we're finding that off-the-shelf robot arms with 7 degrees of freedom, that can handle 10 kilograms per arm, that have impedance control, so torque sensors that detect how much force they're exerting, are more than adequate. The third thing is the AI model itself, probably the most important. That's the part that has never existed before. The final thing is the software that pulls everything together, the bridge software that gets data from the sensors, sends it to the model, gets data back, and uses it to control the actuators. All four require work, but none of it, in our view, is rocket science that hasn't already been shown. The part that was rocket science was the AI model, and that part we think is in good shape.

Jay: Of these four, which do you think is going to be the hardest or least predictable in the coming years?

Jagdeep: Ironically, if you'd asked me a year and a half ago when we started, I'd have said: is there enough demand for this kind of automation, and can the AI model handle the tasks? Both of those we've gotten tremendous validation on. The part I wouldn't have listed high is the software that pulls it all together. This is old-school software engineering, and it turns out there's a lot of work to be done there. This can't just be a research prototype where you fire up two or three computers with command-line interfaces. It has to be push-button. One button, you say go, and the model runs.

Jay: You're one of the few founders in physical AI who has scaled through many cycles. For the first-time founders listening, what are most founders overweighting versus underweighting when they build today?

Jagdeep: There are probably two pieces of advice I'd give your founders. The first is to focus on four things. These things are not invented, they're discovered. You have to find an opportunity that has all four. One, there has to be a really big unsolved problem. Two, you have to have a differentiated technology. If you're a commodity doing what everybody else does, you'll be in a commodity business with commodity margins. Three, you need an exceptional team. And four, you want early customer validation, particularly for deep tech. I look for customers who say not just this is interesting, but this is so interesting that if you could build it, I want to help you get it to market. The second point is this idea of being contrarian and right. You want to be contrarian, because if you're conventional the value is already priced in. But if you're contrarian and wrong, you're still wrong. So how do you make sure you're contrarian and right? The key is to think of your startup as simply a collection of hypotheses. Do not start drinking your own Kool-Aid. Think of your job as being a chief risk officer. List out the key hypotheses in order of risk, attack them one at a time, and try to de-risk the opportunity. As you de-risk those hypotheses, value will accrue.

Jay: One thing we didn't talk about enough is the gap between those who can have these solutions and those who cannot. One of the founders we had on the show recently runs one of the largest robotic fleets, focused on deployment in manufacturing. He told me less than 10% of factories in America currently have robotic solutions.

Jagdeep: Yeah. And many more want them.

Jay: But there's this massive gap between want and can actually implement. How does Rhoda help solve that?

Jagdeep: Deploying robots today does require a lot more than the robot itself. There's a whole infrastructure you need to put in place. But this is exactly what we're trying to solve. It's why we're going after generalist robots. The idea of a generalist robot is that you can pretty much wheel it in and it can do the task with 10 to 20 hours of training. You don't need a lot more infrastructure. We don't need to build mechatronic systems like conveyor belts around the robot. The robot can work in the same environment humans already work in.

Jay: So said another way, the deployment challenge is that it takes so long to make the thing useful that you're getting fewer and fewer people adopting, because they don't have the patience for it.

Jagdeep: Precisely. The more specialized your robot is, the more infrastructure is required to run it. The more generalist it is, the easier it is to roll out.

Jay: Can we close on form factor and where you see this going? Humanoids are obviously a big piece of the conversation. You've got naysayers who say, why do these things need feet? Why not wheels?

Jagdeep: Why not six arms?

Jay: Right, other form factors. What's the future you're building for?

Jagdeep: We don't have a fixation on the humanoid form for its own sake. Our focus is on solving real problems our customers have. It just turns out many of these use cases are being addressed by humans today. The environments, the objects, the whole setups are built for humans. So if you're building a machine to address these use cases, you'll likely need 7-degree-of-freedom arms to manipulate the objects. You'll need a vision system. You need some place to run the brain, the compute. And it starts to look a bit humanoid. Now, one thing we do not need, in our environments at least, is legs. So we're not building legs. We build robots that run on wheel bases, because our environments already have AMRs and AGVs running around. In fact, legs can be an issue. If you have to e-stop the robot, which is a requirement for manufacturing safety, a robot on legs is going to collapse and either damage itself or, worse, injure a human. Our robots look humanoid because that's the most functional form to build a generalist machine, but they will not conform to human dimensions purely for the sake of appearances.

Jay: I'd love to close on the year ahead. You just announced a major capital raise. What does the year ahead look like for Rhoda?

Jagdeep: It's an incredibly exciting year. We have multiple customer POCs that have already been successful in the field. The next logical step is to convert those POCs into real deployments and real revenue. If you look at the hierarchy of what robot development efforts have done over the years, a lot of people have made robots work in a lab setting and nothing more. Some have taken robots into the field and done some field testing, like a POC. Very few have gone from POCs into actual deployment. We're now at the phase where that's our main goal. If we can succeed at that, we think we're on track to change the world.

Jay: I'm so excited for it, not just as an investor but as someone who sees the massive labor shortages and opportunities across so many industries. Manufacturing, logistics, supply chain, operations across the world.

Jagdeep: Jay, a pleasure talking to you. You obviously have a deep understanding of robotics, which is always very nice to see.

Jay: Thank you so much for being here.

Pull quotes

  1. "The ability of AI models to manipulate atoms is effectively zero right now."

  2. "A lifetime spent teleoperating robots is not going to be enough. It's a drop in the bucket compared to what you need to truly generalize."

  3. "Even scenes with no humans have something to teach the model about the physics of how the world works."

  4. "Do not start drinking your own Kool-Aid. Realize the whole idea you have is just a bunch of hypotheses."

  5. "If you're contrarian and wrong, you're still wrong. So how do you make sure you're contrarian and right?"

Source

From CLIMB Episode 097 with Jagdeep Singh (Rhoda AI). Watch the full episode: https://youtu.be/QTRbW1YPrM4

Guest appearance

Share episode

Stay tuned for every episode

[ DIRTY JOBS Starts in: ]

46 : 07 : 59 : 03

[ section ]

become a sponsor

All rights reserved

DIRTY JOBS SUMMIT 2026

[ DIRTY JOBS Starts in: ]

46 : 07 : 59 : 03

[ section ]

become a sponsor

All rights reserved

DIRTY JOBS SUMMIT 2026

[ DIRTY JOBS Starts in: ]

46 : 07 : 59 : 03

[ section ]

become a sponsor

All rights reserved

DIRTY JOBS SUMMIT 2026