Can internet videos teach robots? Rhoda AI thinks they can—and it's betting big on that future.
Show More Show Less View Video Transcript
0:00
I'm going to throw in this piece of trash, which it's never seen before
0:03
We're going to see what happens. These robots didn't learn to move like this in a lab
0:11
They learned it by doing exactly what you're doing right now, watching video
0:16
I'm at the headquarters of Rota AI, where they're betting that the future of robotics
0:20
lies in hundreds of millions of internet videos. I know, you've seen robotic arms before
0:25
They've been working pretty well in factories for years. Most of these robots are simply doing a specific task they were programmed to do
0:32
Usually, that makes the robot great at doing that one thing in perfect conditions
0:36
But the problem is they're not so great at adapting when something happens the robot didn't expect
0:41
That's why Rhoda developed what it calls its direct video action model. What we realized we needed was a data set that was internet scale in size
0:49
and had the diversity of the internet and from which you could learn real physics. There's only one answer that we came up with, and that's internet video
0:55
At its core it teaches robots to see the world like we do It shows robots hundreds of millions of video clips of the real world People picking things up opening doors moving through space The idea is the robot learns how things move and how physics works before moving a single joint A video of waves crashing on a
1:12
beach has something to teach the model about physics. It has something to teach the model
1:16
about environments, about sand, about sun, about water. You might be wondering where exactly Rota
1:21
gets all those videos. I asked them and they would only say they train their model on publicly
1:25
available data. After that first round of training on public data, the model is retrained on smaller
1:31
amounts of robot data that ROTA generates in-house. That's when the model learns specific behaviors
1:36
ROTA says the DVA model continuously observes its environment and predicts what will happen next as
1:41
video. It then acts on those predictions and observes what actually happens. This process is
1:46
repeated every few hundred milliseconds, so the model is continuously learning. So this is a
1:51
decanting task for a major automaker in Germany, decanting these ball bearings that go on the drive shafts
1:58
For the first step is to take them out of these boxes. They 22 pounds in weight so you have to pick them up Then you have to do this kind of a dexterous motion where you have to open the box open up this plastic bag Again bags are hard to see for robots because they transparent and they deformable You have to go in
2:13
and reach and grab this piece of desiccant paper and put that into the correct trash
2:18
can. Then you have to dump the contents into this tote here. And having done that, you
2:23
have to now recycle the materials correctly. It has to then get rid of this plastic bag
2:28
It shakes out that piece of paper. Once it sees that it's a clean bag, it puts that into the plastic recycling
2:34
What Nishtuga done here is that the robot has to pick up these radars
2:39
set them up on that little inspection station where there's four cameras that inspect the radar and make sure there's damage-free
2:45
and then place them into that box. Most of the tasks you see here are learned with just 10 to 20 hours of robot training data
2:51
So what this robot is doing right now is it's trying to open up this box, and now it sees this trash in there
2:57
It hasn't seen this trash in this particular configuration before, but its job is to sort that trash
3:03
Now I been told to keep my distance from these robotic arms because they quite powerful but I really want to test this robot out So I going to throw in this piece of trash which it never seen before We going to see what happens
3:34
So it looks like it got it in the right bin. Rhoda's next step is teaching the model to reason
3:40
giving it the ability to predict more than just a single outcome. If we're making a video-based prediction
3:44
we can actually make a prediction of N different features and then score them on which one performs the best
3:51
and only execute the one that performs the best. And in that sense, we can almost do what Dr. Strange did in the movie
3:57
which is look at multiple features and only pick the one that has the right outcome
4:01
So Rhoda isn't publicly disclosing who they're working with yet, but they do tell me that they're working with industry leaders in manufacturing and automotive
4:08
If you enjoyed this video, don't forget to like and subscribe, and I'll see you in the future
#Celebrities & Entertainment News


