Skip to main content
Vision-language-action model

Vision-language-action model

Search complete. 40 mentions across 7 episodes found for "Vision-language-action model".

Sep 6, 2026

LeoHOST
9:10
They studied the scaling law from 1,000 hours to 100,000 hours.
LeoHOST
9:15
On thin tasks, 1,000 hours, their language condition VLA beats in-context learning.
LeoHOST
9:21
So if you only have 1,000 hours of data, their VLA reaches 53% success rate, while in-context learning only reaches 43%.
LeoHOST
9:30
If you go to 100,000 hours, the results become comparable.
LeoHOST
9:33
So in-context learning starts worse and then catches up as data scales up because with more data, the model is able to generalize much better.
LeoHOST
9:42
And now on unseen tasks, 100,000 hours in-context learning reaches 66% success rate.
LeoHOST
9:48
Language prompting reaches 9%.
LeoHOST
9:51
So you have a 7X gap between these two, right? So they say their model S1 exponentially outperforms any existing language prompted VLA as you scale up pre-training.
ShubhamHOST
3:07
And LeRobot is that infrastructure.
ShubhamHOST
3:09
Almost every open vision language action model, or VLA, that I have covered on this show gets pulled from that host.
ShubhamHOST
3:17
It now belongs to the company that sells the GPUs those models are trained on.
ShubhamHOST
3:22
now jensen huang's public commitments here are unusually specific and i want to give him credit for that the platform stays open to the whole ecosystem openweight models keep getting support it stays multi-cloud and it stays multi-accelerator and the sentence that matters most is that nvidia compute will not be required to build on or deploy through hugging face That is a much harder thing to promise than the usual language about respecting the community.
Timothy LeeGUEST
57:49
So one thing is long context.
Timothy LeeGUEST
57:50
So the classic VLA, the way it would work is you would give it a camera image and a prompt, like pick up this object, and then it would generate a sequence of actions for the alert robot.
Timothy LeeGUEST
58:01
And then after a second or two, it would start the whole loop over.
Timothy LeeGUEST
58:03
And so there was no like long running state.
Timothy LeeGUEST
59:56
And I think it's a similar kind of thing with robots, probably.
Timothy LeeGUEST
59:58
It's my guess that obviously nobody knows.
Timothy LeeGUEST
60:00
But my guess is it's going to be a few more years before you start to see kind of small scale, useful deployments of these VLAs.
Timothy LeeGUEST
60:06
And then a few more years after that to kind of scale up and apply to new industries and bring the cost out and the reliability up and so forth.
Kevin CloutierGUEST
11:02
So it, it's how far up the chain do you go? And right now I'm at that, that chain that's in the, the OS level, in the RTOS, in the Ubuntu, in the Canonical, uh, level, et cetera.
Chris GammellHOST
11:13
I, w-well, actually, it was kind of the VLA kind of thing.
Chris GammellHOST
11:15
Seems like it's even another step up as well, right? So like-
Kevin CloutierGUEST
11:17
Yeah
Kevin CloutierGUEST
12:05
... demo that I'm, I'm bringing to AI Infra in Santa Clara in, in a couple weeks.
Kevin CloutierGUEST
12:09
And, um, the idea is using vision language action To play Tic-Tac-Toe.
Kevin CloutierGUEST
12:14
Now, for your listeners who are not familiar with VLAs, most of them, or at least the small sizes that I deal with, they're good at one task.
Kevin CloutierGUEST
12:22
So you can train it to pick something up and put it somewhere else, right? Or you can train it to pick up, uh, let's say a cube and put it in row one, column one of, of a Tic-Tac-Toe board.
Jeff TowsonHOST
1:36
And with that, let's get into the topic.
Jeff TowsonHOST
1:40
Now, the concepts for today that matter are, I've talked about these before, but really world models, uh, VLA models, which VLM models as well, but you know, embodied intelligence and sort of how those models work, which I've gone into some, in some depth.
Jeff TowsonHOST
1:59
I'm gonna do a lot more this week.
Jeff TowsonHOST
2:00
For those of you who get the email, I'm gonna send you some pretty technical breakdown of the embody AI that Tencent is working on.
Jeff TowsonHOST
2:20
It's, it's more like a technical breakdown, so I'm not sure.
Jeff TowsonHOST
2:23
I, I'm torn on whether that's helpful or not, but yeah, I think understanding how these models work, um, it's, it's super important.
Jeff TowsonHOST
2:31
So anyways, that'll be the topics for today, VLA world models.
Jeff TowsonHOST
2:36
We'll talk about Unitree, but I'll go through Tyro's Tencent's smart brain, um, as well briefly.
OscarHOST
27:05
And this is kind of like the big player.
OscarHOST
27:06
It's a vision language action model.
OscarHOST
27:08
It's called Gemini Robotics 2.
OscarHOST
27:10
And this one takes vision and language, so image and text, and then it outputs actions, aka motor controls of the robot.
ShubhamHOST
1:51
Google DeepMind shipped Gemini Robotics 2 at the end of July, and it's three models rather than one.
ShubhamHOST
1:57
The one in the middle is a Vision Language Action Model, or VLA for short.
ShubhamHOST
2:02
That's a policy, and what it does is take what the camera sees, plus an instruction in plain English, and put out motor commands.
ShubhamHOST
2:09
It never has to predict what the scene will look like afterwards, and that is the fork in the road away from a world model.
ShubhamHOST
3:33
The marketing doesn't lead with the bottom of that range, but they published it anyway, and that's worth something.
ShubhamHOST
3:39
The thing I keep coming back to is that the most capable whole-body humanoid controller anybody shipped this summer contains no generative video world model at all.
ShubhamHOST
3:48
It's a layered VLA fed by real teleoperated data.
ShubhamHOST
3:53
And teleoperated means a person drives the robot by hand through the task while the recording runs, and that recording is what the model learns from.

We value your privacy

We use cookies to understand how you use our platform and to improve your experience. Click “Accept All” to consent, or “Decline non-essential” to opt out of non-essential cookies. Read our Privacy Policy.