
Vision-language-action model
40
MENTIONS
7
EPISODES
5
PODCASTS
Search complete. 40 mentions across 7 episodes found for "Vision-language-action model".
Sep 6, 2026
XPENG Just Build A Humanoid Robot That Moves Almost Too Human | XPENG Iron
L
9:10LeoHOST
They studied the scaling law from 1,000 hours to 100,000 hours.
L
9:15LeoHOST
On thin tasks, 1,000 hours, their language condition VLA beats in-context learning.
L
9:21LeoHOST
So if you only have 1,000 hours of data, their VLA reaches 53% success rate, while in-context learning only reaches 43%.
L
9:30LeoHOST
If you go to 100,000 hours, the results become comparable.
L
9:33LeoHOST
So in-context learning starts worse and then catches up as data scales up because with more data, the model is able to generalize much better.
L
9:42LeoHOST
And now on unseen tasks, 100,000 hours in-context learning reaches 66% success rate.
L
9:48LeoHOST
Language prompting reaches 9%.
L
9:51LeoHOST
So you have a 7X gap between these two, right? So they say their model S1 exponentially outperforms any existing language prompted VLA as you scale up pre-training.
NVIDIA Bought Hugging Face and Never Said Robots
S
3:07ShubhamHOST
And LeRobot is that infrastructure.
S
3:09ShubhamHOST
Almost every open vision language action model, or VLA, that I have covered on this show gets pulled from that host.
S
3:17ShubhamHOST
It now belongs to the company that sells the GPUs those models are trained on.
S
3:22ShubhamHOST
now jensen huang's public commitments here are unusually specific and i want to give him credit for that the platform stays open to the whole ecosystem openweight models keep getting support it stays multi-cloud and it stays multi-accelerator and the sentence that matters most is that nvidia compute will not be required to build on or deploy through hugging face That is a much harder thing to promise than the usual language about respecting the community.
AI:AM — Soft Robotics: How Materials Sense and Adapt · September 4, 2026
T
57:49Timothy LeeGUEST
So one thing is long context.
T
57:50Timothy LeeGUEST
So the classic VLA, the way it would work is you would give it a camera image and a prompt, like pick up this object, and then it would generate a sequence of actions for the alert robot.
T
58:01Timothy LeeGUEST
And then after a second or two, it would start the whole loop over.
T
58:03Timothy LeeGUEST
And so there was no like long running state.
T
59:56Timothy LeeGUEST
And I think it's a similar kind of thing with robots, probably.
T
59:58Timothy LeeGUEST
It's my guess that obviously nobody knows.
T
60:00Timothy LeeGUEST
But my guess is it's going to be a few more years before you start to see kind of small scale, useful deployments of these VLAs.
T
60:06Timothy LeeGUEST
And then a few more years after that to kind of scale up and apply to new industries and bring the cost out and the reliability up and so forth.
Hands-on Physical AI with Kevin Cloutier
K
11:02Kevin CloutierGUEST
So it, it's how far up the chain do you go? And right now I'm at that, that chain that's in the, the OS level, in the RTOS, in the Ubuntu, in the Canonical, uh, level, et cetera.
C
11:13Chris GammellHOST
I, w-well, actually, it was kind of the VLA kind of thing.
C
11:15Chris GammellHOST
Seems like it's even another step up as well, right? So like-
K
11:17Kevin CloutierGUEST
Yeah
K
12:05Kevin CloutierGUEST
... demo that I'm, I'm bringing to AI Infra in Santa Clara in, in a couple weeks.
K
12:09Kevin CloutierGUEST
And, um, the idea is using vision language action To play Tic-Tac-Toe.
K
12:14Kevin CloutierGUEST
Now, for your listeners who are not familiar with VLAs, most of them, or at least the small sizes that I deal with, they're good at one task.
K
12:22Kevin CloutierGUEST
So you can train it to pick something up and put it somewhere else, right? Or you can train it to pick up, uh, let's say a cube and put it in row one, column one of, of a Tic-Tac-Toe board.
5 Take-Aways from the Unitree Filing (293)
J
1:36Jeff TowsonHOST
And with that, let's get into the topic.
J
1:40Jeff TowsonHOST
Now, the concepts for today that matter are, I've talked about these before, but really world models, uh, VLA models, which VLM models as well, but you know, embodied intelligence and sort of how those models work, which I've gone into some, in some depth.
J
1:59Jeff TowsonHOST
I'm gonna do a lot more this week.
J
2:00Jeff TowsonHOST
For those of you who get the email, I'm gonna send you some pretty technical breakdown of the embody AI that Tencent is working on.
J
2:20Jeff TowsonHOST
It's, it's more like a technical breakdown, so I'm not sure.
J
2:23Jeff TowsonHOST
I, I'm torn on whether that's helpful or not, but yeah, I think understanding how these models work, um, it's, it's super important.
J
2:31Jeff TowsonHOST
So anyways, that'll be the topics for today, VLA world models.
J
2:36Jeff TowsonHOST
We'll talk about Unitree, but I'll go through Tyro's Tencent's smart brain, um, as well briefly.
Google Just Gave Robots a Massive Intelligence Upgrade | Gemini Robotics
O
27:05OscarHOST
And this is kind of like the big player.
O
27:06OscarHOST
It's a vision language action model.
O
27:08OscarHOST
It's called Gemini Robotics 2.
O
27:10OscarHOST
And this one takes vision and language, so image and text, and then it outputs actions, aka motor controls of the robot.
Your Simulator Has Never Seen a Robot Fail
S
1:51ShubhamHOST
Google DeepMind shipped Gemini Robotics 2 at the end of July, and it's three models rather than one.
S
1:57ShubhamHOST
The one in the middle is a Vision Language Action Model, or VLA for short.
S
2:02ShubhamHOST
That's a policy, and what it does is take what the camera sees, plus an instruction in plain English, and put out motor commands.
S
2:09ShubhamHOST
It never has to predict what the scene will look like afterwards, and that is the fork in the road away from a world model.
S
3:33ShubhamHOST
The marketing doesn't lead with the bottom of that range, but they published it anyway, and that's worth something.
S
3:39ShubhamHOST
The thing I keep coming back to is that the most capable whole-body humanoid controller anybody shipped this summer contains no generative video world model at all.
S
3:48ShubhamHOST
It's a layered VLA fed by real teleoperated data.
S
3:53ShubhamHOST
And teleoperated means a person drives the robot by hand through the task while the recording runs, and that recording is what the model learns from.