Skip to main content
Ren Wang

Ren Wang

Jun 21, 2026

24:16
So- Overall, is there a sense or is there an idea, or do you believe that with physical AI as well, we'll get to a certain point where the robot will be good for a certain task that is w-widely used enough that will then-- So that will get good enough, and then with some RLHF of different sorts, whether in a, in a simulator or in a controlled environment, for one task, it will get good enough that will kick off a data flywheel?
24:44
Yeah.
24:44
So I think what you described is, in my head, i-in many ways, like the only path to actually realizing physical AI.
24:54
It will not happen in a simulator, though.
24:55
So, so if we come back to simulator land for a second, you know, it is true that RLHF enabled GPT-3 to go to GPT-3.5, and, and, and ChatGPT was basically helped by, by this particular approach, and you could say the same for the current wave of agentic systems as well.
25:17
Basically, there were really, really good, um, environments for software development and, and so on that, that people built, and then you could just do RL at scale with, with these language models inside these environments.
25:29
And if you look at the domains where this approach has been particularly successful, in particular coding and, uh, math, all of these domains share this property of being verifiable, which means that you, you can know very easily, very quickly if your answer was, was right or, or wrong.

7 MINS LATER

32:56
Does that even scale? Or like if you have a factory and a bunch of robots, are you gonna have like a GPU right there with them, or is it like really streaming somewhere else in someone else's cloud and coming back?

We value your privacy

We use cookies to understand how you use our platform and to improve your experience. Click “Accept All” to consent, or “Decline non-essential” to opt out of non-essential cookies. Read our Privacy Policy.