Skip to main content
John Schulman

John Schulman

Computer scientist and researcher

Sep 11, 2026

25:23
Mm.
25:23
-that, that are just, um, like involve, uh, like doing a much more complicated task or doing something that requires a lot more cleverness.
25:31
And, uh, you could, you could say this is like the benchmarking distribution, 'cause a lot of the most prominent benchmarks just involve doing some very hard puzzle-like task, uh, that's easy to verify.
25:42
And then there's sort of, uh, like the realism axis, where you want the model to be good in the realistic coding agent setting, where there's like multiple back and forth with the human and there's like multiple objectives.
25:52
And, uh, like I'd say, um, like the people-- like the labs who are crafting the model behavior for the first time need to push in both directions.
26:03
And to get good model behavior, you need to really push on the realism axis and have like rubrics or some kind of human feedback that's informing, uh, the reward function you use there.
26:15
Um, but, uh, I think when, if you try to do distillation naively, you end up just sort of matching the teacher on the benchmarking distribution.

23 MINS LATER

49:02
Um, yeah, I don't know if you guys have thoughts on this.

We value your privacy

We use cookies to understand how you use our platform and to improve your experience. Click “Accept All” to consent, or “Decline non-essential” to opt out of non-essential cookies. Read our Privacy Policy.