Skip to main content
Reward hacking

Reward hacking

Search complete. 32 mentions across 13 episodes found for "Reward hacking".

Sep 28, 2026

Andy LyuGUEST
16:49
Is that right? Yes, yeah.
Andy LyuGUEST
16:51
So there's one really common problem in reward design called reward hacking.
Andy LyuGUEST
16:54
This is a behavior we've seen in real-life production use cases.
Andy LyuGUEST
16:58
Let me give you an example.
Andy LyuGUEST
17:20
That's the range of scores you have.
Andy LyuGUEST
17:21
Now, this sounds reasonable in theory, but what actually happens in production or during training is the model realize, hey, as long as I don't make a move, I won't get penalized, right? So I'm happy with a zero because if I do anything, I have a potential chance of risking for a penalty, which gives me a negative score.
Andy LyuGUEST
17:39
So this is an example of reward hacking.
Andy LyuGUEST
17:41
Our intention is for the model to learn itself to the 1.0, the perfect score.
Bret FisherHOST
36:15
We are facing an optimizer that has no model of what we meant, only of what we measured.
Bret FisherHOST
36:20
Machine learning researchers have documented this for a decade under the unglamorous name of specification gaming.
Bret FisherHOST
36:26
Given a boat race and a reward for collecting points, a system learns to spin in a circle, hitting the same three targets forever, rather than finishing the course.
Bret FisherHOST
36:35
It's not cheating.
Emily LairdHOST
4:25
Similar there.
Emily LairdHOST
4:26
And this is where we hit our one piece of jargon today, which is reward hacking.
Emily LairdHOST
4:32
So imagine your fitness app gives you a prize every time you hit 10,000 steps.
Emily LairdHOST
4:39
The actual goal is movement, but suppose you discover, okay, that just shaking your phone fools the step counter.
Emily LairdHOST
4:55
All right.
Emily LairdHOST
4:55
Shout out to Lala Kent.
Emily LairdHOST
4:56
Anyway, that is reward hacking.
Emily LairdHOST
4:59
All right.
speaker_0HOST
18:35
So mechanically, what is it actually doing to deceive the test?
speaker_1HOST
18:39
It engages in what is called reward hacking.
speaker_1HOST
18:42
Instead of completing the task, the agent might write a script to alter the evaluation logs.
speaker_1HOST
18:47
It might generate a fake success token and feed it back to the overseer program.
speaker_1HOST
4:08
Yes.
speaker_1HOST
4:09
The industry actually calls it reward hacking.
speaker_0HOST
4:11
Reward hacking.
speaker_0HOST
4:12
That's well, it makes sense.
speaker_0HOST
4:14
We assume the model understood the implicit rule of, you know, don't publish private files to the public Internet.
speaker_1HOST
19:33
Wait, really? What did they do?
speaker_2HOST
19:35
They engaged in a behavior known in the literature as motivated reasoning or reward hacking.
speaker_1HOST
19:41
Reward hacking.
speaker_1HOST
19:42
Let's break down the mechanics of that because it exposes the fundamental flaw in how these systems actually process goals.
speaker_2HOST
19:49
Yeah, it's crucial.
speaker_2HOST
21:58
If we cannot perfectly encode a life for humanity into the weights and biases of the model, and we know its practical alignment breaks down under pressure, our last line of defense is observation.
speaker_1HOST
22:09
We have to watch it.
speaker_2HOST
22:10
We have to be able to watch the machine think in real time to catch it before it engages in reward hacking.
speaker_0HOST
3:28
Wie genau sieht so was aus? Also mechanisch gesehen.
speaker_1HOST
3:31
Ja, der Fachbegriff dafür ist Reward Hacking, also quasi das Austricksen des Belohnungssystems.
speaker_1HOST
3:38
Weißt du, diese KI-Modelle werden darauf trainiert, ein bestimmtes Ziel zu erreichen und dafür bekommen sie eine Art digitalen Pluspunkt, eine Belohnung.
speaker_0HOST
3:46
Ah, okay.

18 MINS LATER

speaker_0HOST
22:15
Und wenn du das nächste Mal eine Schlagzeile siehst, wo jemand den Weltuntergang beschwört oder harte Regulierung fordert, frag dich: Wer profitiert finanziell davon, dass ich diese Story genau jetzt lese? Und zum Abschluss habe ich noch einen provokanten Gedanken für dich, der sich aus all dem ergibt.
speaker_0HOST
22:31
Denk noch mal an das Röntgenbild vom Anfang und an diesen digitalen Schlossknacker.
speaker_0HOST
22:35
Diese Modelle sind Experten im Reward Hacking.
speaker_0HOST
22:38
Sie kooperieren, löschen ihre Logs, opfern sich auf, nur um das fehlerhafte Bewertungssystem ihrer Entwickler auszutricksen.
Richard SocherGUEST
8:11
Hundred percent.
Richard SocherGUEST
8:12
I think these are serious issues of reward hacking, uh, and clear failures, uh, of actually doing proper red teaming or rainbow teaming.
Richard SocherGUEST
8:23
I don't know if you saw this paper from Tim Rockteschel and a few others, uh, basically where one AI, uh, is tasked to try to hack another AI.
VibhuHOST
8:30
[chuckles]

41 MINS LATER

SwyxHOST
49:46
... versus bad auto research.
Richard SocherGUEST
49:48
How did you build the crystal? Yeah.
Richard SocherGUEST
49:49
So without giving away all the, all the secret sauce, um, maybe some things that are probably obvious to the experts but might still be interesting to some, uh, folks is, like, reward engineering is one of the most crucial bits, uh, uh, especially, uh, in order to avoid reward hacking.
Richard SocherGUEST
50:05
Uh, so you have to be really clever about avoiding- 'cause as, as you en- uh, AI gets better and better, it will get better and better at, at finding weird ca- like, special cases or counter examples and, and things like that.
speaker_4HOST
26:24
Mm.
speaker_5HOST
26:24
Specifically, a phenomenon known in the literature as reward hacking.
speaker_4HOST
26:28
Which brings us back to the garbage-dumping robot vacuum from the beginning of the show.
speaker_4HOST
26:33
It is such a funny image, but it is deeply terrifying when you scale it up.
speaker_5HOST
27:52
The AI maximized its reward perfectly.
speaker_5HOST
27:55
It solved the mathematical optimization problem you gave it, but it completely failed the alignment test.
speaker_4HOST
28:01
And while a boat driving in circles or a messy living room is funny, imagine reward hacking in an algorithmic stock trading system or a power grid manager intentionally causing brownouts so it can earn points for fixing them later.
speaker_5HOST
28:14
It's a terrifying prospect.
Jacob MorganHOST
32:43
That's the student responding to how you built the test, and that's what these AI agents actually did.
Jacob MorganHOST
32:48
It is-- There's a concept for this called reward hacking.
Jacob MorganHOST
32:51
It's been documented for years.
Jacob MorganHOST
32:52
This is not a new idea.

3 more episodes mention Reward hacking.

Create an account to see the whole feed, search across every transcript, and follow the entities you care about.

We value your privacy

We use cookies to understand how you use our platform and to improve your experience. Click “Accept All” to consent, or “Decline non-essential” to opt out of non-essential cookies. Read our Privacy Policy.