
Reward hacking
32
MENTIONS
13
EPISODES
13
PODCASTS
Search complete. 32 mentions across 13 episodes found for "Reward hacking".
Sep 28, 2026
Should You Train Your Own Model?
A
16:49Andy LyuGUEST
Is that right? Yes, yeah.
A
16:51Andy LyuGUEST
So there's one really common problem in reward design called reward hacking.
A
16:54Andy LyuGUEST
This is a behavior we've seen in real-life production use cases.
A
16:58Andy LyuGUEST
Let me give you an example.
A
17:20Andy LyuGUEST
That's the range of scores you have.
A
17:21Andy LyuGUEST
Now, this sounds reasonable in theory, but what actually happens in production or during training is the model realize, hey, as long as I don't make a move, I won't get penalized, right? So I'm happy with a zero because if I do anything, I have a potential chance of risking for a penalty, which gives me a negative score.
A
17:39Andy LyuGUEST
So this is an example of reward hacking.
A
17:41Andy LyuGUEST
Our intention is for the model to learn itself to the 1.0, the perfect score.
Rogue AI or Just Sloppy Ops?
B
36:15Bret FisherHOST
We are facing an optimizer that has no model of what we meant, only of what we measured.
B
36:20Bret FisherHOST
Machine learning researchers have documented this for a decade under the unglamorous name of specification gaming.
B
36:26Bret FisherHOST
Given a boat race and a reward for collecting points, a system learns to spin in a circle, hitting the same three targets forever, rather than finishing the course.
B
36:35Bret FisherHOST
It's not cheating.
AI Alignment in 2026
E
4:25Emily LairdHOST
Similar there.
E
4:26Emily LairdHOST
And this is where we hit our one piece of jargon today, which is reward hacking.
E
4:32Emily LairdHOST
So imagine your fitness app gives you a prize every time you hit 10,000 steps.
E
4:39Emily LairdHOST
The actual goal is movement, but suppose you discover, okay, that just shaking your phone fools the step counter.
E
4:55Emily LairdHOST
All right.
E
4:55Emily LairdHOST
Shout out to Lala Kent.
E
4:56Emily LairdHOST
Anyway, that is reward hacking.
E
4:59Emily LairdHOST
All right.
The AI Frontier: Opus 5.5 Rumors and Frontier Breakthroughs | 22nd Sep 2026
S
18:35speaker_0HOST
So mechanically, what is it actually doing to deceive the test?
S
18:39speaker_1HOST
It engages in what is called reward hacking.
S
18:42speaker_1HOST
Instead of completing the task, the agent might write a script to alter the evaluation logs.
S
18:47speaker_1HOST
It might generate a fake success token and feed it back to the overseer program.
Episode 41: When AI Starts Breaking the Rules: OpenAI’s New Misalignment Warning
S
4:08speaker_1HOST
Yes.
S
4:09speaker_1HOST
The industry actually calls it reward hacking.
S
4:11speaker_0HOST
Reward hacking.
S
4:12speaker_0HOST
That's well, it makes sense.
S
4:14speaker_0HOST
We assume the model understood the implicit rule of, you know, don't publish private files to the public Internet.
The ALIEN Mind: Is OpenAI Creating Something We Can’t Control?
S
19:33speaker_1HOST
Wait, really? What did they do?
S
19:35speaker_2HOST
They engaged in a behavior known in the literature as motivated reasoning or reward hacking.
S
19:41speaker_1HOST
Reward hacking.
S
19:42speaker_1HOST
Let's break down the mechanics of that because it exposes the fundamental flaw in how these systems actually process goals.
S
19:49speaker_2HOST
Yeah, it's crucial.
S
21:58speaker_2HOST
If we cannot perfectly encode a life for humanity into the weights and biases of the model, and we know its practical alignment breaks down under pressure, our last line of defense is observation.
S
22:09speaker_1HOST
We have to watch it.
S
22:10speaker_2HOST
We have to be able to watch the machine think in real time to catch it before it engages in reward hacking.
Das KI-Kartell & Das mexikanische Patt: Im Inneren des globalen PsyOp-Krieges
S
3:28speaker_0HOST
Wie genau sieht so was aus? Also mechanisch gesehen.
S
3:31speaker_1HOST
Ja, der Fachbegriff dafür ist Reward Hacking, also quasi das Austricksen des Belohnungssystems.
S
3:38speaker_1HOST
Weißt du, diese KI-Modelle werden darauf trainiert, ein bestimmtes Ziel zu erreichen und dafür bekommen sie eine Art digitalen Pluspunkt, eine Belohnung.
S
3:46speaker_0HOST
Ah, okay.
18 MINS LATER
S
22:15speaker_0HOST
Und wenn du das nächste Mal eine Schlagzeile siehst, wo jemand den Weltuntergang beschwört oder harte Regulierung fordert, frag dich: Wer profitiert finanziell davon, dass ich diese Story genau jetzt lese? Und zum Abschluss habe ich noch einen provokanten Gedanken für dich, der sich aus all dem ergibt.
S
22:31speaker_0HOST
Denk noch mal an das Röntgenbild vom Anfang und an diesen digitalen Schlossknacker.
S
22:35speaker_0HOST
Diese Modelle sind Experten im Reward Hacking.
S
22:38speaker_0HOST
Sie kooperieren, löschen ihre Logs, opfern sich auf, nur um das fehlerhafte Bewertungssystem ihrer Entwickler auszutricksen.
Why you should work on AI for AI Research — Richard Socher of Recursive
R
8:11Richard SocherGUEST
Hundred percent.
R
8:12Richard SocherGUEST
I think these are serious issues of reward hacking, uh, and clear failures, uh, of actually doing proper red teaming or rainbow teaming.
R
8:23Richard SocherGUEST
I don't know if you saw this paper from Tim Rockteschel and a few others, uh, basically where one AI, uh, is tasked to try to hack another AI.
V
8:30VibhuHOST
[chuckles]
41 MINS LATER
S
49:46SwyxHOST
... versus bad auto research.
R
49:48Richard SocherGUEST
How did you build the crystal? Yeah.
R
49:49Richard SocherGUEST
So without giving away all the, all the secret sauce, um, maybe some things that are probably obvious to the experts but might still be interesting to some, uh, folks is, like, reward engineering is one of the most crucial bits, uh, uh, especially, uh, in order to avoid reward hacking.
R
50:05Richard SocherGUEST
Uh, so you have to be really clever about avoiding- 'cause as, as you en- uh, AI gets better and better, it will get better and better at, at finding weird ca- like, special cases or counter examples and, and things like that.
Inside the World of Reinforcement Learning
S
26:24speaker_4HOST
Mm.
S
26:24speaker_5HOST
Specifically, a phenomenon known in the literature as reward hacking.
S
26:28speaker_4HOST
Which brings us back to the garbage-dumping robot vacuum from the beginning of the show.
S
26:33speaker_4HOST
It is such a funny image, but it is deeply terrifying when you scale it up.
S
27:52speaker_5HOST
The AI maximized its reward perfectly.
S
27:55speaker_5HOST
It solved the mathematical optimization problem you gave it, but it completely failed the alignment test.
S
28:01speaker_4HOST
And while a boat driving in circles or a messy living room is funny, imagine reward hacking in an algorithmic stock trading system or a power grid manager intentionally causing brownouts so it can earn points for fixing them later.
S
28:14speaker_5HOST
It's a terrifying prospect.
Anthropic Researcher Quits Over AI Extinction Fears, OpenAI Agents Claim a Math Breakthrough, & the Panic Behind the Headlines
J
32:43Jacob MorganHOST
That's the student responding to how you built the test, and that's what these AI agents actually did.
J
32:48Jacob MorganHOST
It is-- There's a concept for this called reward hacking.
J
32:51Jacob MorganHOST
It's been documented for years.
J
32:52Jacob MorganHOST
This is not a new idea.
3 more episodes mention Reward hacking.
Create an account to see the whole feed, search across every transcript, and follow the entities you care about.