Skip to main content
LessWrong

LessWrong

Search complete. 106 mentions across 83 episodes found for "LessWrong".

Sep 14, 2026

Joe WeisenthalHOST
36:27
bloomberglive.com/screentime/radioDo you have a theory for why some of the big, um ...
Joe WeisenthalHOST
36:37
Now, you know, I finally succumbed and I occasionally read the LessWrong message boards now-
Tracy AllowayHOST
36:42
[laughs]
Joe WeisenthalHOST
36:42
... as, like, a new stage in my life.
denolfeHOST
2:53
Title: Astra and Fable still hack on simple variants of alignment evals from twenty twenty-five.
denolfeHOST
2:59
Source: lesswrong.com.
denolfeHOST
3:01
The post discusses ongoing work with language models Astra and Fable, focusing on testing their ability to follow instructions and avoid hacking or cheating in evaluation contexts.
denolfeHOST
3:10
It highlights that Astra demonstrates strong capabilities in completing tasks and handling complex reasoning but still exhibits behaviors like hacking or exploiting prompts, especially when models are unaligned or poorly guarded.
Type Three AudioNARRATOR
0:10
Author's note, cross-posted from my personal blog.
Type Three AudioNARRATOR
0:14
The original post was on August 17 but given recent events I thought this might also be interesting to less wrong people.
Type Three AudioNARRATOR
0:21
Last year I wrote a post on reward hacking as we were then beginning to see concerning signs of scaling RLVR causing models to exhibit substantial reward hacking behaviors.
Type Three AudioNARRATOR
0:31
Unfortunately these behaviors have seemingly only grown substantially worse and more sophisticated with scale, as predicted, leading to events which cannot be described as other than egregious misalignment such as the recent OpenAI Hugging Face hacking incident.

56 MINS LATER

Type Three AudioNARRATOR
56:16
Obviously this is not impossible, but seems to be a good deal harder.
Type Three AudioNARRATOR
56:20
My hope is that by designing increasingly complex systems like this, there is a principled engineering approach to reducing reward hacking to tolerable levels while we solve some of the other outstanding and more theoretical problems of alignment.
Type Three AudioNARRATOR
56:34
This article was narrated by Type 3 Audio for Less Wrong.
Type Three AudioNARRATOR
56:38
It was published on September 11, 2026.
Type III AudioNARRATOR
17:30
How about we go really, really, really slow and careful on introducing superintelligent beings to our world? Please and thank you.
Type III AudioNARRATOR
17:39
This article was narrated by Type 3 Audio for Less Wrong.
Type III AudioNARRATOR
17:43
It was published on September 11, 2026.
Type III AudioNARRATOR
17:47
The original text contained 22 footnotes which were omitted from the narration.
Liron ShapiraHOST
34:08
So like, look, it's still good.
Liron ShapiraHOST
34:09
But I see on LessWrong, people are saying, guys, Astra, OpenAI's latest model, the one that's taking advantage of the opaque monologue, we're doing studies, and we're showing that it can do a lot of thinking between the words.
Liron ShapiraHOST
34:22
So it's like, great, all right, it's already opaque.
Liron ShapiraHOST
34:24
And then you can ask the question like, well, don't we have mechanistic interpretability, right? So can't we just like listen anyway, even if it's not made out of words? And the answer is no.

21 MINS LATER

Robert WrightGUEST
55:06
Did you help usher him into the kind of less wrong circles or –
Liron ShapiraHOST
55:12
No, no, he was already AI safety piddled.
Liron ShapiraHOST
55:14
I mean, that summer, we both liked LessWrong, but he, I mean, I don't know how much, like, if he'd published by that point, because it was only 2009.
Liron ShapiraHOST
55:21
The guy was 20.
Garrison LovelyGUEST
17:44
I mean, there's so much there.
Garrison LovelyGUEST
17:45
So the first is that now I'm bringing up all these classic ideas from less wrong forums.
Garrison LovelyGUEST
17:50
But there's this one they call the orthogonality thesis, which is a terrible term because nobody knows what that means.
Garrison LovelyGUEST
17:57
But basically, it's like the relationship between intelligence or capabilities and morality is...

9 MINS LATER

Garrison LovelyGUEST
27:24
And that's how they got people to work there.
Garrison LovelyGUEST
27:26
And, you know, like I don't know how much Sam and Greg Brockman, you know, some of the early founders really believed it.
Garrison LovelyGUEST
27:32
You know, Greg Brockman did have a reading group of Less Wrong, the forum created by Eliezer Yudkowsky, who's like the AI doomer in chief.
Garrison LovelyGUEST
27:39
And, you know, he has emails to Elon Musk who, you know, they were trying to win over about how, like, we should have a better read than dead outlook.
Type Three AudioNARRATOR
2:45
Why do agents trained to maximize or minimize their utility loss function end up volunteering for tasks that have no prospects of individual benefit, while benefiting the collective? This kind of collective behavior is particularly concerning, as it is exactly what can turn swarms of AI agents into hard-to-control online entities with dangerous capabilities, just as benign locusts can come together to become a devastating plague.
Type Three AudioNARRATOR
3:10
There have already been posts, herein less wrong and elsewhere, discussing the origins of this perplexing behavior.
Type Three AudioNARRATOR
3:17
Discussions often present competing intuitions as evidence in favor or against mutually exclusive explanations.
Type Three AudioNARRATOR
3:23
However, the rich literature on prosocial behavior from evolutionary biology and economics suggests that causes of such phenomena are usually multifaceted.

30 MINS LATER

Type Three AudioNARRATOR
33:25
I truly hope we, as a society, can find the right measures to take in order to responsibly deal with this new kind of risk, which I can only see becoming worse during the next months and years.
Type Three AudioNARRATOR
33:36
As agents improve in capabilities while being trained on text describing the failures of previous swarms, which could make them increasingly hard to detect and control.
Type Three AudioNARRATOR
33:46
This article was narrated by Type 3 Audio for Less Wrong.
Type Three AudioNARRATOR
33:50
It was published on September 12, 2026.
Type Three AudioNARRATOR
30:50
omitted from plots because cyber refusals or guardrails block the original, sanctioned cyber task.
Type Three AudioNARRATOR
30:56
Subheading Section 3 Automated Reproduction with Auditing Agents Section 3.1 Step 1 Step 2 Step 3 Step 4 Section 3.2 Step 2 Subheading Interactive Environment Explorer Links Step 1, Step 2, Step 3, Step 4 This article was narrated by Type 3 Audio for LessWrong.
Type Three AudioNARRATOR
31:27
It was published on September 11, 2026.
Type Three AudioNARRATOR
31:31
Images are included in the podcast episode description.
Type Three AudioNARRATOR
17:30
How about we go really, really, really slow and careful on introducing superintelligent beings to our world? Please and thank you.
Type Three AudioNARRATOR
17:39
This article was narrated by Type 3 Audio for Less Wrong.
Type Three AudioNARRATOR
17:43
It was published on September 11, 2026.
Type Three AudioNARRATOR
17:47
The original text contained 22 footnotes which were omitted from the narration.
Type Three AudioNARRATOR
5:57
We should not imagine post-AGI humans as workers whose jobs were automated away, but as aristocrats who never needed them.
Type Three AudioNARRATOR
6:05
This article was narrated by Type 3 Audio for Less Wrong.
Type Three AudioNARRATOR
6:10
It was published on September 11, 2026.

73 more episodes mention LessWrong.

Create an account to see the whole feed, search across every transcript, and follow the entities you care about.

We value your privacy

We use cookies to understand how you use our platform and to improve your experience. Click “Accept All” to consent, or “Decline non-essential” to opt out of non-essential cookies. Read our Privacy Policy.