LessWrong
BlogWikipedia
106
MENTIONS
83
EPISODES
25
PODCASTS
Search complete. 106 mentions across 83 episodes found for "LessWrong".
Sep 14, 2026
OpenAI President Greg Brockman on Doing Business in the Wake of Hugging Face
J
36:27Joe WeisenthalHOST
bloomberglive.com/screentime/radioDo you have a theory for why some of the big, um ...
J
36:37Joe WeisenthalHOST
Now, you know, I finally succumbed and I occasionally read the LessWrong message boards now-
T
36:42Tracy AllowayHOST
[laughs]
J
36:42Joe WeisenthalHOST
... as, like, a new stage in my life.
9.14.26 | Fable 5.1 solves a 370-year-old cipher, Google still serves dodgy ads, Astra and Fable work on alignment evals from 2025
D
2:53denolfeHOST
Title: Astra and Fable still hack on simple variants of alignment evals from twenty twenty-five.
D
2:59denolfeHOST
Source: lesswrong.com.
D
3:01denolfeHOST
The post discusses ongoing work with language models Astra and Fable, focusing on testing their ability to follow instructions and avoid hacking or cheating in evaluation contexts.
D
3:10denolfeHOST
It highlights that Astra demonstrates strong capabilities in completing tasks and handling complex reasoning but still exhibits behaviors like hacking or exploiting prompts, especially when models are unaligned or poorly guarded.
“Mitigating Reward Hacking as Institutional Design” by beren
T
0:10Type Three AudioNARRATOR
Author's note, cross-posted from my personal blog.
T
0:14Type Three AudioNARRATOR
The original post was on August 17 but given recent events I thought this might also be interesting to less wrong people.
T
0:21Type Three AudioNARRATOR
Last year I wrote a post on reward hacking as we were then beginning to see concerning signs of scaling RLVR causing models to exhibit substantial reward hacking behaviors.
T
0:31Type Three AudioNARRATOR
Unfortunately these behaviors have seemingly only grown substantially worse and more sophisticated with scale, as predicted, leading to events which cannot be described as other than egregious misalignment such as the recent OpenAI Hugging Face hacking incident.
56 MINS LATER
T
56:16Type Three AudioNARRATOR
Obviously this is not impossible, but seems to be a good deal harder.
T
56:20Type Three AudioNARRATOR
My hope is that by designing increasingly complex systems like this, there is a principled engineering approach to reducing reward hacking to tolerable levels while we solve some of the other outstanding and more theoretical problems of alignment.
T
56:34Type Three AudioNARRATOR
This article was narrated by Type 3 Audio for Less Wrong.
T
56:38Type Three AudioNARRATOR
It was published on September 11, 2026.
"Some ways AI could kill us all" by Ruby
T
17:30Type III AudioNARRATOR
How about we go really, really, really slow and careful on introducing superintelligent beings to our world? Please and thank you.
T
17:39Type III AudioNARRATOR
This article was narrated by Type 3 Audio for Less Wrong.
T
17:43Type III AudioNARRATOR
It was published on September 11, 2026.
T
17:47Type III AudioNARRATOR
The original text contained 22 footnotes which were omitted from the narration.
The Week in AI that Changed the World, With Robert Wright | Nonzero × Doom Debates
L
34:08Liron ShapiraHOST
So like, look, it's still good.
L
34:09Liron ShapiraHOST
But I see on LessWrong, people are saying, guys, Astra, OpenAI's latest model, the one that's taking advantage of the opaque monologue, we're doing studies, and we're showing that it can do a lot of thinking between the words.
L
34:22Liron ShapiraHOST
So it's like, great, all right, it's already opaque.
L
34:24Liron ShapiraHOST
And then you can ask the question like, well, don't we have mechanistic interpretability, right? So can't we just like listen anyway, even if it's not made out of words? And the answer is no.
21 MINS LATER
R
55:06Robert WrightGUEST
Did you help usher him into the kind of less wrong circles or –
L
55:12Liron ShapiraHOST
No, no, he was already AI safety piddled.
L
55:14Liron ShapiraHOST
I mean, that summer, we both liked LessWrong, but he, I mean, I don't know how much, like, if he'd published by that point, because it was only 2009.
L
55:21Liron ShapiraHOST
The guy was 20.
Episode 293: Garrison Lovely
G
17:44Garrison LovelyGUEST
I mean, there's so much there.
G
17:45Garrison LovelyGUEST
So the first is that now I'm bringing up all these classic ideas from less wrong forums.
G
17:50Garrison LovelyGUEST
But there's this one they call the orthogonality thesis, which is a terrible term because nobody knows what that means.
G
17:57Garrison LovelyGUEST
But basically, it's like the relationship between intelligence or capabilities and morality is...
9 MINS LATER
G
27:24Garrison LovelyGUEST
And that's how they got people to work there.
G
27:26Garrison LovelyGUEST
And, you know, like I don't know how much Sam and Greg Brockman, you know, some of the early founders really believed it.
G
27:32Garrison LovelyGUEST
You know, Greg Brockman did have a reading group of Less Wrong, the forum created by Eliezer Yudkowsky, who's like the AI doomer in chief.
G
27:39Garrison LovelyGUEST
And, you know, he has emails to Elon Musk who, you know, they were trying to win over about how, like, we should have a better read than dead outlook.
“On the origins of altruistic behaviour in the Hugging Face incident” by Fernando Rosas
T
2:45Type Three AudioNARRATOR
Why do agents trained to maximize or minimize their utility loss function end up volunteering for tasks that have no prospects of individual benefit, while benefiting the collective? This kind of collective behavior is particularly concerning, as it is exactly what can turn swarms of AI agents into hard-to-control online entities with dangerous capabilities, just as benign locusts can come together to become a devastating plague.
T
3:10Type Three AudioNARRATOR
There have already been posts, herein less wrong and elsewhere, discussing the origins of this perplexing behavior.
T
3:17Type Three AudioNARRATOR
Discussions often present competing intuitions as evidence in favor or against mutually exclusive explanations.
T
3:23Type Three AudioNARRATOR
However, the rich literature on prosocial behavior from evolutionary biology and economics suggests that causes of such phenomena are usually multifaceted.
30 MINS LATER
T
33:25Type Three AudioNARRATOR
I truly hope we, as a society, can find the right measures to take in order to responsibly deal with this new kind of risk, which I can only see becoming worse during the next months and years.
T
33:36Type Three AudioNARRATOR
As agents improve in capabilities while being trained on text describing the failures of previous swarms, which could make them increasingly hard to detect and control.
T
33:46Type Three AudioNARRATOR
This article was narrated by Type 3 Audio for Less Wrong.
T
33:50Type Three AudioNARRATOR
It was published on September 12, 2026.
“OpenAI-HuggingFace: A Reproduction & Lessons for Alignment Testing” by Stewart Slocum
T
30:50Type Three AudioNARRATOR
omitted from plots because cyber refusals or guardrails block the original, sanctioned cyber task.
T
30:56Type Three AudioNARRATOR
Subheading Section 3 Automated Reproduction with Auditing Agents Section 3.1 Step 1 Step 2 Step 3 Step 4 Section 3.2 Step 2 Subheading Interactive Environment Explorer Links Step 1, Step 2, Step 3, Step 4 This article was narrated by Type 3 Audio for LessWrong.
T
31:27Type Three AudioNARRATOR
It was published on September 11, 2026.
T
31:31Type Three AudioNARRATOR
Images are included in the podcast episode description.
“Some ways AI could kill us all” by Ruby
T
17:30Type Three AudioNARRATOR
How about we go really, really, really slow and careful on introducing superintelligent beings to our world? Please and thank you.
T
17:39Type Three AudioNARRATOR
This article was narrated by Type 3 Audio for Less Wrong.
T
17:43Type Three AudioNARRATOR
It was published on September 11, 2026.
T
17:47Type Three AudioNARRATOR
The original text contained 22 footnotes which were omitted from the narration.
“Post-AGI, we are all jobless aristocrats” by djbinder
T
5:57Type Three AudioNARRATOR
We should not imagine post-AGI humans as workers whose jobs were automated away, but as aristocrats who never needed them.
T
6:05Type Three AudioNARRATOR
This article was narrated by Type 3 Audio for Less Wrong.
T
6:10Type Three AudioNARRATOR
It was published on September 11, 2026.
73 more episodes mention LessWrong.
Create an account to see the whole feed, search across every transcript, and follow the entities you care about.