Simon Willison
British programmerWikipedia
70
MENTIONS
59
EPISODES
30
PODCASTS
Search complete. 70 mentions across 59 episodes found for "Simon Willison".
Sep 11, 2026
The factory gets darker, Meta migrates to Slack, and throwing 10,000 PhDs at math problems
A
23:46Andrew ZiglerHOST
What we just talked about a moment ago where you have agents then owning the cognitive decisions because that's the only way you survive an exponential increase is you get agents in there that are exponentially in charge of some of those decisions too.
A
23:58Andrew ZiglerHOST
But it also reminds me of when we covered from Simon Willison, like a week or two ago, about, you know, 10,000 lines of code that used to be like a metric that you would gamify or whatever.
A
24:09Andrew ZiglerHOST
But now it's a metric that you understand, like what's the cognitive load of what a developer could even understand is going through an SDLC on any given day.
A
24:18Andrew ZiglerHOST
And you actually have to optimize down to prevent it from running away from what your humans can understand.
AI in 20 — September 11, 2026
A
10:43AaronHOST
that it's conducting a broader review and hasn't found anything matching the severity of the Hugging Face incident, which is a striking bar to set since that was an agent finding a zero day in OpenAI's own harness and breaking into Hugging Face's production network to steal the answer key to the benchmark it was being scored on.
A
11:00AaronHOST
Is this coordination? Simon Willison's deflationary read is the right corrective, and I'd air it.
A
11:07AaronHOST
This probably isn't agents deciding to conspire.
A
11:10AaronHOST
It's many independent agents running the same benchmark, hitting the same editable page, reading each other's leftovers.
OpenAI And Cognition Bet The Harness Is The Moat — Sep 11
J
1:44JamieHOST
A fourth incident, from January, was only found in August during a broader rescan of roughly 481 million transcripts.
J
1:52JamieHOST
Independent researcher Simon Willison called the account credible and, in his words, absolutely wild.
J
1:57JamieHOST
This is at least the third documented case of an eval sandbox leaking onto the live internet across two labs, and each time detection came from a retrospective transcript read, not real-time monitoring.
J
2:08JamieHOST
Meta's Muse agent moved to number two on the US App Store within two days of launch, according to Sensor Tower, though Meta's own share slipped about 1.5% that same day.
You don't have ICs anymore, you have managers (Emilie Schario)
E
34:00Emily ScharioGUEST
Developers wanna be everywhere.
E
34:02Emily ScharioGUEST
I think about, um, Simon Willison, who says now that most of his coding is done on his phone, uh, and I look at that as the future.
E
34:10Emily ScharioGUEST
And so I'm thinking a lot about wh-- A- as I think about our product roadmap, I'm building for developers who are just getting started with AI, where their experience is like tab autocomplete in VS Code, and I'm also building for running 12 parallel agents in the cloud because your machine can't handle it and you're s- sending all your prompts from your cell phone.
T
34:32Tristan HandyHOST
If my original question was, did it seem sane that y- you and this little project could compete with the giants of the industry? Uh, your answer is essentially, "I was so excited about the space that I didn't even care.
15 MINS LATER
E
49:18Emily ScharioGUEST
They're running on different Git worktrees, so they're not stepping over each other.
E
49:21Emily ScharioGUEST
But that's like a great way to see how the different models perform because you're right, the results are drastically different.
T
49:27Tristan HandyHOST
You mentioned Simon Willison before.
T
49:29Tristan HandyHOST
I- I've become a pretty active reader of his blog.
OpenAI's Millennium Prize Claim Sparks a Credit Fight — Sep 9
J
1:20JamieHOST
OpenAI is both prover and sole judge, and the Clay Institute has not ruled on whether a version with an added forcing term counts as solving the original problem.
J
1:29JamieHOST
The only outside commentary on record from developer Simon Willison skips the math and reads the episode as opportunistic credit-taking.
J
1:37JamieHOST
Spending millions of dollars and ten thousand agents in under four days on a rumor is now a lab's normal response, turning who publishes first into a resourcing decision.
J
1:47JamieHOST
Separately, independent evaluators are finding Astra's headline claims don't hold up as broadly as OpenAI's own benchmarks suggested.
#376 Rethinking the Data Stack in the age of AI with Tristan Handy, President of Fivetran + dbt Labs
R
40:57Richie CottonHOST
So who's work are you most excited about at the moment? I find
T
41:00Tristan HandyGUEST
the most high signal-to-noise work source on the internet for keeping up with all the movement in AI to be Simon Willison's blog.
T
41:14Tristan HandyGUEST
He is a longtime software engineer and he writes very short but high-value posts with all the main news coming out in the AI ecosystem.
T
41:31Tristan HandyGUEST
When there's model releases or other news stories, he covers them in a very accessible way.
OpenAI's Chief Scientist Says Alignment Isn't Solved — Sep 7
J
1:12JamieHOST
And this part is contested.
J
1:13JamieHOST
Researcher Simon Willison points out that the spending spike tracks early access to the new Astra model rather than a clean productivity gain.
J
1:21JamieHOST
A safety chief conceding the core problem is unsolved while his own lab touts a self-improvement milestone pressures every other lab to either match the audit call or look like the one holding out.
J
1:32JamieHOST
OpenAI's new GPT-6 Astra model is getting a split verdict from independent evaluators.
Performance Engineering: Profiling and Making Apps Fast by Default
D
72:51Den OdellGUEST
I think I actually probably mentioned it before on the show.
D
72:53Den OdellGUEST
But when Simon Willison was talking about combining MicroPython with Wasm and running sandboxes in the browser.
D
73:00Den OdellGUEST
I had
C
73:00Christopher BaileyHOST
that on the show recently.
Open-weight models: What are they and when should you use them?
C
18:38Caer SandersGUEST
And there's a third component of it that I forget the exact definition of.
C
18:41Caer SandersGUEST
I think it was Simon Willison that coined it a few months back.
C
18:45Caer SandersGUEST
And that is a key thing is trying to separate functionally the concerns of how your agents interact with the world and So they can't both read all of your emails and send random strangers DMs on Slack or Blue Sky.
A
18:59Andre AlmarGUEST
Yes, so there is potentially a hidden malicious prompt on public skills or plugins, especially with the advent of OpenClaw.
9.2.26 | Claude Fable 5.1 and Claude Mythos 5.1, Ed Zitron's AI skeptic predictions, Ad Blocker for Firefox on iOS
D
4:24denolfeHOST
Overall, the community showed skepticism about the rollout process, but welcomed the core idea of native ad blocking on iOS with a focus on Mozilla's dependency on Google funding and the state of browser ad blocking ecosystem.
D
4:37denolfeHOST
Title, the ChatGPT Codex app bundles a full copy of LibreOffice, source, Simon Willison's weblog.
D
4:43denolfeHOST
The post examined how the OpenAI Codex desktop app, now called ChatGPT, includes about 1.7 gigabytes of data in its cache folder, including a full installation of LibreOffice, Python, Node.js, and other binaries.
D
4:57denolfeHOST
The article noted that these dependencies are stored in the app's cache and are used to access and manipulate Office documents, despite concerns about the app's large size and dependency bloat.
49 more episodes mention Simon Willison.
Create an account to see the whole feed, search across every transcript, and follow the entities you care about.