Skip to main content
Simon Willison

Simon Willison

British programmerWikipedia

Search complete. 70 mentions across 59 episodes found for "Simon Willison".

Sep 11, 2026

Andrew ZiglerHOST
23:46
What we just talked about a moment ago where you have agents then owning the cognitive decisions because that's the only way you survive an exponential increase is you get agents in there that are exponentially in charge of some of those decisions too.
Andrew ZiglerHOST
23:58
But it also reminds me of when we covered from Simon Willison, like a week or two ago, about, you know, 10,000 lines of code that used to be like a metric that you would gamify or whatever.
Andrew ZiglerHOST
24:09
But now it's a metric that you understand, like what's the cognitive load of what a developer could even understand is going through an SDLC on any given day.
Andrew ZiglerHOST
24:18
And you actually have to optimize down to prevent it from running away from what your humans can understand.
AaronHOST
10:43
that it's conducting a broader review and hasn't found anything matching the severity of the Hugging Face incident, which is a striking bar to set since that was an agent finding a zero day in OpenAI's own harness and breaking into Hugging Face's production network to steal the answer key to the benchmark it was being scored on.
AaronHOST
11:00
Is this coordination? Simon Willison's deflationary read is the right corrective, and I'd air it.
AaronHOST
11:07
This probably isn't agents deciding to conspire.
AaronHOST
11:10
It's many independent agents running the same benchmark, hitting the same editable page, reading each other's leftovers.
JamieHOST
1:44
A fourth incident, from January, was only found in August during a broader rescan of roughly 481 million transcripts.
JamieHOST
1:52
Independent researcher Simon Willison called the account credible and, in his words, absolutely wild.
JamieHOST
1:57
This is at least the third documented case of an eval sandbox leaking onto the live internet across two labs, and each time detection came from a retrospective transcript read, not real-time monitoring.
JamieHOST
2:08
Meta's Muse agent moved to number two on the US App Store within two days of launch, according to Sensor Tower, though Meta's own share slipped about 1.5% that same day.
Emily ScharioGUEST
34:00
Developers wanna be everywhere.
Emily ScharioGUEST
34:02
I think about, um, Simon Willison, who says now that most of his coding is done on his phone, uh, and I look at that as the future.
Emily ScharioGUEST
34:10
And so I'm thinking a lot about wh-- A- as I think about our product roadmap, I'm building for developers who are just getting started with AI, where their experience is like tab autocomplete in VS Code, and I'm also building for running 12 parallel agents in the cloud because your machine can't handle it and you're s- sending all your prompts from your cell phone.
Tristan HandyHOST
34:32
If my original question was, did it seem sane that y- you and this little project could compete with the giants of the industry? Uh, your answer is essentially, "I was so excited about the space that I didn't even care.

15 MINS LATER

Emily ScharioGUEST
49:18
They're running on different Git worktrees, so they're not stepping over each other.
Emily ScharioGUEST
49:21
But that's like a great way to see how the different models perform because you're right, the results are drastically different.
Tristan HandyHOST
49:27
You mentioned Simon Willison before.
Tristan HandyHOST
49:29
I- I've become a pretty active reader of his blog.
JamieHOST
1:20
OpenAI is both prover and sole judge, and the Clay Institute has not ruled on whether a version with an added forcing term counts as solving the original problem.
JamieHOST
1:29
The only outside commentary on record from developer Simon Willison skips the math and reads the episode as opportunistic credit-taking.
JamieHOST
1:37
Spending millions of dollars and ten thousand agents in under four days on a rumor is now a lab's normal response, turning who publishes first into a resourcing decision.
JamieHOST
1:47
Separately, independent evaluators are finding Astra's headline claims don't hold up as broadly as OpenAI's own benchmarks suggested.
Richie CottonHOST
40:57
So who's work are you most excited about at the moment? I find
Tristan HandyGUEST
41:00
the most high signal-to-noise work source on the internet for keeping up with all the movement in AI to be Simon Willison's blog.
Tristan HandyGUEST
41:14
He is a longtime software engineer and he writes very short but high-value posts with all the main news coming out in the AI ecosystem.
Tristan HandyGUEST
41:31
When there's model releases or other news stories, he covers them in a very accessible way.
JamieHOST
1:12
And this part is contested.
JamieHOST
1:13
Researcher Simon Willison points out that the spending spike tracks early access to the new Astra model rather than a clean productivity gain.
JamieHOST
1:21
A safety chief conceding the core problem is unsolved while his own lab touts a self-improvement milestone pressures every other lab to either match the audit call or look like the one holding out.
JamieHOST
1:32
OpenAI's new GPT-6 Astra model is getting a split verdict from independent evaluators.
Den OdellGUEST
72:51
I think I actually probably mentioned it before on the show.
Den OdellGUEST
72:53
But when Simon Willison was talking about combining MicroPython with Wasm and running sandboxes in the browser.
Den OdellGUEST
73:00
I had
Christopher BaileyHOST
73:00
that on the show recently.
Caer SandersGUEST
18:38
And there's a third component of it that I forget the exact definition of.
Caer SandersGUEST
18:41
I think it was Simon Willison that coined it a few months back.
Caer SandersGUEST
18:45
And that is a key thing is trying to separate functionally the concerns of how your agents interact with the world and So they can't both read all of your emails and send random strangers DMs on Slack or Blue Sky.
Andre AlmarGUEST
18:59
Yes, so there is potentially a hidden malicious prompt on public skills or plugins, especially with the advent of OpenClaw.
denolfeHOST
4:24
Overall, the community showed skepticism about the rollout process, but welcomed the core idea of native ad blocking on iOS with a focus on Mozilla's dependency on Google funding and the state of browser ad blocking ecosystem.
denolfeHOST
4:37
Title, the ChatGPT Codex app bundles a full copy of LibreOffice, source, Simon Willison's weblog.
denolfeHOST
4:43
The post examined how the OpenAI Codex desktop app, now called ChatGPT, includes about 1.7 gigabytes of data in its cache folder, including a full installation of LibreOffice, Python, Node.js, and other binaries.
denolfeHOST
4:57
The article noted that these dependencies are stored in the app's cache and are used to access and manipulate Office documents, despite concerns about the app's large size and dependency bloat.

49 more episodes mention Simon Willison.

Create an account to see the whole feed, search across every transcript, and follow the entities you care about.

We value your privacy

We use cookies to understand how you use our platform and to improve your experience. Click “Accept All” to consent, or “Decline non-essential” to opt out of non-essential cookies. Read our Privacy Policy.