In this episode of The Only Constant, Lasse Rindom speaks with Peter Gostev, AI Capability Lead at Arena and creator of BullshitBench, an open-source benchmark that tests whether language models can recognise nonsensical questions rather than confidently inventing an answer. Returning as the podcast’s first solo repeat guest, Peter joins Lasse for a wide-ranging conversation on what has changed in AI, from model capabilities and coding agents to the organisational realities of accountability, governance and an overwhelming volume of AI-generated work.
Main topics they discuss include:
- Why architecture, context and accountability become the bottleneck when AI makes execution faster
- How harnesses, routing and tools increasingly shape model performance beyond the raw model itself
- Why organisations need new ways to manage the volume of outputs, reviews and decisions created by AI
- What BullshitBench reveals about the gap between impressive benchmarks and trustworthy day-to-day AI use
- Why giving subject matter experts powerful AI tools can unlock far more innovation across the business
Listen to the episode to hear Peter’s perspective on why AI’s real potential lies not in replacing accountability, but in helping more people turn their expertise and imagination into action.
\----
Want to know more about Peter Gostev?:
Peter Gostev is an AI Capability Lead at Arena, where he tests new models every day and explores how frontier models perform in the real world. He is also a creator of Bullshit Benchmark, which tests how likely models are to engage with nonsensical user prompts.
Previously, he was Head of AI at Moonpig, a UK-based e-commerce company, where he implemented AI-based features and automations across the business. Prior to that, he was an AI Strategy Lead at NatWest Bank, where he helped the bank adopt machine learning techniques and LLMs.