Supervisor, not overseer
In my post about my Showboat project (https://simonwillison.net/2026/Feb/10/showboat-and-rodney/) I used the term "overseer" to refer to the person who manages a coding agent. It turns out that's a
In my post about my Showboat project (https://simonwillison.net/2026/Feb/10/showboat-and-rodney/) I used the term "overseer" to refer to the person who manages a coding agent. It turns out that's a
An AI-generated report, delivered directly to the email inboxes of journalists, was an essential tool in the Times’ coverage. It was also one of the first signals that conservative media was
Skills in OpenAI API (https://developers.openai.com/cookbook/examples/skills_in_api) OpenAI's adoption of Skills continues to gain ground. You can now use Skills directly in the OpenAI API with
GLM-5: From Vibe Coding to Agentic Engineering (https://z.ai/blog/glm-5) This is a huge new MIT-licensed model: 754B parameters and 1.51TB on Hugging Face (https://huggingface.co/zai-org/GLM-5)
cysqlite - a new sqlite driver (https://charlesleifer.com/blog/cysqlite---a-new-sqlite-driver/) Charles Leifer has been maintaining pysqlite3 (https://github.com/coleifer/pysqlite3) - a fork of the
A key challenge working with coding agents is having them both test what they’ve built and demonstrate that software to you, their overseer. This goes beyond automated tests - we need artifacts
Structured Context Engineering for File-Native Agentic Systems (https://arxiv.org/abs/2602.05447) New paper by Damon McMillan exploring challenging LLM context tasks involving large SQL schemas (up
AI Doesn’t Reduce Work—It Intensifies It (https://hbr.org/2026/02/ai-doesnt-reduce-work-it-intensifies-it) Aruna Ranganathan and Xingqi Maggie Ye from Berkeley Haas School of Business report
Friend and neighbour Karen James (https://www.etsy.com/shop/KarenJamesMakes) made me a Kākāpō mug. It has a charismatic Kākāpō, four Kākāpō chicks (in celebration of the 2026 breeding season
People on the orange site are laughing at this, assuming it's just an ad and that there's nothing to it. Vulnerability researchers I talk to do not think this is a joke. As an erstwhile vuln
Vouch (https://github.com/mitchellh/vouch) Mitchell Hashimoto's new system to help address the deluge of worthless AI-generated PRs faced by open source projects now that the friction involved in
Claude: Speed up responses with fast mode (https://code.claude.com/docs/en/fast-mode) New "research preview" from Anthropic today: you can now access a faster version of their frontier model Claude
I am having more fun programming than I ever have, because so many more of the programs I wish I could find the time to write actually exist. I wish I could share this joy with the people who are
Last week I hinted at (https://simonwillison.net/2026/Jan/28/the-five-levels/) a demo I had seen from a team implementing what Dan Shapiro called the Dark Factory
I don't know why this week became the tipping point, but nearly every software engineer I've talked to is experiencing some degree of mental health crisis. [...] Many people assuming I meant job
There's a jargon-filled headline for you! Everyone's building sandboxes (https://simonwillison.net/2026/Jan/8/llm-predictions-for-2026/#1-year-we-re-finally-going-to-solve-sandboxing) for running
An Update on Heroku (https://www.heroku.com/blog/an-update-on-heroku/) An ominous headline to see on the official Heroku blog and yes, it's bad news. Today, Heroku is transitioning to a sustaining
When I want to quickly implement a one-off experiment in a part of the codebase I am unfamiliar with, I get codex to do extensive due diligence. Codex explores relevant slack channels, reads related
Mitchell Hashimoto: My AI Adoption Journey (https://mitchellh.com/writing/my-ai-adoption-journey) Some really good and unconventional tips in here for getting to a place with coding agents where
Two major new model releases today, within about 15 minutes of each other. Anthropic released Opus 4.6 (https://www.anthropic.com/news/claude-opus-4-6). Here's its pelican