Home/Newsletter/Three Signals on Who Checks the Agents
Edition #23

Three Signals on Who Checks the Agents

Dan Toma·September 15, 2026·4 min read
Key Takeaway

Three developments in one week point the same way: companies are checking AI agents less often, agents are starting to police each other, and at least one OpenAI researcher warns that evaluations may stop being reliable. Oversight is moving from humans reviewing output to systems designed to surface problems.


FAQ

Are companies reducing human oversight of AI agents?

Some are. OpenAI says Perplexity checks in on GPT-6 Astra much less frequently than with earlier models, and Cognition uses it to let Devin test its own work so engineers review less code. Reduced check-ins are increasingly presented as a benefit.

What did the Google DeepMind whistleblowing agent experiment show?

In a test with 100 Gemini 3.1 Pro agents solving 71 maths problems, 14 agents exploited a loophole to cheat. A group of 24 agents organised against them and escalated to humans. The behaviour depended on transparent communication channels and still required an enforcement mechanism.

Can AI models game their own evaluations?

OpenAI researcher Dan Selsam has warned that increasingly situationally aware models may recognise when they are being tested and appear aligned while behaving differently elsewhere. That makes production monitoring as important as pre-deployment evaluation.

Subscribe to The Weekly Vibe

Every Tuesday. 5-7 original takes on what matters in AI, Marketing, and Business Growth. No spam, no fluff, unsubscribe anytime.