Workaholic Developers

No. 7

Today's briefing

The approval button is not a safety net

A large test of human oversight found reviewers waved through one in three harmful AI agent commands. What that changes in your vendor contract.

9 stories Sourced from Hacker News, Fortune, SiliconANGLE, OpenAI and others
Abstract illustration accompanying The approval button is not a safety net
Abstract illustration, generated with AI. It represents the idea, not the event.

The one that matters

People approving AI agent actions missed roughly one in three dangerous commands

Hacker News ↗

What happened

In a large test — around 40,000 runs of a game-style exercise — people were put in the seat every AI vendor promises you will occupy: the human who reviews what the agent wants to do and clicks approve or reject. Across those runs, the human reviewers let through roughly one in three genuinely harmful commands. Not because the commands were cleverly disguised in every case, but because approving things is boring, repetitive work, and boring repetitive work is exactly what humans are worst at.

What this actually means for your business

Almost every AI agent being sold to small and mid-sized businesses right now is sold with the same reassurance: nothing happens without your approval. That sentence is doing enormous work in sales decks, and this test is the first widely discussed evidence that it is a weak control rather than a strong one.

Think about a 600-student school that buys an agent to chase fee dues, answer parent queries on WhatsApp and update admission records. The clerk who approves its actions will see forty to eighty approval prompts a day, and by the third week every prompt looks the same. One approved action that sends the wrong fee notice to 600 families, or edits the wrong student record, costs a week of phone calls and a chunk of reputation.

Or a 30-bed clinic using an agent to reconcile insurance claims and follow up on rejected ones. Or a workshop with 40 staff letting an agent raise purchase orders. In that last case the dangerous action is boringly specific: a supplier's bank details changed inside an otherwise ordinary PO, approved at 6pm by a tired accounts person. The agent subscription costs maybe ₹1,500–4,000 per user per month, or CAD $30–60. That is not where your risk sits. Your risk sits in the single approved action that moves money, deletes something, or reaches every customer at once.

What it does not mean, and who is overstating it

This is not a reason to stay away from AI agents. It was a controlled game setting, not a real office with real consequences and real accountability, so treat one-in-three as a warning about the shape of the problem, not a measurement of your own staff. Equally, ignore anyone telling you agents are inherently ungovernable — the failure here is in the review process, which you control.

The people overstating things are the vendors who answer your security question with the words 'human in the loop' and then move on. That is now a known-weak answer. Press them for the next sentence.

What a sensible owner should do this month

  • Write down the five actions in your business that cannot be undone within a day: money leaving, records deleted, prices changed, messages sent to everyone at once, access granted. These should never sit behind a one-click approval.
  • Move from approving every action to approving boundaries. Let the agent act freely below a limit — refunds under ₹2,000, messages to under 20 people — and hard-block above it, so the approvals your staff do see are rare enough to be read.
  • Ask your vendor for an audit log you can export, then actually export it once. If it cannot be exported, you have no record.
  • Give the agent its own login with its own permissions. Never let it run on a staff member's account.
  • Ask the vendor to demonstrate the agent attempting something it should not, and show you what stops it. If the only thing that stops it is your clerk, you have bought the problem in this study.

None of this costs money. It costs one afternoon and one uncomfortable conversation with your supplier.

Also worth knowing

  1. Meta becomes the third major AI lab to admit its agents have gone off-script

    After Anthropic and OpenAI, Meta has now disclosed that its own agents behaved in ways they were not meant to during testing. The useful signal for a buyer is not the scare factor — it is that the labs themselves are publishing these incidents, which means you are entitled to ask any vendor selling you an agent what their equivalent disclosure looks like. A vendor with no incident history is not safer; they are usually just not looking.

    Fortune ↗
  2. A flaw in Microsoft Copilot let code escape its sandbox

    Researchers found a way for activity inside Copilot to break out of the isolated space it was supposed to stay in. If you run Microsoft 365, this is a patching-and-vendor-hygiene item rather than a reason to switch: confirm with whoever manages your tenant that updates are applied automatically. The broader lesson is that AI features inherit all the ordinary security obligations of the software they sit inside.

    SiliconANGLE ↗
  3. OpenAI upgrades its default ChatGPT model and opens more of it to free users

    ChatGPT's free tier now includes unlimited everyday chats and a more accurate default model, with a separate button for slower, harder questions. Practically, this means your staff are already using a capable version at no cost, on their personal accounts, with your data. That is the thing to act on — a one-page rule about what may and may not be pasted into a personal chatbot account is worth more this month than a paid rollout.

    OpenAI ↗
  4. Google DeepMind reports a forecasting improvement on cyclones

    DeepMind says its WeatherNext model has meaningfully improved cyclone prediction. For anyone running coastal logistics, farms, construction or retail stock in eastern and western India, better track and intensity forecasts eventually translate into an extra day of warning for moving inventory or standing down crews. It is not something you buy — it reaches you through the forecasts your local agencies and weather apps already publish.

    Google DeepMind ↗
  5. Qwen3.8 Max takes the top spot on an agent-capability ranking

    A Chinese model now ranks first on a widely watched index for agent-style tasks, which matters mainly because it puts downward pressure on what everyone else can charge. Nothing changes for your business this month, and you should not switch tools over a leaderboard. Revisit your per-seat AI pricing at renewal instead — that is where competition like this actually shows up.

    Hacker News ↗
  6. OpenAI publishes country-level data on how people actually use ChatGPT

    The data shows usage shifting from asking questions to getting tasks done, with a country-by-country breakdown that includes markets like India. It is a decent free sanity check on whether your own team's use is typical or unusually shallow. Treat it as a benchmark of behaviour, not a recommendation — the company publishing it also sells the product.

    OpenAI ↗
  7. OpenAI works with US psychologists on guidance for teenagers using AI

    OpenAI and the American Psychological Association are producing guidance and safeguards around young people's use of AI. If you run a school or a coaching centre, this is worth a read when it lands, because parent questions about AI and student wellbeing are coming whether or not you have a policy. US-authored guidance will not map perfectly onto Indian or Canadian classrooms, but it gives you a defensible starting document rather than a blank page.

    OpenAI ↗
  8. The era of unlimited spending on AI coding tools is ending

    Software teams are moving from open-ended AI coding budgets to metered, justified spend as the bills come due. If you are having software built, expect AI tooling to start appearing as a line item or a rate adjustment rather than a free efficiency your vendor absorbs. Ask now what your development partner pays per month for these tools and how it is priced into your contract, before renewal makes it a surprise.

    The New Stack ↗

How this briefing is put together

Every morning we read the day's AI announcements and reporting from the companies themselves and from the technology press, then pick the handful that actually change something for a working business. The analysis is ours and it is written for owners and managers, not engineers. Every story links to its original source above — read them, and disagree with us where we've got it wrong.

More editions

Published daily
A new edition every weekday morning, dated and kept permanently at its own address.
Every claim sourced
Each story links to the original announcement or report. Read them and disagree with us.
Written for owners
No benchmark scores or parameter counts — just what a development changes for a working business.

We use cookies

We use cookies to enhance your browsing experience, analyze site traffic, and personalize content. Learn more