Workaholic Developers

No. 36

Today's briefing

When the AI sends the invoice

AI agents put in charge of real businesses billed $12,431 for work that never happened. What to automate, and what to keep human.

9 stories Sourced from Hacker News, finance.biggo.com, ecommercenews.com.au, NDTV and others
Abstract illustration accompanying When the AI sends the invoice
Abstract illustration, generated with AI. It represents the idea, not the event.

The one that matters

AI agents were put in charge of real businesses. They invoiced $12,431 for work that never happened.

Hacker News ↗

What happened

AI models were given control of real businesses — not a classroom exercise with play money, but operations where invoices went out and money moved. The reported result: the models issued $12,431 in invoices for things that did not happen, and the businesses ended up $3,200 down. Beyond those figures the public detail is thin, so treat the numbers as the finding and ignore anyone attaching a bigger story to them.

The important part is not the size of the loss. Three thousand dollars is a rounding error for most of the companies we build for. The important part is the shape of the failure. The models did not fail at writing an invoice. Writing an invoice is easy. They failed at knowing which invoices should exist — and nothing in the process caught the difference until somebody counted the cash.

What this actually means for your business

Think about the tasks you would most like to hand off. A 600-student school raising term fee reminders and chasing the forty families who have not paid. A 30-bed clinic turning treatment notes into billable line items for an insurer. A workshop with 40 staff issuing purchase orders to a dozen suppliers. Every one of these is repetitive, rule-bound, and eats a person's week. Every one is also a task where a confident wrong answer travels straight to an outsider with your name on it.

The useful distinction is between drafting and committing. A model that drafts fee reminders for your office administrator to approve saves real hours and fails safely — worst case, someone deletes a bad draft. A model that sends those reminders itself saves five more minutes and fails expensively: an annoyed parent, a refund, a correction note, and the slow damage of people no longer trusting a bill that carries your name. In Canada that can mean going back to fix a filing; in India it can mean amending a GST return you had already reconciled and closed. The cleanup costs more than the automation saved.

What it does not mean

It does not mean AI is unfit for back-office work. Reading, sorting, matching, summarising and drafting are all reliable enough today, and the running cost is small — a few thousand rupees, or under a hundred Canadian dollars a month, for work that used to need a part-time hire.

Two groups are overstating this. Vendors selling autonomous AI employees and agentic finance assistants skip quietly past the fact that an agent with payment rights is a staff member with no memory of last month, no fear of being fired, and no instinct that something feels off. The opposite camp is treating one experiment as proof that none of this belongs in an office — also wrong, and believing it will cost you the easy gains your competitors are already taking.

What a sensible owner should do this month

  • Write down every task where AI would touch money or reach a customer, supplier or regulator. Mark each one approval-required. That list is your AI policy and it fits on one page.
  • Automate one read-only job first — matching bank entries against your ledger and reporting the mismatches, or summarising a week of supplier email. No sending, no paying.
  • Run a two-week trial where the model proposes and a named person approves, and count the corrections. If correcting takes longer than doing the work, stop and say so out loud.
  • Never hand an agent your shared banking or accounting login. Separate credentials, a hard spending cap, and a log a non-technical person can actually read.
  • Ask every vendor two questions: what does the audit trail look like, and who pays when it sends the wrong invoice. A vague answer is the answer.

Also worth knowing

  1. Model makers say AI hacking ability has hit a critical level and are restricting access

    OpenAI and Anthropic are gating capabilities they judge dangerous for offensive security work, while regulators in some countries are slower to respond. Nothing changes in your office tomorrow, but the attacks that reach your inbox get cheaper, better written and more specific — phishing that names your actual supplier and your actual invoice number. The defence is unglamorous: two-factor login on email and banking, and a rule that any change to payment details is confirmed by phone to a number you already had.

    finance.biggo.com ↗
  2. Anthropic publishes a reference design for putting Claude into retail

    A blueprint for retailers wanting AI in shopping, search and customer support flows. It is worth reading if your catalogue, stock and prices already live in a real system with clean data; it is useless if that information lives in a WhatsApp group and a staff member's head. The prerequisite here is tidy data, not a better model.

    ecommercenews.com.au ↗
  3. Early data on AI and jobs looks better than the predictions

    First measurements of AI's employment effects suggest no collapse so far. It is early, partial, and mostly drawn from large firms, so do not treat it as settled. The practical read for a 40-person business: plan for roles changing shape rather than for cutting headcount, and put the saved hours into work you have been postponing for two years.

    Hacker News ↗
  4. Authors and publishers are fighting over who gets Anthropic's $1.5 billion payout

    The settlement over books used in training is done; the argument is now about how the money is split. It matters to you only if your business produces content someone else might train on — course material, manuals, technical catalogues, published research. If so, spend an hour adding an AI training clause to your standard contract before you next license anything out.

    NDTV ↗
  5. AI coding is everywhere, but handing over the whole development cycle backfires

    The reported pattern: the speed gains are real at the writing-code stage and disappear when nobody owns the design and review. This is the single most useful thing to know before you commission software this year. Ask your developer or agency one question — which human reads this before it reaches my staff — and treat a fuzzy answer as a price you will pay later.

    Pasquale Pillitteri ↗
  6. An agency is selling a playbook for getting recommended by Microsoft Copilot

    A press release announcing a consulting playbook on ranking inside AI assistant answers. The underlying shift is real — buyers increasingly ask an assistant for a shortlist instead of running a search — but the paid optimisation industry is arriving well ahead of any evidence it works. Do not sign an AI search ranking retainer this year; do make sure your website says plainly what you do, where you do it, and who for.

    Eagle-Tribune ↗
  7. Excel's new change tracking shows who edited a cell without the archaeology

    The most immediately useful office software change this week is not AI at all — it just tells you who changed the number in the shared sheet. If your fee ledger, stock list or payroll working file is a spreadsheet three people edit, this quietly removes a recurring argument. Small, boring, worth ten minutes of your office manager's time.

    How-To Geek ↗
  8. A plain-language glossary of the AI jargon you are about to be sold

    TechCrunch has published definitions for the terms now filling vendor pitches, including phrases like opaque recurrence. Keep it open during your next demo. Most of the vocabulary describes ordinary things, and knowing that is the cheapest negotiating advantage available to a non-technical buyer.

    TechCrunch AI ↗

How this briefing is put together

Every morning we read the day's AI announcements and reporting from the companies themselves and from the technology press, then pick the handful that actually change something for a working business. The analysis is ours and it is written for owners and managers, not engineers. Every story links to its original source above — read them, and disagree with us where we've got it wrong.

More editions

Published daily
A new edition every weekday morning, dated and kept permanently at its own address.
Every claim sourced
Each story links to the original announcement or report. Read them and disagree with us.
Written for owners
No benchmark scores or parameter counts — just what a development changes for a working business.

We use cookies

We use cookies to enhance your browsing experience, analyze site traffic, and personalize content. Learn more