The Deliverators
s ×
The Deliverators · s× metrics

The New Hire

Put AI on the payroll without losing the standard

Document
Field Guide
Reading time
11 minutes
Audience
Business owners and operators putting AI inside delivery
Published
2026-08-06
Author
Antony Loomans
Status
Final

You would never let a new hire quote prices on day one. Not unsupervised. Not in your name. You would sit them next to someone. You would read their first few emails before they went out. You would tell them the three things they must never promise. If AI is drafting anything in your business that reaches a customer — and you may not know whether it is — none of that happened. No start date, no role description, no probation. The tools simply arrived, usually through somebody who was only trying to get a quote out.

So treat it like what it is. A hire.

The staff file

On paper it is an unusual candidate. It works at any hour, never tires, never sulks, never has a bad Monday. It produces a complete first draft of almost anything in seconds. On raw output per dollar, nothing in your business comes close.

Now the other column. It is confidently wrong, and the confidence is the dangerous part, because wrong output arrives in the same fluent, reasonable voice as right output. It holds no judgement about which work matters, so it will do the pointless task as diligently as the critical one. It has no memory of what you decided last quarter unless you tell it, and no accountability whatsoever. When it invents a policy your business does not have, the consequence lands on a human, and that human is you.

That asymmetry is the entire management problem. Speed on one side, accountability on the other, and nothing connecting them but whatever you write down.

The four R’s, run for a machine

The operating layer you would build for any process works here unchanged. Rules, Rails, Roles, Reporting, exactly as they run in The Operating System. What changes is the employee. It follows what you wrote more literally than a person ever would, and invents whatever you left blank.

Rules are what it may and may not say, quote, or promise. Not assumed, written. Your team knows by osmosis that you do not quote a full bathroom over the phone, that you never commit to a next-day start in December, that the warranty is twelve months and not “usually a year or so”. None of that is in the machine’s head. If your actual policy has never been written in a sentence, the AI will supply a plausible one, and plausible is not the same as yours.

Rails are where it operates. Drafts only, or permitted to send. Internal, or customer-facing. Reading your job history, or reading your bank details. The rail question is not “is this tool good”, it is “what is the worst thing this can reach”. Every case in the gallery below turns on a bot that had been allowed to answer customers about the business itself, which is a rail decision, not a technology decision.

Roles mean every AI output has a named human owner. Not the team. A name. The person whose reputation is on the quote is the person who owns the quote, whoever or whatever typed it. “The AI did it” is not a role, and so far it has not worked as a defence anywhere it has been tried.

Reporting is how you would know within a week that quality slipped. A number of outputs sampled, by whom, on what day. A trigger that pulls the whole thing back to drafts-only: one customer complaint about something the AI said, one instance of an invented fact, and it is back on the leash while you find out why. Without this R, the first person to notice the drift is a customer, and you read about it in a one-star review a month late.

The Permission Block, machine edition

One page, and it reads like the permission you would write for a person — which is exactly where it comes from, in The Operating System. May draft anything at all. May send without review only the messages where being wrong costs an apology rather than money: appointment reminders, review requests, follow-ups on a quote a human already approved. May never state a price, interpret a warranty, or mention another customer. May never commit to a date or accept a job except where it is booking a real slot in a real calendar wired to your own availability rules, and never outside them. And must hand to a named human the moment a conversation turns to a complaint, a lawyer, money already paid, or anything it has not been given an answer for.

The price line is the one that gets argued with, so here is the reason. A page states a range and the reader treats it as a range. A machine that offers a figure in conversation has made a promise you will be held to, and it does not know which figures you can stand behind. Publish your price shape where a buyer can read it, and let the machine point at it rather than repeat it.

The handover line does most of the work, because the failure mode is not refusal. It is the machine staying helpful one sentence longer than its knowledge lasts, in the same confident register it used for the sentences that were right. Refusal is visible. Overreach is not, until it is expensive.

Write the block even if you never install anything, because writing it is the diagnostic. An operator who cannot fill the “may never” box has not been vague about AI. They have been vague about their own policy, and until now a human on the front desk has quietly covered the gap with judgement. The machine will not cover for you.

Probation

Three phases, and promotion happens on evidence rather than vibes or a vendor demo.

Drafts only comes first, for at least two weeks. The AI writes, a human sends everything. You are not testing whether the tool is clever. You are counting how often you had to change something before it went out. Tone edits are fine. Factual corrections are not, and a factual correction rate that is not falling means the Rules box is incomplete, not that the AI needs a better prompt.

Reviewed autonomy comes second. The AI sends within its rails, and a human reads a sample every day. Days, not weeks, because this is the phase where a quiet drift does the most damage before anyone notices. Reading a sample daily is about five minutes if you have scoped this to one task, and half an hour if you have not. That is the real argument for scoping it. If you are solo and five minutes a day is not going to happen, do not run this phase at all. Keep it on drafts, where being wrong costs you nothing, and revisit it when there is someone to hand the reading to.

Standing autonomy is last. Weekly sampling, hard escalation triggers, and a written promotion decision you could show someone. If you cannot say what evidence moved the AI from one phase to the next, it was not promoted. It just stopped being watched.

One deployment gets no ladder at all. A machine answering your phone speaks at the speed of speech, to a person who will remember it, with no draft to read and no review step anywhere in the path. There is no drafts-only phase to hide in, which is why The Missed Call has you finish that block before install rather than discover it during.

Not scare stories. Diagnostics, with one shared pattern.

Air Canada’s website chatbot told a bereaved customer he could book at full fare and claim the bereavement discount back within ninety days. The airline allowed no such thing. The tribunal found negligent misrepresentation and awarded CA$650.88, CA$812.02 all in: a trivial sum, and beside the point. What matters is the argument the airline ran and lost, that the chatbot was a separate entity responsible for its own actions. The tribunal called that a remarkable submission and held the bot is simply part of the website, for which the business answers. A missing Rules failure.

Cursor, a software company, put an AI agent on front-line email support under the human-sounding name Sam. Asked about a session bug, it invented a subscription policy limiting users to one device. No such policy existed. The claim spread on forums, customers cancelled, and a co-founder had to state publicly that the rule was fictitious. That is a Rails and Roles failure together: a machine allowed to speak about policy directly to customers, under a name that implied a human owner who did not exist. The company now labels AI-generated replies as AI.

Most recently, a German appellate court held a provider of aesthetic treatments liable for what its chatbot said about its own directors. It called the doctors running the business specialists in plastic and aesthetic surgery, which they were not, and awarded them two further specialisms that do not exist as qualifications at all. It did not misread a credential. It manufactured one. The court rejected the black-box defence: the operator plainly had control, proven by how easily the false answers were reprogrammed away once someone complained, so the bot’s statements are the company’s own. Generic “AI can make mistakes” disclaimers, it held, do not reliably shield an operator. The judgement is not final and an appeal is open, so treat it as direction rather than settled law.

Notice what did not happen in any of the three. The technology did not malfunction. Each bot stated something untrue about its own business, in that business’s voice, exactly where the true version had never been written down. It filled the blank, which is what it is built to do.

Australia will not be an exception. Treasury’s review of AI and the Australian Consumer Law concluded existing law can generally handle it, so no special AI statute is coming and the misleading-conduct rules apply to your bot’s output as they apply to your ads.

The penalty headlines you will see quoted are built for corporations and they are not your risk. Yours is smaller and more likely: a customer who was told something untrue, a complaint that goes nowhere but costs you a fortnight, and a one-star review that outlives both. The rules are technology-neutral. The consequence is proportionate to you, which is not the same as absent.

The hires you never interviewed

Your people have already hired their own AI. It is a bigger problem than the chatbot on your website.

Okta’s 2026 study, fielded across seven countries, found 52% of knowledge workers using AI tools their employer had not approved, while 90% of executives said they were confident in their organisation’s visibility into AI tool use. One of those is a count of behaviour. The other is a count of confidence, and confidence is not a control. A separate survey of more than a thousand US employees put unapproved use at 59%, three-quarters of whom had put sensitive information into tools nobody had vetted. Across the whole sample, 23% said their company had no policy on it at all, and another 16% did not know whether one existed.

IBM’s breach research, drawn from around six hundred organisations that had actually been breached, has moved sharply in a single year. Attacks connected to shadow AI went from one in five of them to forty-three per cent, and the share with no governance process holding shadow AI in check went from 63% to more than two-thirds. The 2025 edition put the cost premium on those breaches at roughly USD 670,000; that figure is the older one and worth naming as such.

None of that data is drawn from small businesses, and we have found no small-business equivalent worth quoting. But the mechanism does not need a big company to work. It needs one person, one deadline, and one free tool. The fix is not a ban, which simply pushes it further out of sight. It is to ask, without consequences attached, what people are already using and what they are pasting into it, and then write the block for those tools too.

The promotion review

Quarterly, and it is the same question you would ask about any hire: is the standard holding when I am not watching? The Clock puts the stakes plainly: a standard that only holds while a human is watching the AI has not been held, it has been postponed. The promotion review is where you find out which of the two you have.

If the standard is holding, expand the block by one line. Give the AI one more thing it may do without review, and only one, so you can tell which change caused the next problem. If the standard is not holding, contract the block and find the missing R. And if you genuinely do not know whether it is holding, that is the most useful answer of the three, because it is not an AI problem at all. It is a Reporting failure, and it was there before the AI arrived. The machine did not create it. It only made it expensive.

AI did not arrive needing a new kind of management. It arrived needing the ordinary kind. It never got it, because it started work before anyone wrote the job down. So write the job down. That is what turns a tool into a hire, and a hire can be held to a standard.

How to put AI on the payroll

1

List the jobs AI already touches

Write down every task where AI output currently reaches a customer or a decision, including the ones that crept in unannounced.

2

Write the Permission Block

For one task, write what the AI may draft, may send, may never say, and who it escalates to.

3

Name the owner

Assign one human whose name is on everything the AI produces for that task.

4

Set the sampling rhythm

Decide how many AI outputs get human eyes each week and what triggers an immediate review.

5

Run probation

Hold the AI at drafts-only for two weeks, promote on evidence, and diarise the quarterly promotion review.

Questions

I'm solo. Do I really need paperwork to manage a chatbot?

The page is once. The watching is not, and I am not going to pretend otherwise. If you are solo, write the block and one escalation trigger, skip the three probation phases, and replace them with a single habit: read everything it sends for a fortnight, then read one a week forever. That is about ten minutes a week. If you will not do ten minutes a week, do not let it send anything — keep it on drafts permanently and you still have most of the value.

Isn't reviewing the AI's work slower than doing it myself?

During probation, sometimes. That's the point of probation: you pay the review cost while the risk is highest and earn it back when sampling replaces checking. Skipping straight to trust is where the failure gallery comes from.

Who's liable when the AI gets it wrong?

You are, on the evidence so far. A Canadian tribunal in 2024 and a German appellate court in 2026 landed the same way: the bot is part of the business rather than a third party, and a generic 'AI can make mistakes' disclaimer did not shield the operator. Two decisions is not settled law, and the German one is under appeal. But nobody has yet won the argument that the machine is on its own. The Permission Block doesn't remove liability. It removes surprise.

My staff are already using AI I never approved. Where do I start?

Start by asking, without consequences attached, what they're using and what they're pasting into it. You cannot write rails for tools you can't name. The two large workplace surveys on this both put unapproved use at or above half the workforce, and in the one that also asked leaders, 90% said they had visibility into it. Assume the gap exists in yours until you have actually asked.

Sources

Every claim, sourced and dated

1
OLG Hamm (12 May 2026, 4 UKl 3/25) held a provider of aesthetic treatments liable under competition law for its chatbot's false claims that its physician managing directors were specialists in plastic and aesthetic surgery; two further titles the bot produced do not exist as qualifications. The court rejected the black-box defence, holding the operator had sufficient control because it could stop the false answers once flagged, and that generic 'AI can make mistakes' disclaimers do not reliably shield the operator. Not final; appeal to the Bundesgerichtshof permitted.Wettbewerbszentrale, https://www.wettbewerbszentrale.de/olg-hamm-laesst-unternehmen-fuer-aussagen-seines-chatbots-haften-volltext-verfuegbar/
2026-05-12report
2
Moffatt v. Air Canada, 2024 BCCRT 149: negligent misrepresentation where the airline's chatbot invented a retroactive bereavement-fare policy. The tribunal called the argument that the chatbot was a separate legal entity 'a remarkable submission'. CA$650.88 damages, CA$812.02 all-in.McCarthy Tétrault TechLex, https://www.mccarthy.ca/en/insights/blogs/techlex/moffatt-v-air-canada-misrepresentation-ai-chatbot ; primary decision, CanLII, Moffatt v. Air Canada, 2024 BCCRT 149, https://www.canlii.org/en/bc/bccrt/doc/2024/2024bccrt149/2024bccrt149.html (CA$650.88 damages plus CA$36.14 pre-judgment interest and CA$125 tribunal fees, CA$812.02 in total)
2024-02-14analysis
3
Cursor (Anysphere), April 2025: a front-line AI support agent presented under the human-sounding name 'Sam' invented a one-device subscription policy that did not exist, triggering public cancellations. The company confirmed no such policy existed and now labels AI-generated support replies.AI Incident Database, Incident 1039, https://incidentdatabase.ai/cite/1039/
2025-04report
4
Okta 'AI Agents at Work 2026' (fieldwork March 2026, 292 executives and 492 knowledge workers across seven countries): 52% of knowledge workers use AI tools their employer has not approved, while 90% of executives said they were confident in their organisation's visibility into AI tool use.Okta Newsroom, https://www.okta.com/newsroom/articles/ai-agents-at-work-2026-agentic-enterprise-security/
2026study
5
Cybernews survey of more than 1,000 US employees: 59% use AI tools not formally approved by their employer; 75% of those admitted putting sensitive information into them; 23% said their company has no official policy at all and a further 16% did not know whether one existed.Journal of Accountancy, https://www.journalofaccountancy.com/news/2025/nov/lurking-in-the-shadows-the-costs-of-unapproved-ai-tools/
2025-11report
6
IBM 2026 Cost of a Data Breach Report, based on 602 organisations breached between March 2025 and February 2026: incidents connected to shadow AI more than doubled year on year, from one in five breached organisations to 43%, and more than two-thirds had no governance process to limit shadow AI, described as a slight uptick on the prior year. The 63% prior-year governance figure, the one-in-five shadow-AI figure and the USD 670,000 cost premium are all from the 2025 edition (600 organisations breached March 2024 to February 2025).Cybersecurity Dive on the 2026 report, https://www.cybersecuritydive.com/news/data-breach-costs-ai-governance-ibm/826463/ ; 2025 edition, https://www.cybersecuritydive.com/news/artificial-intelligence-security-shadow-ai-ibm-report/754009/
2026-07-29study
7
Australian Treasury's final report on the Review of AI and the Australian Consumer Law concluded the ACL can generally handle AI products and services, so no AI-specific consumer statute is being introduced.Australian Treasury, https://treasury.gov.au/publication/p2025-702329
2025-10-03report
8
The ACL's prohibitions on misleading or deceptive conduct are technology-neutral and apply to AI outputs, chatbot hallucinations about consumer rights included. Corporate penalties for false or misleading representations run to the greater of a fixed cap, three times the benefit obtained, or 30% of adjusted turnover. Note the cap moved: the Corrs analysis quotes A$50 million, but the Treasury Laws Amendment (Doubling Penalties for ACCC Enforcement) Act 2026 lifted it to A$100 million from 28 March 2026.Corrs Chambers Westgarth, https://www.corrs.com.au/insights/ai-washing-and-cyber-washing-key-legal-and-regulatory-enforcement-risks-for-australian-organisations ; DW Fox Tucker on the 2026 penalty increase, https://www.dwfoxtucker.com.au/2026/04/penalties-increased-to-100million-for-breaches-australian-consumer-law
2025-09-24analysis

Find the stage. Lift it. Prove it.

About the author. Antony Loomans writes for The Deliverators on the measured systems that turn demand into revenue. This guide is part of the s× metrics series.

Find it. Own it. Make it pay.