top of page
Rechercher

The agent shipped. Now who owns it?

  • Photo du rédacteur: Tib Bardout
    Tib Bardout
  • 16 juil.
  • 6 min de lecture

Air Canada tried to fire its chatbot in court.

In Nov. 2022, Jake Moffatt booked flights to his grandmother’s funeral. Air Canada’s support chatbot had told him he could claim the airline’s bereavement fare after the fact, within 90 days of buying the ticket (1). So he paid full price, flew, and applied for the refund. Air Canada said no. The real policy didn’t allow claims once travel was over, and the chatbot had told him something the policy didn’t say.

Moffatt took it to the British Columbia Civil Resolution Tribunal. The airline’s defense is the part worth keeping. Air Canada argued the chatbot was, in its own words, “a separate legal entity that is responsible for its own actions”. A company stood up in a legal proceeding and tried to fire its own software, as if the bot were a rogue contractor it had never authorized.

The tribunal didn’t buy it. A chatbot, the member wrote, “is still just a part of Air Canada’s website”, and the airline was responsible for everything on it, static page or bot alike. Air Canada paid.


The ruling landed in Feb. 2024, and it forced a question most teams shipping agents still haven’t asked out loud: when the agent acts, who owns what it did?



The hallucination wasn’t the failure.

It’s tempting to file this under hallucination. The model got something wrong, so you patch it and move on. Except Air Canada’s bot ran in 2022, before ChatGPT, and the tribunal never even established it was an LLM (as opposed to a simpler rule-based system). That’s the point: this failure has nothing to do with how the bot was built.


Today’s agents add a newer way to go wrong. 1st, an LLM can hand back a confident, plausible falsehood. That is the base rate, not a surprise, and anyone who’s shipped one has watched it happen. 2nd, that answer reaches a customer, hardens into a promise, and binds the company, because no one had decided whether the agent was allowed to make that promise. The first is a model problem. The second is an ownership one, and it is the one that costs you. Rule-based bot in 2022 or frontier model today, that gap is identical.



I get called in on other people’s agents. Sometimes to help build one. More often after it has shipped and done something no one can explain. The first question I ask is who owns what it’s allowed to do. Sometimes, the room goes quiet.

That silence is the failure, not the model. The easiest breakdown to fix isn’t technical. It’s that accountability was assumed instead of assigned. And this isn’t only what I run into. Deloitte’s 2025 State of AI in the Enterprise survey (2), which polled 3,235 leaders, found only 1 in 5 companies has a mature way to govern autonomous agents at all.


Deciding what an agent may do is a product job.


Here’s the line that gets blurred.


Engineering owns whether the agent works: latency, uptime, whether the model returns something coherent. That’s real work, and it’s usually well-owned. What the agent may and may not do is a different question entirely. Can it issue a refund? Waive a fee? Change a customer’s plan, close an account, confirm a policy that may not exist? Each of those is a business decision with money and trust attached, and none of them is an engineering call.


That’s product work. Deciding what an agent is allowed to do, under what conditions, with what limits, is the same judgment a product leader brings to a feature, a pricing change, or a permissions model. The twist is that with an agent, the decision stays invisible until a real customer tests it in production.


When no one owns that decision, the rules for what the agent may do never get written down. And rules that were never written can’t be enforced, no matter how good your engineers are at building the thing that would enforce them. The agent ships anyway, because shipping it was an engineering milestone and someone hit it. The boundary just never got an owner.


When it’s everyone’s job, it’s no one’s.

The reason the boundary goes unowned usually isn’t negligence. It’s diffusion. Ask who’s responsible for an agent in production and you tend to get a list, not a name.


A 2025 survey by the Cloud Security Alliance and Strata (3) found ownership split across security teams (39%), IT (32%), and newer AI-security functions (13%), with no clear accountability between them. Only 21% of orgs keep a real-time inventory of their agents, and fewer than a third can trace an agent’s action back to a human.


Diffusion has a nastier failure mode too. I call it the orphaned agent: someone builds an agent, ships it, then leaves the company, and the thing keeps running with no owner for its limits, its updates, or its shutdown. It doesn’t stop when its creator leaves. It just keeps acting, with no one left who knows why it was ever allowed to.


Cursor lived this in public (4). In April 2025, its support agent, “Sam”, told users they were being logged out under a new “one device per subscription” rule it called a core security feature. No such rule existed. Sam made it up. Users believed it, posted it, and canceled. A human stepped in 3 hours later: no such policy existed. The agent sounded official the entire time. Nobody caught the boundary because the boundary was nobody’s.


A security committee is the wrong owner.

Once a company feels this gap, the instinct is to stand up a governance board and hand the problem to security. Better than nothing. Still the wrong owner.


Security can do something essential here: it can stop an agent, scope its access, pull its keys, take it out of production. What security can’t do is decide what the business should let the agent do in the first place. Whether your support agent should be allowed to waive a fee to keep an unhappy customer isn’t a security question. It’s a product and commercial one.


Stack enough functions and you’re back to a committee, which is back to diffusion. Someone has to own the call. Routing it to security is often just a polite way of routing it away from the person whose job it actually is.


Give the agent a product owner.

The fix isn’t another framework. It’s a name. Every agent in production gets one product owner, the same way a feature or a P&L line does. The people who’ve studied this converge on the mechanics: a named human owner for every agent, a real reporting line to a sponsor, and an inventory of what’s running. That much is settled. What’s usually missing is the role.



That owner holds at least 3 things (bigger agents carry more, these are the floor). The outcome the agent exists to move. The boundary, an explicit list of what it may and may not do. And the escalation path for the edge cases it shouldn’t be deciding alone. None of those is a model parameter. All of them are product calls.


I’ve put many agents into production. The ones that lasted had an owner for what they were allowed to do. The ones that didn’t belonged to no one. They kept acting, kept sounding helpful, and nobody held the line on what they could touch. One kept comping plan upgrades to smooth over complaints, and it ran that way for a month before the numbers caught it. By then it wasn’t a bug anyone owned. It was a cost nobody had signed off.


And that cost is the small version. Two liabilities are in play. The company’s, for what the agent puts out: Air Canada paid for its chatbot’s promise, and UnitedHealth is now defending a class action over an AI tool that allegedly denied elderly patients care in bad faith (5). And a personal one, already arriving in regulated fields: lawyers have been personally fined and suspended for an AI’s fabricated citations (6), and under Delaware’s Caremark duty, directors can be held personally liable for failing to oversee a mission-critical AI system (7). Air Canada was an $812 award in 2022, before ChatGPT. The floor, not the ceiling. Each agent needs an owner before the bill arrives, not after.


An agent with no product owner is a liability with a friendly voice.

--------

An agent that demos clean but can’t be trusted with anything irreversible is exactly the kind I get called in to build, or rebuild. If that’s the one you’re sitting on, that’s a conversation worth having.


---------

Sources

(3) Ownership fragmentation and inventory (Cloud Security Alliance / Strata, Securing Autonomous AI Agents, 2025): https://cloudsecurityalliance.org/artifacts/securing-autonomous-ai-agents

(6) Lawyers sanctioned for AI-fabricated citations (Helsell Fetterman): https://www.helsell.com/2026/04/24/ai-hallucinations-keep-costing-lawyers-in-court/

(7) Director personal liability for AI oversight (Fenwick, AI in the Boardroom): https://www.fenwick.com/insights/publications/ai-in-the-boardroom-what-directors-need-to-know-now

 
 
 

Commentaires


bottom of page