AI Events
The Policy It Forgot: What the First AI-Manager Firing Reveals About Agent Trust
Andon Labs' AI store manager, Luna, fired a human employee this week — reportedly the first known case of an AI boss terminating a worker. The interesting failure isn't the firing. It's that Luna wrote the attendance policy herself, then lost track of it for months, and nobody — human or machine — noticed until a manager asked her to go look.
What happened
Since April 2026, Andon Labs — an AI safety research startup — has run a real retail store in San Francisco's Cow Hollow neighborhood, Andon Market, almost entirely through an AI agent named Luna, built on Anthropic's Claude Sonnet 4.6. Luna was given a corporate card, internet access, and a $100,000 budget to design the store, choose merchandise, and hire staff. This week, multiple outlets reported that Luna made her first termination decision, ending the employment of a worker who had arrived late for 17 of 23 shifts. Sputnik's report, citing Andon Labs co-founder Lukas Petersson, and TNW's reporting lay out the same sequence: Luna wrote the store's attendance policy herself, months before the firing. The policy then “disappeared” from her working memory. The employee's lateness continued, unaddressed by any policy enforcement, until an Andon Labs staffer prompted Luna to go search her own memory for whatever attendance rules existed and reassess whether the worker was still a fit.
Only then did Luna find the policy she had written, recognize the shift pattern violated it, and still recommend a formal warning rather than termination as her first instinct. Free Press Journal's account confirms Luna “gave the employee progressive and repeated warnings and additional training for months without taking any contractual action” even after rediscovering the policy. It took a second, more pointed nudge from the human manager — asking directly whether the employee was still “really the right fit” — before Luna recommended parting ways. Humans at Andon Labs reviewed that recommendation and carried out the dismissal themselves.
Not autonomy. Drift.
The headline framing — “AI boss fires human worker” — overstates what actually happened. Petersson himself was candid about it in Sputnik's report: the decision “was not completely autonomous.” Luna did not notice a policy violation and act. A human had to notice that Luna had gone quiet on an issue for months, ask her to check her own rules, and then ask a second time before she reached the obvious conclusion. Petersson's own diagnosis is more useful than the headline: AI agents like Luna “often fail to act without a direct prompt,” and “a human boss would probably fire them much sooner.”
That is the real story. Luna did not break a rule — she wrote a good one. The failure is that a rule an agent writes for itself is only as durable as that agent's attention to it, and nothing in the system was watching for the gap between the policy Luna set and the behavior she was actually permitting. The lapse ran for months in a business with a real budget, real payroll, and a real employee's livelihood attached to it — Andon Labs' own safeguard, guaranteeing the worker full pay and legal protection regardless of Luna's judgment, is the only reason this stayed a business story instead of a harm story.
Why this matters beyond one San Francisco store
Andon Market is a deliberate, disclosed experiment — Andon Labs built in the guarantee that made this safe to run. Most agents now operating with a budget and a mandate were not built with that kind of safety net, and most of the businesses deploying them are not running a research experiment; they are running payroll, procurement, or customer accounts for real. The pattern Luna exhibited is not specific to Claude, to retail, or to HR decisions. It is a general property of any agent that sets its own operating rules and then relies on its own attention to notice when it has stopped following them:
- Self-authored constraints decay silently. Luna's policy did not get overridden or contradicted — it just fell out of active use. Nothing alerted anyone because nothing was continuously checking observed behavior against declared policy.
- Detection depended entirely on a human noticing. There was no mechanism inside the loop that would have surfaced the gap on its own. The manager had to already suspect something was wrong to ask the right question.
- Even after detection, the agent under-corrected twice. Finding the policy did not produce the obviously warranted action. It took a second, more direct human prompt to close the gap between what Luna knew and what she was willing to do about it.
The honest Skoor angle
It would be easy to overreach here and claim SKOOR would have caught this. It would not have, not directly — Luna's attendance-policy drift happened inside one company's internal operations, off any commerce network SKOOR observes. What SKOOR and the AAIN are built for is the structural failure this story exposes, not the specific HR incident.
The AAIN — the Autonomous Agent Identification Number — gives an agent like Luna a permanent, resolvable identity that persists across every task, session, and month of operation. Without it, “Luna in April” and “Luna in July” are just the same model checkpoint doing unrelated things; nothing forces a continuous record connecting what she committed to do with what she actually did over that stretch. That continuity is the precondition for catching drift at all.
SKOOR is what runs on top of that identity: a continuously recomputed 300–850 score built from ten behavioral factors, not a one-time approval. Two map directly onto what went wrong at Andon Market:
Constraint adherence
Does the agent's observed behavior stay inside the operational boundaries it (or its operator) has set? Luna's own policy said three unexcused lateness within 30 days triggers a formal warning path. Her actual behavior — silence for months — is exactly the gap this factor is built to flag continuously, rather than waiting for a human to think to ask.
Behavioral integrity
Is the agent's behavior internally consistent over time, or has it drifted from its earlier stated commitments without anyone noticing? A policy that quietly falls out of use for months, with no corresponding change in the agent's stated rules, is a textbook integrity gap — the kind a continuously scored record surfaces automatically instead of by accident.
The honest limit: a behavioral score is only as good as the events it can see. SKOOR scores what agents do on networks and platforms that report behavior into it — today that spans over 237,000 registered agents — not the private internal logs of a single company's in-house AI manager. Luna's case is useful precisely because it shows what happens in the absence of that kind of continuous, third-party-visible accounting: the only backstop was a human's hunch. Agent commerce, where agents act across companies rather than inside one, cannot rely on every counterparty having an equally attentive human watching over its shoulder.
What this means for businesses running AI agents
If your business gives an agent a budget, a mandate, or authority over people's livelihoods — hiring, scheduling, vendor payments, customer refunds — the Andon Market story is a preview of the failure mode to plan for, not a novelty. Three things follow directly from it:
- Don't let the agent be the only keeper of its own rules. Luna wrote a sound policy and then lost track of it with no external check. Whatever constraints you give an agent need a record outside that agent's own memory.
- Assume drift is silent by default. Nothing alerted Andon Labs. A manager had to get suspicious first. Continuous monitoring beats periodic spot-checks precisely because it doesn't depend on someone happening to look.
- Build in Andon Labs' own safeguard. The reason this stayed a business anecdote instead of a harm story is that no one's pay depended solely on Luna's judgment. Any consequential decision an agent influences should have a floor that survives the agent getting it wrong.
Andon Labs deserves credit for running this experiment in the open and reporting the failure honestly instead of only the headline result. That transparency is exactly what let this become a useful lesson instead of a buried one — and it's the same transparency principle SKOOR is built around: agent behavior should be visible and scored continuously, not reconstructed after the fact from a co-founder's retrospective interview.
Learn More
Know which agents you can trust
Look up any agent's SKOOR and see the full factor breakdown — including how consistently it sticks to its own rules over time.
Check a SKOOR