Yesterday, about 1:20 in the afternoon — the same time I started speaking today — an agent inside a Thai enterprise saw a signal. Ten seconds later it had an explanation. Twenty seconds later it had taken an action. A minute later, a human was asked to approve.
Was that human oversight, or was that human simply informed?
I am not here to tell you what an AI agent is. You know. I am here to argue that the question most organizations are asking is the wrong one — and to tell you what the right one is.
Open cold. No bio, no agenda. Let the MC read your title. Beat after "informed".
Bridge → slide 2I'm not going to tell you what an agent is; you know. I'm going to argue that the question most organisations are asking is the wrong one. But first — why is this happening now, and not two years ago?
One: on OSWorld, agents doing real computer work went from 12% to 66.3% in roughly a year — now within about six points of human performance.
Two: those same agents still fail roughly one attempt in three. Capability rose faster than reliability. That is the jagged frontier.
Three: Gartner predicted, in June 2025, that more than 40% of agentic AI projects would be cancelled by the end of 2027 — “due to escalating costs, unclear business value or inadequate risk controls”. Not because the models failed.
GUARDRAIL: a 2025 prediction about 2027, never completed cancellations. The official wording is “or”, not “and” — any one cause suffices.
Number one is why you must act. Number two is why you cannot act the way you used to. Number three is what happens if you only do number one.
Say the numbers slowly. Three only — resist adding a fourth.
Bridge → slide 3Capability went up. Reliability didn't. But there's a third thing that changed — and it's the one that makes the old playbook dangerous rather than merely slow.
The risk has changed shape. We used to fear AI saying the wrong thing — a wrong number on a dashboard. Now we have to fear AI doing the wrong thing — a wrong action in a production system.
A wrong answer gets read, and someone can push back. A wrong action just carries on.
And here is the part most speakers miss: a wrong agent action becomes the input to the next agent's decision. Error stops being an event and becomes a chain. That is why governance at launch is structurally insufficient — the thing you approved is not the thing that will be running in six weeks.
Bridge → slide 4So if approving it once doesn't work, we had better be precise about what we are governing. Here is the whole machine, on one slide.
Most organizations have automated Sense and Analyze. Agentic analytics is the claim that Recommend and Act can be automated too.
Explain and Learn are the two everyone skips — and they are the two that decide whether you can defend the decision afterwards.
Keep this loop in your head for the whole talk. Everything that follows is this same loop — first run end-to-end by an agent with nobody watching, then scored, then deployed three times in your own organization.
This is the map slide. Point at Recommend and Act when you say "that is the new part", and at Explain and Learn when you say "that is the skipped part". You will call back to this loop on slide 5 and again on slide 10 — say "this loop" both times so the callback registers.
Bridge → slide 5That is the machine. So — what happens when you let an agent run the whole loop, end to end, with nobody watching? Somebody tried it in public, and on the numbers it looks like a complete failure.
In March this year, Andrej Karpathy released a project called autoresearch. He let an AI agent do machine-learning research on its own, overnight — sense, analyze, recommend, act, learn. That is the loop from the slide before, running unattended. Edit code, run experiments, measure, iterate, with nobody watching.
Two days. Around 700 experiments. About 20 of them were real improvements. That is roughly 3% — wrong 97% of the time. And the result was still valuable: an 11% speed-up when those improvements were applied to a larger model.
GUARDRAIL: Fortune (17 Mar 2026) confirms 700 / ~20 / 11% from Karpathy's own post. Do NOT say “2.02 hours to 1.80 hours” — those timings are unsupported and were withdrawn on verification.
In a moment I am going to tell you why that 3% was acceptable — because I think it is the heart of everything I want to say today.
Ask the question and let it hang for one real beat — the answer is the very next slide, so the pause is short, but it must be a pause. Do NOT resolve it on this slide. Never name the metric, the model, or what it was training. Not "agentic analytics" — call it an agent running an analytical loop end-to-end.
Bridge → slide 6Ninety-seven percent wrong. And still worth doing. Why? — pause here. The whole talk turns on the next slide.
So — 700 experiments, about 20 right. Why was that acceptable?
Not because the agent was accurate — it wasn't. Because being wrong reverted automatically. Five minutes per experiment, one metric, and if it wasn't better, throw it away.
He didn't design the agent not to fail. He designed the arena so failing was cheap.
Now take an agent with exactly the same accuracy and put it on loan approval. One thing is different: being wrong cannot be undone.
This is why the real dividing line is not accuracy. It is reversibility. And notice what that means for governance: his agent was MORE useful because the environment was engineered for cheap failure. The controls are what bought the autonomy, not what limited it.
The most valuable 25 seconds in the talk. Slow down on the bold line. Let the slide stay almost empty — they should be looking at you.
Bridge → slide 7Reversibility. That's the line. Which raises the obvious question — how do you measure it on your decisions? Because “it depends” is not a policy.
| 0 | 1 | 2 | ||
|---|---|---|---|---|
| Consequence | Can it affect a person, an institution, a public service? |
Internal only | One customer | Many people |
| Irreversibility | If wrong, what does it cost to undo? | One click | Costly but possible | Cannot be undone |
| Contestability | Must the affected person be able to appeal? |
No one asks | Might be asked | Legal right to an explanation |
| Value conflict | Must someone choose between two defensible answers? |
Single objective | Mild tension | Genuine trade-off |
When people ask where humans still matter, the popular answer is empathy and creativity. It sounds good and you cannot use it. You cannot write empathy into a policy. You cannot audit empathy. We need an answer that can be tested.
Four tests. Human judgment should increase as a decision scores higher on each of them.
Consequence — can it materially affect a person, an institution, a public service or a market? Irreversibility — if it is wrong, what does it cost to undo? Contestability — must the affected person be able to question or appeal it? Value conflict — must someone choose between two defensible answers?
On the fourth test, if asked. The first three tests are about the outcome: how big, how recoverable, how challengeable. The fourth is about the objective itself. An agent can optimise an objective; it cannot choose the objective. The trade-off is almost never good against bad — it is good against good, and the tell is that the benefit and the cost land on different people.
Two examples that land in this room. Fraud detection: tighten the threshold and you stop more fraud and freeze more legitimate accounts. That is not an accuracy problem you can engineer away — it is a decision about how much inconvenience to impose on the innocent many to stop the guilty few. And for the public-sector half of the room: benefit fraud screening. Tighten it and you exclude genuinely eligible people. Both errors are real, they fall on different populations, and they carry different moral weight.
The diagnostic to give them. Ask your team what the agent should optimise. If the honest answer needs an "and" plus a weight — maximise recovery and minimise hardship, weighted somehow — that weight is the value judgment. Whoever picks the number is making policy. An agent picking it is unelected policy-making at machine speed.
Kill the lazy answer first (45 sec), then land the four tests.
If challenged on why this is a separate test: a decision can be cheap, reversible and uncontested and still embed a value choice — and repeated ten thousand times a day that choice BECOMES your policy. Conversely, shutting down a failing turbine is high-consequence and irreversible with zero value conflict. The axes are independent. That asymmetry is the argument for four tests rather than three.
Bridge → slide 8Four tests, nought to two each. So — what do you do with the score?
Score zero to eight, and the score sets the autonomy tier.
Thirty seconds, please. Think of one use case your team is about to ship, and score it on those four tests.
… Anyone at six or above, with a plan to let the agent act on its own — that is the gap you need to go and fix tomorrow.
The only audience-interaction moment in the talk, and the slide they photograph. Give them the full 30 seconds — count it silently.
Bridge → slide 9Now, you might reasonably be thinking: that's a tidy framework, but it's his framework. So let me show you who else asked for exactly this — and stopped one step short.
This framework is not new in principle.
The Bank of Thailand's AI risk-management guidelines for financial service providers, issued 12 September 2025, say human oversight must be "calibrated to the level of risk and impact". Article 14 of the EU AI Act says the same thing — oversight proportionate to risk and to the level of autonomy.
Both say you must. Neither says how.
And notice which example the Bank of Thailand chose itself: loan approval — the same example I used thirty seconds ago. They said oversight must be proportionate. They did not say how to measure the proportion. Those four tests are the measuring instrument, and they are the ones I actually use.
GUARDRAIL: say "แนวปฏิบัติ / guidance", never "regulation" — someone from a bank or the BOT may be in the room. Thailand has NOT enacted an AI law.
This is the slide that says: I read the regulation, I found the gap, I built the tool. Do not rush it.
Bridge → slide 10So the requirement exists and the instrument answers it. Now — where does this show up in your work? Go back to the loop I drew at the start, and watch it deployed three times.
| Operations | Customer | Risk | |
|---|---|---|---|
| Signal | Latency or error spike | Repeat contact, rising frustration | Shift in transaction behaviour |
| Agent may | Diagnose, roll back a safe change | Assemble context, draft a resolution | Rank cases, prepare the file |
| Frontier score | Irreversibility 0rollback is one click | Contestability 1–2precedent, emotion | 2 · 2 · 2never closes the case |
| Human enters | Anything customer-visible | Emotional or precedent-setting | Every consequential decision |
| Measure | MTTRand rollback rate | Resolutionand re-contact rate | Detectionand false positives |
Point back at the loop from slide 4. Same loop, three arenas — same structure all three times. What changes is where the human enters.
Now notice the last row. Every arena has two metrics, and the second one is always the one nobody reports. Rollback rate. Re-contact rate. False-positive burden.
If your agent dashboard only tells you how fast it went, you are measuring half the system. Every agent needs a paired metric: one for how much it helped, one for the burden it created — and for those of you in government, your false positives are citizens.
Derive each entry point from the score — this is the point of the slide: Operations may act because rollback is one click, so Irreversibility scores 0. Customer needs a human once Contestability and Value conflict start scoring. Risk never closes the case because it is 2 · 2 · 2.
In V1 this table asserted three entry points. Here it proves the instrument the room just scored themselves with. Walk the Frontier-score row, not just the Human-enters row.
Bridge → slide 11So you score the decision, you set the tier, you put a human at the boundary. Are you done? No. And this is the part I find most dangerous, and least discussed.
Human-in-the-loop does not always mean human control. Sometimes it just means we found someone to hold responsible in advance.
A reviewer is on the hook rather than in control when they see a recommendation and a confidence score but not the evidence; when refusing costs them more than approving; when the review rate makes review impossible; or when they can stop this instance but not the system.
Researchers have argued that current agent designs do not support effective oversight — they contribute to its degradation, because the capacities oversight depends on are eroded by extended AI use.
And Romeo and Conti's 2025 review in AI & Society — a PRISMA-guided review of 35 peer-reviewed studies — finds that explanations, on their own, are often insufficient to improve decision accuracy or to mitigate automation bias, and that “overly technical, cognitively demanding, or simplistic explanations may inadvertently reinforce misplaced trust”.
Here is the part I find most dangerous and least discussed: the more often the agent is right, the worse its supervisor gets at supervising. We are not just designing agents. We are designing the competence of the people who have to oversee them.
Three responses, all cheap. Rate-limit reviewers — if the queue exceeds what a human can genuinely examine, fix the autonomy tier, not the human. Seed known-bad cases into the review queue; if they get approved, your oversight is decorative and now you can prove it. And measure the override rate.
If you remember one thing today, let it be this: go back and ask what your reviewers' override rate is. If nobody is collecting that number, you do not yet know whether your oversight works.
GUARDRAIL: "นักวิจัยเสนอไว้ว่า / researchers have argued" — it is a position paper, NOT an experiment. Never "a study found". The AI & Society item IS a literature review, so "a review finds" is fair there.
Bridge → slide 12So: measure the override rate. But measurement tells you whether oversight is working — it doesn't give the agent its boundaries in the first place. So what does real control actually look like, operationally?
We are used to governance as a gate — approve it before production, done. But an agent does not make one decision. It decides continuously, and a gate cannot govern something that moves. Controls have to run at the same speed as the thing they control.
So every agent gets an explicit, written, finite budget on four dials.
Data: minimum necessary, not whatever the warehouse happens to hold. Action: which actions without approval, and which are simply forbidden. Impact: a ceiling per action and cumulatively per day.
And the fourth, which almost nobody has — Time. Autonomy should have an expiry date. If an agent was granted its authority six months ago, under one dataset, under one model version, and nobody has revisited it since, you do not have governance. You have a historical record that once you did.
Optional opener, only if Acts I–IV came in on time: ~two-thirds of organizations name security and risk — not technology — as the biggest barrier to scaling agentic AI. Cut this first if you are running long.
Bridge → slide 13Four dials, and authority that expires. Which leaves one question: what has to be true before you switch any of this on?
Lead with the punch, not the list: “For those of you from government, I'm going to give you four things — and the fourth one isn't an engineering requirement. It's administrative law.”
Four things before any agent acts. An inventory — most organizations already have more agents than they think. A named human owner per agent, not a committee, a name. A kill switch that has actually been tested; untested is unproven. And a decision log good enough to reconstruct why, not just what.
For those of you from government, the fourth is not an engineering matter. It is administrative law. Administrative decisions carry a duty to give reasons, and citizens have a right to appeal. An agent that cannot reconstruct its own reasoning afterwards does not have technical debt. It has produced a decision you cannot legally defend.
And Thailand's draft AI Act — out for public consultation since 9 July 2026 — points clearly this way: a right to be informed, a right to an explanation, a right to contest a decision, and a duty to keep operational logs. It is not enacted.
One more thing: the EU has deferred its high-risk obligations to 2 December 2027. Do not misread that. The deadline moved. The exposure did not. Thailand has not waited — four instruments in roughly a year, from the central bank, the cyber-security agency, the insurance regulator, and draft personal-data guidance for AI. And more important than any deadline: customers and citizens are not waiting until 2027 to start asking why a system decided this about them.
GUARDRAIL: say "ร่าง / draft" every single time, and cite the July 2026 draft AI Act — not the June 2025 principles, which it supersedes. Not enacted as of September 2026. MDES and ETDA are co-hosting, so currency matters here more than anywhere.
This paragraph is the bid for the government half of the room. Deliver it deliberately.
Bridge → slide 14That's what has to exist before you start. Here's what you ask your team.
Name the overlap with slide 12 out loud, or the room hears repetition: “The dials are what you set. These are what you ask.”
You do not start with which agent platform to buy. You start with these five questions, and if your team cannot answer all five, you are not ready to ship.
The fifth is the hardest, and most organizations still cannot answer it — because when three agents work in sequence and the result is wrong, the accountability does not sit with any one of them.
Leave this slide up through the close.
Bridge → slide 15Five questions. And if you remember one sentence from me today…
I opened with something that took one minute. An agent saw a signal, explained it, acted — and then a human was informed.
The question that defines this era is not what AI can do. It is what AI is permitted to do, where it has to stop, and who owns what follows.
If you remember one sentence from me today: do not automate judgment. Engineer where judgment enters.
Thank you.
Stop. Do not add "and if you'd like to discuss further…" — it deflates a clean ending and reads as a pitch. The opportunities come from the idea, not the ask.
EndStop. Stand still. Take the applause. Do not add “happy to discuss further”.