Working with Agentic Analytics · Where Human Judgment Still Matters

ONE MINUTE.

Was that oversight — or just a notification?
SIGNAL0s
EXPLANATION10s
ACTION20s
HUMAN TOLD60s

THREE NUMBERS

66.3%
Agents on real computer tasks.
Up from 12% in a year.
1 in 3
Attempts still fail.
Reliability did not follow.
40%+
Agentic projects cancelled
by 2027.

A WRONG
ANSWER
WAITS.

A WRONG
ACTION
KEEPS GOING.

Error stops being an event. It becomes a chain.

THE DECISION LOOP

Sense
Analyze
Explain
Recommend
Act
Learn
Two of these are the ones everybody skips.
700
~20
3%
Two days. An agent running that entire loop, unattended.
Hold that number.
700
~20
3%
He didn't design the agent not to fail.
He designed the arena so failing was cheap.
Same accuracy. Different arena. Loan approval.

THE JUDGMENT FRONTIER

Score each 0–2. Human judgment rises with the score.
012
Consequence Can it affect a person, an institution,
a public service?
Internal onlyOne customerMany people
Irreversibility If wrong, what does it cost to undo? One clickCostly but possibleCannot be undone
Contestability Must the affected person be able
to appeal?
No one asksMight be askedLegal right to
an explanation
Value conflict Must someone choose between
two defensible answers?
Single objectiveMild tensionGenuine trade-off

SCORE 0–8. THE SCORE SETS THE TIER.

0–2
ACT
Agent acts. Human audits a sample afterwards.
3–5
RECOMMEND
Agent proposes. A human approves first.
6–8
ASSIST
Agent assembles evidence. The human decides.

“CALIBRATED TO RISK
AND IMPACT”

Bank of Thailand
AI risk management guidelines for financial service
providers, 12 Sep 2025 — non-binding guidance
EU AI Act, Article 14
Oversight proportionate to risk and to the level
of autonomy
Both say you must.
Neither says how.
Those four tests
are the how.

ONE LOOP, THREE ARENAS

OperationsCustomerRisk
Signal Latency or error spikeRepeat contact,
rising frustration
Shift in transaction
behaviour
Agent may Diagnose, roll back
a safe change
Assemble context,
draft a resolution
Rank cases,
prepare the file
Frontier score Irreversibility 0rollback is one click Contestability 1–2precedent, emotion 2 · 2 · 2never closes the case
Human enters Anything
customer-visible
Emotional or
precedent-setting
Every consequential
decision
Measure MTTRand rollback rateResolutionand re-contact rateDetectionand false positives

IN THE LOOP ≠ IN CONTROL

Sometimes it just means we pre-assigned someone to blame.
No evidence,
only a score
No real authority
to say no
200 approvals a day
is not review
A reviewer with a 0% override rate is a rubber stamp. Measure it.

EVERY AGENT GETS A FINITE BUDGET

Data
What can it reach?
Action
What can it do
unapproved?
Impact
How much, per action
and per day?
Time
When does its
authority expire?

BEFORE ANY AGENT ACTS

1
An inventory — you cannot govern what you cannot list
2
A named human owner — a name, not a committee
3
A kill switch that has been tested — untested is unproven
4
A decision log that reconstructs why  — not just what
For government, number four is not engineering. It is administrative law.

FIVE QUESTIONS

1
Outcome — what is this agent allowed to pursue?
2
Evidence — what data may it use?
3
Authority — what may it do without approval?
4
Boundary — where must it stop and hand over?
5
Accountability — who owns the outcome, including damage across several agents?
Do not automate judgment.
Engineer where judgment enters.
QR code linking to linkedin.com/in/kirati-th Scan to connect
Act I · The Shift · The Delegation Moment
25:00 target
1 / 15  ·  press ? for keys

Keys

→ ↓ Space
Next slide — also PageDown / Enter (clicker forward)
← ↑
Previous slide — also PageUp / Backspace (clicker back)
. or B
Blank the screen (clicker middle button)
H
Hide the bottom bar for full-bleed projection
1–9
Jump to slide
Home / End
First / last slide
S
Speaker notes
T
Start / pause the 25-minute timer
R
Reset the timer
O
Overview grid
F
Fullscreen
Esc
Close panels