The weekly half hour that keeps an agent honest

Every post about running an agent ends with "read a sample of threads weekly", and almost nobody does it, because it is not a task with a clear shape. This is that shape: what to pull, what to look for, and what to do with the findings so the half hour compounds instead of evaporating.

4 min read

Fifteen threads pulled from a week and read end to end
Thirty minutes, fifteen threads, one honest question each.

Sample deliberately

Random sampling is the intuitive choice and the wrong one, because the interesting threads are rare. Pull fifteen, weighted:

  • Five ordinary ones. The bread and butter, to see whether the common case is still good.
  • Three the agent escalated. Was the escalation right, and would you have wanted it sooner.
  • Three that ran long, measured in turns. Long threads are where an agent is failing politely.
  • Two a human corrected. The most valuable two in the set.
  • Two flagged by any signal you have: a reopened thread, a complaint, a bounce, or an unusual recipient.

If a category is empty, that is information: no escalations at all in a week usually means the rules are too loose rather than that everything went well.

Read the whole thread, not the reply

The reply in isolation always looks reasonable, which is exactly the trap. Read from the customer's first message, in order, as they experienced it. The failures that matter only appear that way: an answer to the wrong question, a fact that contradicts message two, a tone that hardens after a complaint.

For each thread, one question: would I have sent this? Not "is it defensible", which everything is.

Then sort into the four buckets from the first week with an email agent: would have sent it, tone fix, grounding fix, should not have answered at all. The distribution is the finding, not any individual thread.

BucketMeansGoes to
Would have sent itWorkingNothing
Right answer, wrong tonePrompt driftPrompt, if it happened twice
Wrong answerGrounding too looseA tool, a source, or an escalation rule
Should not have answeredRule missingTooling, not prose

The numbers, after the reading

Look at the metrics second, not first, because numbers without threads produce confident wrong conclusions. Three that matter, defined in evaluating an email agent:

Correction rate. Trending down is the whole game. Trending up after a prompt change means the change was wrong, whatever the samples looked like.

Escalation precision and recall. Precision falling means the agent is dumping; recall falling means it is answering things it should not.

Reopen rate. The one that catches the confident wrong answer nothing else sees.

Deliverability gets its own weekly glance: complaints, hard bounces, and per-provider placement, per knowing where your agent's mail actually lands.

Threads first, then the numbers, then one change
Numbers without threads produce confident wrong conclusions.

Change one thing

The discipline that makes this compound: one change per week, and it goes into the golden thread suite before it ships.

More than one and you cannot attribute next week's movement. Zero, week after week, means either the agent is genuinely stable, which happens, or the review has become a ritual without teeth, which is more common. If two weeks pass with nothing worth changing, widen what the agent handles rather than continuing to inspect a solved problem.

Fixes go where they belong: tone in the prompt, grounding in tools and sources, and anything that should never happen into an enforced limit rather than an instruction, per designing the tools your email agent calls.

Who does it

The agent's owner, and it should be a named person rather than a rota, because the value is in noticing drift between weeks. A rotating reviewer sees fifteen threads; a consistent one sees that the fifteen are worse than last month's.

Half an hour, same slot every week. If it needs an hour, sample ten instead of fifteen rather than skipping a week, because the habit is what produces the pattern recognition.

Write down four lines

Keep a running log, one entry a week:

Week 32: 15 read. 11 fine, 2 tone, 1 grounding, 1 should not have answered.
Correction rate 1.8% (was 2.4%). Escalation precision fine, recall low on refunds.
Change: refund threshold moved into the tool rather than the prompt.
Added to golden set: the partial-refund thread from Tuesday.

Four lines, and after three months it is the most useful document about your agent that exists: what changed, when, and what happened next. It is also what makes handing the agent to someone else possible, which is the thing that otherwise never happens.

Questions

How many threads should I read each week?
About fifteen, sampled deliberately rather than randomly: ordinary ones, escalations, long threads, corrected ones, and anything a signal flagged.
Why read the whole thread?
Because a reply in isolation always looks reasonable. The failures that matter appear only in sequence, as the customer experienced them.
Metrics or threads first?
Threads. Numbers without threads produce confident wrong conclusions, and the threads tell you what the numbers mean.
How many changes per week?
One, added to the golden thread suite before it ships. More than one and next week's movement is unattributable.
What if nothing needs changing?
For a week or two, fine. Beyond that, either widen what the agent handles or accept that the review has become a ritual without teeth.
Who should do it?
The agent's named owner, consistently. The value is in noticing drift between weeks, which a rotating reviewer cannot see.

Give your agent an address it can answer from.

Create an inbox