ChanlChanl
Blog/Testing & Evaluation

Testing & Evaluation

Browse 48 articles in testing & evaluation.

Testing & Evaluation Articles

48 articles · Page 1 of 4

Two glowing documents sit in a dim interrogation room: a machine-printed receipt stamped with a red error seal beside a cheerful summary claiming success, while an examiner traces the mismatch with a pen light, in a Blade Runner-style warm-plum palette
Testing & Evaluation·15 min read

How to Test Agents That Call the Right Tool and Still Get It Wrong

Tool-use benchmarks check whether an agent picks the right tool. They miss agents that call correctly, then mishandle the result. Here's how to test for it.

Read More
A wood-paneled mission control room in amber light; an operator magnifies two punched instruction cards that both fit the same slot but one row of holes sits offset; gauges and a trembling docking sequence behind (Interstellar style, golden-amber palette)
Testing & Evaluation·15 min read

How to Test Tool Argument Correctness in AI Agents

The most common production agent bug isn't picking the wrong tool. It's picking the right tool and passing the wrong arguments. Here's how to catch it.

Read More
A calibration curve chart showing the gap between LLM judge scores and human expert ratings across a sample of CX agent conversations
Testing & Evaluation·14 min read

LLM judges grade your agents wrong. Here's how to fix them.

LLM judges are the only practical way to evaluate agents at production scale, but they have predictable biases that give you false confidence. Here's how to calibrate them so the scores you see reflect the quality you actually have.

Read More
Two agent nodes connected by an arrow, representing a customer conversation flowing between a routing agent and a specialist agent
Testing & Evaluation·12 min read

Testing AI agent handoffs before they break in production

Multi-agent CX systems fail at the seam between agents. Here's how to build handoff tests that catch context loss, duplicate actions, and loop failures before your customers feel them.

Read More
A circular diagram showing production conversation traces flowing into an evaluation dataset and back into a CI testing pipeline
Testing & Evaluation·11 min read

The trace-to-dataset loop: turning live conversations into eval cases

Your best eval cases are already in your production traces. Here's how to automatically curate the interesting ones into a test suite that gets better every week without manual effort.

Read More
A production monitoring dashboard showing incomplete agent trace coverage with gaps in tool call visibility
Testing & Evaluation·13 min read

Why most AI agents in production are flying blind

57% of organizations have AI agents in production. Only a third are satisfied with their observability. Here's what's wrong with single-call LLM logging and how to build a monitoring stack that actually works.

Read More
Code terminal showing eval output with passing and failing agent test cases side by side
Testing & Evaluation·16 min read

Build evals before you build the agent

Eval-Driven Development (EDD) means your grader ships before your prompt. Here's the complete methodology for building CX agents with evals as the first deliverable, not an afterthought.

Read More
Terminal showing agent test scenarios with injected network faults, rate limit errors, and partial API responses highlighted in red against a dark background
Testing & Evaluation·15 min read

Fault Injection Is the Missing Layer in Agent Testing

ReliabilityBench found rate limiting is the most damaging fault for AI agents in production -- more damaging than wrong answers. Here's how to inject chaos into your test suite before your users do.

Read More
Side-by-side showing an agent conversation claiming a refund was processed next to a tool execution log confirming the actual transaction
Testing & Evaluation·16 min read

Tool receipts: verifiable proof of what your agent actually did

When your AI agent claims it called a tool and got a result, can you verify that? Tool receipts are structured execution proofs that catch fabricated tool outputs before they damage customer trust.

Read More
Dashboard showing three-tier evaluation sampling metrics for production AI agents
Testing & Evaluation·14 min read

Eval sampling for production agents: the 100/10/1 playbook

Running LLM-as-judge on every production conversation costs more than the agent itself. Here's the three-tier sampling system -- 100% lightweight heuristics, 10% LLM judge, 1% human review -- that keeps quality high without multiplying your eval bill.

Read More
A terminal showing a multi-step agent trace with one failing tool call highlighted in red
Testing & Evaluation·14 min read

Tracing AI agent failures across multi-step tool chains

When a production CX agent returns the wrong answer, the bug rarely lives in the last LLM call. Here's how to trace failures back to their root cause across multi-step tool chains.

Read More
Dashboard showing a high containment rate alongside a declining customer satisfaction score, illustrating the gap between the two metrics
Testing & Evaluation·14 min read

Your CX agent hits 95% containment. Why is CSAT tanking?

Containment rate tells you how many conversations your agent handled without escalation. It says nothing about whether they went well. Here's the quality-weighted scorecard that tells the real story.

Read More

El briefing de Signal

Un email por semana. Cómo los equipos líderes de CS, ingresos e IA están convirtiendo conversaciones en decisiones. Benchmarks, playbooks y lo que funciona en producción.

500+ líderes de CS e ingresos suscritos