Articles tagged “agent-eval”
2 articles

Testing & Evaluation·11 min read
The trace-to-dataset loop: turning live conversations into eval cases
Your best eval cases are already in your production traces. Here's how to automatically curate the interesting ones into a test suite that gets better every week without manual effort.
Read More

Testing & Evaluation·16 min read
Tool receipts: verifiable proof of what your agent actually did
When your AI agent claims it called a tool and got a result, can you verify that? Tool receipts are structured execution proofs that catch fabricated tool outputs before they damage customer trust.
Read More
Learn Agentic AI
Weekly. Patterns for shipping agents that work. MCP, scorecards, regression tests, prompts, model comparisons.