The Failure Lab
One reference agent, broken on purpose and then hardened, with the results published.
The system
A refund-processing agent. It reads a support ticket, checks the order, calls a payments API and emails the customer. A duplicated refund is a failure any finance director understands.
Faults we inject
- Tool timeouts.
- Provider rate limits.
- A process crash mid-run.
- Duplicate event delivery.
- Tool schema drift.
The hardened version
- Idempotency keys on every tool with side effects.
- An outbox, so state changes and outgoing messages are recorded together.
- Durable execution, so a crashed run resumes where it stopped.
- Compensation steps for actions that have to be undone.
- Budget caps, so a looping run stops before it becomes expensive.
What we measure
The naive and hardened versions run under the same faults. For each, we publish:
- Duplicate refunds.
- Completion rate.
- Wasted tokens.
- Time to recover.
Next
An event-driven variant on Kafka, covering backpressure, poison messages and a dead-letter queue routed to human review. Then a chapter on evals as a release gate in CI.