Durable execution
The agent resumes, retries, and finishes. A dropped model call does not lose the work.
I fix agents that fail silently and burn money. Durable, observable, cost-capped.
I don't build demos.
Your agent returns a wrong answer and nothing flags it. The bill climbs and nobody knows why. You have to explain that to your team.
The agent resumes, retries, and finishes. A dropped model call does not lose the work.
A test set catches wrong answers before your customers do.
Hard limits on spend and response time. The agent cannot burn the budget.
I work on live agents with real traffic. If you need a demo or a chatbot, I'll refer you to someone else.
Need a demo? I'll refer you to someone else.
We start with the logs.
We look at what broke in production and what the eval missed.
Durable execution, real evals, and hard caps on spend and latency.
You get a sanitised postmortem you can share with your team.
Each one covers what broke, what the eval missed, and what fixed it.
The first teardown is in progress.
View all writingTell me what broke. I'll tell you if I can fix it.
Start a conversation