Taking on one reliability engagement
I keep AI agent systems reliable after they ship.
For small AI teams with a critical browser agent or tool-using workflow already in use. I find the failures hiding behind the demo, build evidence around them, and make recovery part of the system.
20-minute technical conversation · paid diagnostic first · no sales theatre
- 01Goal acceptedscope and evidence target fixedpass
- 02Tool callbrowser state changed as expectedpass
- 03Result driftpage changed; expected control missingcaught
- 04Recovery pathbounded retry preserved prior workrecovered
- 05Independent checkoutcome verified outside agent reportpass
This is for you if
Your agent has real users, real tool calls, and failures that cost engineering time or customer trust.
Where the work starts
The failures that appear after the clean demo.
Selectors, permissions, rate limits, and third-party behavior change underneath the workflow.
A retry starts over, repeats side effects, or discards the evidence needed to understand what happened.
The agent reports success without a source, artifact, or independent check that proves the task completed.
A workflow still finishes, but uses more model calls, takes longer, or becomes operationally expensive.
The low-risk starting point
Start with a paid reliability diagnostic.
Before a retainer, I map the system around the workflows that matter most. The result is a bounded baseline: what fails, how to reproduce it, what is already protected, and what should change first.
Describe your system →Critical journey mapUp to three workflows, their dependencies, and their actual success conditions.
Failure reproductionConcrete examples, logs, and the smallest repeatable cases behind the instability.
Reliability baselineRegression gaps, recovery gaps, and observable cost or latency risks.
Prioritized planA decision-ready sequence of fixes, with retainer scope only if ongoing work makes sense.
Ongoing engagement
A narrow retainer, built around evidence.
Regression coverage for the agreed critical workflows.
A concise failure, drift, and operating-risk report.
One agreed reliability improvement shipped and verified.
Bounded investigation, recovery support, and a useful postmortem.
Operator documentation and checks your team can keep using.
What the retainer is
Focused ownership of one agreed reliability surface, with business-hours collaboration and explicit monthly priorities.
What it is not
Unlimited feature development, 24/7 on-call support, a replacement CTO, or a promise that no failure will ever happen.
Public evidence
The offer comes from systems I already build and test.
The public proof shows the engineering approach. Private client systems stay private.
Stateful web work, resumable progress, separate outcome checks, and thousands of public tests.
case study →A public benchmark built around measured degradation and externally verified outcomes.
methodology →A focused way to freeze expected behavior and expose meaningful change after models or tools move.
inspect →This is a newly focused service offer. The examples above are public systems work, not disguised client logos. A diagnostic establishes the evidence and scope for your system before either side commits to an ongoing engagement.
One useful first message
Tell me what is running, where it breaks, and what success needs to mean.
Remote · Toronto · one new engagement available