Agent reliability retainer
A paid diagnostic followed by bounded ongoing ownership of regression coverage, drift, recovery, and operating evidence.
see the reliability offerToronto-based systems builder working across browser automation, agent evaluation, recovery, and local-first infrastructure. There are two clear ways we can work together.
A paid diagnostic followed by bounded ongoing ownership of regression coverage, drift, recovery, and operating evidence.
see the reliability offerAgent infrastructure, browser automation, evaluation tooling, and local-first Linux software with explicit milestones and handoff.
start a conversationThe difficult part of an agent system starts after the clean demonstration. Real websites change, tools fail, recovery repeats side effects, and confident outputs arrive without enough evidence. My work focuses on making those failures observable, reproducible, and less expensive to repeat.
The retainer is intentionally narrow. It is not unlimited development or 24/7 emergency support. We agree on the critical workflows, establish a baseline, and make one meaningful reliability improvement at a time.
| Project | What it demonstrates |
|---|---|
| Blackreach and Huginn | Browser automation and self-hosted research infrastructure backed by public code and extensive tests. |
| Lethe | A public benchmark for measuring behavioral drift in long-running agents. |
| Rigr | Small, focused regression tooling for freezing expected agent behavior and detecting change. |
| claude-voice | A polished local voice tool used by other developers, with no cloud API required. |
| AMP Discovery | An end-to-end ML experiment with leakage-aware evaluation and documented results. |
I use multiple coding agents when parallel work helps, but I own the result. Agent output is reviewed against the repository, exercised in the real environment, and described honestly. A tool saying its tests passed is not the same as independent verification.
I prefer bounded milestones, working slices, and visible acceptance criteria. For contract work, the handoff includes the documentation and checks needed for someone else to operate the result.
Some of my strongest recent systems work is private. I can provide a sanitized walkthrough of the problem, outcome, and verification method without disclosing private architecture, operational details, or data.
My primary consulting offer is the agent reliability retainer. I am also open to software engineering roles where AI infrastructure, developer tooling, evaluation, browser automation, or Linux systems are central, plus contained contract work with a clear acceptance boundary.
Available for one reliability engagement, remote roles, and selected contract work. Toronto-based, open to relocation and US timezone alignment.
contact@phnix.dev · résumé ↓ · github ↗
A useful first message tells me what you are building, where it is getting stuck, and what a successful outcome would look like.