Hourly
Contract engineering
- Task and verifier authoring inside your pipeline and tooling
- Part-time or full-time
Hourly part-time or full-time
Home / AI agent evaluation
For AI labs, environment companies and evaluation teams. I design software-engineering tasks and verifiers built from real distributed-systems failures: race conditions, duplicate delivery, sync drift and ordering bugs. They are hard for the right reasons and can't be passed by guessing.
What I build
Why distributed systems
Most coding tasks can be solved by reading one function. Distributed-systems bugs can't. The cause sits in the interaction between processes, timing and state: a commit that happens after a side effect, a cache filled from a stale read, an event applied twice. To fix one, an agent has to build a model of the whole flow.
They are also unusually easy to verify well. Correctness can be stated as an invariant ("every order is charged exactly once", "both replicas converge") and checked by replaying the same scenario under many interleavings. That gives a strong, hard-to-game reward signal.
I bring 11+ years of building and debugging these systems in production, which is where the realistic failure modes come from.
Sample task
An illustrative task written for this site, showing how I structure one.
| Instruction | Customers report occasional double charges after deploys. Find the cause and fix it without changing the public API. |
|---|---|
| Environment | A small order service: an HTTP API, a Kafka-style consumer that charges payments, a Postgres database, a fake payment provider, and a test harness that can trigger consumer restarts. |
| Planted cause | The consumer charges the payment, then commits the offset. A restart between the two replays the message and charges again. |
| Verifier | Hidden tests replay 500 orders while forcing restarts at random points across 20 seeds. Pass only if every order is charged exactly once and none is lost. The agent cannot see or edit these tests. |
| Common wrong answers | Adding sleeps, increasing timeouts, committing before charging (which loses orders instead), or deduplicating in memory (which forgets on restart). All of them fail the verifier. |
| Accepted fixes | An idempotency key stored in the same transaction as the charge record, or an idempotent request to the payment provider keyed by order ID. |
See a verifier work
Fixes an agent might submit for the sample task above, each run against 20 seeds with random restarts. Plausible fixes pass most seeds and still fail. Illustrative results, simulated in your browser.
Engagement
Hourly
Hourly part-time or full-time
Fixed scope
Per sprint fixed number of tasks
Review
Per task set fixed scope
Questions
It is a sandboxed software project, usually a repository running in a container, that an AI agent works in to complete a task. A verifier then checks the result automatically, which produces the reward signal used in reinforcement learning or the score used in an evaluation.
It checks behaviour rather than matching a reference diff, runs outside the agent's reach, fails every known wrong-but-plausible fix, and gives the same result every time it runs. For timing bugs, that means fault injection and fixed seeds.
Hidden tests run in a clean step after the agent finishes, the agent's diff is checked for edits to tests and fixtures, the answer is kept out of git history and comments, and shortcuts such as sleeps, retries or hard-coded outputs are confirmed to fail.
Yes. I can author tasks and verifiers in your tooling and formats, under your confidentiality terms, part-time or full-time.
Tell me the kind of tasks you need and your timeline. I reply within one business day.