Home / AI agent evaluation

Hard, verifiable coding tasks for AI agents

For AI labs, environment companies and evaluation teams. I design software-engineering tasks and verifiers built from real distributed-systems failures: race conditions, duplicate delivery, sync drift and ordering bugs. They are hard for the right reasons and can't be passed by guessing.

What I build

Tasks, environments and verifiers

  • Task environmentsRealistic multi-service repositories in containers, with the conventions, history and technical debt real codebases have.
  • VerifiersHidden behavioural tests, invariant checks and stress runs across many seeds, so a race condition fails reliably instead of one time in fifty.
  • Difficulty calibrationMultiple agent attempts per task, a target pass-rate band, and a read of every failure to confirm the agent failed for the intended reason.
  • Shortcut and leak reviewChecks that a task can't be passed by editing tests, special-casing inputs or finding the answer elsewhere in the repo.
  • Task QAReviewing other authors' tasks for ambiguity, unfair hidden requirements and verifiers that accept wrong answers.
  • Evaluation of agentic codingComparing how agents and setups perform on realistic engineering work, with failure analysis you can act on.

Why distributed systems

These bugs test reasoning, not recall

Most coding tasks can be solved by reading one function. Distributed-systems bugs can't. The cause sits in the interaction between processes, timing and state: a commit that happens after a side effect, a cache filled from a stale read, an event applied twice. To fix one, an agent has to build a model of the whole flow.

They are also unusually easy to verify well. Correctness can be stated as an invariant ("every order is charged exactly once", "both replicas converge") and checked by replaying the same scenario under many interleavings. That gives a strong, hard-to-game reward signal.

I bring 11+ years of building and debugging these systems in production, which is where the realistic failure modes come from.

Sample task

"Orders are sometimes charged twice"

An illustrative task written for this site, showing how I structure one.

InstructionCustomers report occasional double charges after deploys. Find the cause and fix it without changing the public API.
EnvironmentA small order service: an HTTP API, a Kafka-style consumer that charges payments, a Postgres database, a fake payment provider, and a test harness that can trigger consumer restarts.
Planted causeThe consumer charges the payment, then commits the offset. A restart between the two replays the message and charges again.
VerifierHidden tests replay 500 orders while forcing restarts at random points across 20 seeds. Pass only if every order is charged exactly once and none is lost. The agent cannot see or edit these tests.
Common wrong answersAdding sleeps, increasing timeouts, committing before charging (which loses orders instead), or deduplicating in memory (which forgets on restart). All of them fail the verifier.
Accepted fixesAn idempotency key stored in the same transaction as the charge record, or an idempotent request to the payment provider keyed by order ID.

See a verifier work

Five fixes, twenty seeds, one real answer

Fixes an agent might submit for the sample task above, each run against 20 seeds with random restarts. Plausible fixes pass most seeds and still fail. Illustrative results, simulated in your browser.

Engagement

How we can work

Hourly

Contract engineering

  • Task and verifier authoring inside your pipeline and tooling
  • Part-time or full-time

Hourly part-time or full-time

Fixed scope

Task sprints

  • An agreed number of calibrated tasks with verifiers
  • Delivered with notes on difficulty and failure modes

Per sprint fixed number of tasks

Review

Task and verifier audit

  • A review of an existing task set for leaks, ambiguity and gameable verifiers
  • Findings with concrete fixes

Per task set fixed scope

Questions

Frequently asked

What is an RL environment for a coding agent?

It is a sandboxed software project, usually a repository running in a container, that an AI agent works in to complete a task. A verifier then checks the result automatically, which produces the reward signal used in reinforcement learning or the score used in an evaluation.

What makes a good verifier for a coding task?

It checks behaviour rather than matching a reference diff, runs outside the agent's reach, fails every known wrong-but-plausible fix, and gives the same result every time it runs. For timing bugs, that means fault injection and fixed seeds.

How do you stop agents from gaming a task?

Hidden tests run in a clean step after the agent finishes, the agent's diff is checked for edits to tests and fixtures, the answer is kept out of git history and comments, and shortcuts such as sleeps, retries or hard-coded outputs are confirmed to fail.

Can you work inside our existing harness and formats?

Yes. I can author tasks and verifiers in your tooling and formats, under your confidentiality terms, part-time or full-time.

Building environments or evals?

Tell me the kind of tasks you need and your timeline. I reply within one business day.

Start a conversation