Home / Writing / AI evaluation

Designing verifiable coding tasks for AI agents

A coding task is only as good as its verifier. Here's how I build tasks that are hard for the right reasons, using distributed-systems bugs as the raw material, and verifiers that a capable agent can't talk its way past.

Short answer

A good coding task for an AI agent has three parts: a realistic environment (a repository running in a container), a clear instruction, and a verifier that checks behaviour the agent can't see or edit. It should be solvable by a strong engineer, fail weaker attempts for the intended reason, and resist shortcuts such as editing tests or special-casing inputs. Bugs from distributed systems make strong tasks because the cause spans processes and timing, and correctness can be checked as an invariant across many replayed runs.

The anatomy of a task

PartWhat it isWhat goes wrong
EnvironmentA repository and its services in a container, with tools the agent can useToy code with no conventions, so the task tests nothing realistic
InstructionWhat a teammate would write in a ticketHidden requirements the verifier checks but the instruction never mentions
VerifierTests and checks run after the agent finishes, outside its reachChecks that accept wrong fixes, or reject correct ones written differently

Seven properties of a good task

  1. Realistic. It looks like work someone gets paid to do, in a codebase with history and some mess.
  2. Solvable. A strong engineer can finish it with only what's in the environment. I solve every task myself before shipping it.
  3. Unambiguous. Everything the verifier checks is implied by the instruction or by the code's existing contracts.
  4. Verifiable. Success is decided by behaviour, not by matching a reference diff.
  5. Shortcut-resistant. No pass by deleting tests, hard-coding outputs, or adding a sleep.
  6. Leak-free. The answer isn't sitting in git history, a comment, or a sibling file.
  7. Calibrated. It separates agents. A task every agent passes, or none does, tells you little.

Why distributed-systems bugs work so well

Most benchmark-style tasks can be solved by reading one function. Distributed bugs can't. The cause lives in the interaction: a side effect that happens before a commit, a cache filled from a read that was already stale, two replicas applying the same change in different orders. An agent has to build a mental model of the whole flow to fix it, which is the skill you want to measure.

They also produce a strong reward signal. Correctness can be written as an invariant, such as "each order is charged exactly once" or "both sides converge to the same state", and checked under many different timings. A fix that only looks right fails.

Building verifiers that hold up

Test behaviour, not implementation

Assert on outcomes the user would notice: rows in the database, calls to the fake payment provider, the final state of both replicas. Avoid asserting on function names or internal structure, or you will reject correct fixes that take a different route.

Make timing bugs fail reliably

A race that shows up once in fifty runs makes a useless verifier. Two techniques help. First, add fault-injection hooks to the environment (restart the consumer after the side effect, delay one replica, drop an acknowledgement) and trigger them at chosen points. Second, run the scenario across many fixed seeds, so the same interleavings are tested every time and results are reproducible.

Keep the verifier out of reach

Hidden tests are copied in after the agent finishes and run in a clean step. Visible tests can stay in the repository as hints, but they must never be the thing that decides the reward. Also check the agent's diff for edits to test helpers, fixtures or the fault-injection hooks themselves.

Check for the wrong-but-plausible fixes

For each task I write down the fixes a hurried engineer would try and confirm the verifier fails every one. For a duplicate-processing bug, that list includes adding retries, raising timeouts, committing before processing (which loses data instead), and deduplicating in memory (which forgets on restart).

Calibrating difficulty

Run several attempts per task with the agents you care about and look at the pass rate. Then read the failures. A failure only counts if the agent failed for the reason the task is about. If agents fail because a tool is broken, a dependency won't install, or the instruction is unclear, the task is measuring the environment, not the agent. Fix it and run again.

A worked example

The instruction says customers occasionally get charged twice after deploys and asks for a fix without changing the public API. The environment holds an order service with an HTTP API, a queue consumer that charges payments, a Postgres database and a fake payment provider. The planted cause is a consumer that charges first and commits its offset second.

The hidden verifier replays 500 orders while restarting the consumer at random points across 20 seeds, and passes only if every order is charged exactly once and none is lost. Accepted fixes include an idempotency key stored in the same transaction as the charge, or an idempotent call to the provider keyed by order ID. You can read more about the underlying bug in why Kafka consumers process messages twice.

Checklist before shipping a task

  • I solved it myself inside the environment, from the instruction alone.
  • Every verifier check traces back to the instruction or an existing contract.
  • Known wrong-but-plausible fixes all fail.
  • Timing-dependent checks run across fixed seeds and pass or fail reproducibly.
  • The answer isn't in git history, comments or neighbouring files.
  • Agent failures were read, and they fail for the intended reason.

Frequently asked questions

What is a verifier in AI agent evaluation?

A verifier is the automated check that decides whether an agent completed a task, for example a hidden test suite, an invariant check or a comparison of system state. Its result becomes the reward in reinforcement learning or the score in an evaluation.

Why not grade agents against a reference solution?

Matching a reference diff rejects correct fixes that take a different route and can accept wrong ones that look similar. Checking behaviour, such as what ends up in the database or what calls were made, measures what matters.

How do you make race-condition tasks reproducible?

Add fault-injection hooks to the environment, such as restarting a consumer after a side effect, and run the scenario across a fixed set of seeds so the same interleavings are tested every time.