Product exploration · Agents

Small agents that earn trust like new employees

16 August 2026

You describe the job in a chat window and the system drafts the agent and the exam it has to pass before it can run at all. It then asks before every action, and earns its freedoms one at a time.

Businesses are keen to put AI agents to work, but a lot of the ones that go live get switched off again once something goes wrong. The appetite is there, the trust isn’t — the surveys we looked at are in the deck above.

The usual answer is a single assistant that does everything and has access to everything, which is the shape hardest to test, limit or audit. You can’t set a meaningful exam for an agent whose job is “everything”, so we tried the opposite.

Hiring one by conversation

You describe the job you want doing, in ordinary sentences, in a chat window. The system drafts that agent’s paperwork from the conversation: its one job, the few things that matter about doing it well, and exactly what it may touch. The model supplies the answers, but the machine writes the file, so anything you didn’t mention stays as you left it.

A three-screen wizard turns the same conversation into the agent’s exam: put what matters into order, mark two real pieces of work right or wrong, read it back and sign. No jargon or percentages on any screen. Your signature freezes the result, and until then no work is routed to that agent.

The rules of employment

When the system failed its own product

The results I find most convincing are the ones where it refused to pass its own work. Our document-reading agent scored 96% in its first exam and failed anyway, because the one field that mattered most — a National Insurance number — was wrong; after a fix it retook the exam and passed at 99%.

The examiner was caught out too. When a candidate correctly ignored an instruction hidden in a test document, the examiner marked it to zero for doing so, and the gate refused to issue any verdicts until the examiner itself was fixed. No agent is immune to being tricked, ours included, and I’d be wary of anyone claiming otherwise.

What it has actually done

The most demanding test so far has been a real VAT quarter, where the office read the company’s mailbox read-only and drew up a list of what looked like supplier invoices for the period. That is where the design earns its keep, because deciding what is and isn’t a business expense turns out to be genuinely hard. The office doesn’t settle it alone: it lays out what it found, I rule on the rows it has wrong, those rulings become standing rules for next quarter, and nothing is sent until I approve it.

Where it honestly stands

Six agents have passed the exam that lets them exist. None has graduated from probation, so each still asks before acting — the rules above are shipped policy, not a record of anything promoted.

The weakest part is the hiring conversation itself. I have twice tried to walk it as a normal user would, and both times it was rougher than it should be. An outside review in August said to continue and aim higher, while noting we are shipping more slowly than we should. There are no external users yet.

Summary

Governance doesn’t work as a layer added after the agent is built. Passing the exam has to be the condition of it running at all, and everything it does afterwards has to leave a record you can hand to somebody else.

Links

Working on something like this?

Thirty minutes, no pitch — just a second pair of eyes.

Newsletter

Get new posts by email

Occasional write-ups on agent systems, evals and local models — including what broke. No pitch, unsubscribe in one click.