How a Tiny Team Ships Like a Big One

Humans own the spec: how a retro complaint becomes shipped code

Every complaint, customer question, and retro item lands in one intake pipeline that sorts work into mechanical, design, and human lanes - and design work can't start until a person approves the spec.

7 min readagentsprocessclaudespec-driven

At our sprint retro back in July, four different people typed some version of "we should really fix this" into the thread. A missing soak window, a CI gate we'd been promising ourselves for weeks, that kind of thing. In most retros I've sat in, that's where sentences like those end.

By the end of that day all four were GitHub issues, labeled and prioritized, each one checked against the open issues first so we weren't filing a duplicate. One had become two issues, because whoever read it noticed the two halves could be fixed independently. Except nobody read it. An agent had picked up the retro thread, done the filing, and posted a note saying what it had done.

Most of what we've written about our agents is the part that writes code. This is the other half, the one that moves work around the company, and the thing we had to settle before we'd run any of it unattended was which calls stay ours.

Where those sentences used to go

Nowhere, mostly. Someone nods, the thread scrolls, and by Thursday it's gone. We lost things the same way after incidents and on customer calls: somebody says "we should add a guard for that", everybody agrees, nobody files it. At six people you feel that more than a bigger team would, because the ideas were usually good and there's nobody whose job it is to catch them.

We already had agents fixing production bugs (the SRE fleet). The intake pipeline points the same machinery at everything else.

Three lanes, one gate

Any GitHub issue labeled needs-triage gets picked up by a triage sweep that runs every hour, whether the issue came from a retro, from customer feedback, or from a human jotting down a thought. A classifier sorts it into one of three lanes. The classifier turned out to be the least interesting part of the design; the two rules we put around it are what made us comfortable leaving it running.

  1. It fails safe. When the classifier is unsure, the verdict is "human". Some checks run before any model call at all, so tracking issues and epics route straight to a person without a model ever seeing them.
  2. Some work is un-automatable by construction. The classifier's instructions carry a hard rule: anything involving a production data migration, a backfill, a credential change, or a cross-account infrastructure decision is never mechanical. If the core of the job is deciding something about state that already exists, a person gets it, even when a code change would help.

Then the lanes:

  1. Mechanical - a well-understood fix. Low-severity items dispatch straight to the fixer agent; anything higher waits for a human to apply an approval label first. A daily cap keeps the fleet from flooding the review queue.
  2. Design - the interesting one. A spec agent drafts exactly one spec file and writes no code. The spec goes up as its own small PR. If a human merges it with a ready-for-dev label, a scope-confined implement agent picks it up and builds only what the spec declares. Merge the spec without the label and you've accepted the thinking without committing to the build.
  3. Human - the pipeline writes a verdict comment with a suggested owner and then touches nothing.

When implementation does happen, the PR gets a reviewer assigned automatically, from CODEOWNERS plus the file's actual git history, and that person gets pinged on Slack. There's a nag module too, which comes back when an item sits untouched. We added that after watching a few things stall politely for a week.

The intake pipeline: three sources, three lanes, human gates · click to enlarge

The same loop, pointed at customers

The sources aren't only internal. A watcher reads our customer-feedback Slack channel every 17 minutes. Bug reports route to the fixer, feature requests become tracked issues, and either way a confirmation goes back into the customer's thread so they know it didn't vanish.

The one that stuck with me: a customer asked why their scores looked different from what they'd seen in a colleague's demo. The loop traced it to a per-org feature flag, checked that explanation against live production data before saying anything, and posted the answer in the same thread. They'd asked a question in the morning and had a root cause back, rather than a promise that someone would look into it.

Real numbers

  1. 4 retro items → 4 triaged issues, same day, priorities assigned, duplicates searched first - one item split in two because the fixes were independent.
  2. Every design-lane item produces exactly one spec file before any code exists.
  3. The spec agent runs on our strongest models. A spec gets read by every implementation that follows it, so it's the last place we'd go looking for savings.
  4. Customer questions answered with verified root causes in-thread, feature requests confirmed in-thread, on a 17-minute ingestion loop.

Where the humans sit

Mostly at the spec. The agent drafts it, a person reads it, argues with it on the PR, sharpens it, and only a human's merge-with-label sends anything to implementation. As Michael put it while we were building this: make sure the spec is clear, because that's where humans should sit. If it isn't clear, iterate on the spec, not on the code.

Below that gate there are more. Higher-severity mechanical work waits for an explicit human approval, the human lane is left alone entirely, and the never-mechanical rules mean whole categories of risk can't get into the automated path no matter what the classifier decides.

Steal this

Two things, both cheap. Make "one spec file, no code" the mandatory first deliverable for anything that smells like design, because the spec is where you can still argue cheaply, and an agent that writes code first takes that away from you. Then write your never-automate list before you write your classifier prompt. Ours took about ten minutes and it's done more work than anything else in the pipeline: migrations, backfills, credentials, cross-account changes. Yours will look different. Write it down anyway.

Next Thursday: none of these agents could do their jobs if they couldn't see - the protocol that gives them hands, from internal tool to product feature. Next Tuesday, the product track: the 44 use cases nobody had to write.


This post is part of How a Tiny Team Ships Like a Big One, a series on how six builders run a production AI company. Building at Aithon - if this is how you want to work, talk to us.