How a Tiny Team Ships Like a Big One · The Product Track
One deal, 468 facts, and nothing regenerated twice
Naive AI features re-analyze everything on every change and die of cost and latency. Our deal analysis is event-sourced: changes invalidate only the sections they touch.
The first version of every AI analysis feature works the same way: something changes, so you regenerate the analysis. All of it. It's the obvious design, it demos beautifully, and it fails in production for three compounding reasons: cost (you're re-buying conclusions you already paid for), latency (users watch a spinner while the world is re-derived), and churn - the subtlest one - where re-generation rewrites sections whose inputs never changed, so users see conclusions wiggle for no reason and quietly stop trusting the product.
We hit all three. A deal view like the 468-element one from the opening of this track synthesizes hundreds of facts across dozens of sources into scored, written analysis - and deals change constantly: an email lands, a call is logged, a CRM field moves. Regenerate-everything at that scale isn't a product, it's a furnace.
Change as an event, not a trigger
The redesign treats change the way an accountant treats money: nothing moves without a record.
- Detection. Source data is fingerprinted; when a fingerprint moves, the pipeline writes an evidence event - a durable, deduplicated record that this specific input changed. Not "something happened, rerun everything"; a ledger entry saying exactly what happened.
- Resolution. Each evidence event resolves into section invalidations: which sections of which analyses does this change actually touch? A new stakeholder email might invalidate the stakeholder map and one score, and nothing else. Every section of the analysis is versioned, so an invalidation is a precise mark against a precise version - and events targeting sections that don't exist yet resolve to nothing, wasting no work.
- Regeneration. A warm consumer works the invalidation queue and regenerates only the marked sections. Everything else stays exactly as it was - same words, same scores, because nothing about their inputs changed.
The user-visible result: a deal view that's never more than about half an hour behind reality, updates that touch only what the change touched, and a cost curve that scales with change volume instead of with deal size. The stable sections don't wiggle. Trust survives.
The failure mode we monitor for
An event-sourced pipeline has one classic way to lie to you: the consumer hangs while the infrastructure reports it healthy. The process is up, the health check passes - and the queue quietly grows. So the monitoring watches the queue, not the process: dead-letter depth and, more tellingly, the age of the oldest unprocessed message. A hung consumer that ECS still calls healthy shows up as aging messages within minutes, in the same alert channel everything else lands in. We added that alarm after the silence, not before. Most good monitoring is scar tissue.
Real numbers
- ~30 minutes: the refresh sweep cadence, and so the bound on how far a deal analysis can lag the evidence behind it.
- One section, not one deal: the typical regeneration triggered by a single change - the difference between a furnace and a feature.
- 468 data elements on the deal from the opening of this track - the scale at which regenerate-everything stops being an option.
- Two queue alarms - DLQ depth, and oldest-message age firing at 15 minutes - because a hung consumer looks healthy to everything except its backlog.
Where the humans sit
The section boundaries - what counts as an independently-generated unit of analysis - are a human design decision, and they're the entire trick: draw them wrong and every change invalidates everything, draw them right and most changes touch one section. Engineers own that map. The pipeline just obeys it.
Steal this
Version your AI outputs at the section level and make invalidation an explicit, recorded step between "input changed" and "regenerate" - even if your first resolver is crude. The moment invalidation is a first-class record, you get three things for free: a cost model (count the invalidations), an audit trail (why did this section change?), and an end to conclusion-wiggle. And monitor your queue's oldest message age, not your consumer's health check. The health check is the last to know.
Next week the build track tours our inner loop: a TUI and twenty-two markdown files. Then back here: the data model underneath all of this, and the argument we had with ourselves about whether "context graph" is a design or a buzzword.
This post is part of the product track of How a Tiny Team Ships Like a Big One, a series on how six builders run a production AI company. Building at Aithon - if this is how you want to work, talk to us.