← All case studies

AI-Native Engineering: 950+ Reviewed PRs in Four Months

CTO at Assetera · 2026 - Present

How I ran the Assetera rebuild with AI agents doing the reading and the typing, while architecture, acceptance criteria, and review stayed with people. The workflow, the guardrails that keep it safe in a regulated firm, and the open-source tooling that came out of it.

AI Engineering
Claude Code
Multi-Agent
Developer Productivity
Open Source

Problem

The rebuild at Assetera needed far more code than a small team could write in four months: about 15 services, six frontends, an infrastructure repository, and the contracts. Hiring a large team would have taken longer than the rebuild itself.

AI coding agents can write that much code. The open question was how to use them in a regulated firm, where every change must be reviewed, traceable, and safe to roll back, and where "the model said it works" is not evidence.

Constraints

The review bar does not move. Every change is a pull request with CI, a named code owner, and a person who has read the diff.

The agent's report is not evidence. What changed is whatever git says changed. Whether it works is whatever the tests say.

Parallel work must not collide. Several agents in the same repositories at the same time must not overwrite each other or each other's assumptions.

Cost must stay sane. Using the most capable model for every keystroke is expensive and slow.

Architecture

People own judgement, agents own volume. I frame each task: the decision already taken, the files in scope, the tests to write, and the exact verification commands. The agent implements. I review the diff and the test run, then merge or send it back.

Isolation by default. Every task runs on its own branch in its own git worktree. A change that spans repositories is one branch and one pull request per repository. Nothing touches a shared checkout.

A cheap tier for bounded work. I built flash-agents, a Claude Code plugin that hands bounded tasks (a feature slice with tests, a mechanical refactor, a read-only code map) to a much cheaper model. Each writing job runs in a disposable copy-on-write clone and comes back as a git-computed patch. Claude keeps the architecture and the review. The worker spends its own tokens.

Memory that outlives the session. Facts learned the hard way (a platform default that is unsafe, a migration trap, a vendor quirk) are written into a shared memory and into ADRs, so the next task starts from them instead of rediscovering them.

Meetings into work items. I record and transcribe meetings locally with PrivateGoat, my own macOS app, and turn the action items into Jira tasks. Nothing leaves the machine until I send it.

My role

I designed the workflow, built the tooling, and ran it daily. I personally reviewed and merged 950+ pull requests across 29 repositories in four months. The team works with the same branching, review, and AI-assisted practices.

Outcome

Tradeoffs

Review is the bottleneck, on purpose. Agents can produce more code than a person can review carefully. I chose to merge less and read everything, rather than merge more and hope.

Green alone, red together. Changes developed in parallel can each pass CI and still break each other when merged. Merging in order and re-running CI on the combined result is slower, and necessary.

More code means more CI. Test suites grow quickly when writing tests is cheap. Keeping pipelines fast became its own piece of work.