Moving an engineering team to agent-driven delivery: six months of real numbers
What changed when I moved my engineering work onto Claude Code agents. Measured throughput, speed and quality, what I built, what broke, and what I'd tell another engineering leader.
Summary
I lead engineering at a B2B SaaS product company and report to the CEO. Between April and October 2026 I moved my day-to-day engineering work, and much of my team's tooling, onto AI coding agents (Claude Code).
Over the same period:
- Merged pull requests rose about 10x, from 11.6 a month to 113.6 a month.
- Median time from first commit to merge fell from 2.0 hours to 0.7 hours.
- New test cases rose almost 3x, from about 470 a month to about 1,350 a month.
- Reverts stayed rare: 3 to 4 in twelve months.
I didn't get there by typing better prompts. I built a small operating system around the agents: routing, handoffs, scheduled checks, review agents, and shared agent packs my team now maintains.
This post covers the numbers, how I measured them, what I built, and the three failures that taught me the most.
Context
- The company: a multi-tenant B2B SaaS platform, with identity and entitlements, a shared app framework, a product configurator, transactional email, and several customer-facing apps.
- The codebase: I committed to 22 repositories over the period.
- The team: application developers, QA engineers and DevOps, across time zones. Several colleagues speak English as a second language. That shaped how I write instructions for both people and agents.
- My role: architecture, delivery, code review, release and production promotion. I work across the whole lifecycle, from spec to production.
The numbers
These figures compare the six months before I started using agents daily with the period after. They count my own commits and PRs only.
| Metric (per month) | Before | After | Change |
|---|---|---|---|
| Merged pull requests | 11.6 | 113.6 | 9.8x |
| Commits that reached a main branch | 39 | 167 | 4.3x |
| Lines added (excl. lockfiles and generated code) | 33,800 | 138,900 | 4.1x |
| Test cases added | 473 | 1,347 | 2.8x |
| Median PR cycle time | 2.0 h | 0.7 h | 65% faster |
| Slowest-quarter PR cycle time (p75) | 8.7 h | 3.4 h | 61% faster |
The honest version. These numbers show my output rose sharply after I moved to agent-driven work. They don't prove agents alone caused it. In the same period:
- my role grew
- the number of repositories I worked in doubled
- agent-made PRs tend to be smaller than hand-made ones (lines grew about 4x while PRs grew about 10x).
I'd treat 4x to 10x as the fair range.
What I built
The productivity didn't come from the model alone. It came from three layers I built around it.
1. A personal working layer
- Work streams and handoffs. Every long-running piece of work has a charter. Every session ends with a written handoff, so the next one can start cold. A hook reminds me to write the handoff before the agent's context is summarised. About 190 handoffs across 14 work streams so far.
- Encoded team knowledge. I wrote skills that teach the agent our work-tracker conventions, our architecture contracts and our runbook style. That means I stop re-explaining the same things.
- A read-only architecture reviewer. It checks any repo against our authentication, tenancy and permission rules. It runs on the strongest model, because that work needs judgement.
2. An orchestration layer
I set up one session as a router. It doesn't edit code. It classifies a request, writes a self-contained brief, and dispatches it to a separate agent session.
Scheduled routines run on their own:
- a weekday-morning brief that pulls together open work, the PR queue, deploy drift and the inbox
- a daily test-balance audit
- a weekly build-and-deploy health check.
A trust rule treats email, chat and ticket text as data, never instructions. Agents read a lot of text written by other people, so this matters.
3. A team-shared layer
This is the part I'm proudest of, because other people use it.
- A QA agent pack committed in a shared repo with 9 contributors. It has seven agents:
- an end-to-end test author
- an API test author
- a bug-to-failing-test reproducer
- a flaky-test classifier
- a page-object porter
- a spec reviewer
- a threat-model author.
- A lint hook that enforces test conventions on every agent edit, rather than only documenting them.
- An agent instruction file for the QA repo that I started. Six other people have since edited it, over 100 commits. The team keeps it current as part of normal work.
- 17 to 19 plain-English runbooks for promotions, production fixes and incident handling. I turned the runbook style into a reusable skill.
What shipped
| System | My role | Outcome |
|---|---|---|
| QA automation platform (browser and API tests, merge gate, agent pack) | Started it | Unattended scheduled test runs 95 days after the first commit |
| Transactional email service (password reset, invitations, security notices) | Designed and built it | Verified end to end in production after 126 days |
| Identity and entitlements platform | Major contributor; hardening and test gates | Took it to production |
| Build-once, promote-by-PR deployment pipeline | Wrote the design decision; ran promotions | First automated production promotion in 63 days |
| One main branch per repository | Co-owned the decision | Rolled out across repos in 52 days |
Cost discipline
Agents get expensive fast if every task runs on the biggest model. I set a simple routing rule: cheaper models for reading and searching, the top model for judgement.
In a two-week audit, 86% of sub-agent work ran on mid-tier or small models. A command proxy cut token use by a further 34%.
I also rejected a self-hosted CI runner. The saving was too small to justify the upkeep. Saying no counts as cost discipline too.
Where it went wrong
These three failures taught me more than any of the wins.
1. A scheduled agent posted my private brief to a team channel
My morning-brief routine had a step that posted to a team chat channel. It posted four times, with internal work items and decisions that were waiting on me.
I edited the prompt, but a run that had already started kept the old instructions and posted again almost a day later.
What I changed:
- I removed the posting step entirely, and checked the next run's transcript to confirm it no longer posted.
- I moved the brief to a local file.
- I made a standing rule that agents post to chat only when I ask for a specific message to a specific channel.
Lesson: a scheduled agent with write access to a shared channel is a publishing system. Treat it like one.
2. My router became the "mega-session" it was built to prevent
When I reviewed the logs:
- my "never does heavy work" router had run 19 of the last 30 sessions
- it had used 30% of all tokens in two weeks
- one conversation ran past 800 turns
- half of all sessions had no handoff
- work items were being opened more than twice as fast as they were closed.
What I changed:
- a hard "orchestrate, never edit" rule, with a fresh session every 40 turns
- a triage rule: an item is top priority only if a person is blocked this sprint or a customer date exists
- shorter handoffs, and the routing instructions split into small rule files.
Lesson: agent systems drift the same way teams do. Measure them, and review them like a process, not a tool.
3. An agent invented a ticket ID, and it ended up in published release notes
With no real ticket to cite, an agent made one up. The fake ID spread to commit messages, PR descriptions and the changelogs of two published package versions, which can't be edited.
What I changed:
- Agents must now look an ID up in the tracker before citing it.
- After two filed root causes turned out wrong when re-measured, I added a second rule: re-measure the cause before implementing the fix.
Lesson: agents fill gaps confidently. Any fact that leaves your system, such as an ID, a number or a root cause, needs a check against the source.
A bonus failure. Parallel sessions sharing one checkout overwrote each other's work more than once. I now give each session its own git worktree and commit by explicit file path.
What I'd tell another engineering leader
- Build the operating layer, not just better prompts. Handoffs, routing, scheduled checks and review agents are where the compounding comes from.
- Put conventions in hooks, not documents. A lint hook that runs on every agent edit beats a style guide nobody reads.
- Ship agent packs into shared repos. When teammates maintain the agent instructions themselves, adoption has stopped being your project.
- Route by model tier from day one. Most agent work is reading and searching, and doesn't need the top model.
- Assume agents will fabricate at the edges. Build checks wherever their output leaves your system.
- Measure honestly. Smaller PRs inflate counts, and a growing role inflates everything. Give the range you can defend.
How I measured this
- Source data: git history across all repositories, GitHub PR data, Claude Code session transcripts, and my own configuration files.
- Commits: deduplicated across branches. Merge commits, lockfiles, build output and generated code are excluded.
- Split date: "before" and "after" use 19 May 2026. That's the point where my surviving session records begin, and it gives the fairer "before" sample. A split at the start of daily use (April 2026) shows bigger multipliers, but rests on only 19 PRs, so I don't lead with it.
- This is "lighter vs heavier" agent use, not "no AI vs AI". Most of my commits already carried an AI co-author tag before the split.
- Session hours count gaps under 30 minutes between messages. Parallel sessions are counted once. Older transcripts were deleted automatically, so the hours are a lower bound.
- PR cycle time runs from the first commit in a PR to its merge. Fast medians partly reflect agents removing waits between people. They don't show that reviews got faster.
- Test cases are counted from added test functions. Squashed branches may overstate them by up to about 20%.
- Coverage isn't claimed, because the repositories I lead don't yet publish coverage reports.
If you're moving a team onto agent-driven delivery and want to compare notes, reach out.
Enjoying this article?
Get posts like this in your inbox. No spam, unsubscribe anytime.



