Case study: internal engineering tooling

Ticket to UAT

An agentic development pipeline built on Claude Code. Six engineers run it across three repositories. Tickets that used to take two or three days close the same day.

Role
Designed it, including where people stay in the loop
Where
QsrSoft, internal engineering tooling
Scale
6 engineers, 3 repositories
Built on
Claude Code, on the CI we already had
Measured
2–3 day tickets → same-day close

What this page does and doesn't say

This is internal work at my employer. What follows is the shape of the workflow, the reasoning behind each guardrail, and the turnaround number. There is no employer code here, no repository names, no ticket contents, and no detail about internal tooling. If that seems like a strange thing to point out on a portfolio, it is the same judgment I would be applying to your codebase.

The problem

A ten-minute change, surrounded by two days of overhead.

“This report should show net sales, not gross.” “Add average check as a metric in this widget.” The change itself is ten minutes of work.

Everything around it is not. Load the context. Find the code. Branch. Make the change. Work out what to test. Write the tests. Open the pull request. Write the UAT email. Then remember, three days later, to check whether anyone approved it.

That fixed cost is why a ten-minute change took two or three days. Not because it was hard, but because it was surrounded. And the smaller the ticket, the worse the ratio, which is exactly backwards from how it should feel to fix something small.

The shape of it

Eleven steps, three of which are people.

The pipeline reads and understands the issue on a ticket, creates a branch, and attempts the change. Once a developer has verified the behavior, it writes the tests, opens a separate pull request for each testing branch, merges them into the testing environments from the command line, files new tickets for anything it found outside the original scope, and sends the UAT email.

From there it is people. UAT approval comes first; then, once the guardrails pass, a person moves the change to master. Where those three human steps sit is the entire design, and everything below is an argument about their placement.

The pipeline, from ticket to masterA vertical flow of 11 steps. 1: Read the ticket. 2: Create a branch. 3: Implement the change. 4: A developer verifies the behavior, which is a human step. 5: Tests, against what was verified. 6: One PR per testing branch. 7: Merge into the testing environments. 8: Send the UAT email. 9: UAT approval, which is a human step. 10: Guardrails before master. 11: A person moves it to master, which is a human step. Anything found outside the ticket's scope becomes a new ticket rather than a commit.
Steps 04, 09 and 11 are people. None of them is a hedge about model quality. Each is a step where a decision gets made that no agent is in a position to make.

What I decided,
and what I rejected

Three calls, and what each one cost.

Call 01Tests come after a human has verified the behavior

This is the decision I get argued with about most, so it goes first. It is deliberately not test-first.

The pipeline implements the change. A developer then verifies the new behavior is actually correct. Only then are tests generated, against that verified behavior.

The reasoning: a model writing tests before the implementation is guessing at intent. It will produce tests that pass, that look thorough, and that encode the wrong requirement. Now the wrong requirement is locked in and defended by a green suite. Test-first works because a human holds the intent while writing the test. That property does not transfer to an agent holding only the ticket text.

The exception is characterizing existing behavior, where writing the test first is exactly right, because the intent genuinely is “whatever it does today.”

Human verification sits between implementation and test generation. That position is the whole design, and it is the first thing I would defend in a review.

Chose
Implement → a human verifies → generate tests against verified behavior.
Rejected
Test-first. It is the received best practice, and here it is the wrong one.
The cost
You give up the design pressure TDD puts on an interface. What you buy is tests that assert the thing that was actually wanted.

Call 02Out-of-scope findings become tickets, not commits

An agent opening a file to fix one thing will find three others. Left alone, the default outcome is a pull request that began as a one-line fix and arrives as forty files, which is unreviewable and therefore gets rubber-stamped.

So the pipeline files new tickets for anything outside the original scope and leaves the branch alone. Same reasoning behind one pull request per testing branch rather than one large one: a reviewer should be able to hold the whole change in their head.

Scope discipline is the thing careful senior engineers do by hand and almost nobody encodes in tooling. Encoding it is what made this safe to hand to five other people.

Chose
File a ticket, keep the diff small, one pull request per testing branch.
Rejected
“Fix it while you're in there.”
The cost
Real problems wait in a backlog instead of being fixed while the file was already open. Some of them will not get picked up.

Call 03Merge freely into test; master is a person's call

We run several testing environments, and moving a change through them is exactly the repetitive work the pipeline exists to remove. So it merges into them directly from the command line, the way a developer would by hand, just every time, and without forgetting one.

Master is different. Before anything moves there, guardrails check that every test passes, submodules are up to date, and the change has already been merged into the testing environments. UAT approval is a business decision about whether this is the change that was wanted, and nobody in the loop can make it except the person who asked for it. After that approval, a person, not the pipeline, moves it to master.

Underneath all of it is the belief the pipeline is built on: AI-written code should be surrounded by tests that check for explicit behavior. Speed without that just moves the defects downstream.

Chose
Merge into the testing environments from the command line; gate master behind UAT approval, guardrails and a person.
Rejected
Stopping at a pull request and hand-merging every environment, or letting a green suite promote straight to master.
The cost
More guardrails to maintain, and throughput is capped by how fast UAT comes back. The pipeline can finish quickly and the ticket still waits on UAT. That is the correct cap.

Outcome

Six engineers, three repositories, same-day close.

  • Tickets that took two or three days now close the same day. The effect is largest on small tickets, which is where the overhead-to-work ratio was worst.
  • Review stayed reviewable as throughput went up, because the pull requests did not get bigger.
  • Out-of-scope findings became a backlog instead of pull-request sprawl.
  • It is used by five engineers besides me, on code I do not review. That is a different bar from a workflow that only has to work for its author.

I also maintain the rules and context files that make those repositories legible to an agent in the first place. That turned out to be more of the work than the pipeline, and it is the part that does not show up in a demo.

It keeps getting tuned. The latest pass rewrote its commands to use fewer tokens, shortened the commit messages it writes, and narrowed test generation to the new behavior. Regression is already covered by the existing suite, so re-testing it was cost without coverage.

What I'd do
differently

Speed is measured. Quality is next.

The number I have is turnaround: tickets that took two or three days now close the same day. What I don't have yet is a defect-rate comparison: bugs that escape, before and after.

So far it's going well. But “faster” is only half the story for any AI tooling, and escaped defects are the number I want next, even if it comes back inconvenient.