AI Made Coding Faster. So What’s the Bottleneck Now?

A candid conversation between Pandium CTO Shon Urbas and Justworks Senior Staff Engineer Cullen MacDonald about what changes when writing code stops being the hard part.
Written by
Bronwen Malloy, Marketing Coordinator
Last updated
October 8, 2026

Most AI takes land in one of two buckets. Either engineers are about to be replaced, or every team is about to ship ten times more stuff. Put two people who have spent years shipping production software in a room for an hour and you get a messier and far more useful picture.

Cullen MacDonald has been at Justworks for over five years, most of that as director of engineering for the benefits teams. Recently, he's moved into an IC role as a senior staff engineer. (He also spent a year at Pandium, so this one was a bit of a reunion.) Justworks is an all-in-one HR, benefits and payroll platform for small businesses.

Cullen describes himself as a former AI skeptic. Then something shifted last November. Here's how he put it.

"The models got better, the harnesses around those models got better. And then the patterns on how to actually use these tools started being blogged about in a way that made it a lot easier to pick it up and build something."

His first real test was an infrastructure migration on systems he had set up himself five years earlier. With Claude Code, he wrote the Helm charts, the bash migration scripts and the runbook. He read every line and ran every step himself. Work that would have taken a long time to write, review and think through moved fast.

So if writing code got that much faster, why didn't everything else? That question runs through this whole episode.

The Framework, Your Slowest Step Sets the Pace

Cullen borrowed the framing for this conversation from Eliyahu Goldratt's The Goal. The core lesson is simple. A system's throughput equals the speed of its slowest part. Speed up any other step and the system moves exactly as fast as it did before.

Shon's shorthand was "you're only as strong as your weakest link." Same idea, applied to shipping software.

Where the bottlenecks actually live

AI made the build step faster. The rest of the product development lifecycle (PDLC) stayed roughly where it was. Cullen named three spots where work still piles up.

  • Code review. Opening a PR gives you a little dopamine hit, and then you grab the next ticket while the PR sits. Cullen thinks any manager who convinces their team that a day spent reviewing code is a great day has done their job well.
  • Prioritization. You have ten good ideas and capacity for three. Picking those three, and deciding how big a slice of each to take, is still a negotiation between sales, support, product and engineering.
  • Enablement. Sales needs to know what shipped. Support needs an answer when a customer asks about it. Ops needs training.

"Building is even easier now, but that doesn't solve those other problems."

Why "build more" is the wrong lesson

Justworks is hiring engineers right now, because engineers have never been more productive. Cullen was clear that the goal is still to ship the right things, since every feature has to be maintained, observed, monitored and sold.

"We could build a new feature every day and our sales team won't know about it, our customer support team won't know it. Somebody will ask about it and customer support wouldn't know how to answer."

What does change is the size of the first bite. Teams take a bigger swing at a problem up front, and that trickles all the way down to ticket size and PR size. Shon is seeing the same thing at Pandium. A ticket turns into a PR in a day or two, and then the review stretches.

Stacked PRs and tools like jujutsu help organize the changes, but the underlying question stays the same. How do you get comfortable with a big pile of changes landing at once? Cullen's old rule still holds up well here. Work on one thing at a time, and actually sit on your PR once it's open instead of opening four more.

Build the Tool, Skip the Subscription

One of the first places Cullen pointed AI was internal tooling. Justworks used to be customer number one on its own platform. Then the company grew to roughly 1,700 people and outgrew a product built for small businesses. The people team moved to Workday, but that's a tool for HR, and engineering managers still needed answers to basic questions.

So Cullen's team built HQ, an internal tool for the engineering, product and design org. It tracks who's on which pod (Justworks' version of the Spotify squad, usually five engineers, an engineering manager and a product manager), who reports to whom, Jira hygiene, and DORA-style metrics like PR cycle time. That's a whole category of reporting products teams normally pay real money for.

HQ does a second job that might matter more. It's a sandbox.

"It is our sandbox. It is not customer facing. There's real data... but it is a much safer environment for us to really test the boundaries of some of the patterns or tooling or models."

There's a nice side effect too. Justworks asks every team to follow a fairly rigid structure at the epic and initiative level in Jira, while sprints and story points are up to each team. Follow the structure and the rest of the org gets solid reporting for free. An off-the-shelf tool built for every company probably wouldn't map to that process. A homegrown one fits it exactly.

The integration angle

Shon is seeing the same pattern at Pandium. Once he noticed the documentation platform had a Git integration, docs became easy. Claude reads the code, drafts the documentation, and it gets posted. Documentation went from an afterthought to a first-class part of the work.

Jira works the same way. Shon hands Claude an API key and asks it to write a ticket, or to pull the status of every open ticket, or to answer some one-off question he would have built a report for in the past. In his words, it works like a real-time integration with a third-party system.

The SaaS that stays

None of this means ripping out your stack. Both teams still run on Jira and GitHub. Shon's read is that core systems of record are sticking around, and the tools layered on top of them are the ones getting displaced. For anyone thinking about integrations, that's a pretty big deal. The value moves toward clean APIs on the core systems, because that's what people are now building on.

Trust, Verification, and Testing the Tests

Bigger PRs bring the real question to the surface. How do you trust a large set of changes when an agent wrote a lot of it? At Justworks the answer depends on what the code touches. Customer-facing and money movement code stays closer to what Cullen called "artisanally handcrafted." Elsewhere, trust comes from stacking several layers.

  • Pair programming. One of Justworks' directors came up through the extreme programming world, and many teams pair by default. Two people watching the coding agent work means two people reviewing its output in real time.
  • Agentic review bots. Cullen says the Claude Code reviewer bots have gotten really good. The catch is context. They often won't look across your other repos.
  • Deterministic checks. Linters, tests and other non-agentic ways to verify the code.
  • Observability. Datadog (or New Relic, or whatever you run) with real dashboards and real alerting.

"There's a layer of trust you can get to where it feels the same and almost even better, because now we have even more test coverage."

The asteroid problem

Shon was quick to point out the other side. More output also means more noise. He described Claude warning him about the rare case where a power outage happens at the exact moment an asteroid hits, then proposing extra defensive code to handle a risk that doesn't exist.

Tests are the bigger issue. Shon keeps seeing tests that are tautologically true, meaning they pass no matter what. Cullen added tests that check the framework itself. (Django already has its own test coverage, thanks.) Shon was also refreshingly honest about how he reviews them.

"As a reviewer, I'm gonna be super honest. I look at those tests, I sort of glare over them and I hope they cover the things I need to cover."

That's exactly why the rule from our Explore/Exploit episode with Lizzy matters so much. Make sure you see a test fail before you trust it to pass.

Agents still think they're people

A couple of quirks got a laugh from both of them. Ask for a plan and the agent estimates the work in weeks, because it learned from humans on Reddit. Cullen's reaction was basically "that's cute, you're doing it now and it'll take an hour and a half."

The other one shows up on weekend projects. Change your mind about how something works and the agent assumes millions of users depend on the old version. It keeps a thread back to the old API route or code path and never fully finishes the migration.

Specs Are the New Prompts

For the last couple of years, Justworks has written detailed product briefs and technical specs for new work. Which internal APIs will we call? What's the data model? What do we need to account for? Humans still do most of that thinking on a whiteboard, with an AI asking questions along the way.

That habit paid off. Those same docs now work as the input for the agent. Teams hand them over and ask for help breaking the work into Jira tickets, or use them directly as the prompt to build. We've landed in a similar place at Pandium. Tight constraints and existing, reliable code make the output far easier to trust, which is the approach Lizzy laid out in Building AI into Integration Workflows and the thinking behind our AI-Powered Integration Code Generator.

If that sounds familiar, it should. Shon and Cullen both flashed back to Cucumber and Gherkin and the behavior-driven testing era. The big idea back then was writing plain human language and getting predictable results. Cullen was careful about what "deterministic" means here. The code can come out different on every run, and the behavior still matches the spec.

"If you write a set of specifications, hey, keep going until all of these specifications pass, then in theory you will get deterministic behavior."

A Rails to Go rewrite on autopilot

One Justworks team is putting that to work right now. They have a small Rails app with a lot of code and very well-specified behavior. They turned the spec into pass/fail tests, then pointed four different models at the rewrite in parallel. Claude Code's /goal command keeps the agent working until the goal is met, and the whole thing runs on an EC2 instance so it can go all day.

Shon's first question was the obvious one. What about tokens? Justworks routes everything through Bedrock and has been on API pricing for about a year, so everyone there has a clear sense of what model usage actually costs. Running several models side by side is partly an experiment in which ones earn their keep.

That cost awareness shows up in the culture too. Justworks celebrates people who get solid, measured output from a cheaper model, and the team A/B tests models in product features like document parsing. Cullen is getting results from OpenAI's GPT 5.6 Luna and Terra that he says feel on par with Opus for his day-to-day work, and he's nowhere near his budget.

Pi, Model Routing, and the Case Against Getting Comfy

Cullen's tool of choice lately is Pi, an SDK for building your own agent loop. It also ships as a command line tool that looks a lot like Claude Code, Codex or OpenCode at first glance. Out of the box it has almost nothing. No permission prompts, no subagents, no to-do list. What it does have is a platform for extensions and hooks that can modify the agent itself, whereas skills are plain text the model reads.

Here's what Cullen has built on top of it.

  • Dumb but effective model routing. Writing a commit message goes to a subagent on a really cheap model. Code review goes to a different model than the one that wrote the code.
  • An Obsidian journal. A slash command looks at every session he ran that day, drops the empty threads, and writes a journal entry of what he worked on and shipped, plus to-dos for anything left open.
  • A weekly review. On Monday morning he opens his laptop and sees exactly what he was in the middle of, including the loose threads he swore he wouldn't leave.

His advice for everyone else is to build one yourself, with Claude's help.

"If you don't know that you can describe how those things are not the same thing, then it's probably worth you actually going through the effort of writing your own little agent loop."

"Those things" are the agent, the harness and the model. Shon's quick version is that the harness is the software that runs the model and manages its inputs and outputs. Writing a loop that calls the API, gets tokens back and exposes a tool or two (the same basic idea behind an MCP server) is the fastest way to make that distinction stick.

Explore, exploit, and Shon's honest answer

Cullen's bigger warning was about comfort. He's watching non-engineers at Justworks build software in ways that look foreign to traditional engineers. The docs and to-do lists live as markdown inside the repo, and every internal tool ends up as its own mini Jira and Confluence.

"Anyone getting comfy on Cursor or Claude and being like, great, this is my tool I'll use for the next 10 years, I think you're going to do yourself a disadvantage when the actual job changes four years from now."

Shon, who mostly lives in Claude Code, owned up to sitting firmly in exploit mode. His reasoning is practical. He doesn't expect this iteration of tools to last much longer, so rather than hopping between harnesses today, he plans to pick up whatever comes next when it arrives. If you read our Explore/Exploit piece, you'll recognize the tension.

Verified vs. Vibe Coded, and the Rise of the Tiny Company

Shon shared a story every CTO using these tools will recognize. Auth is one of the most core pieces of Pandium, and he has rewritten the auth service with Claude twice in the last eight months. Both times he threw it away. The code was there. What was missing was a way to validate it that felt safe enough to ship.

That's the bottleneck framework again. Shon can spike out big ideas faster than ever, and he also discards them in a way he never would have before. Fast building with no way to verify just produces more drafts.

He also noticed that prompting Claude feels a lot like managing an engineer. You describe the problem, it comes back with code, you review it and send notes, and occasionally you get something nobody asked for. Shon and Cullen both remember an engineer at a past company who spent six weeks adding Gmail-level keyboard shortcuts to an integration platform. They had to talk him out of merging it. Cullen's response was that adding them would be trivial today.

The companies that win will verify

Cullen sees the trust problem as the real competitive line. He doesn't review the payroll team's code at Justworks, and payroll still runs correctly every day, because that team built the feedback loops and verification that earn his trust. The same standard will apply to agents.

"There are gonna be companies that figure out how to get to a place where they can verify and can trust this in the same way that you can verify and trust a human. And humans make mistakes and ship bugs all the time."

The split looks something like this.

The gamble + skip verification and hope the agent got it right = fast now, with risk nobody can see.

The verified team + tests, observability and review that catch what the agent misses = speed you can actually ship.

Fewer SaaS products, more tiny companies

Why are those verified teams going to win? According to Cullen, it comes down to headcount. Justworks' working hypothesis is that we'll see a lot of new companies that hit five to ten people and simply stay there because they don't need to grow.

And most of them won't be tech companies. Think plumbers, accountants, and mom and pop shops building their own software stack to run the business.

That also answers a question people keep asking. If these tools are so good, where's the next billion-dollar startup? Cullen thinks a lot of the output is going into tools people build for themselves, tools they once would have paid a SaaS company for or simply gone without.

"It's a lot easier to build a thing that doesn't need to scale and that doesn't have to have a multitude of features and configurability because it just needs to work for me and my three coworkers."

Shon pushed back a little. A ten-person company still has to deliver serious value to stay profitable, and that's where vibe-coded tools tend to fall apart. Cullen agreed it doesn't take a billion-dollar business to make that math work.

The AGI Question, and Why Nobody's Losing Sleep

The episode wrapped on the big one. What happens when software can build software, and eventually build everything else?

Shon's take was refreshingly grounded. If that runaway moment were coming, the AI labs would have hit it first. Anthropic, Google and Meta have had years of head start with large language models. A friend at Google describes its internal platform, Borg, as a near-perfect observability system that can optimize memory allocation down to the sub-process level, and it has been doing that for years. Google kept its position, and nothing ran away. Shon's honest read is that the models are sort of tapped out where they are.

Cullen pointed to a recent announcement where AGI got described as more of a spiritual concept than something you can measure, and shrugged. If the little coder bot on his Mac Mini that turns Datadog alerts into PRs ever puts in its two weeks' notice, he'll figure out where else to plug it in.

Find your slowest step

So where does that leave a team trying to move faster with AI? Back at the framework. Code got cheap, which means the build step is rarely your bottleneck anymore. Look at code review, prioritization, enablement and, above all, verification. The teams that invest there are the ones that will actually ship faster, with fewer people and fewer surprises. If integrations are on your roadmap, our Integration Strategy and Design posts are a good next read.

Watch the full interview

Originally published on
October 8, 2026
Latest

From the Blog

Check out our latest content on technology partnerships, integration and APIs. Access research, resources, and advice from industry experts.

The Five Questions to Ask Before (and After) You Build an MCP Server

MCP adoption is growing fast, but most of what's been announced isn't actually live. Five questions, backed by new data, for deciding whether to build one, checking if yours works, and knowing when you still need a traditional integration.

What Dozens of Vertical SaaS Reviews Reveal About QuickBooks Integrations (And How to Build a Better One)

A look at the QuickBooks integration issues vertical SaaS users report most often, why the API makes them so common, and how to build one that holds up.