It's difficult not to fall into constant FOMO these days. The speed of change in software engineering feels like a once-in-a-lifetime shift.

You try out Claude. You build three startup MVPs over a single weekend.

On Monday you go back to work and your boss/clients/stakeholders keep asking what your plan is to move faster with AI.

You open the same Claude on your large work project. You write less code now, but the work isn't moving that much faster.

Online, you see everyone talking about their harnesses, multi-agent setups, skills, MCPs, how AI ships all their features... OK, that's just LinkedIn influencers who are probably farming engagement.

But you still look at your project and wonder - why isn't it moving faster?

Then you see DHH announcing in his Rails World keynote that he doesn't see himself as a software engineer anymore. Pencils down, start making things instead of coding. We can move so much faster. It's going to be awesome.

And that hurts even more - regardless of what you think of DHH, he did create a few great things.

Then your teammate says "Yeah, but DHH is probably just working on some greenfield applications. He couldn't do that with our system."

So - what's wrong with our system?

An existing system carries a lot of baggage. Some of it sits in the codebase, some in the organization around it. Together, they make it hard to speed things up beyond having the AI agent write parts of the new code for you.

This post focuses on the codebase. For a good part of this year I've been working on implementing agentic engineering practices across Upside's clients, and I've noticed a few common impediments that show up quite consistently across the applications that have been built pre-AI.

Knowledge That Lives in People's Heads

There's obviously a large gap between a prototype and a large, mature application. What makes the latter difficult to copy is the depth of the problems that it solves. Over the years of development, a system accumulates a huge number of choices that are made by its creators when discovering the domain and as a response to users' feedback.

While AI agents can do a good job of understanding some workflows and implementation details directly from the code, the difficulty scales with the size of the codebase. At some point, the agent can still get the overall picture of what the codebase does, but it will miss many of the small details.

Which is not much different from humans. Just think about how long it often takes for a new developer to get familiar with an existing codebase and understand why certain features work the way they do.

To work efficiently and make the right implementation decisions, AI coding agents need good-quality documentation that's also easy for them to navigate.

If you work with a pre-AI codebase, that's usually missing. Sometimes it lives in a Confluence site, where outdated and current docs are mixed together. Other times, the important bits may just be in the heads of people like your product manager, your senior developer or the CEO. And that gets in the way of AI moving fast.

The documentation needs to live in a repository, or in an MCP-enabled tool that allows the agent to access and update it. On the technical side it's rather easy, and there's plenty of tools that support MCP now, even Confluence. The tricky part is how to fill it with up-to-date requirements and how to structure them. You can use AI to synthesize documentation from the code and the tests, or to extract it from the existing sources, but you'll still need to verify the result with people who are familiar with certain requirements.

Lack of Automated Validation

The speed of AI-powered iteration can only be as fast as the speed at which we are able to verify the changes. That means both that what has been built for us works as we want it to work, and that it doesn't break any of the existing behaviors of the system.

Many legacy systems lack extensive automated test suites. They used to be costly to build, then costly to maintain, and sometimes even costly to run. And in some industries it didn't even matter that much because you could always move fast, break things, make more money, and then have the customer support reps calm down a few unhappy customers.

But that equation changes drastically when using AI for coding. It's hard to go far without proper tests. Most harnesses will try to generate at least the basic unit tests, or maybe even low-effort integration tests, just to make their own work possible. Without that, they will get lost in an endless loop of fixes and regressions. And once you work in an agentic engineering process, you will notice very fast that you need to build an even more extensive test infrastructure.

For most older codebases, however, you will need to revisit your test strategy. It's tempting to crunch more features with AI, but in the long run it may be better to spend some time on improving the coverage first. Together with clear specs, it's a rather large investment, but it removes the two biggest bottlenecks.

Inconsistent Coding Style and Architecture

If you look at an average older codebase, you'll be able to spot different epochs of programming in it. For example, when looking at a ten-year-old Rails codebase, you'll likely see:

  • A bunch of controllers where someone was trying to use some paradigms from functional programming in an object-oriented language. That was hot in 2016.
  • A connector to Apache Kafka that receives just a few event types from the application, and then the same application reads those events and reacts to them. That's because your Tech Lead had a plan to go into microservices and handle millions of requests per second.
  • Some Rails-based views that use Hotwire instead of React, because it was considered the future of Rails frontend in 2022.

You get the picture. Unless you have a very generous budget for dealing with tech debt, your system probably also has some of those annoyances that confuse everyone, but there's never a good moment to sort them out.

Every single one of those inconsistencies makes it difficult for the AI coding agent to navigate the codebase, to understand the meaning of its different parts and to write new code.

The more standardized the codebase is, the easier it is for the AI agent to perform code lookups and the fewer iterations it needs to reason about functional and architectural decisions. When the AI is able to do that correctly, it produces better-quality output. And by being influenced by the standardized code it reads, the AI ends up creating code that conforms to the same standards more easily.

Shopify spent a good part of last year standardizing their codebase and infrastructure, and there is a good reason for that.

Dead Code

Another seemingly harmless problem with a large impact is code that's no longer used.

Sometimes it's a small leftover function. Sometimes it's an implementation of a process that's not connected to any trigger. Sometimes it's a deprecated API endpoint that someone forgot to remove from the codebase.

In any of these situations, the dead code confuses the AI agent and leads to it crafting an overcomplicated implementation. Very often that implementation will support edge cases that don't really happen in real life, or rely on events or data that's not there.

Fortunately, it's easy to use a harness to prune dead code, or at least dead code that's clearly unused. It gets a bit trickier with API endpoints that no one calls anymore, or for example with features that are only activated based on configuration.

And once you get the initial cleanup done, keep in mind that it's not a one-time thing. As the code evolves, AI coding agents will keep leaving leftovers behind. A periodic cleanup will need to become part of your process.

A Monolith That's Too Large

Large codebases have always been a pain to work on if too many developers were making changes to them at the same time. Too many merge conflicts, too many dependencies between team members.

A common approach was that once your team grew beyond "two pizzas", you had to start thinking about splitting the team and the codebase. Microservices, microfrontends, modules - whatever worked. The core idea is simple: if you have isolated parts of the codebase with clear interfaces between them, engineers can stick to the negotiated interfaces and work within the boundaries of their own part of the system.

That is still a good idea, but the "two pizzas" threshold has changed. A single developer can now have multiple agents working on different branches, on different features in parallel, and they generally create more code faster. If each developer can create as much code as two developers could before, a 5-person team already goes past the "two pizzas" mark.

This is one of the areas where I see a lot of change going on. Making the codebase smaller and running it as independent modules or services with clear interfaces makes it more manageable. It does come with an infrastructure overhead, but a smaller codebase is easier to keep constraints in and to test at its boundaries.

Lack of Automation Around the Codebase

The easy thing for a developer is to use an AI coding agent where they would traditionally type code in their IDE. This is what you see in established teams quite often - the focus is on "how do we prompt" and "how do we build skills that write the code according to our conventions". It surely speeds up writing boilerplate code, but it doesn't make the work go that much faster in the long run.

Now that we're able to generate code using AI, the focus should be elsewhere. Part of that is the domain knowledge, part quality check. But there's more to it - all the pipelines around the system that support code generation and maintenance.

For example: it doesn't have to be a human who creates a pull request with a bugfix - it can easily be a callback from a telemetry system based on recurring error logs. And that's just the low-hanging fruit.

The real challenge lies in identifying the processes and the data that can be used to further improve the product. This can be defining some benchmarks for the test suite, or creating processes that collect data about product usage and feed it back to the AI for improvements.

This is the hardest part, because there are no "industry standard" practices yet. It's also the most exciting one, because it leaves the most room for creativity and experiments.

What's Next

All of this baggage is what keeps you at the point where the AI agent writes parts of the code and nothing more. Clearing it, even if not fully, in my experience leads to a great improvement in the AI coding workflows.

You don't have to do it all at once. Tests come first, because without them you can't trust anything the agent changes. Pruning dead code and adding the first automations are cheap and pay off quickly. Writing down the knowledge that lives in people's heads takes longer, and standardizing the architecture or splitting the monolith is something you do piece by piece.

Improving the codebase is not all of it. There are deeper changes that an organization needs to go through to fully embrace what's possible now. But that's a different story that I'll cover another time.

In the meantime, if you're working through the same problems in your own codebase, I'd like to hear how it's going.