One of the common problems in software teams is that if you ask them what exact path a change takes from the moment they decide to make it until the moment a user actually sees it in the product, you probably won’t get a particularly clear answer. It’s not that they don’t have a process; in fact, they usually have plenty of processes. The work gets created, prioritized, planned, developed, reviewed, tested, approved, and eventually deployed. On paper, everything looks organized. But when you follow a real change from beginning to end—in other words, follow the actual flow of the work—you usually discover something else: several waiting points, multiple handoffs, decisions that were never documented anywhere, manual steps that everyone has simply forgotten are still manual, and, most importantly, places where nobody really knows how to tell whether the change actually worked.
This is more interesting to me than many of the discussions we have about software development methodologies. We love giving names to things in this industry. There seems to be a methodology for almost every problem: TDD, BDD, DDD, SDD, and many others. Then, at the management layer, we have Scrum, Sprints, Stories, Points, Velocity, and all kinds of processes around them. Some of these ideas are genuinely useful, and their value can’t simply be dismissed. The problem starts when the method itself becomes the goal. Instead of asking, “What problem does this solve for us?”, we start focusing on, “Are we following this particular methodology correctly?”
Take TDD, for example. The main question for me isn’t whether we should write the test before the code or after it. The real value of testing is that it shortens the distance between our assumption that the software is correct and the evidence that it actually is. If you change something today and find out five minutes later that its behavior is wrong, the cost of that mistake is usually small. If you discover it three weeks later, after several other changes have already been made around the same part of the system, the story is completely different. So what really matters isn’t the test itself; it’s the length of the feedback loop. Testing is one tool for shortening that loop, not the goal itself.
Once we take this perspective a little further, we can see the same problem throughout the entire software delivery cycle. Suppose you’ve implemented a change, its tests are passing, and the code has been reviewed. Is the work done? If you still can’t confidently deploy it, maybe not. If you can’t tell what effect it had on the system after deployment, still not. If something goes wrong and you don’t know how to quickly limit its impact or roll the change back, definitely not. In other words, “development being finished” doesn’t necessarily mean “the change is finished.” A change is really finished when it has reached the real system, production, and you can get feedback from its actual behavior.
This is where I think the DevOps perspective becomes important. If we strip DevOps away from all the tools and slogans surrounding it, there is a fairly simple idea underneath: the distance between making a change and understanding the result of that change should be as short as possible. You change the code, build it, test it, deploy it, and then you should be able to understand what actually happened in the real world. That “after” isn’t a small part of the work; it is a continuation of the same work.
The problem is that we often treat the test environment as a substitute for reality. No matter how good a test environment is, it still isn’t the real environment. In the real world, we have real data, real users, real request volumes, behaviors nobody wrote tests for, and dependencies that may behave differently at exactly the moment our system changes. So the goal of good engineering isn’t to reach a point before deployment where we can be certain that there are no errors; such a point simply doesn’t exist. The goal is to discover what we don’t know at the lowest possible cost and with the smallest possible blast radius.
That’s why progressive delivery, to me, isn’t just a deployment technique; it’s an engineering decision about managing uncertainty. If we deploy a change to the entire system all at once, we’re effectively accepting all the risk at once. If we first deploy the same change to a small part of the system, we’re accepting the possibility that we might be wrong while limiting the size of the mistake. Then, if the system behaves as expected, we gradually increase the scope. That’s an important difference: we’re no longer trying to eliminate risk; we’re trying to make it controllable.
Of course, this only has value if we can actually understand what happened. That’s where observability comes in, and I think this concept is often confused with simply having a few charts, graphs, and alerts. You can have dozens of dashboards and still have no idea why a user can’t complete a purchase. CPU and memory look normal, the services are up, and the error rate isn’t particularly high, but the user still can’t perform the main action the product is supposed to provide. From an infrastructure perspective, the system looks healthy; from a business perspective, it’s broken.
After something like that happens, the more interesting question isn’t “Why did it break?” but “Why didn’t we know sooner?” The answer usually reveals something about the system itself. Maybe we measured the wrong thing. Maybe an important user behavior isn’t observable at all. Maybe we didn’t have the right alert. Maybe the system was designed in a way that makes failures difficult to diagnose. An incident, in other words, isn’t always just a problem in the software; sometimes it’s a problem with our ability to see the software and with our understanding of what observability actually means.
This creates an interesting loop. An incident happens, we find the cause, and we fix it. But if that’s where the work ends, we haven’t really learned much. We may need to add a new test so the same failure doesn’t come back, add a new metric so we can detect it earlier next time, change the deployment strategy so the impact of a similar failure is smaller, or even change the design of the system itself. In this way, an incident becomes an input into the next development cycle—and ultimately a source of strength.
I think this is one of the important differences between a team that simply fixes bugs and a team that actually learns from its system. In the first kind of team, an incident is an unpleasant event that needs to be closed as quickly as possible. In the second, an incident is an opportunity to identify a permanent weakness in the system and add another layer of defense. That layer might be a test, an alert, a constraint in the code, an architectural change, or even a change in the way the system is deployed. What matters is that the knowledge gained from the incident makes its way back into the system.
If we look at software from this perspective, perhaps even our definition of software quality starts to change. High-quality software isn’t simply software with fewer bugs. It’s software where, when something goes wrong, we can quickly understand what happened, limit its impact, and turn what we learned into something that makes the system a little better next time. That definition is much more interesting than simply saying, “Our test coverage is 80%!”
Now this becomes even more interesting with the arrival of AI agents. If an agent can produce code much faster than a human, a new question appears: okay, what happens next? If code generation becomes ten times faster while review, testing, deployment, and feedback remain at the same speed, we’ve simply increased the speed of producing something that the rest of the system still can’t process at the same rate. The bottleneck may simply have moved from writing code to somewhere else.
That’s why I don’t think the most important value of AI agents is simply that they can “write code faster.” The much bigger opportunity may be their ability to shorten parts of this feedback loop as well. For example, they could make the change, run the necessary tests, put it into an isolated environment or sandbox, observe the system’s behavior, and based on the results decide whether the change is ready to move forward. At that point, the agent isn’t simply a faster programmer; it becomes part of our engineering system.
But there is a serious trap here. If our flow is broken, an agent can execute the same broken flow faster. If the requirements are ambiguous, an agent can produce the wrong solution faster. If review is the bottleneck, more code simply piles up in the review queue. If deployment is risky, an agent can produce more risky changes. A change might require only a few hours of actual engineering work and still spend three days waiting for review, approval, or deployment. In that situation, cutting coding time in half doesn’t change much. Or deployment itself might take only five minutes, while it takes two days to realize that the change caused a problem. Making deployment faster isn’t the real issue there either. What we need to shorten is the distance between “making a change” and “understanding the result of that change.”
This is why, the more I think about this, the less I care about what we call the methodology we’re using. TDD can be useful, Scrum can work well for some teams, domain modeling can be genuinely valuable in a complex system, and other approaches can have their place too. But none of them should replace the fundamental question: does our team have a short, observable, and reliable cycle for changing the system and learning from the result? If we do, many of these techniques naturally find their place. Testing is useful where fast feedback is needed. Progressive delivery is useful where uncertainty is high. Observability becomes critical where system behavior can’t be fully predicted in advance. Incidents become inputs into development, and AI agents can be introduced where they actually reduce a bottleneck.
Ultimately, software engineering is about building a good system that can learn from its own changes. You make a change, try it with limited risk, see what happened, find out quickly if it was wrong, fix it, and make sure the system is a little better the next time. If we have that kind of cycle, perhaps it doesn’t matter quite so much what we call it.