Why faster AI coding can mean harder engineering
InfoWorld ·

We keep asking whether AI can make developers more productive, but that’s the wrong question. We should instead be asking what happens when it does. All those productivity gains can be for good, of course, but they could also be for ill. Simon Willison made this clear in a recent post. “The more time I spend working with coding agents, the more convinced I am that they make software engineering even harder.” Wait, what? Why would he say that? Because getting the most from them requires “extraordinary discipline and knowledge.” This wasn’t a denunciation of AI coding. If you’ve read Willison for more than a millisecond, you know he’s a fan. Indeed, in his post, he explicitly celebrated what the tools make possible. One respondent apparently missed that, however, warning that without deep expertise, it’s difficult to distinguish correct output from plausible nonsense. Willison’s response ? “You’re making the exact same point as me.” This isn’t just retreading familiar ground (that AI still needs experienced engineers to impose that discipline). It’s something different. After all, faster implementation doesn’t automatically create extra human capacity. Yes, it can change the work people attempt, the decisions they must make, and the expectations attached to their jobs. But a team can deliver more and still find the work more demanding. Both can be true. Companies should be careful about budgeting as though AI has already given them both a larger road map and a smaller need for engineering expertise. It’s by no means clear that faster coding is an unalloyed good. Faster isn’t the same as easier In the replies, one reader, Hillel, put it neatly : “It doesn’t get easier, you just get faster.” Geoffrey Huntley concurs , describing the work as “incredibly taxing” even for experienced operators. He reported roughly 16-hour days over the preceding 12 days, explicitly noting that this was his choice, not a requirement at work. That qualification matters because these comments aren’t representative measurements of developer productivity, nor are they evidence that agents inevitably exhaust everyone. They’re simply observations from people using the tools enthusiastically . The enthusiasm is part of the story. As more becomes possible, people attempt more. I’ve argued before that AI coding agents need good software engineers . AI doesn’t eliminate the need for engineering discipline. It raises the price of not having it. But having the expertise and having enough capacity to apply it are different things. Consider a team that uses faster implementation to take on a migration it had postponed. That’s a perfectly reasonable use of AI. With AI the team can afford to attempt something valuable, but that doesn’t change the fact that the migration still introduces decisions about compatibility, customer disruption, and which old behaviors must survive. The engineers may now spend more of their day resolving those questions, even if they spend less time writing the changes. In other words, the gains are real, but so is the additional responsibility. Comparing the new, more ambitious workload with the old one and concluding that engineering has become easy misses what the team actually bought with its saved time. Nor does this depend entirely on models making mistakes. Better agents can reduce correction and verification work. They can also make previously impractical projects feasible. How much of the resulting capacity becomes breathing room and how much becomes another commitment is a choice the organization has to make. The extra work becomes the job There’s some research behind this concern, though it’s maybe not as incontrovertible as it appears on the surface. UC Berkeley researchers Aruna Ranganathan and Xingqi Maggie Ye studied AI use at a 200-person technology company over eight months. Their work included observation and more than 40 interviews across functions. They saw employees expand their responsibilities, fit prompting into former pauses, and keep more work running simultaneously. Much of this was voluntary, as people were excited by what they could accomplish. The problem was how that excitement reset expectations. As Ye explained, “what was once extra effort becomes standard performance.” This was qualitative, in-progress research at one company, not proof that AI makes every workplace more exhausting. But it identifies a plausible way productivity gains can become difficult to sustain: A burst of experimentation becomes the ordinary workload against which everyone is measured. The management question is what gets counted as success. If the goal is more ambitious work, admit that the team is spending its gains on ambition. If the goal is lower costs, measure the full cost of completing and maintaining the work before assuming the savings. If people are working longer to hit the new targets, some of the apparent improvement may be coming from additional labor. I recently wrote about the tools developers are building to manage coding agents. Better tools can help people recover context and finish work. They can’t decide how much work management should expect. A well-organized queue can still contain an unreasonable amount of work. Spend the gains deliberately Willison’s own work offers a useful example of what productive AI use can look like. In September, he described a security audit of Datasette using several coding agents. External vulnerability reports prompted the audit, and successive rounds uncovered more problems. Willison and Alex Garcia then divided the work. For most issues, one wrote tests demonstrating the problem while the other implemented the fix. Two humans examined each issue alongside agents using different models. Fixes went into the main development branch, with selected changes also applied to the stable release. AI helped them discover useful work, and they supplied an explicit process for completing it. The additional findings were valuable precisely because the maintainers acted on them. This account doesn’t tell us how many hours AI saved, but it does show why counting findings alone would miss much of the accomplishment. A longer list of vulnerabilities to fix can be evidence of a better audit, making it absurd to call that a failure because it created work. It would be equally absurd to assume the people fixing them had suddenly become less necessary. For an engineering leader, that suggests a more useful conversation than demanding another increase in the percentage of code written by AI. Ask the team where the time went. Did a change reach customers sooner? Did the saved implementation time pay for a deeper audit? Did review spill into evenings? These are quite different outcomes, even if the coding agent looked equally impressive in all three. Then give engineers permission to spend some of the gains on making the next task less demanding. A test that reliably catches a recurring failure, creates a clearer interface, or removes an unnecessary dependency can reduce the decisions someone has to revisit. That work needs room in the plan. Otherwise, every improvement in implementation speed risks being consumed by new features while the cost of understanding the system keeps accumulating. An enthusiastic experiment also shouldn’t automatically become next quarter’s staffing assumption. Before turning a burst of output into a standing commitment, find out whether the team sustained it within its normal workday, including review and maintenance. People can choose to throw themselves into an interesting project, but that doesn’t establish how much work they can routinely absorb, and it shouldn’t become an obligation for colleagues who didn’t volunteer for the experiment. Willison is describing demanding work that can be worth doing. That’s a much more credible case for AI than promising that software engineering will become effortless. Enterprises should pursue the additional capability, then budget honestly for the people and time needed to use it. The productivity gain belongs in the plan once. If the bigger road map only works because engineers extend their workday, some of that gain is coming from the engineers working more, with the AI simply serving as taskmaster.
We keep asking whether AI can make developers more productive, but that’s the wrong question. We should instead be asking what happens when it does. All those productivity gains can be for good, of course, but they could also be for ill. Simon Willison made this clear in a recent post. “The more time I spend working with coding agents, the more convinced I am that they make software engineering even harder.” Wait, what? Why would he say that? Because getting the most from them requires “extraordinary discipline and knowledge.” This wasn’t a denunciation of AI coding. If you’ve read Willison for more than a millisecond, you know he’s a fan. Indeed, in his post, he explicitly celebrated what the tools make possible. One respondent apparently missed that, however, warning that without deep expertise, it’s difficult to distinguish correct output from plausible nonsense. Willison’s response ? “You’re making the exact same point as me.” This isn’t just retreading familiar ground (that AI still needs experienced engineers to impose that discipline). It’s something different. After all, faster implementation doesn’t automatically create extra human capacity. Yes, it can change the work people attempt, the decisions they must make, and the expectations attached to their jobs. But a team can deliver more and still find the work more demanding. Both can be true. Companies should be careful about budgeting as though AI has already given them both a larger road map and a smaller need for engineering expertise. It’s by no means clear that faster coding is an unalloyed good. Faster isn’t the same as easier In the replies, one reader, Hillel, put it neatly : “It doesn’t get easier, you just get faster.” Geoffrey Huntley concurs , describing the work as “incredibly taxing” even for experienced operators. He reported roughly 16-hour days over the preceding 12 days, explicitly noting that this was his choice, not a requirement at work. That qualification matters because these comments aren’t representative measurements of developer productivity, nor are they evidence that agents inevitably exhaust everyone. They’re simply observations from people using the tools enthusiastically . The enthusiasm is part of the story. As more becomes possible, people attempt more. I’ve argued before that AI coding agents need good software engineers . AI doesn’t eliminate the need for engineering discipline. It raises the price of not having it. But having the expertise and having enough capacity to apply it are different things. Consider a team that uses faster implementation to take on a migration it had postponed. That’s a perfectly reasonable use of AI. With AI the team can afford to attempt something valuable, but that doesn’t change the fact that the migration still introduces decisions about compatibility, customer disruption, and which old behaviors must survive. The engineers may now spend more of their day resolving those questions, even if they spend less time writing the changes. In other words, the gains are real, but so is the additional responsibility. Comparing the new, more ambitious workload with the old one and concluding that engineering has become easy misses what the team actually bought with its saved time. Nor does this depend entirely on models making mistakes. Better agents can reduce correction and verification work. They can also make previously impractical projects feasible. How much of the resulting capacity becomes breathing room and how much becomes another commitment is a choice the organization has to make. The extra work becomes the job There’s some research behind this concern, though it’s maybe not as incontrovertible as it appears on the surface. UC Berkeley researchers Aruna Ranganathan and Xingqi Maggie Ye studied AI use at a 200-person technology company over eight months. Their work included observation and more than 40 interviews across functions. They saw employees expand their responsibilities, fit prompting into former pauses, and keep more work running simultaneously. Much of this was voluntary, as people were excited by what they could accomplish. The problem was how that excitement reset expectations. As Ye explained, “what was once extra effort becomes standard performance.” This was qualitative, in-progress research at one company, not proof that AI makes every workplace more exhausting. But it identifies a plausible way productivity gains can become difficult to sustain: A burst of experimentation becomes the ordinary workload against which everyone is measured. The management question is what gets counted as success. If the goal is more ambitious work, admit that the team is spending its gains on ambition. If the goal is lower costs, measure the full cost of completing and maintaining the work before assuming the savings. If people are working longer to hit the new targets, some of the apparent improvement may be coming from additional labor. I recently wrote about the tools developers are building to manage coding agents. Better tools can help people recover context and finish work. They can’t decide how much work management should expect. A well-organized queue can still contain an unreasonable amount of work. Spend the gains deliberately Willison’s own work offers a useful example of what productive AI use can look like. In September, he described a security audit of Datasette using several coding agents. External vulnerability reports prompted the audit, and successive rounds uncovered more problems. Willison and Alex Garcia then divided the work. For most issues, one wrote tests demonstrating the problem while the other implemented the fix. Two humans examined each issue alongside agents using different models. Fixes went into the main development branch, with selected changes also applied to the stable release. AI helped them discover useful work, and they supplied an explicit process for completing it. The additional findings were valuable precisely because the maintainers acted on them. This account doesn’t tell us how many hours AI saved, but it does show why counting findings alone would miss much of the accomplishment. A longer list of vulnerabilities to fix can be evidence of a better audit, making it absurd to call that a failure because it created work. It would be equally absurd to assume the people fixing them had suddenly become less necessary. For an engineering leader, that suggests a more useful conversation than demanding another increase in the percentage of code written by AI. Ask the team where the time went. Did a change reach customers sooner? Did the saved implementation time pay for a deeper audit? Did review spill into evenings? These are quite different outcomes, even if the coding agent looked equally impressive in all three. Then give engineers permission to spend some of the gains on making the next task less demanding. A test that reliably catches a recurring failure, creates a clearer interface, or removes an unnecessary dependency can reduce the decisions someone has to revisit. That work needs room in the plan. Otherwise, every improvement in implementation speed risks being consumed by new features while the cost of understanding the system keeps accumulating. An enthusiastic experiment also shouldn’t automatically become next quarter’s staffing assumption. Before turning a burst of output into a standing commitment, find out whether the team sustained it within its normal workday, including review and maintenance. People can choose to throw themselves into an interesting project, but that doesn’t establish how much work they can routinely absorb, and it shouldn’t become an obligation for colleagues who didn’t volunteer for the experiment. Willison is describing demanding work that can be worth doing. That’s a much more credible case for AI than promising that software engineering will become effortless. Enterprises should pursue the additional capability, then budget honestly for the people and time needed to use it. The productivity gain belongs in the plan once. If the bigger road map only works because engineers extend their workday, some of that gain is coming from the engineers working more, with the AI simply serving as taskmaster.