TL;DR:
- AI coding assistants are neither uniformly good nor uniformly bad. Task complexity decides. Bounded work (tests, prototypes, documentation, scripting) produces reliable gains. Complex work (architectural judgment, multi-file changes, tacit domain knowledge) pays the saved time back in verification overhead and rework.
- This is why aggregate productivity measurements often show “no effect”: gains and friction are of similar magnitude and cancel out in the average. Not because nothing is happening, but because both happen at once.
- The task complexity boundary is movable: experienced developers invest more time upstream in specification and requirement formulation, converting complex tasks into bounded ones.
A while ago, I burned an entire afternoon together with an AI assistant. In our microservice setup with micro frontends, Tailwind styles that worked fine in the individual frontends broke once they were loaded into the host application. The assistant and I went in circles: suggestion, test, failure, new suggestion. The next day a colleague took one look and fixed it with a one-liner. He knew the setup and had the mental model. The AI and I had neither.
The same gap shows up at scale. Vendors promise 26–56% time savings based on controlled experiments, but production data tells a different story: a study across 400 companies found that a 65% increase in AI usage produced only a 7.76% increase in pull request throughput. A METR study showed experienced open-source developers were 19% slower with AI while feeling 20% faster. I wanted to understand this gap between promise, perception, and measurement. Not in a lab, but here at inovex.
What I studied
My master’s thesis (Hochschule der Medien in Stuttgart) is a mixed-methods case study at inovex with three independent data strands:
- Telemetry: 177 days of organization-level GitHub Copilot usage data (~138,000 code suggestions across 72 languages)
- Survey: a questionnaire built on the SPACE framework, 36 responses (28 AI users, 8 non-users)
- Thematic analysis of the free-text responses (61 statements, 8 themes)
I intentionally did not measure “productivity” as a single number, because research shows this is unreliable for AI-assisted coding and in general. Instead, I examined where three independent strands converge on the same conclusion; confidence rises. Where they contradict each other, the contradiction itself is a finding.
Where AI coding assistants help, and where they don’t
All three data strands converge on the same pattern. In the telemetry, Python (structured, self-contained tasks) has a 36% acceptance rate, while TypeScript React (logic, markup, and state in the same file) sits at just 19%. In the survey and free-text responses, the dominant benefit lives in bounded tasks: tests, prototypes, documentation, and scripts. The frustration lives on the complex side, wherever architectural judgment, multi-file coordination, or tacit knowledge is required.
The DORA report calls AI an “amplifier”: it strengthens existing competence as well as existing uncertainty. Task complexity is the mechanism underneath. My Tailwind micro frontend afternoon sat exactly on the complex side of that task complexity boundary. The problem lived in the integration between systems, not in any single file the AI could see.
A null result that isn’t one
The efficiency comparison between AI users and non-users is not statistically significant. That sounds like no effect, but the opposite is true. At the item level, strong forces face off: 96% agreement on prototyping, 86% on documentation search, and 79% on task speed against the 61% who say post-processing takes longer than expected. The average across all tasks is the signature of two opposing forces, not the absence of an effect.
What this means in practice
Use AI where it reliably delivers. Tackle bounded tasks actively with AI: tests, scaffolding, prototypes, conversions, and documentation. The gain here is real and reproducible.
Secure the complex side instead of overselling it. Promising AI gains for complex architectural work undermines the credibility of the real gains. This side needs consistent reviews and enablement, meaning the time to build usage competence rather than expecting instant wins.
Shift from writing to specifying. The most interesting adaptation in the data is experienced developers invest significantly more time in use-case description and requirement engineering before letting the AI work. One calls it “spec-driven development.” They are moving the boundary itself, turning complex tasks into bounded ones. With the shift toward agentic workflows, this turns from a trick into a prerequisite. The friction doesn’t disappear; it moves upstream into specification and verification overhead.
Therefore, the craft of software development isn’t being replaced. It’s relocating: less into the keystrokes that produce code and more into the upstream intent that directs it and the downstream judgment that decides what gets to stay.