An article with a blunt title hit the Hacker News front page over the last few days: AI Coding will Prevent Expertise. As I write this, the discussion thread sits at 497 points and 485 comments.
The author, Lars Faye, is an engineer who has written code for decades, and the core of his argument is a catch-22: AI coding tools demand expert judgment to use well, and they bypass exactly the friction that produces expert judgment. In his words: “If these tools demand expertise, yet the tools can actively circumvent the friction that cultivates expertise, then what is the path for one to become an expert so they can effectively use these tools?” Junior developers are pushed to adopt AI to stay competitive before they have built the judgment that wielding it requires, judgment that used to come from writing code by hand and getting stuck on it. The result, he says, is finishing with an “illusion of competence” rather than true understanding.
None of this is a new worry, and it has usually stayed at the level of senior-engineer intuition; the comment section split into the usual camps. This time I want to take it seriously, because there is now experimental data aimed directly at the question, and one piece of it comes from a maker of the tools.
The hunch now has a control group
In late January, Anthropic published a randomized controlled trial: 52 software engineers, most of them junior, were randomly split into two groups and asked to learn Trio, a Python async library none of them had used before, and implement two features with it. One group had access to an AI assistant and the other did not; because assignment was random, the average difference in outcomes can be read as the effect of AI access. Right after the coding task, everyone took a quiz covering four skills: debugging, code reading, code writing, and conceptual understanding.
The AI group averaged 50%. The hand-coding group averaged 67%. That is a 17-point gap, close to two letter grades, statistically significant with a sizable effect (Cohen’s d = 0.738). What the AI group bought in exchange was about two minutes of speed, and even that was not significant. Of the four question types, the widest gap was in debugging.
One detail gives this study unusual weight: it comes from Anthropic, the developer of Claude Code. When the shovel seller runs the test and reports that the shovel hurts your hands, evidence against interest of that kind is usually more credible than a third-party attack.
Learning happens at the moment you get stuck
Why would this happen? Anthropic’s report is direct about the mechanism: cognitive effort, “even getting painfully stuck,” is “likely important for fostering mastery.” When you learn something new, pulling things out of memory, trying, failing, and correcting is itself how memory and intuition form; memory research has well-established evidence that actively retrieving information from memory is a key step in building long-term retention. AI takes that step over wholesale: the code arrives on schedule, and the learning never happens. Faye reaches for the German word Fingerspitzengefühl, “fingertip feeling,” the accurate but hard-to-articulate intuition that veterans grind out over years of hands-on work. Remove the friction and that intuition has no soil to grow in.
What makes it worse is that the loss is invisible from the inside. METR’s experiment last year had 16 experienced open-source maintainers work on 246 real issues in the large projects they maintain, with AI access randomized per issue. With AI they were actually 19% slower, yet beforehand they forecast a 24% speedup, and afterwards they still believed AI had sped them up by 20%. That study measured speed, not learning, but it establishes the same point: subjective impressions of AI’s effect and objective measurement can diverge by forty percentage points. The sense of competence that comes from shipping working code does not necessarily correspond to real understanding.
There is a second layer in Anthropic’s data: variance inside the AI group was large. The high scorers used the assistant in visibly different ways. Some asked “why is it written this way” after code was generated; some had the AI produce explanations alongside the code; some asked only conceptual questions and fixed every error themselves. The lowest scores came from delegating the whole task to the AI and waiting for the result, and the other heavy-reliance patterns all averaged under 40%. To be clear, this part is a post-hoc grouping by observed usage, and Anthropic states explicitly that causal conclusions cannot be drawn from the interaction patterns. The randomization establishes the causal effect of the AI-access condition; how much each usage pattern hurts is, for now, correlational evidence.
One more dataset points in the opposite direction. Two researchers exploited the staggered timing of Claude Code adoption across developers to run a quasi-experimental analysis of monthly panel data on 5,346 GitHub developers. The authors note that because adoption is voluntary, the results are event-time associations, not definitive causal effects. After adoption, the number of programming languages a developer actively used rose by 2.5 on average, against a baseline trend of 0.9; first attempts at new languages rose by 1.2; and those first attempts concentrated among developers whose prior technical range was narrowest. So on one side, the compressed immediate understanding in Anthropic’s trial; on the other, an expansion of range: territory people previously didn’t dare enter, they now do. The two results measure different things, and “used a new language” is not the same as “learned a new language.” This study measures usage behavior, not understanding.
Finally, the boundary of the evidence. Anthropic’s quiz was taken minutes after the coding task ended; it measures immediate understanding, and it does not answer whether experts will fail to form over the next decade. Faye’s claim is long-term; the experimental evidence so far covers the short term. Getting from “17 points less learned in one task” to “a generation’s expertise collapses” requires a compounding assumption: that every time you learn something new you use the delegate-everything pattern, and the losses stack. The extrapolation is reasonable, but I could not find a longitudinal study that tests it directly, and Anthropic itself lists whether immediate quiz performance predicts long-term skill development as a question this study leaves open.
Why calculators never ran up this debt
The whole time I was writing this, one object kept coming to mind: the calculator. We handed arithmetic over to it entirely. I can still do it by hand, just slowly, and nobody mourns the decay of that craft; we moved on to other things. Will AI coding turn out the same, one more tool transition that looks like a false alarm twenty years from now?
Lining the two up, my judgment is that the analogy is half right, and the wrong half is the part that matters.
The right half is offloading itself. Hand arithmetic is mechanical labor; giving it up costs nothing worth keeping, and boilerplate you have already written a hundred times is the same. The wrong half has two pieces. First, ordering. At least in the math education I went through, the calculator arrived after arithmetic was learned: years of doing it by hand came first, the friction stayed intact through the learning period, and the tool took over something already mastered, which falls squarely on Faye’s “offloading” side. Math education treats that ordering as consequential: the US National Mathematics Advisory Panel cautioned that “to the degree that calculators impede the development of automaticity, fluency in computation will be adversely affected”. A junior developer entering the industry today typically meets the tools in the reverse order: the AI is there from day one, before the judgment it demands has had a chance to form. Second, verification. A calculator performs basic operations deterministically; the same input gives the same result every time, and checking it is trivial. AI-generated code can compile and pass the tests it was given while still being wrong: when researchers expanded the test suites of a standard code-generation benchmark, the added tests caught enough previously undetected wrong LLM-generated code to cut pass rates by up to 19.3-28.9%. And checking AI code requires precisely the judgment you can only build by writing code yourself. The two tools also take over work of different magnitudes: a calculator handles single fixed-procedure operations, while people now hand agents the whole decision chain from technology selection down to implementation detail, and every link in it can be wrong. So the skill the calculator retired (hand arithmetic) and the skill needed to use a calculator well (knowing what to compute) are two different skills, while the skill AI coding retires and the skill needed to use it well are the same one. Faye’s catch-22 never existed for calculators. For AI coding it is structural.
The analogy raises a bigger question: what if one day AI-written code genuinely no longer needs a human to verify it? Then the analogy holds completely, and people can put programming down the way we put down hand arithmetic. Push one step further: programming and spreadsheet work are jobs an industrial society invented, not obviously “natural” for humans; maybe silicon really is better suited to them than carbon, and people should be released to the outdoors, to the physical world. I don’t find that vision absurd. But it describes an endgame, and every number in this post is about the transition. The awkwardness of the transition is exactly this: AI is not yet good enough to skip verification, as METR’s slowdown and the debugging gap keep reminding us, and Anthropic says plainly that humans still need the skills to catch errors and provide oversight for AI deployed in high-stakes environments. Society still needs a supply of engineers with judgment as the backstop, and the pipeline that supplies judgment is the very thing AI has already rewritten. Even if the endgame really is humans exiting programming, getting through the transition takes a generation that still has judgment. As for when the endgame arrives and what people should do once it does, I don’t know, and I won’t pretend to.
Keep the friction where you intend to grow
The most useful cut in Faye’s piece, to me, is the distinction between cognitive offloading and cognitive debt: “Cognitive debt is abdicating your judgment and decisions, whereas cognitive offloading is delegating the mechanical or tedious.” Handing over mechanical labor you have already mastered is offloading, and it’s fine. Handing over judgment you have not yet built is debt, and it comes due. In practice:
- When learning a new framework, language, or domain, don’t let AI write the code. Use it as a tutor: ask for explanations, ask follow-up questions, fix your own errors. The high scorers in Anthropic’s data used exactly these patterns. Faye’s version is more extreme: “the most productive learning that can happen with an AI coding tool is when it isn’t used to generate much of any code at all.”
- Do your own debugging. It was the category with the largest gap between groups, and it is also where the case for keeping friction is strongest: in the passage Faye quotes, François Chollet calls LLMs “interpolation engines,” while software engineering “is an exercise in adaptation and novel problem-solving.”
- In domains you have already mastered, delegate freely, and use the tool to widen your range. Stacks you have never touched are especially worth trying: in the GitHub panel data, first attempts at new languages concentrated among developers with the narrowest prior range.
- If you manage a team, design for this deliberately. Anthropic’s advice to managers is to “consider systems or intentional design choices that ensure engineers continue to learn as they work.” My two concrete suggestions: keep AI-free exercises in new-hire training, and in code review require the author to explain every generated block. No explanation, no merge.
Expertise will not collapse automatically just because AI exists. What the current data can say is this: at least for immediate understanding, how much you lose tracks how you use the tool, a pattern Anthropic itself reports only as correlational, and no study I could find answers the long-term question. But delegating the whole task and waiting for the result, the pattern that takes the least effort, is exactly the pattern that lost the most points in the experiment. Friction is no longer programming’s default configuration. It has become something you must actively choose to keep, and where you choose to keep it determines how much irreplaceable judgment you will still hold five years from now.
References
- AI Coding will Prevent Expertise — Lars Faye — the piece under discussion; source of the core paradox, the cognitive debt vs. offloading distinction, and all Faye quotes
- Hacker News thread — discussion volume (497 points / 485 comments as of this writing) and the arguments on both sides
- How AI assistance impacts the formation of coding skills — Anthropic — RCT design and all quantitative results (52 participants, 50% vs. 67%, non-significant two-minute speedup, largest gap in debugging, high-scorer usage patterns, advice to managers)
- Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity — METR — 16 developers / 246 issues, 19% slowdown, 24% forecast speedup, 20% post-hoc perceived speedup
- Agentic Delegation and the Language Frontier of Software Developers: A Model and Evidence from Claude Code on GitHub — arXiv:2605.25438 — 5,346-developer panel, +2.5 active languages (0.9 baseline), +1.2 new languages, concentration among previously specialized developers; the authors’ event-time-association caveat
- The critical role of retrieval practice in long-term retention — Roediger & Butler, Trends in Cognitive Sciences — retrieval practice as a key mechanism of long-term memory formation
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation — arXiv:2305.01210 — expanded test suites (HumanEval+) catch previously undetected wrong LLM-generated code, reducing pass@k by up to 19.3-28.9%
- Foundations for Success: The Final Report of the National Mathematics Advisory Panel — U.S. Department of Education, 2008 — the caution that calculator use which impedes the development of automaticity adversely affects computational fluency