The cover letters got better. As evidence about the people who wrote them, they got worse.
On a large online labor market, an AI writing tool helped job seekers tailor their applications more closely to vacancies. The biggest improvements appeared among people whose letters had previously been weakest. There was some evidence of a modest increase in callbacks, but the effect faded after two months.
Then something stranger happened. After the tool became available, the relationship between how well a letter was tailored and whether its author received a callback fell by 51 percent. Employers began relying more heavily on applicants’ previous work histories.
The letters had improved at the task they were meant to perform: presenting a candidate persuasively. But they had become less useful at another task we had quietly assigned them: revealing something about the candidate.
That distinction reaches far beyond hiring.
We usually judge a technology by what it helps us produce. Does the answer become more accurate? Does the report become clearer? Does the program run? Does the student solve the problem? Does the worker finish sooner?
Those are sensible questions. But many human activities have never produced only the visible result. They also leave something behind in the person doing them: understanding, skill, judgment, confidence, commitment, habits of attention, or sometimes dependence. And because the work bears traces of those things, other people use it to decide whom to trust, hire, admit, promote, or place in charge of the next problem.
Generative AI can loosen the relationship among these outcomes. It can improve the artifact, alter what the person learns from producing it, and change how much the artifact tells us about that person—all at the same time.
The argument over whether AI makes people more capable or less capable therefore starts too late. Before asking whether the person has improved, we need to notice that the work itself has been doing more than one job.
What Work Leaves Behind
For some tasks, only one outcome matters.
If software converts a photograph from one file format to another, the converted file is the point. Few people care whether its owner could perform the conversion manually. If automation removes a tedious operation that nobody expects a human to perform again, losing practice at that operation may be a gain rather than a loss.
Other activities are different.
Writing produces sentences, but the act of writing can also produce understanding. Practice produces a performance and, sometimes, a skill. Deliberation produces a decision while shaping the judgment brought to the next decision. Planning creates a sequence of intended actions, but the act of making the plan can also create ownership of it.
The visible product is one result of the work. The other is what the person can see, understand, remember, question, verify, or do afterward.
That second result is not necessarily good. Repetition can produce bad habits. Easy success can produce misplaced confidence. A badly designed activity can consume effort without building anything useful. The point is not that manual struggle is inherently virtuous. It is that a process can change its participant, and sometimes that change is part of what the process is for.
Learning research has long distinguished performance from learning. Conditions that make an exercise easier in the moment do not necessarily produce stronger retention or transfer later; difficulty during practice can sometimes produce more durable learning.
Generative AI makes the distinction unusually consequential because it can move through a task rather than merely perform one fixed operation inside it.
A calculator does arithmetic. A spelling checker flags words. A search engine retrieves pages. A general-purpose AI system can frame a question, propose an explanation, generate alternatives, retrieve information, draft an answer, attack that answer, revise it, summarize the evidence, and decide that the task is finished.
The boundary between what the person does and what the tool does can shift several times before a single piece of work appears.
That flexibility is enormously useful. It is also why the same finished product can now conceal very different kinds of human activity.
One Tool, Many Routes
Imagine three analysts asked to write the same memorandum.
The first researches the problem, develops an argument, writes a draft, and asks AI to tighten the prose.
The second asks the model for competing explanations, searches for evidence that might defeat each one, rejects most of what the system suggests, and writes a conclusion the model never proposed.
The third delegates the research, outline, first draft, and revision, then checks the finished memorandum for obvious errors.
A reader might judge all three documents excellent.
But the analysts have not performed the same cognitive work. They have had different opportunities to discover which evidence matters, notice contradictions, revise beliefs, struggle with uncertainty, and build a mental model of the problem.
Recent experiments show why the location of assistance matters.
In a field experiment involving nearly 1,000 high-school students, researchers compared ordinary instruction with two forms of AI-assisted mathematics practice. Both AI systems improved performance while help was available. But students using the relatively unrestricted chatbot later performed worse than the control group when assistance was removed. A tutor designed to guide students without simply handing over answers largely avoided that penalty.
The underlying technology was not the whole intervention. What mattered was what the system left for the student to do.
Other evidence runs in the opposite direction, which is just as important. In experiments on professional writing, people who practiced with an AI editor later wrote better cover letters when the AI was removed. Simply studying a strong AI-revised example produced a similar benefit, suggesting that the tool could sometimes function as a teacher rather than a substitute. The result depended on how assistance entered the work.
Workplace evidence points in the same direction. A study of 5,172 customer-support agents found that an AI assistant increased productivity by about 15 percent on average, with particularly large gains among less experienced workers. When the software occasionally became unavailable, workers with prior exposure retained some of their improvement—evidence consistent with at least some learning rather than pure dependence. The assistant had not merely completed work around them; some of its practices appear to have been absorbed.
The evidence points to no single effect on human capability.
“Using AI” is too crude a description.
Asking for an answer, asking for a hint, generating objections, studying an example, debugging a model’s mistake, revising one’s own draft, and delegating an entire task can involve the same underlying system while producing very different forms of human participation.
The division of cognitive labor is part of the treatment.
And once that division becomes variable, a second problem appears.
A Better Result, a Worse Clue
Much of everyday life depends on an inference we rarely state explicitly:
I can learn something about you from the quality of what you produce.
A school reads an essay partly as evidence of understanding. An employer reads a work sample partly as evidence of judgment. A client sees a strong analysis and becomes more willing to trust its author with a harder problem. A promotion committee treats past performance as evidence of readiness for future responsibility.
These inferences were never perfect. Editors, teachers, colleagues, templates, search engines, and unequal access to help long predated generative AI.
What has changed is how many cognitive routes can now lead to the same outward result.
A polished memorandum remains a polished memorandum whether AI corrected a few sentences, supplied the decisive counterargument, found the evidence, or generated almost everything.
That matters because three questions that once traveled together can now move apart.
How good is the output?
What is the person capable of?
How much does the first tell us about the second?
The distinction is not theoretical. In experiments involving job pitches and startup proposals, access to ChatGPT improved evaluators’ ratings of the pitches while making the writers’ underlying expertise harder to identify. Screening errors increased by an estimated 4 to 9 percent. Yet the direction was not universal: among some writers from non-English-speaking countries, AI benefited experts more strongly and actually improved the ability to distinguish expertise. The information carried by performance changed according to who benefited and by how much.
This produces a possibility that much of the debate over AI misses:
Performance can become less revealing even when nobody becomes less capable.
Suppose AI helps everyone improve, but raises weaker performers more sharply and compresses the visible differences among them. Average capability can rise while a finished product becomes a poorer guide to who is best prepared for an unfamiliar problem.
The reverse can happen too. If highly skilled people extract greater gains from AI, assisted work may become more, not less, informative.
Output quality, human capability, and the evidentiary value of performance are therefore different variables. Improvement in one does not settle what happened to the others.
Nor does disclosure solve the problem.
An “AI-assisted” label tells us that a tool was involved. It does not tell us what the person delegated, what they verified, what they understood, or what they could reproduce under changed conditions. A complete transcript would reveal more about the route, but even that would not directly tell us what remained in the person at the end of it.
The problem is not simply authorship.
It is interpretation.
When the Signal Moves
Institutions do not stop making decisions because an old signal becomes noisy. They look for another one.
That is exactly what appeared to happen in the cover-letter study. As tailored letters became less informative about applicants, employers placed more weight on prior employment histories.
The substitution matters.
In many settings, generative AI can democratize access to high-quality production. Someone with weak prose can communicate a strong idea more clearly. A novice can draw on practices once available only through expensive coaching. A second-language writer can lose some disadvantages that had little to do with the substance of the work.
But equalizing one signal does not eliminate selection. It changes where selection happens.
If polished writing distinguishes candidates less effectively, institutions may turn toward live interviews, supervised exercises, credentials, employment history, professional reputation, process monitoring, or some new form of testing.
One tempting replacement is exhaustive process visibility: prompt logs, screen monitoring, keystroke histories. But a process record can show how a result was assembled without showing what the worker can do next. Turning uncertainty into surveillance would shift the cost of a broken signal onto the person being judged while leaving the underlying measurement problem intact.
Some substitutes will be better than the old signal. Others may simply reward a different set of advantages.
A technology that democratizes polish, for example, could inadvertently increase the value of pedigree if employers fall back on institutional history when the work sample no longer separates applicants clearly. That outcome is not inevitable. It is precisely why the redesign of assessment is not a minor technical question.
When an old proxy weakens, someone chooses the new one.
The important question is therefore not how to restore every pre-AI signal. Some deserved to disappear. Elegant prose has often been mistaken for intelligence; fluency for expertise; expensive preparation for merit.
The opportunity is larger.
If AI disrupts the old relationship between performance and capability, institutions can stop pretending that the relationship was ever automatic and ask a more demanding question:
What capability are we actually trying to observe?
Test What Remains
The answer will differ by setting.
If the real job consists of working with AI and the system will normally be available, banning the tool from an assessment may produce false reassurance. Test the combined human-machine system instead. Give people realistic, difficult cases with the tools they will actually use.
But then test the part for which the human remains responsible.
If a doctor, engineer, analyst, programmer, or manager is expected to catch a system’s unusual error, an assessment should include a plausible error. Can the person detect it? Explain why it matters? Recognize uncertainty? Decide when intervention is warranted?
If those abilities are what “human oversight” means, clicking approve is not evidence of them.
The issue is not how much human participation remains, but what that participation is for.
Hiring can use the same principle. A polished take-home assignment should carry less weight as a stand-alone proxy when the route to producing it is highly variable. A realistic AI-assisted exercise can be paired with a short oral defense, an unexpected variation, or a deliberately flawed input. The aim is not to catch candidates using AI. It is to observe the judgment the institution expects them to supply.
Education presents a different problem because formation is part of the purpose. A good assignment is not merely one that produces a good answer. It should leave the student better prepared for the next answer.
That does not require preserving every pre-AI exercise. It means measuring beyond the moment of assistance. Ask students to explain a concept later, transfer it to a new problem, detect a subtle mistake, or reconstruct enough of the reasoning to show what has been internalized.
A tool can remain available during much of learning while some carefully chosen moments reveal what the learner can carry forward.
The same logic applies to work. If AI generates routine code while allowing programmers to spend more time on architecture, losing some first-draft fluency may be a sensible trade. If those programmers can no longer diagnose unfamiliar failures, the trade looks different.
There is no universal quantity called “human capability” that must always be maximized.
We need to know which capabilities the new system eliminates, which it creates, which it merely hides, and which will still be needed when conditions change.
That is a more useful standard than preserving human effort for its own sake.
Automation should remove operations that are merely costs. It should be treated more carefully when it removes the very participation that produces a capability the larger system still depends on.
The Person After the Work
The cover letters had genuinely become better letters.
They were answering one question increasingly well: How good is this letter?
Employers had been using them to answer another: What kind of worker is behind it?
Generative AI widened the distance between the questions.
That distance will grow as AI moves beyond drafting isolated outputs and begins to participate in longer chains of research, planning, execution, revision, and decision-making. The more stages a system can occupy, the more ways a successful result can be produced—and the less the result alone can tell us about the human role inside it.
This need not be a story of human decline. AI may leave some users more knowledgeable, some more dependent, some better at verification, some weaker at first drafts, and many simply different from the people the old assessments were designed to measure.
The mistake would be to look only at the finished work and assume we know which future we are seeing.
The next generation of AI will be judged by what it can produce, solve, and finish. Those benchmarks matter.
But wherever the human still matters, another question matters just as much:
After the work is finished, what has the work produced in the person?
ZNetwork is funded solely through the generosity of its readers.
Donate
