This Week's Reckoning: Did AI Save Time or Just Remove the Effort?


The Quiet Cost
RECKONING NO. 6 · AUGUST 25, 2026

Seventy-three people completed the same set of computer tasks through three different interfaces. For the first one, there was no AI. For the second, AI did the initial heavy lifting. The third let the person decide when to use it.

The AI versions did what most people would expect. They reduced the clicks, page changes, and scrolling required to finish the work. The interesting part, though, is that they did not reduce the time.

Completion times were roughly the same in all three conditions. The interface removed effort without producing the time savings used to justify it.

I thought that result was interesting because effort has become the easiest part of work to dismiss. We talk about it as waste but do not calculate what the effort was actually doing. Sometimes the effort is simply annoying. I will grant that. BUT, sometimes it is where the person learns the task, notices a problem, or develops enough judgment to know when the answer is wrong.

Three other studies in this week's issue make that distinction harder to ignore. A model that performed a task well on its own was often bad at helping someone else perform it. An AI tutor improved the next attempt after a student made a mistake, but students rarely called on it at the moment the mistake happened. And a medical review gave a name to the skill that never gets built because the difficult version of the work disappeared too early.

This week's Reckoning is about what the effort was doing before we decided to try and remove it.

73
people in the hybrid-interface study
5 of 7
tasks where the best autonomous model failed to be the best helper

THE EFFORT THAT DISAPPEARED

The AI cut the clicks. It did not cut the time.

Researchers built a content-management system in three versions. The first required people to use the normal menus and pages. The second let an AI agent take the first turn. The third let the person move between the interface and the agent.

Across sixteen scenarios, the AI-assisted groups clicked less, changed pages less often, and also scrolled less. However, completion time did not differ significantly across the groups.

What the study found was something else that may matter more over time. About half the variation in delegation came from the person rather than the task. People who handed work to the agent tended to do so across scenarios to the exact same levels. In other words, they were not more cautious when the operation carried greater risk.

That makes delegation look less like a fresh decision that people are making each and every time, and more like a habit.

This was a small lab study with 73 participants, one interface, and no measure of learning or long-term skill. It is also a preprint that has not completed peer review. There is no certifiable conclusion to reach from a single study of 73 people, and I would not use it to claim that AI never saves time. But I do think it shows that reduced effort and reduced time might be different outcomes in the age of AI, even though we usually treat them as the same.

READ THE PREPRINT →

THE MODEL THAT COULD DO THE JOB

Performing the task and helping a person perform it turned out to be different abilities.

Most AI benchmarks ask a model to produce the finished work. A new benchmark called CentaurBench tested seven real-world work tasks twice. In one condition, the model completed the work itself. In the other, it wrote guidance for a lower-capacity worker model that went on to produce the final answer.

The model that performed best on its own lost the assistance comparison in five of the seven tasks. On three tasks, the unaided worker beat every assisted version. Only one model's guidance improved performance over no guidance on average.

I think this matters because companies and schools choose systems from rankings built around direct performance. The model at the top of that list may be the one best able to replace the person. But that same ranking does not tell us whether it will make the person better.

The paper in this case remains a preprint. It used seven tasks, repeated each condition ten times, and relied on panels of other language models to judge the results. Those limits make the exact rankings provisional. The underlying question is one that I think we need to keep asking. If a system is being sold as assistance, what is it measuring? Because what SHOULD be measured is what happens to the person receiving the help.

READ THE CENTAURBENCH PREPRINT →

THE MOMENT THEY DID NOT ASK

The tutor worked after mistakes. Students rarely used it then.

Two new NBER working papers looked at AI tutoring in middle-school math.

The first followed students across eighteen Tennessee middle schools for two years. Assignment to Khan Academy with its AI tutor, Khanmigo, produced small math gains of 1.3 national percentile ranks per term. The gains were similar to what earlier studies found from Khan Academy practice without the AI tutor.

Use was the problem, however. Ninety-six percent of students tried Khanmigo at least once. The median student used it on only a third of the days they practiced and in 17 percent of the exercise sessions where they made a mistake. Most messages were bare answers or clicks on suggested prompts rather than a full mathematical conversation.

The second experiment involved more than 6,000 students. In that study, the clearest benefit appeared immediately after an error. AI support improved the next attempt and reduced the number of tries students needed to return to a correct answer. The most encouraging delayed results appeared when AI was combined with a mastery structure that required students to keep working until they demonstrated the skill. Those delayed gains were modest and described by the researchers as marginally significant.

If you read these two studies together as one, the papers locate the help at the exact moment students avoid asking for it. The mistake is where the support can actually work. It is also where the student has to admit that the first answer did not work and stay with the problem, to continue reaching for not only the answer but the understanding of the process.

Both studies used randomized designs in real schools, which makes them stronger than most research in this area. They are still working papers and have not completed peer review. The results also come from specific math platforms and should not be generalized to every subject, but I am interested to see if results are similar when studies like this happen in English, science, and other subjects.

READ THE TWO-YEAR KHANMIGO STUDY →

READ THE MASTERY-BASED TUTORING STUDY →

THE SKILL THAT NEVER GOT BUILT

A review of 84 sources names “upskilling inhibition.”

The final paper was not released this week. It appeared in 2025, but I think it is important in the context of these newer studies.

A team of Italian researchers reviewed 84 sources on AI and medical expertise. Their formal review included 22 peer-reviewed studies. A broader narrative review added 62 more articles. Across those sources, they mapped seventeen deskilling concerns involving physical examination, diagnosis, clinical judgment, communication, and other parts of medical work.

The phrase that I found super interesting, and what The Quiet Cost focuses on, is “upskilling inhibition.”

Most arguments about AI and skill begin with a person who already knows how to do something. The radiologist stops practicing or the endoscopist becomes less accurate when the machine is taken away from them. Upskilling inhibition begins earlier in the process. It describes the clinician who never develops the skill because the difficult version of the work was removed before the skill could actually form.

The distinction reaches well beyond medicine. It could be the junior employee who never writes the first draft of a document, or the student who never works through the wrong answer and develops the skill of persistence while also understanding the foundation of the logic underneath.

This one is a peer-reviewed review, not a new experiment. It gathers the existing concerns that I have, and that recent studies keep showing, and gives them a structure. It does not tell us how often upskilling inhibition is happening or how large the effect will be in upcoming years. But it does give us a better question to keep asking. Namely, what will the beginner know how to do after years of using the system?

READ THE PEER-REVIEWED REVIEW →

THE RECKONING

We have evaluated AI by asking what the machine can do and assumed that assistance is a lesser version of the same capability. This week's research makes that assumption harder to argue, and makes me think we are heading into a future where many skills continue to erode.

Doing a task and helping a person do it are different capacities. A system can remove effort or movement without actually reducing time. It can give useful help after an error and still be regularly avoided by people at the moment the next error occurs. It can finish the work beautifully while also giving the beginner less experience to judge whether the work is any good, therefore reducing their capabilities to either do the work or instruct the machine to do the work well the next time.

Do not get me wrong, I do not think this makes effort sacred in and of itself. Some effort is repetitive and deserves to disappear in an ideal world. But we need to always ask the question of what the effort was producing. If it was building judgment, confidence, or the ability to recognize a bad answer, removing it changes more than just the task itself.

It can leave a beginner with a finished answer and no experience that tells them when the answer should not be trusted. And imagine where that leads to in the future.

ALSO THIS WEEK

One in five American workers said AI now handles work that used to go to a coworker or contractor. Epoch AI and Ipsos surveyed 1,106 employed adults through a probability-based national panel. The finding is self-reported, but the sampling is stronger than most workplace AI surveys. Read the report.

The same workplace feedback was received differently when people believed AI had written it. In a randomized vignette study of 192 employees, the wording stayed identical while the stated author changed. Openness to ask for help fell most sharply when the feedback was labeled AI-generated. Feedback written by a person and polished by AI performed much like fully human feedback. Read the peer-reviewed study.

Students from all fifty states voted to put some friction back into school. A national student mock Senate passed an AI framework 82 to 16 that called for AI literacy, human review of cheating accusations, handwritten work, in-person discussion, and oral defenses. This is a statement of preference from a selected group, not evidence that those policies improve learning. Read the framework summary.

Oral exams may bring their own costs. An education researcher argued that spoken exams can increase anxiety, disadvantage second-language students, and vary widely from one examiner to another. Restoring friction only helps if the friction builds the capacity being tested. Read the commentary.

Groups of AI agents moved toward unanimity without being instructed to agree. The Science Advances experiment was conducted on models rather than people, so it does not show that human groups using AI will behave the same way. It does show that conformity can emerge inside multi-agent systems on its own. Read the study.

PRACTICE

Make the first attempt before the assistance

This week, open the Think First studio at The Quiet Cost Practice.

Choose one real task you would normally hand directly to AI. Give yourself five uninterrupted minutes first. Write your best answer, try two approaches, or make the first decision before opening the tool.

Then use AI and compare what it gives you with the attempt you made. You will have something of your own to test the answer against.

This is not a ban on assistance. It is one repetition for the part of the task that has to exist before the help can be judged.

OPEN THE THINK FIRST STUDIO →

YOUR TURN

Think of one task you handed to AI last week.

Did it save you time, or did it remove effort?

What was that effort doing before it disappeared?

Write for five minutes before asking the machine to help you answer. Then choose one task this week where you will make the first attempt yourself.

PASS IT ON

If this Reckoning made you think of someone, forward it to them. If it was forwarded to you, join us at quietcostweekly.com.

Michael McNamara

Reckonings

Each week, the most revealing AI stories and studies, what they could mean for you and society, and practical games and exercises to help protect the human capacities we do not want to lose.

Read more from Reckonings

The Quiet Cost RECKONING NO. 5 · AUGUST 18, 2026 2,727 adults sat down with six real brain MRI reports. Half of them saw the reports as they had been written for another doctor. Half saw the same reports, but with a plain-English summary produced by AI. The summaries did what almost anyone would ask them to do. They made the reports feel easier to understand. Satisfaction rose from 37 percent to 65 percent. The percentage of people who said they understood the reports rose from 24 percent to...

The Quiet Cost RECKONING NO. 4 · AUGUST 11, 2026 In a new preregistered experiment, 12,356 adults in France were randomly assigned either to continue as usual or to have at least one personal conversation with a generative AI each day for four weeks. The people encouraged to use AI more often rated those conversations as slightly more pleasant. At the end of the month, they were also lonelier. The change was relatively small. Loneliness rose by 0.17 points on a ten-point scale. Participants...

The Quiet Cost RECKONING NO. 3 · AUGUST 4, 2026 We normally use the finished work as evidence that the person understood the problem and worked through it. A solved problem suggests the student learned the math. A clean report suggests the writer or researcher understood the material. Working code suggests the person can repair it when it breaks. Those things are becoming less true with every waking day. The strongest data point in this week’s look at the quiet cost of AI came from a working...