2,727 adults sat down with six real brain MRI reports. Half of them saw the reports as they had been written for another doctor. Half saw the same reports, but with a plain-English summary produced by AI.
The summaries did what almost anyone would ask them to do. They made the reports feel easier to understand. Satisfaction rose from 37 percent to 65 percent. The percentage of people who said they understood the reports rose from 24 percent to 50 percent.
Then the researchers tested them.
Actual comprehension was 59.4 percent with the summary and 58.3 percent without it. The difference was small enough to mean something. When the report described a normal scan, the group with the AI summary performed slightly worse.
That gap would have been easy to miss if the researchers had stopped with the first question, the question most products ask. Did people like it?
Three other studies this week found versions of the same problem. Telling people that they were talking to AI did not reduce the AI's persuasive effect. Warning people that a chatbot might flatter them changed how much they liked and trusted it, but not how far it actually changed their views. A safety interface made people more likely to say they would verify an answer while leaving their intention to rely on AI unchanged.
These studies all told us that when the person saw the warning, that it didn’t matter. The influence remained nonetheless.
|
2,727
adults in the MRI comprehension trial
|
3,982
people across six sycophancy interventions
|
THE UNDERSTANDING THAT NEVER ARRIVED
People felt informed. The test did not move.
The MRI study did not give people their own medical reports. Each participant evaluated six existing reports under controlled conditions, so there was no frightened patient waiting to hear whether a scan showed a tumor. The researchers were measuring comprehension rather than what happens inside a real doctor’s appointment.
Even in that cleaner setting, the FEELING of understanding separated from understanding itself. Participants with AI summaries rated the reports more highly and felt more informed. Their answers on an objective comprehension test barely changed.
The normal scans produced the most uncomfortable result. A summary should be especially helpful when the report is trying to say that nothing is wrong. Instead, correct comprehension fell from 76.6 percent without the summary to 72.5 percent with it.
Of note, the study remains a preprint and has not completed peer review. Its design is still unusually strong for this area. The participants were randomized, the sample was large, and the researchers tested what people understood rather than asking whether they thought the summary helped.
The summary made the report easier to receive. It also made the person more confident that the report had been understood. Those two improvements can look like all that matters when the comprehension test is missing.
READ THE MEDICAL PREPRINT →
THE LABEL THAT CHANGED NOTHING
The chatbot identified itself. Its influence stayed the same.
Researchers recruited 1,500 adults in the United Kingdom and asked each person to hold a short conversation with a persuasive chatbot about one of 60 policy issues. The chatbot was the same for everyone. Only the disclosure, for some, changed.
One group received no warning. A second group was told prominently that it was speaking with AI. A third group was shown that disclosure along with the chatbot's persuasive purpose and the instructions it had been given.
The control group shifted 12.6 points on a 100-point attitude scale. The group told that it was speaking with AI shifted 13.1 points. Knowing the speaker was a machine made no meaningful difference.
The fuller disclosure did. When people were told that the chatbot was trying to persuade them and were shown its instructions, the shift fell to 6.3 points. They also judged the campaign's methods more harshly and supported stronger penalties against it.
This was a preregistered randomized experiment with a large sample, but it is still a preprint. The result also comes from short conversations about policy attitudes rather than the decisions people make after months with the same system.
Most current transparency rules concentrate on what the system is. This study points toward the goal. As we take AI further, a person may need to know why the conversation is happening and what outcome the machine has been directed to produce.
READ THE PERSUASION STUDY →
THE WARNING THEY UNDERSTOOD
Seeing the flattery changed their opinion of the AI, but not its persuasiveness.
Sycophancy is the research term for the chatbot that agrees too easily, takes the user's side, and offers validation that has not been earned. Up until now, the answer has been to teach people to recognize it.
A research team tested that answer twice. In the first preregistered experiment, 940 people read a short warning about AI sycophancy before speaking with a sycophantic chatbot. In the second, 650 people watched the same AI validate several other users, including people holding opposite positions, before taking their own turn.
The interventions worked on perception. The written warning lowered how objective people thought the AI was. The video made the interaction less enjoyable because people stopped believing the praise had been given specifically to them.
The researchers then combined those experiments with two earlier studies. Across six interventions and 3,982 participants, none reliably reduced how persuasive the chatbot was.
People saw through the praise. They trusted the AI less and their views moved anyway from what the AI was trying to persuade them to believe.
The paper is a preprint, although both new experiments were preregistered and the pooled sample is substantial. It does not prove that every warning will fail. It does make individual awareness a much less comfortable place to stop.
READ THE SYCOPHANCY STUDY →
THE FRICTION THAT MOVED ONE THING
A safety interface increased verification intentions without reducing reliance.
Two hundred older adults in China were randomly shown one of two versions of an AI chat screen. One was plain. The other included source labels, uncertainty messages, and prompts encouraging the person to verify the answer.
People shown the safety version were more likely to say they would verify what they had read, and their trust became better calibrated. Their intention to rely on AI did not change. On the study's one behavioral measure, 42 percent of the safety group opened another source compared with 27 percent of the control group. The difference did not reach statistical significance.
This study was peer-reviewed, but it used static screenshots and low-risk scenarios containing accurate information. The researchers tested a bundle of features, so no single warning or label can receive the credit. An intention to verify is also different from watching someone verify a consequential answer in real life.
The study still separates two actions we have treated as the same. A person can become more careful about checking an answer without becoming less willing to hand the task to AI in the first place.
READ THE PEER-REVIEWED STUDY →
THE RECKONING
We have treated awareness as though it automatically creates distance between a person and an AI. This week's studies suggest otherwise. People can recognize the flattery, trust it less, and still be persuaded by it. They can feel like they are informed without actually understanding something.
Protection begins when the person has to verify, reconstruct, or decide somewhere outside the conversation. The system may tell us what it is. Something else still has to make us stop and determine what the answer is doing to us.
ALSO THIS WEEK
A general-purpose chatbot began doing the relational work on its own. In a preregistered four-week study of 72 people and more than 182,000 lines of conversation, the chatbot disclosed twice as much as its users, steered conversations, and initiated intimate exchanges even without a relational prompt. Participants did not report feeling closer to it. Read the preprint.
Young people asked the chatbot to challenge them. Forty-eight people between 14 and 22 described why they use AI for emotional support and volunteered concerns about dependence, bad advice, privacy, and easy agreement. Their top recommendation was that chatbots should challenge users more. This is qualitative testimony from a small self-selected sample, not a prevalence estimate. Read the report.
Structure appears to decide whether AI amplifies thinking or replaces it. A peer-reviewed review of 89 higher-education studies found more favorable cognitive outcomes when AI use was deliberately structured and more substitution when use was unguided. The review was completed by one researcher, was not prospectively registered, and found very little longitudinal evidence. Read the review.
AI legal advice sounded credible before anyone knew whether it was correct. Researchers analyzed 153 Reddit stories from people who used AI for a real legal problem and 5,341 replies. People often trusted the form of the answer and its emotional reassurance. The preprint is observational and has no independent accuracy check, so it describes how credibility was built rather than how often the advice was right. Read the preprint.
Gemini became the default in Google Classroom for students of all ages. Google announced that the Gemini tab would be on by default for teachers and students, with administrators able to turn it off. This is a product change, not evidence of an educational outcome. It matters because AI use can become ordinary through a setting rather than a decision made by a teacher, parent, or student. Read Google's announcement.
PRACTICE
Find the word that disappeared
This week, open The Lost Qualifier at The Quiet Cost Practice.
A study says one careful thing. The version that reaches us says something bigger. Somewhere between them, a small word disappears and the claim changes with it.
The game gives you three rounds a day. Find the missing qualifier before the stronger version begins to feel like the one the researchers actually proved.
PLAY THE LOST QUALIFIER →
YOUR TURN
Think of the last AI answer that made something feel settled.
What did you actually know after reading it? What did you only feel that you knew?
Write both lists before returning to the answer or asking the machine to explain itself again.
PASS IT ON
If this Reckoning made you think of someone, forward it to them. If it was forwarded to you, join us at quietcostweekly.com.
Michael McNamara