Your AI Summary Says More Than Your Notes Do: The Over-Generalization Bias
Ask an AI to summarize your careful notes and it will likely say more than you did. Not less — more. In a controlled study, large language model summaries were nearly five times more likely than human ones to widen a scoped finding into a sweeping claim. The caveats you wrote do not survive.
That five-times figure is precise. Uwe Peters and Benjamin Chin-Yee tested ten leading models across 4,900 machine summaries; in a direct comparison against human-written summaries, the machine versions carried an odds ratio of 4.85 for broad generalizations, with a 95% confidence interval of [3.06, 7.70] and p < 0.001 1 2. The models did not merely drop your hedges. They replaced them with confidence you never claimed. A line that read "in these two interviews, most participants said" comes back as "participants said." The scope quietly inflates. The authors name the pattern flatly: a "strong bias in many widely used LLMs towards overgeneralizing scientific conclusions" 1.
What we think a summary is
We treat a summary as a faithful shorter version of the original: the same claims, in fewer words. That belief is reasonable. It is also how the tool is sold, and how it behaves most of the time on plain, unqualified prose. The trouble begins exactly where your notes are most careful.
Careful notes are full of scope. You write down who said it, how many times, under what condition, and how sure you are. That hedging is not clutter. It is the difference between a record and a rumor. Strip the scope and you do not get a smaller truth. You get a bigger claim wearing the smaller claim's clothes.
Consider a note you might write after a meeting:
In our two customer calls this week, both people said onboarding felt slow — though one was mid-migration, so this may not generalize.
Now the summary a chatbot hands back: Customers say onboarding feels slow. Every qualifier is gone. The sample of two, the "this week," the migration caveat, the "may not generalize" — all removed, and in their place a clean, quotable sentence that reads like a finding. It is not a finding. It is two calls and an honest doubt, stripped of the doubt.
The summary says more, not less
The distortion has a direction. It is not random noise; it is a consistent widening of scope. Across ten leading models and 4,900 summaries, generalizations grew broader than the source in a way human summarizers did not match — the odds of a broad overgeneralization ran nearly five to one against the original 1.
This is not folklore about "AI being vague." The study tested a named field: ChatGPT-4o, ChatGPT-4.5, DeepSeek, LLaMA 3.3 70B, and Claude 3.7 Sonnet among ten models, each summary graded against the paper it came from 1.
The point is not that the models are careless. They are fluent, and fluency is the trap: a broadened claim reads more like knowledge than the cautious original it came from, so it wins on the surface exactly while it loses on the facts.
There is already a companion argument that a summary is a lossy view, so you should keep the original bytes rather than trust the compression 1. You can read that case in let the AI summarize, but keep the original. Over-generalization is its darker twin. Omission drops what you wrote; over-generalization asserts what you did not. One leaves a gap. The other fills the gap with a claim you never made and now appear to stand behind.
Asking the AI to be accurate makes it worse
The obvious fix fails. Telling the model to be accurate does not reduce over-generalization — it increases it. Peters and Chin-Yee found that "prompts that included direct requests to avoid inaccuracy increased algorithmic overgeneralizations," with several models overstating in 26 to 73 percent of cases even under that instruction 1.
Newer is not safer either. In the same study, "newer models tended to perform worse in generalization accuracy than earlier ones" 1. Capability and caution point in opposite directions.
The cause is training, not malfunction. The authors trace it to the reward signal: because "reinforcement learning from human feedback (RLHF) enhanced models' helpfulness, it often led them to express unwarranted confidence or reduced their ability to hedge claims to indicate uncertainty" 1. A confident, fluent, universal sentence scores well with a human rater. A hedged one reads as evasive. So models "may learn to prioritize confident fluency over caution and precision, increasing their tendency to produce overgeneralized statements" 1.
The polish is the problem. The smoother the sentence, the further it has often traveled from what your note actually supports.
Where the honesty lives, and where it does not
Two caveats keep this in scope. First, the bias was not universal: in that study Claude models "did not significantly differ from the original texts" 1. Second, the research measured summaries of scientific papers, not personal notes. The mechanism is general; the exact five-times figure belongs to its domain, and that boundary is worth stating plainly.
State it, and the finding still lands where you live. Any hedged source text, whether a meeting note, a reading note, or a field observation, carries the same scope words the reward signal has learned to shave off.
Why does the direction of the error matter? Because people believe the confident version. Neil Rathi, Dan Jurafsky, and Kaitlyn Zhou found that "LLMs are linguistically overconfident in English, leading users to overrely on confident generations," and that "overreliance risks are high across languages" 3. An inflated summary is not merely wrong; it is persuasively wrong, and the persuasion travels into the next document, the next decision, the next quote of the quote. That confidence is also unreliable on its own terms — the model is just as confident when it is wrong about your notes.
The failure is, at least, a known and describable defect. Work on teaching models to signal uncertainty observes that "generated responses are typically unhedged or hedged in ways that do not reflect this variability" 4 — a faithfulness gap that even strong models show. Naming it is the first step to reading around it.
What to do with your own notes
The defense is not a better prompt; it is a kept original. Write your hedges into the note and keep that note as the source of truth. Then treat every summary as a draft to be checked against it — and check the universal claims first, because those are exactly where the inflation lands 1.
Five habits make that concrete:
- Keep the hedged original. Your scope words ("roughly," "in one case," "we think") are the data, not decoration. The summary is a view; the note is the record.
- Ask for preserved scope, not "accuracy." Since requesting accuracy backfires 1, instead ask the model to keep every qualifier, sample size, and condition, and to quote the source's own stated limits.
- Trace each universal claim back to the note. For any sentence that says "always," "customers," or "the data shows" without an exception, search your original for the qualifier it dropped. If the note hedged and the summary did not, the summary is wrong.
- Lower the temperature. The paper's own suggested mitigations include "lowering LLM temperature settings" — a less adventurous model over-generalizes a little less 1.
- Keep the caveat in your hands. A model can compress your prose; it cannot be trusted to carry your uncertainty. The doubt is yours to store, in plain text you own and can search.
Frequently asked questions
These are the questions people actually ask when a chatbot hands back a tidy version of their own writing. The short answer runs through all of them: the summary is likelier to inflate your claim than to lose it, and the fix lives in the original you keep, not in the prompt you write.
Does an AI summary exaggerate what my notes say? Often, yes. In a controlled study of 4,900 summaries, large language models were nearly five times likelier than humans to turn a scoped observation into a broad generalization (odds ratio 4.85). They tend to drop your qualifiers and replace them with confidence the note never carried 1.
Why does ChatGPT overstate what my document says? The likeliest cause is training. Reinforcement learning from human feedback rewards confident, fluent answers, so models learn to prioritize confident fluency over caution and hedging. A qualified sentence reads as evasive to raters; a sweeping one reads as helpful. The result is a systematic tilt toward broader claims than the source supports 1.
Does asking the AI to be accurate make summaries better? No — it can make them worse. Peters and Chin-Yee found that prompts explicitly requesting accuracy increased over-generalization rather than reducing it, with some models overstating in 26 to 73 percent of cases even under that instruction. Ask the summary to preserve every qualifier and condition instead 1.
Are newer AI models more accurate at summarizing? Not for this problem. In the same study, newer models tended to perform worse in generalization accuracy than earlier ones. Capability and caution moved in opposite directions. A more advanced model may write a smoother summary while being more likely, not less, to overstate the scope of your original 1.
How do I stop an AI summary from over-generalizing my notes? Keep the hedged original as your source of truth, and check the summary against it. Ask the model to preserve scope and quote the source's limits rather than to "be accurate." For any universal claim in the summary, search your note for the qualifier it dropped. Lowering the model's temperature can help too 1.
Do AI summaries just drop the caveats in my writing? That is only half of it. Dropping detail is omission — a summary is a lossy view, so keep the original. Over-generalization is the opposite and worse: the summary adds certainty and reach your note never had, replacing "in one case" with a confident universal you now appear to endorse 1.
A summary is a convenience, not a witness. It can tell you roughly what a page said; it cannot be trusted to remember how carefully you said it. The confidence in the output belongs to the model, not to you — and the only copy that still holds your doubt is the one you wrote down and kept.
MNMNOTE keeps your notes as plain Markdown on your own device, where the hedges you wrote stay exactly as you wrote them and stay searchable when you check a summary against them — mnmnote.com.
Footnotes
-
Uwe Peters and Benjamin Chin-Yee, "Generalization bias in large language model summarization of scientific research," arXiv:2504.00025, 28 March 2025. https://arxiv.org/abs/2504.00025 — accessed 2026-07-27. ↩ ↩2 ↩3 ↩4 ↩5 ↩6 ↩7 ↩8 ↩9 ↩10 ↩11 ↩12 ↩13 ↩14 ↩15 ↩16 ↩17 ↩18 ↩19
-
Uwe Peters and Benjamin Chin-Yee, "Generalization bias in large language model summarization of scientific research," Royal Society Open Science 12(4):241776, April 2025. https://doi.org/10.1098/rsos.241776 — accessed 2026-07-27. ↩
-
Neil Rathi, Dan Jurafsky and Kaitlyn Zhou, "Humans overrely on overconfident language models, across languages," COLM 2025 (arXiv:2507.06306), 8 July 2025. https://arxiv.org/abs/2507.06306 — accessed 2026-07-27. ↩
-
Bryan Eikema et al., "Teaching Language Models to Faithfully Express their Uncertainty," arXiv:2510.12587, 14 October 2025. https://arxiv.org/abs/2510.12587 — accessed 2026-07-27. ↩