
Ask a model for a homepage headline and you get ordinary copywriting.
Know your market before your competitors know they've lost it.
Win the argument before you walk into the room.
Ask it for ten options, which is what anyone actually does, and somewhere down the list this turns up.
Litigation Intelligence, Grounded in Authority.
Confident Clinical Decisions, Grounded in Patient Context
Every Employee Question, Answered and Sourced.
I wanted to know how often that happens, so I ran 400 headlines through two models. Ten analytical products, ten with no data angle at all, ten options a time.
My first version of this test asked for one headline per product, and sourcing language showed up zero times in 20. That result is what made me think headlines were safe, and it was an artifact of the question, because nobody asks a model for a single headline. Ask for ten and it turns up in 6 of 200, spread across 5 of the 20 times I asked.
Three percent of headlines sounds like nothing until you remember that you read all ten and keep one. A quarter of the time, one of the ten in front of you is about sources.
The control was dog walking, karaoke rooms, car washes, meal kits and tattoo studios. Zero out of 200, not one.
I expected GPT to be the offender, since that's where I'd seen this discussed. GPT came in at 2% and Claude at 4%. There's no interesting difference between them, which kills the tidiest explanation I had.
Headlines get argued over in a way the rest of the page never does.
Asked for a full page, a hero and a subheadline and three feature blurbs, sourcing language appeared on 5 of 15 analytical products at 5.6 instances per thousand words. Control arm: zero again.
One honest note on that. The control did produce "verified" three times, and I excluded all three, because each attaches to a person rather than a claim. Verified photographers. Verified renters. Healed-and-verified tattoos. That's a trust badge, a different sense of the word.
Here are two things one model wrote in the same session, off prompts of identical length.
Pricing Decisions, Backed by Data
Every suggestion links back to NICE, CKS, and primary literature, so you can see the reasoning, not just the recommendation.
The first names nothing. It's a price tracker, all it has is data, and the line would survive being moved to any product on earth that touches a number. The second names three things you could go and check by lunchtime.
This is the part I didn't expect. The model can clearly write the good version, and does it unprompted, about as often as it writes the bad one. It just shows no sign of knowing the difference.
What separated them was the domain. Clinical medicine has NICE and CKS sitting there waiting to be named. A price tracker's data is its own, there's no outside authority to point at, so the model reached for the shape of evidence instead. The empty version turns up precisely where nothing nameable exists to back the claim.
Inside an agent, provenance is load-bearing and nobody is being precious about it. An answer assembles itself out of a user message, an email, a retrieved PDF, a web search, another agent's summary and a tool call, and then the thing says "the policy expires Friday" and the only question worth asking is: according to what?
Without it you can't debug a wrong action, and you can't separate what the user said from what a web page said, which is most of what stands between you and prompt injection. It settles arguments too. Salesforce says 2 million, an email says 2.5, and provenance is how you work out which one to believe.
The root of it is that language models flatten types. A row in Postgres carries a type, a source table and an updated_at. The sentence "the policy limit is around 2 million" carries none of that. Provenance is an attempt to staple identity and lineage back onto tokens that lost both.
So engineers have spent three years writing about grounding and traceability, constantly and for good reason. That writing is a meaningful share of what a model has ever seen described as rigorous.
Instruction tuning spends enormous effort teaching a model not to claim what it can't support. Qualify. Ground. Cite. Flag uncertainty. Correct when someone asks a question, and most of why these things are usable at all.
The trouble is that it becomes a register the model can't drop when the genre changes. Reinhart and colleagues measured this in PNAS: instruction-tuned models write in a noun-heavy, informationally dense style, with GPT-4o using nominalisations at 2.1 times the human rate. Their conclusion is that instruction tuning narrows the range of styles a model can imitate at all.
Here's a guess, and I'm calling it a guess because the paper doesn't cover it. Provenance, traceability, verification, grounding, attribution: every one is a verb wearing a noun's clothes. My guess is that a model reaching for a nominalisation lands on these more often than chance, because they're the nominalisations that surround discussions of rigour. Reinhart shows density. He doesn't show a preference for credibility nouns.
There's a commercial loop underneath too. The founding paper on generative engine optimisation, out of Princeton and Georgia Tech, tested roughly 10,000 queries and found that adding statistics, quotations and citations measurably raised how often generative engines cited a page. The advice to stuff pages with visible sourcing is advice that works, and it produces more evidence-shaped pages for the next model to read.
All 30 generated pages contained at least one em dash. Thirty out of thirty, 58 across 3,903 words, which works out at 149 per 10,000.
Pew's analysis of the 2026 web puts em dashes at 11.19 per 10,000 words, roughly double their 2023 rate. Those two numbers measure different things, since Pew is general web text and mine is marketing copy, so I'm not printing a ratio. But 30 out of 30 is worth sitting with, from a model that was only asked to sell a car wash.
Humans have written this way forever. Clinically proven. Doctor recommended. Backed by science. Trusted by 10,000 teams. Category-empty credibility claims are the oldest move in advertising, and if that's all this is, then a model is faithfully reproducing an old register and none of the above is needed to explain it.
The thing that would settle it is whether the vocabulary changed. Trusted and proven are ancient. Grounded, traceable, provenance and auditable were engineering words five years ago, and if those crossed into consumer-facing copy after 2022, that's a shift rather than a habit.
I tried to test it against archived analytics homepages from 2017 and 2021 and couldn't get a sample out of the Wayback Machine big enough to count. So I don't have that number and I'm not going to argue around the hole. What I can say is that the words are on shipped pages today. Glean, You.com and ThoughtSpot all use "grounded". Elicit uses "traceable" and "auditable". Klue says "trusted source". That's six of the twelve AI product homepages I could read as plain HTML. I ran the same check on ten consumer sites and only three returned readable markup, which is too thin to compare against.
That's an observation. The measurement needs the Wayback Machine back.
Given the subject, this is worth saying out loud. The post started with six studies and ships with two.
The four I cut were about biomedical abstracts, news writing across 34 languages, and consumer reactions to disclosed AI authorship. All real, all well conducted, none of them about marketing copy. They would have sat here looking like evidence while measuring something adjacent, which is the exact move I'm complaining about. Two more turned out not to exist at all, which is its own lesson about where I got the idea.
A citation that supports a named claim is doing work. "Research-backed" in a headline names nothing. That distinction is the whole post, and it applies to me.
Provenance should be a feature, not a personality.
For insurance, legal and compliance buyers it genuinely is the product. "Every answer cites the policy paragraph" is why someone signs, and the test is whether that buyer has been burned yet. If they have, lead with it.
Everyone else should cut it, and not by making it more specific. "Backed by your Salesforce data" is "backed by data" with more words. The shape is the problem, so filling the blank does not fix it.
Look again at what the model wrote when it got this right. Not "backed by NICE and CKS". It wrote every suggestion links back to NICE, CKS, and primary literature. That is a verb. It describes something the reader can do, and there is no slot left to stuff a noun into.
That is the whole rule. Write what the reader does, not what you have. "Click any number and see the paragraph it came from" is a sentence about them. "Source-backed" is a sentence about you.
The people who lose here are the engineers who started it.
Provenance was a precise word. It meant this value came from that source at that time and here is the chain. Spend it on enough homepages with no chain to show and it stops carrying the meaning. Then a system that genuinely does track every claim back to a policy paragraph will have no way left to say so, because the sentence that says it will sound like everyone else's.
In a few years nobody will be able to explain why their homepage says grounded in trusted sources. It will just be what homepages say.

Everything you need to know about how LLMs break text into tokens - and why it explains most of their weird behaviors.
AI
Why the same model name can behave differently across providers, covering routing layers, quantization differences, fallbacks, and hidden wrappers with a practical debugging checklist.
AI
Understanding how bit-rates and quantization shape LLM deployment, from precision trade-offs to practical quantization methods like GPTQ, AWQ, and SmoothQuant.
AI