For two years the argument about machine written content has been conducted entirely in adjectives. Flooded. Drowning. Overrun.
Pew Research supplied a number instead. Roughly 10% of webpages show signs of AI authorship or editing, based on a sample of about 500,000 pages measured this August.
Among pages published after ChatGPT launched, the figure is 35%.
The detector is counting punctuation
Here is the part that should slow you down. The markers this kind of analysis relies on are stylistic residue, not substance.
Em dash frequency doubled, from 5.79 to 11.19 per 10,000 words. Oxford comma use rose 63% since early 2023.
I have a standing rule against the em dash in everything I publish. That preference now reads as a machine avoidance tactic rather than a house style, which is the weakness of the method in one sentence.
A classifier counting punctuation is measuring a model's default settings. It is not measuring whether the page is accurate, useful, or worth anyone's time.
Which is why the estimates scatter so badly. Graphite put AI generated articles at 49.9% of new publications in Q1 2026. Pew says 10% of the web overall.
Both are honest pieces of work. They count different things, on different samples, at different thresholds. The distance between 10% and 50% is not a disagreement about reality, it is a disagreement about where you draw the line.
Authored and edited are not the same activity
One detail in the framing deserves more attention than it got. The category is signs of AI authorship or editing.
Those are different activities with very different implications, collapsed into one bucket.
A researcher who writes 2,000 words from their own fieldwork and then runs it through a model to tighten the prose lands in the same category as a page generated wholesale from a keyword list.
The first is a person using a tool. The second is a machine filling a slot. A punctuation based classifier cannot separate them, and neither can anyone quoting the 10% figure from a conference stage.
This matters commercially, because most serious teams now sit in the first category. Nearly every marketing department I work with runs drafts through a model at some stage of production.
If your internal policy treats that as contamination, you have written a rule against your own productivity while your competitors quietly ignore theirs. I wrote about the arriving cost of this in the AI content bill coming due, and the bill is being itemised now.
The split that actually tells you something
Ignore the headline percentage for a moment and look at the domain breakdown.
Commercial .com domains sit around 9.35%. Educational and government domains sit under 2%.
That is a four to five times difference, and it has nothing to do with technical capability. Universities have access to the same models as everyone else.
The difference is incentive. Machine writing concentrates exactly where publishing volume converts into money, and stays away from where it converts into liability.
If you operate in a commercial category, that ratio is your operational reality. Your competitors are producing at machine speed, and they started before you noticed.
Volume stopped working as a differentiator somewhere around the moment the second company in your category bought a seat. The Pew split is simply the receipt for that.
Why a measurable thing becomes a ranked thing
Google's public position has been consistent and reasonable. Quality matters, authorship does not. Helpful, reliable, high quality content wins regardless of what produced it.
I believe they mean it. I also think that position is stable only while authorship stays expensive to detect.
Pew just demonstrated that a research team can classify half a million pages using linguistic markers. That is not an expensive operation. That is a weekend of compute and a decent classifier.
Anything that can be counted cheaply and repeatedly ends up as an input somewhere. Not necessarily as a penalty. Possibly as a weighting, a confidence adjustment, or a tiebreaker between two pages answering the same question.
Building your content architecture on the assumption that authorship stays invisible is building on a temporary condition. We already saw one version of this when a third of the internet turned out to be machine generated and nothing visible happened for a year.
The only position that holds
The wrong response is to optimise for undetectability. Strip the em dashes, vary the sentence length, hand edit until the classifier shrugs.
You lose that race. Detectors update faster than your style guide, and you will have spent the budget on camouflage instead of substance.
The position that holds is built on inputs a model cannot produce from a prompt.
Proprietary data from your own accounts. A client outcome with a real number attached. An opinion with a date on it, published before the result was known. A measurement you took yourself.
None of that registers as machine writing, because none of it is available to the machine.
A recent piece on Search Engine Journal covering the Pew findings framed this as an authorship question. From where I sit, it is a sourcing question.
The model can write. It cannot know what happened inside your business last quarter.
The web just got its first proper census. The interesting part is not how many pages the machines wrote. It is that we can now count, that we will keep counting, and that every content strategy is now being run in public.