Updated September 6, 2026 · 4 minute read

AI writing statistics, 2026 edition.

Numbers on AI writing, rework, trust, and detection.

By Manav Mishra

We cannot put a credible global percentage on bad AI writing. Research can answer narrower questions about the time spent repairing it, the effect on colleagues, and changes in published language. Those findings become useful only when the population and method stay attached to the number.

Weak AI work shifts the burden

In a September 2025 survey by BetterUp Labs and Stanford’s Social Media Lab, 40% of 1,150 full-time US desk workers said they had received “workslop” in the previous month. The term describes AI-generated work that looks complete but leaves the recipient without enough information to move the task forward.

Recipients recalled losing an average of 1 hour 56 minutes per incident, according to BetterUp’s research summary. That is roughly two hours spent repairing work that appeared ready to use. The study measured respondents’ recollections rather than timing the repairs.

In the same survey, 42% of recipients saw the sender as less trustworthy. Fifty percent saw the sender as less capable, and 32% said they would not want to work with that person again. The survey records those reactions; it does not establish how long they last.

Count the rework in time savings

In Workday’s January 2026 report, 85% of respondents said they saved one to seven hours a week. Hanover Research surveyed 3,200 active AI users at large employers in November 2025, across workplace tasks beyond writing.

Workday estimated that rework absorbed nearly 40% of those reported savings. As an illustration, five hours saved with 40% spent on rework would leave three hours. That calculation is not a prediction of what any individual worker saves.

Fourteen percent consistently achieved what Workday called clear positive net outcomes. That finding does not establish that the other 86% broke even or lost time. The report distinguishes consistent gains from the broader range of experiences.

Gallup asked workers how AI affected their productivity and efficiency. Sixty-five percent of US employees at organizations using AI reported a positive effect in Gallup’s May 2026 data. This is a probability-based survey of workers’ own assessments.

METR’s randomized experiment assigned AI access for 246 tasks completed by 16 experienced developers in familiar repositories. They expected a 24% speedup; with early-2025 tools, completion took 19% longer than without AI. The measured result applies to that coding setting and generation of tools.

A 2026 follow-up reported selection problems: developers were less willing to submit work that might have to be done without AI. That limited the comparison.

A later survey of 349 technical workers found self-reported gains, but used a convenience sample.

In Stack Overflow’s 2025 survey, 66% of those answering the AI-frustrations question selected nearly correct answers as a frustration, and 45% selected extra debugging effort. The question drew 31,476 answers and allowed multiple choices.

AI is changing published prose

At least 13.5% of 2024 biomedical abstracts were estimated to have received LLM assistance in Kobak and colleagues’ 2025 Science Advances paper. The study analyzed more than 15 million abstracts and inferred assistance from shifts in word frequency.

An August 2026 preprint by Holzwarth, González-Márquez, and Kobak estimates that 89% of open-access biomedical papers showed excess LLM-associated vocabulary by late 2025. The preprint examines full text using a different method, so the two percentages are not directly comparable. Neither measures writing quality.

The 2024 RAID benchmark contains over six million generations. Its evaluation of 12 detectors found weaknesses under attacks, unfamiliar models, and different generation settings. Those tests concern authorship detection; they do not tell us whether a passage needs editing.

What Zero Slop has measured

Our editing test used our AI Slop test corpus of drafts with deliberately obvious flaws. Their average score was 76.3 out of 100. Because the sample was designed to be bad, it is not representative of AI writing. Twenty-five is our review threshold, not a measured point where readers begin to notice slop.

After GPT-5.4 edited those drafts using Zero Slop, the mean was 12.8. Every rewrite passed the automated source-detail check. That check protects names, numbers, quotes, and links; it cannot prove every claim or catch every shift in meaning. See the full comparison.

In our RAID+ audit, 24.9% of 7,627 nonempty abstracts scored at or above 25. That is the share that triggered our threshold; the dataset’s labels record machine origin and cannot certify writing quality.

Readers are pushing back

Over one million people clicked “seems like AI slop” in the feature’s first two weeks, according to LinkedIn’s Chief Product Officer. LinkedIn also reported an aggregate 40% reduction in views of content it classified as slop compared with a few weeks earlier. That does not establish a 40% penalty for an individual flagged post.

Hari Srinivasan also reported that LinkedIn was catching hundreds of thousands of automated comment attempts daily. That figure counts attempted abuse, not the proportion of ordinary posts that are poorly written.

Before reusing a statistic, check when the data was collected, who was measured, and how. If one of ours is wrong, send the source through the feedback form.

← All posts