Updated September 6, 2026 · 4 minute read

Anatomy of the slop score.

What changes the number, and what the number can tell you.

By Manav Mishra

An 80 on Zero Slop’s scale means a draft contains more of the writing patterns it tracks than a draft scoring 20. It does not mean there is an 80% probability that AI wrote it. Confusing those two meanings turns an editing aid into an unsupported accusation.

The scoring code is public. With the same version, settings, and rules, the same text produces the same number. Whether a flagged passage should change depends on what it says and who it is for.

What raises the score

The scorer looks for 294 weighted patterns and 96 watchlist terms. It also checks 26 words that count only when nearby language makes them suspicious. These include stock phrases and vocabulary associated with generic AI prose. Length adjustments and repeated-pattern rules prevent document length alone from deciding the result.

Sentence variety, repetition, readability, and formatting also contribute. Instructions can benefit from consistent rhythm, and technical papers need technical terms. The formal-writing setting disables the rhythm-uniformity and formality penalties for those uses.

With no phrase or vocabulary hits, and no emoji or hashtags, the remaining checks can push the score only as high as 20.2. That ceiling keeps a formal register or ordinary punctuation from triggering the default review threshold on its own.

Why the score moves quickly in the middle

The code combines its measurements into a total called E. It then applies this curve and rounds to one decimal place:

score = 100 / (1 + exp(-(E - 9) / 4))

A total of zero produces 9.5; a total of nine produces 50. The score moves fastest around the middle and more slowly near either end. The curve controls how the number responds to additional evidence; it does not measure how readers react.

25: editorial thresholdevidence 9 = score 50curve is steepest hereminimum: 9.5score100924evidence
Calculated from the scoring equation. The threshold is a product choice; the curve is not fitted to reader judgments.

We use 25 as the review threshold. It is a product choice, with no evidence that readers universally start noticing bad writing at that number. Our human reference passages help catch false alarms without defining a ceiling for human prose.

What the public tests show

Beemo contains model answers, expert edits, and independent human answers to the same prompts. In our audit of 2,187 sets, model output averaged 30.2, expert edits 25.3, and independent human answers 20.0. These are group averages; individual edits can move in either direction.

Expert editing lowered the score in 52.2% of pairs; the remaining scores stayed unchanged or rose. An editor may need to explain more or keep a phrase that the scorer flags. The average improvement does not tell you whether any particular edit is good.

30.225.320.0raw model outputexpert-editedhuman answers
Group means for 2,187 Beemo sets. Beemo records authorship and editing history; it does not label slop quality.

Our RAID+ audit scored 7,627 abstracts from four model families in formal-writing mode. The share at or above 25 ranged from 10.1% for DeepSeek V3 to 41.7% for Llama 3.3 70B. Those percentages tell us how often our threshold triggered; RAID+ labels identify model origin and cannot establish a ranking of writing quality.

A separate editorial review used two LLM raters to assess 72 passages from 12 drafts without seeing which method produced them. The raters agreed 77.8% of the time, but only 38 passages received a shared, definite judgment. Borderline cases and disagreements stayed unresolved, and no human validation is implied.

The editing replay used GPT-5.4 to apply four sets of editing instructions to our AI Slop test corpus, which contains deliberately obvious flaws. The table reports performance against Zero Slop’s own score and release checks. It does not establish which editor is best on ordinary drafts.

Instructions usedMean scoreDrafts passing release checks
Original drafts76.30%
Zero Slop12.8100%
avoid-ai-writing23.383.3%
no-ai-slop28.466.7%
humanizer35.450%

Use the flags to find passages worth reviewing. Keep a phrase when it says what you mean, even if it costs a point. Then compare the edit with your source and read it aloud before accepting it.

← All posts