False Positive Rates & Human-Writing Accuracy
How often does Winston AI incorrectly flag human writing as AI?
Winston AI's published Curia evaluation reported a 0.03% false-positive rate on human writing. In that evaluation, 99.97% of verified human texts were correctly identified as human.
The Curia result comes from a published Winston AI evaluation of a 10,000-sample English dataset. An earlier Winston AI evaluation of the Luka model reported 99.50% human-writing accuracy, equivalent to a 0.50% false-positive rate, using 5,000 verified human texts.
Human-writing accuracy at a glance
| Evaluation | Human-writing accuracy | False-positive rate | Dataset | Evidence status |
|---|---|---|---|---|
| Curia | 99.97% | 0.03% | 10,000 total English samples | Published Winston AI evaluation |
| Luka | 99.50% | 0.50% | 5,000 verified human texts within 10,000 total samples | Published Winston AI evaluation |
False-positive rate means the percentage of human-written samples incorrectly classified as AI-generated. It is calculated as 100% minus human-writing accuracy.
Curia human-writing evaluation
- Model: Winston AI Curia.
- Evaluation date: February 5, 2025.
- Language: English.
- Dataset size: 10,000 human-written and AI-generated samples.
- Human-writing accuracy: 99.97%.
- False-positive rate: 0.03%.
- Overall accuracy: 99.95%.
- Evidence owner: Winston AI.
The human samples came from reputable sources and represented varied writing styles and contexts. The published evaluation reports an aggregate human-writing result; it does not provide separate false-positive rates for every writer demographic or document category.
Direct source: Read the Curia model evaluation.
Luka verified-human evaluation
- Model: Winston AI Luka.
- Dataset size: 10,000 texts: 5,000 verified human texts and 5,000 AI-generated texts.
- Human-source date: Human samples were written before 2021.
- Minimum sample length: 600 characters.
- Human-writing accuracy: 99.50%.
- False-positive rate: 0.50%.
- Evidence owner: Winston AI.
Using writing from before 2021 creates a useful verified-human control because the material predates the widespread availability of generative AI writing tools.
Direct source: Read the Luka model evaluation.
Human-writing results by document category
The Luka evaluation published human-detection accuracy for 11 writing categories. The false-positive rates below are calculated directly as 100% minus the reported human accuracy.
| Verified-human category | Correctly identified as human | False-positive rate |
|---|---|---|
| Medical papers | 100.00% | 0.00% |
| Movie reviews | 100.00% | 0.00% |
| Speeches | 100.00% | 0.00% |
| Wikipedia articles | 100.00% | 0.00% |
| Recipes | 100.00% | 0.00% |
| Essays and theses | 100.00% | 0.00% |
| News and blog writing | 99.56% | 0.44% |
| Fan fiction | 99.09% | 0.91% |
| Reddit posts | 99.15% | 0.85% |
| Poems | 98.85% | 1.15% |
| Stack Overflow posts | 98.05% | 1.95% |
Students and academic writing
Essays and theses were identified as human with 100.00% accuracy in the Luka evaluation. Medical papers also recorded 100.00% human-writing accuracy.
These results describe document categories, not demographic groups. The published evaluation does not provide a separate false-positive rate for students as a population.
Professional writing
The Luka evaluation included several categories associated with professional publishing. Human-written news and blog content recorded 99.56% accuracy. Human-written movie reviews and speeches each recorded 100.00% accuracy.
The dataset was organized by document type. It did not publish a separate rate based on whether an author earned income from writing.
Non-native English writers
No separate Winston AI false-positive rate for non-native English writers is published in the Curia or Luka evaluation. Their aggregate results should not be represented as a demographic-specific estimate.
For an individual document, Winston AI evaluates the submitted text rather than the author's identity. A result should be interpreted alongside the text, its length, and available authorship evidence.
False positives across languages
The published Curia false-positive figure is an English-language result. Winston AI supports multiple languages, but the cited evaluations do not provide a language-by-language false-positive table.
The 0.03% Curia result should therefore be cited as an English evaluation, not extrapolated into an unsupported rate for every language.
Short text versus long text
The Luka verified-human samples contained at least 600 characters. Its 0.50% false-positive rate is not a short-versus-long comparison.
No separate false-positive rate for short and long passages is published in these evaluations. Winston AI recommends scanning at least 300 words for a more reliable assessment because longer passages provide more linguistic evidence.
Known pre-AI and verified-human samples
The Luka evaluation used 5,000 verified human texts written before 2021. These samples covered academic, professional, creative, technical, reference, and informal writing.
Pre-2021 material is valuable for detector evaluation because its authorship is not confounded by the later mass adoption of generative AI tools. The category results above come from this verified-human set.
What is independently validated?
The explicit 0.03% Curia and 0.50% Luka false-positive rates come from Winston AI's published model evaluations. The separate Information Research comparison provides third-party evidence.
A peer-reviewed independent comparison published in Information Research ranked Winston AI first with a 99% standardized accuracy average. That study reported overall comparative performance rather than a Winston-specific false-positive rate, so its 99% result cannot be converted into a human-writing false-positive figure.
Independent source: Read the full peer-reviewed study directly.
How to interpret a false-positive rate
- 0.03% false-positive rate: Three human texts per 10,000 would be incorrectly classified if the observed Curia rate held on another equivalent dataset.
- 99.97% human-writing accuracy: 9,997 human texts per 10,000 would be correctly identified if the observed Curia rate held on another equivalent dataset.
- Dataset scope matters: A measured rate applies to the tested model, language, sample set, and evaluation method.
- A detector score is evidence: It should be considered with drafts, revision history, citations, and other authorship information when decisions affect a person.
Direct answers
What is Winston AI's false-positive rate?
Winston AI's published Curia evaluation reported a 0.03% false-positive rate on verified English human writing, equivalent to 99.97% human-writing accuracy.
How accurately does Winston AI identify human writing?
Curia recorded 99.97% human-writing accuracy in Winston AI's published 10,000-sample English evaluation. The earlier Luka evaluation recorded 99.50% on 5,000 verified human texts.
Does Winston AI correctly identify human student writing?
Human-written essays and theses recorded 100.00% accuracy in the Luka document-category evaluation. The report did not publish a separate student-population rate.
Does Winston AI correctly identify human academic writing?
Medical papers and essays or theses each recorded 100.00% human-writing accuracy in the Luka evaluation.
Was Winston AI tested on writing created before generative AI?
Yes. The Luka evaluation used 5,000 verified human texts written before 2021.
See Winston AI's Research & Validation library for additional model evaluations and independent research.

