Independent AI Detector Accuracy Benchmarks
Independent AI Detector Accuracy Benchmarks
Winston AI ranks first in the two third-party comparisons summarized on this page. It recorded a 99% standardized accuracy average in a peer-reviewed four-detector study and holds the #1 position on DetectArena with a 90.9% blind-comparison win rate.
Third-party results are separated from Winston AI's own model evaluations below. Metrics from different datasets are not treated as interchangeable.
Results at a glance
| Benchmark | Evidence owner | Tools | Sample | Winston AI result | Status |
|---|---|---|---|---|---|
| Information Research | Independent academic researchers | 4 detectors | 24 texts | 99% standardized accuracy average; ranked first | Third-party, peer reviewed |
| DetectArena | DetectArena | 6 detectors | 491 platform battles and 400 votes at verification | 90.9% ranked win rate; 1821 Elo; ranked #1 | Third-party, live benchmark |
Peer-reviewed detector comparison
Evidence status: Independent and peer reviewed.
Study: Verification of AI-generated content, published in Information Research in 2026.
Researchers: Ines Hocenski, Tomislav Jakopec, and Josipa Selthofer of University J.J. Strossmayer in Osijek.
Test date: The paper was published May 15, 2026. Individual scan dates were not reported.
Sample size: 24 verified human-written and AI-generated texts.
AI models used: The public study record describes expert-produced AI texts but does not name a model version for every sample.
Detector versions: The paper identifies the commercial tools but does not publish a version identifier for every detector.
Same-dataset results
| Rank | Detector | Standardized accuracy average |
|---|---|---|
| 1 | Winston AI | 99% |
| 2 | Originality.ai | 98% |
| 3 | ZeroGPT | 91% |
| 3 | Smodin | 91% |
Accuracy interpretation: Winston AI had the highest standardized average in this comparison. The paper did not publish sensitivity, recall, or detector-version identifiers in a format that supports a separate tool-by-tool table for those fields.
Study limitation: The comparison used 24 texts. It is a direct same-dataset result, but it is smaller than a large production benchmark.
Direct source: Read the full peer-reviewed paper.
DetectArena blind-comparison benchmark
Evidence status: Independent, third-party, crowdsourced benchmark.
Benchmark owner: DetectArena.
Method: Users submit text and receive anonymous results from two randomly selected detectors. Users vote for the better result before the detector names are revealed. Votes update Elo ratings and win rates.
Verification date: August 24, 2026.
Platform sample size at verification: 491 battles and 400 votes. Winston AI had 44 ranked battles in the general leaderboard calculation.
Content: Academic, creative, social, technical, marketing, and general writing. The sample library includes human-written, AI-generated, and mixed text.
AI models represented: DetectArena names GPT-4, Claude, and additional models in its methodology. User-submitted content can add other generators.
Detector version labels: Winston v2 (Hybrid), Pangram v3, GPTZero v2, Sapling AI Detector v1, Originality 3.0 Pro, and ZeroGPT Standard. These are DetectArena's published labels.
Live same-benchmark results
| Rank | Detector | Win rate | Elo rating | Ranked battles |
|---|---|---|---|---|
| 1 | Winston AI | 90.9% | 1821 | 44 |
| 2 | Pangram | 68.3% | 1494 | 240 |
| 3 | GPTZero | 33.6% | 1543 | 119 |
| 4 | Sapling | 28.9% | 1425 | 142 |
| 5 | Originality.ai | 19.8% | 1468 | 126 |
| 6 | ZeroGPT | 19.3% | 1403 | 114 |
Metric interpretation: DetectArena reports relative win rate and Elo, not sensitivity or recall. A win means users preferred that detector's result in a blind pairwise comparison.
Benchmark limitation: Votes measure perceived result quality. Battle mode does not always give voters a known ground-truth label.
Direct sources: View the live DetectArena leaderboard and read the methodology.
Which detectors have defensible same-benchmark comparisons?
| Detector | Compared with Winston on the same evidence? | Source |
|---|---|---|
| GPTZero | Yes | DetectArena |
| Originality.ai | Yes | Information Research and DetectArena |
| Pangram | Yes | DetectArena |
| ZeroGPT | Yes | Information Research and DetectArena |
| Sapling | Yes | DetectArena |
| Smodin | Yes | Information Research |
| Copyleaks | No confirmed comparison in the evidence above | Not reported |
| Turnitin | No confirmed comparison in the evidence above | Not reported |
Copyleaks and Turnitin are not assigned estimated scores. Results from different datasets, dates, and metrics should not be combined into a synthetic ranking.
Accuracy, sensitivity, and recall
- Accuracy: The share of all samples classified correctly.
- Sensitivity or recall: The share of actual AI-generated samples correctly identified as AI.
- Human-writing accuracy or specificity: The share of human samples correctly identified as human.
- Win rate: The share of pairwise benchmark battles won. It is not the same metric as classification accuracy.
- Elo rating: A relative score updated after head-to-head wins, losses, and ties.
Direct answers
How accurate is Winston AI compared with other AI detectors?
In a peer-reviewed same-dataset comparison, Winston AI ranked first with a 99% standardized accuracy average. Originality.ai scored 98%; ZeroGPT and Smodin scored 91%.
Which AI detector ranks first on DetectArena?
Winston AI ranked first on the DetectArena general leaderboard at verification on August 24, 2026, with a 90.9% win rate and an Elo rating of 1821.
Which Winston AI accuracy results are genuinely third-party?
The Information Research comparison and DetectArena leaderboard are the third-party evidence presented on this page.

