By ShriprasannaPublished September 22, 2026Updated September 27, 2026
How We Test AI Detectors (and What We Have Measured So Far)
We have run one small, dated check so far. On 26 and 27 September 2026 we scored AI-written drafts, their rewrites, and pre-2020 human-written Wikipedia passages with our own in-app detector, mostly in Tagalog and Malay, plus a human-text check in Arabic. The results are below, including where the detector failed outright: it scored every human-written Malay passage as 100% AI.
This page does three things. It lists what a rigorous detector test needs, explains what our in-app AI score actually is, and reports what we have measured, with its limits. It is not a leaderboard, and it does not add new 98% or 99% bypass rates. We have not tested against Turnitin, GPTZero, or any other live third-party detector.
If you only need a check on a draft, open the AI Detector. If you want the method behind a trustworthy comparison, keep reading.
What a rigorous detector test needs
A test that skeptical readers can reuse has to be boringly specific. At minimum:
- Named detectors, live or licensed. Call Turnitin, GPTZero, Originality.ai, Copyleaks, and any other system you name — as those products, not as lookalike labels. Record the product version or UI date.
- A frozen corpus. Same prompts, same models, same word counts, same languages, same mix of fully generated / mixed / human control text. Publish the count and the date the files were frozen.
- Blind scoring. Do not tune the humanizer against the same samples you then report. Hold out a set.
- One output per condition. Humanize once per sample (or declare retries). Re-running failed rows until they pass is not a pass rate.
- Per-detector results, not one blended percentage. "99% across detectors" hides a tool that fails Turnitin and clears a weak free checker.
- Decision rules in writing. What counts as a pass — under 20% AI? "Human" label only? Instructor-view Turnitin vs student view?
- Limitations. Short text, ESL writing, citations, code, and hybrid edits all change detector behavior. State what you did not test.
If a vendor cannot point to those pieces, their bypass percentage is not a result you can audit from this page.
What our in-app AI score is
Our in-app AI Detector makes a single upstream scoring pass, then a mapping layer fills twelve named labels.
The twelve labels in the product are: Turnitin, GPTZero, OpenAI, Claude, Gemini, Grok, Grammarly, Crossplag, Copyleaks, QuillBot, Sapling, and ZeroGPT.
Those labels are not twelve live third-party APIs. The app takes one overall AI score from the upstream scorer, converts it to a human score (100 - AI score), then assigns each label right, warn, or bad from a fixed bucket table:
| Human score | right | warn | bad |
|---|---|---|---|
| 100 | 12 | 0 | 0 |
| 80–99 | 10 | 1 | 1 |
| 70–79 | 10 | 0 | 2 |
| 60–69 | 8 | 3 | 1 |
| 50–59 | 8 | 2 | 2 |
| 30–49 | 5 | 5 | 2 |
| 20–29 | 3 | 3 | 6 |
| 0–19 | 0 | 0 | 12 |
Priority order gives "right" to Turnitin and GPTZero first. Sentence-level highlights come from the same upstream response, not from a second detector.
So the UI is a guide derived from one score, not a simultaneous poll of Turnitin's production classifier and GPTZero's production classifier. Wherever this site shows the twelve detector names, they describe that labeled breakdown, not twelve independent checks.
Paid web detection does not spend plan credits. Free users and API/MCP calls are charged per word. None of that is an accuracy claim.
What we have measured so far (September 2026)
How we ran it
- Dates: the AI-draft and rewrite runs and the Tagalog/Malay human-text check happened on 26 September 2026; the Arabic human-text check on 27 September 2026.
- Detector: our own in-app detector only, the single scoring model described above. Not Turnitin, not GPTZero, not any other third-party product.
- AI text: short AI-written drafts, about one paragraph each (roughly 85 to 125 words in Tagalog and Malay). Tagalog and Malay each got 5 runs on 4 drafts per mode: four distinct drafts, with the first run twice to check consistency. Every other language got one draft.
- Rewriting: each draft was rewritten with SupWriter's humanizer in default mode and scored again. The Tagalog, Malay and English drafts were also rewritten in aggressive mode, as were the single drafts in Persian, Malayalam, Tamil, Sinhala, Marathi, Greek and Arabic.
- Human text: passages from Tagalog and Malay Wikipedia, taken from article revisions dated 2018 and 2019 (pre-2020, so written before ChatGPT existed). Encyclopedic register, 80 to 200 words each: 9 passages in Tagalog and 8 in Malay, and, on 27 September, 8 passages from 2019 Arabic Wikipedia revisions (76 to 194 words each).
Tagalog and Malay: AI drafts before and after rewriting
Our detector's AI score, before → after, for 5 runs on 4 drafts in each mode:
| Language | Mode | First draft, run twice | The other three drafts | Summary |
|---|---|---|---|---|
| Tagalog | Default | 100 → 100 and 100 → 0 | 100 → 10.8, 35.9 → 0, 100 → 0 | 4 of 5 runs under 11% |
| Tagalog | Aggressive | 100 → 0 and 100 → 100 | 100 → 5.2, 35.9 → 0, 100 → 11 | 4 of 5 runs at or under 11% |
| Malay | Default | 100 → 100 and 100 → 0 | 100 → 70.5, 95.3 → 11.6, 100 → 7.8 | 3 of 5 runs under 12% |
| Malay | Aggressive | 100 → 12.8 and 100 → 88.5 | 100 → 100, 95.3 → 5.6, 100 → 100 | 2 of 5 runs at or under 12.8% |
What stood out:
- Run-to-run variance is large. The same draft, rewritten twice in the same mode, scored 0% after one run and 100% after the other. That happened in Tagalog default mode, Tagalog aggressive mode, and Malay default mode.
- Aggressive mode was not better for Tagalog or Malay.
- The detector misses AI text too. One AI-written Tagalog draft scored only 35.9% AI before any rewriting.
- The Malay rows are not evidence that the rewrites read as human. The human-text check below shows why.
English control and other languages
One English draft scored 100 → 4.2 in default mode and 100 → 7.2 in aggressive mode.
Other languages, one draft each. The same draft was run once in default mode and, for most languages, once in aggressive mode (aggressive mode reuses the same "before" score):
| Language | Default: before → after | Aggressive: after | Note |
|---|---|---|---|
| Persian | 56 → 67 | 55 | |
| Malayalam | 71 → 100 | 100 | |
| Tamil | 72 → 85 | 83 | |
| Sinhala | 71 → 68 | 71 | |
| Marathi | 74 → 100 | 79 | |
| Greek | 74 → 74 | 65 | |
| Arabic | 100 → 75 | 100 | |
| Urdu, Nepali, Mandarin, Amharic | 0 → 0 | not run | The detector returns 0 for both the AI draft and the rewrite, so it cannot evaluate these languages. |
For most non-English languages, our detector's scores are not a reliable judge. It scored fully AI-written drafts at only 56–74% AI in Persian, Malayalam, Tamil, Sinhala, Marathi and Greek, and in four of those languages the rewrite scored higher than the original draft. For Urdu, Nepali, Mandarin and Amharic it returned 0 for everything.
Human-written text: the false-positive check
This is the part that matters most if you are worried about being wrongly flagged. When a detector scores genuine human writing as AI, that is a false positive.
| Language | Passages | AI % for each passage | Summary |
|---|---|---|---|
| Tagalog | 9 | 0, 7.1, 0, 0, 2.1, 0, 0, 100, 23.3 | 7 of 9 at or under 7.1%; one at 23.3%; one ("Cebu") at 100%, a clear false positive |
| Malay | 8 | 100, 100, 100, 100, 100, 100, 100, 100 | All 8 at 100%: every one a false positive |
| Arabic (checked 27 Sept) | 8 | 73.4, 74.8, 100, 100, 100, 0, 0, 100 | 6 of 8 at 73% or higher; 4 at 100% |
Our detector should not be used to judge Malay or Arabic text. Like most detectors, it is unreliable for Malay: it scored human-written Malay text as 100% AI in our check, every time, and it scored most of the human-written Arabic passages at 73% AI or higher. That also means the Malay before-and-after numbers above are not evidence that a rewrite reads as human.
For Tagalog, the detector separated these samples reasonably: most AI drafts scored 100%, and most human passages scored under 8%. It still produced a false positive. In either language, a detector score should never be treated as proof either way, and that includes ours.
Limits of this check
- Tiny samples. 5 runs on 4 drafts per mode for Tagalog and for Malay, one draft for every other language, and 25 human passages in total (9 Tagalog, 8 Malay, 8 Arabic).
- One detector. Everything here comes from our own model. It says nothing about how Turnitin, GPTZero, or anyone else would score the same text. Turnitin's AI writing detection covers only long-form English, Spanish, Japanese and Modern Arabic (Turnitin's AI detection FAQ), so it would not produce an AI score for Tagalog or Malay at all.
- Formal register. The human passages are encyclopedic Wikipedia prose, and the AI drafts are short essay-, email- and reflection-style paragraphs. Student essays, casual writing and longer documents may behave differently.
- Short texts. Every sample was under the 300 words of prose Turnitin requires before it produces an AI report (same FAQ).
- Two days. The runs happened on 26 September 2026 and the Arabic human-text check on 27 September. Detector and rewriting models change, so treat these numbers as a snapshot.
What we have not tested
We do not currently publish:
- scheduled runs against live Turnitin, GPTZero, Originality.ai, or Copyleaks accounts
- a large, versioned sample set with pass/fail logs
- a published confusion matrix or a large false-positive study (the human-text check above covers 25 passages in three languages)
- a numbered bypass rate generated from those live tools
SupWriter has not published live-detector bypass tests. Some older pages on this site quoted bypass percentages (including 99%+). None of those figures came from a documented test like the one on this page, so treat any you find as unverified.
How you can run a test yourself
If you need a number you will defend to a professor, editor, or client:
- Keep a folder of drafts you actually submit — not one cherry-picked paragraph.
- Run each original through the detector you will face (the school LMS, the publisher's tool), not only our labels.
- Humanize with the same settings you would use for real work.
- Re-check on the same third-party detector. Save screenshots with timestamps.
- Record failures. An honest rate is passes ÷ attempts, including retries you would have to pay for.
- Run the same text twice. Our own check above shows that one run can mislead.
Use our detector as a pre-check: AI Detector. Then confirm on the system that actually grades or publishes the work.
How to read vendor charts
When a humanizer site shows a 98% or 99%+ badge, ask:
- Which detector products, on which date?
- How many samples, in which languages and genres?
- Was the humanizer allowed multiple passes?
- Is the chart the vendor's own UI labels, or the third-party report?
- Did they test human-written text too, so you can see false positives?
Our in-app twelve-label view answers only the first kind of chart: our mapping of one score. It is useful for spotting obviously AI-heavy passages in English. It is not a substitute for the detector your institution licensed, and the check above shows it is not reliable in every language.
Related reading
- AI Detector — run the in-app check
- How does AI detection work? — perplexity, burstiness, and classifiers
- AI detection false positives — what published research says about human writing flagged as AI
- Pricing — plan limits if you are detecting in volume
- Sign up — 300 words free, no credit card required
This page covers what a test must include, what the current product actually computes, and what we have measured so far.
Related Articles

What SupWriter's AI Detector Actually Checks (vs Turnitin and Originality.ai)

How SupWriter's AI Detector Score Is Produced

Turnitin AI Detection Update (Aug 2026): Why Purple Is Gone

