Free account300 free words·no cardGet started
Now on iPhone & iPadApp Store ↗
Your Text
Style
Type
Lang
60 / 500 words
AI Humanization
11 min read

By ShriprasannaPublished March 17, 2026Updated September 27, 2026

AI Humanization for Non-English Content: What Works

Here's a problem nobody talks about enough: almost every AI humanization tool was built for English first. The marketing pages say "multilingual support," but a tool that accepts Spanish, Arabic or Chinese text isn't necessarily good at rewriting it. AI detectors have a parallel problem. Many support only a handful of languages, and some return no score at all for text outside that list.

This matters because AI-assisted writing isn't an English-only phenomenon. Students in Madrid draft essays with ChatGPT. Marketers in Dubai generate Arabic ad copy with Claude. Researchers in Beijing use DeepSeek to draft papers in Mandarin. The demand for good rewriting in other languages is real, and the tooling on both sides is uneven.

This guide covers what the major detectors say they support, what we saw when we ran non-English drafts through SupWriter's own detector (including where it fails), and how to check a non-English draft yourself. We have not run a multi-tool, multi-detector benchmark across languages, so you won't find competitor bypass percentages here.

The English Bias Problem

The bias starts with training data. The large language models behind today's chatbots are trained mostly on English text, with every other language sharing a smaller slice. GPT, Claude, Gemini and DeepSeek all reflect that imbalance to some degree.

This matters for two reasons:

AI text in non-English languages has different statistical signatures. When a model writes French, it draws on far less French data than it had English data for English. The result is often fluent but subtly different from native writing: not wrong, exactly, but statistically distinguishable. Perplexity profiles, vocabulary distributions and sentence-structure patterns all shift.

AI detectors cover fewer languages than people assume. Turnitin's AI writing detection works only on long-form English, Spanish, Japanese and Modern Arabic; for other languages the indicator shows an empty state instead of a score, and its detection of AI-paraphrased or "bypassed" text is English only (Turnitin AI detection FAQ). GPTZero says it "fully supports" English, German, Portuguese, French and Spanish (GPTZero). Originality.ai's multi-language detector lists 31 languages (Originality.ai help center). Copyleaks claims "over 30 languages" but publishes accuracy figures for only a handful of them (Copyleaks).

There's also a well-documented fairness problem on the English side. In a 2023 Stanford study published in Patterns, seven GPT detectors misclassified more than half of 91 human-written TOEFL essays as AI-generated, with an average false-positive rate of 61.22%, and 89 of the 91 essays (97.8%) were flagged by at least one detector (Liang et al., 2023). If detectors stumble over non-native English, there's little reason to assume they're well calibrated for languages they were barely trained on.

The net effect is a landscape where non-English writers face patchy detection in both directions, and most humanization tools weren't built with their language in mind either.

How AI Detection Works Differently in Non-English Languages

Understanding the mechanics helps explain why detector scores in other languages deserve extra skepticism.

Training Data Gaps

AI detectors learn what "human writing" looks like by analyzing large collections of verified human text. For English, those collections are enormous: academic papers, books, journalism, social media and student submissions. For most other languages far less labeled data exists, and even less of it is labeled specifically for AI detection.

Smaller training sets mean the detector's model of "what human writing looks like" in that language is less nuanced. It has fewer examples of natural variation, which makes it harder to separate "text that looks unusual because a machine wrote it" from "text that looks unusual because the writer has an uncommon style."

Tokenization Differences

Languages that don't use the Latin alphabet present fundamental challenges for detection tools built on English-centric architectures.

Arabic is written right-to-left, uses connected script where letter forms change based on position, and has a rich morphological system where a single root can generate dozens of related words through patterns of vowels and affixes. A tokenizer designed around English can fragment Arabic words in ways that distort the statistical patterns a detector relies on.

Chinese doesn't use spaces between words. Word segmentation, deciding where one word ends and another begins, is itself a non-trivial NLP task. Different segmentation approaches produce different token sequences from the same text, so the patterns a detector analyzes can vary with preprocessing.

Languages with complex morphology present similar challenges: German with its compound words, Turkish with its agglutinative structure, Finnish with its extensive case system. When the tokenizer wasn't designed for the language, the statistical analysis works with a distorted representation of the text.

Stylistic Norms

Every language has its own conventions for academic and professional writing. German academic prose tends toward longer, more complex sentences than English. Arabic formal writing uses more elaborate rhetorical structures. Japanese academic text follows conventions around indirectness and hedging that differ significantly from English norms.

Detectors trained mainly on English norms may read these language-specific conventions as "suspicious" because they deviate from the patterns the detector associates with human writing. That's a structural source of false positives for non-English content.

What We Actually Measured

On 26 September 2026 we ran a small set of AI-written drafts through SupWriter's humanizer and scored the text before and after with SupWriter's own detector. That detector is not Turnitin or GPTZero, and the samples are small, so treat this as a field note rather than a benchmark. Scores below are the detector's AI percentage, before → after rewriting.

Tagalog: 5 runs per mode on 4 drafts (the first draft was run twice)

  • Default mode: 100→100 and 100→0 (same draft, two runs), 100→10.8, 35.9→0, 100→0. That's 4 of 5 runs under 11%.
  • Aggressive mode: 100→0 and 100→100 (same draft, two runs), 100→5.2, 35.9→0, 100→11. Also 4 of 5 runs at or under 11%.

Malay: 5 runs per mode on 4 drafts (the first draft was run twice)

  • Default mode: 100→100 and 100→0 (same draft), 100→70.5, 95.3→11.6, 100→7.8.
  • Aggressive mode: 100→12.8 and 100→88.5 (same draft), 100→100, 95.3→5.6, 100→100.
  • Read these Malay numbers as noise, not as results. When we ran eight human-written Malay passages through the same detector (text from Malay Wikipedia revisions dated before 2020, so written before ChatGPT existed), it scored all eight as 100% AI. A detector that calls human Malay 100% AI can't tell you whether a Malay rewrite "reads human."

English control (one draft): default mode 100→4.2, aggressive mode 100→7.2.

Three things stood out. First, variance: the same draft scored 0% in one run and 100% in another. Second, aggressive mode was not better than default for Tagalog or Malay. Third, the detector behaved very differently by language: on nine human-written Tagalog passages (also pre-2020 Wikipedia text) it scored seven at 7.1% or lower, but it still flagged one at 100%, and for Malay it failed completely, as noted above.

We also ran one AI-written draft in each of several other languages, in default mode:

LanguageBefore → afterNote
Persian56 → 67
Malayalam71 → 100
Tamil72 → 85
Sinhala71 → 68
Marathi74 → 100
Greek74 → 74
Arabic100 → 75
Urdu, Nepali, Mandarin, Amharic0 → 0The detector returns 0 for both the AI draft and the rewrite, so it cannot evaluate these languages.

The honest interpretation: for most non-English languages, our detector's scores are not a reliable judge. It scored fully AI-written drafts at only 56–74% for several languages, rated some rewrites as more AI-like than the originals, and returned 0 for four languages. A low score in those languages tells you very little in either direction, which is why the checks later in this guide lean on native-speaker reading rather than any single number. (How we test)

Results by Language

Here's what the major detectors document for five widely used languages, plus the rewriting problems to watch for in each. Coverage is as stated on each vendor's page (linked above), checked in September 2026.

LanguageTurnitin AI detectionGPTZero "fully supports"Originality.ai multi-languageCopyleaks publishes accuracy
SpanishYesYesYesYes
FrenchNoYesYesYes
GermanNoYesYesYes
ArabicYes (Modern Arabic)Not listedYes (Modern Standard Arabic)Not listed
ChineseNoNot listedYes (Simplified and Traditional)Not listed

"Not listed" means the language isn't in that vendor's stated list, not necessarily that the tool refuses the text. It does mean you shouldn't assume the vendor's headline accuracy applies.

Spanish

Spanish is the best-covered non-English language on the detector side: every detector in the table lists it, including Turnitin. That cuts both ways. If your institution runs Turnitin, a Spanish essay gets a real AI score, so you should hold it to the same standard as English.

For rewriting, meaning preservation is the hidden metric. Aggressive paraphrasing can change a detector score while quietly distorting what the text says. Check the rewrite against the original claim by claim before you check any score.

French

Turnitin's AI indicator doesn't cover French, so a French submission gets no AI score there; GPTZero, Originality.ai and Copyleaks all list it.

French grammar creates plenty of room for rewriting errors: gendered nouns, complex verb conjugations, the subjunctive mood. The mistakes to look for in any rewritten French are the ones a native speaker notices immediately: incorrect gender agreement, the wrong preposition, an awkward subjunctive construction.

German

Turnitin doesn't cover German either; GPTZero, Originality.ai and Copyleaks do.

German is where English-first tools often struggle. The language's compound nouns (Rindfleischetikettierungsüberwachungsaufgabenübertragungsgesetz, anyone?), case system and flexible word order are easy to get subtly wrong. Watch for wrong case endings, compound words split where they shouldn't be, and word order no native speaker would produce. Those errors can make text look machine-processed to a human reader regardless of what a detector says.

Arabic

Turnitin's AI detection now covers Modern Arabic (our explainer: Turnitin's Arabic AI detection), but its paraphrase and bypasser detection remains English only. Originality.ai lists Modern Standard Arabic; GPTZero doesn't list Arabic among its fully supported languages, and Copyleaks doesn't publish an Arabic accuracy figure on its detector page. In our own one-draft check, SupWriter's detector scored an AI-written Arabic draft at 100 and the rewrite at 75, which is not a result to lean on: the same detector scored six of eight human-written Arabic passages at 73% AI or higher.

The connected script, right-to-left direction, root-based morphology and diacritics make Arabic hard to rewrite well. Poorly processed Arabic can end up with unusual token patterns that look strange to any reader, human or machine, because they don't resemble coherent writing. Have a fluent reader check register and morphology before anything else.

Chinese (Simplified Mandarin)

Turnitin doesn't cover Chinese, and GPTZero doesn't list it as fully supported. Originality.ai lists both Simplified and Traditional Chinese. SupWriter's detector returned 0 for both a Mandarin AI draft and its rewrite, so it cannot evaluate Mandarin at all.

Chinese presents the word-segmentation challenge described above, and tools built around Latin-script tokenization can behave erratically with it. One warning sign to watch for: output that reads like translated text rather than native Chinese writing, which suggests the tool translated to English, rewrote, and translated back.

What About Google? AI Content in Non-English Search

You'll see claims that Google is worse at "detecting AI" in other languages, so AI content has a longer runway outside English. We haven't found public evidence that Google Search runs a language-by-language AI-text detector, so that claim is speculation.

What Google does publish is its spam policies, and they don't carve out languages. Google defines scaled content abuse as "when many pages are generated for the primary purpose of manipulating search rankings and not helping users," lists "using generative AI tools or other similar tools to generate many pages without adding value for users" as an example, and says it applies "no matter how it's created" (Google Search spam policies). For more on how Google treats AI-assisted content, see our summary of Google's stance on AI content.

Non-English SEO: Opportunity Without Shortcuts

There's a business angle here that content marketers and SEO professionals should think about carefully.

Publishing for audiences in other languages can be a real opportunity when good content in that language is thin. But the opportunity is being useful where others aren't, not exploiting a supposed detection gap. The spam policy above applies to AI-assisted pages in Spanish or Arabic exactly as it does in English.

A workable approach: use AI as a drafting aid, have a native speaker edit every piece, and add something only you can provide, such as original data, local examples or first-hand experience. SupWriter's editor works in 50+ languages, which helps with the rewriting step, but it doesn't replace that native-speaker edit. For non-native speakers writing in English, the goal is the same: text that reads like you wrote it.

What to Look for in a Non-English Humanizer

If you're working in a non-English language, here's what separates a tool that actually works from one that just claims multilingual support:

Language-specific processing vs. translation pipelines. This is the critical distinction. A tool that processes French text as French will generally do better than one that translates to English, rewrites, and translates back. Ask how the tool handles your language; the output will tell you even if the marketing doesn't.

Native speaker evaluation. Run the output past a native speaker before you trust it. Grammar and meaning preservation matter more than any detector score if the "humanized" text reads like it came out of a malfunctioning translation engine.

Script support. If you're working in Arabic, Chinese, Japanese, Korean or other non-Latin scripts, verify that the tool handles the script natively rather than through romanization. Romanization-based processing destroys the structural information that makes the text read naturally.

Detection tool coverage. Different detectors support different languages, and some return nothing at all for unsupported ones. Check the detector your audience or institution actually uses against its own language list before you read anything into a score.

For a broader comparison of humanization tools, see our best AI humanizer tools roundup.

The Bottom Line

Non-English AI humanization is a mostly unsolved problem, on both sides. Many tools that claim multilingual support were built for English first, and many detectors either don't cover your language or don't publish how well they do. The further a language sits from English, structurally and in its script, the more careful you should be.

Our own measurements point the same way. On our detector, 4 of 5 Tagalog runs (on 4 drafts) scored under 11% AI after rewriting, but the same draft could land at 0% in one run and 100% in another. Our detector can't judge Malay at all: it scored every human-written Malay passage we checked as 100% AI. For most other languages it isn't a trustworthy judge either.

If you're working in a non-English language, don't assume that a tool's English performance predicts its performance in your language. Test it on your own writing. Have a native speaker read the output. And check which detector will actually see your text, and whether it supports your language at all.

Make an AI-assisted draft sound like you

Rewrite stiff, AI-sounding passages in a natural voice.

Free account, no card · 300 free words

Related Articles