Flagged as AI for Writing in Your Second Language: What to Do
It is not your imagination. Detectors score vocabulary range, not authorship, and a peer-reviewed study found over half of non-native English essays wrongly flagged.
If you learned English as a second language and a detector has flagged your writing as AI-generated, the first thing worth saying is that you are not imagining the pattern, and this is not a failure of your English. It is a documented weakness in how these tools work — one that researchers identified and published in a peer-reviewed journal, and one that has caused several institutions to restrict or switch off detector use entirely.
That does not make the situation less stressful. Being suspected of cheating over work you spent weeks on is genuinely awful, and if you are reading this at 2am with a meeting scheduled tomorrow, you have every right to feel wronged. But you are in a much stronger position than you probably realise, and this article is about what to do with that position.
The research
In 2023, Liang and colleagues published a study in the journal Patterns examining how GPT detectors perform on writing by non-native English speakers. They ran essays written by non-native speakers for the TOEFL exam through a set of widely used detectors, alongside essays written by native English speakers.
The finding was stark. The detectors classified a majority of the non-native TOEFL essays as AI-generated. On the native-speaker essays, the same detectors performed accurately — correctly identifying them as human-written nearly all the time.
Same tools. Same task. Human-written text in both cases. The only variable that changed was whether English was the writer's first language, and the error rate went from close to zero to a majority.
The researchers went further and demonstrated something important about the mechanism. When they took the non-native essays and used a language model to rewrite them with richer, more varied vocabulary, the flagging largely disappeared. The detectors were not reacting to anything about how the essays were produced. They were reacting to the linguistic profile of the writing — and specifically to vocabulary range.
Why this happens: the perplexity trap
Understanding the mechanism helps, because it turns an inexplicable accusation into a predictable technical artefact.
Most AI detectors are built around a measure called perplexity. Perplexity asks: given everything up to this point in the sentence, how surprising is the next word? A language model reads your text and, at each word, estimates how likely that word was.
The logic detectors apply is this. Language models generate text by repeatedly choosing likely next words, so their output tends to be highly predictable — low perplexity. Human writers make idiosyncratic choices, reach for unusual words, break their own patterns — high perplexity. So low perplexity gets read as machine authorship.
Now consider how careful second-language writing actually looks. If you are writing in a language you learned rather than absorbed, you tend to:
- Use vocabulary you are confident about rather than reaching for a rare word you might use slightly wrong
- Build sentences with regular, reliable structures instead of unusual constructions
- Avoid idioms, slang, and cultural shorthand where the usage rules are fuzzy
- Follow the essay conventions you were explicitly taught, consistently
- Write more formally overall, because formal registers are what language instruction emphasises
Every one of those is good practice. Every one of them is what a competent writing teacher would advise. And every one of them lowers perplexity.
The detector cannot tell the difference between text that is predictable because a machine generated it and text that is predictable because a careful writer chose words they were sure of.
That is the entire problem in one sentence. Two very different things produce the same statistical signature, and the tool sees only the signature. There is a bitter irony in it: the more disciplined and controlled your English, the more machine-like you look. Writers with messier, more error-prone English often score lower.
Why Grammarly and translation tools can make it worse
Many second-language writers use assistive tools, entirely legitimately. Both common ones tend to push scores up.
Grammar and style checkers work by suggesting the conventional phrasing — the standard construction, the expected word. Conventional means predictable, and predictable means low perplexity. Accept enough suggestions and you have systematically smoothed out exactly the irregularities detectors read as human.
Machine translation is even more direct. If you draft in your first language and translate, the English you get was produced by a neural model trained the same way generative models are trained. It carries their statistical fingerprint because it came from the same kind of system. Detectors were never designed to distinguish "translated by a model" from "written by a model," and mostly they cannot.
This puts second-language writers in a genuinely unfair bind. The tools that help you communicate clearly are the tools that make you look suspicious. Worth knowing before an accusation, and worth explaining calmly if one arrives — using a translator or a grammar checker is not academic misconduct at most institutions, and you should say so plainly rather than hiding it.
What to do before you submit anything
The best protection is evidence of process, gathered as you go. Almost none of this takes extra effort — it is mostly about not deleting things.
- Write in a tool that keeps version history. Google Docs (File → Version history) and Word files saved to OneDrive both keep a full timeline showing your document growing over days, with false starts and rewrites. This is the single most persuasive evidence there is, because generated text does not have a history.
- Do not compose elsewhere and paste in. Pasting a finished document destroys the history and produces a suspicious single-paste timeline.
- Keep your notes. Handwritten pages photographed on your phone, outlines, mind maps, reading notes in your first language — all of it. Notes in your native language are particularly convincing, because they show your thinking happening somewhere a detector never looked.
- Save dated drafts. Even just essay-draft1.docx, essay-draft2.docx. Timestamps do a lot of work.
- Keep a record of your sources. Downloaded PDFs, library records, browser history, annotated readings.
- Email your supervisor a question mid-project. One clarifying question, sent while you are working, creates a timestamped record that you were engaged with the material.
- Note what tools you used. If you used a translator or grammar checker, write down where and how. Being able to explain your process precisely is far better than being asked and improvising.
If you are accused
First, do not panic-rewrite
The instinct is to open the document and change everything immediately. Do not. You will overwrite the version history that proves your case, and a frantic full-document rewrite hours after a flag looks far worse than the original score. Preserve everything first: download the version history, copy your drafts somewhere safe, screenshot anything that might change.
Then work through this
- Reply calmly and briefly. A short message asking to meet and discuss it beats a long defensive email. Long emails read as anxious; meetings resolve things.
- Ask what evidence exists beyond the detector score. Ask it politely and genuinely. This is the most important question you can ask, because detector vendors themselves state their scores are not proof of misconduct. If the only evidence is a percentage, that is worth establishing early and calmly.
- Produce your process evidence. Bring the version history, the drafts, the notes, the sources. Walk through them. Show the document being built.
- Offer to discuss the content. Ask to talk through your argument, why you chose your sources, what you cut and why. Nobody who did not write something can do this convincingly, and everyone who did can.
- Explain your writing process honestly, including tools. "I drafted in Google Docs, I used Grammarly for grammar, I looked up some terms with a translator" is a normal, permitted process at most institutions. Say it plainly.
- Mention the research if it helps. Not as an accusation. Something like: "I understand there is published research in the journal Patterns showing these detectors misclassify non-native English writing at high rates. Could that be a factor here?" Most instructors do not know this and are glad to be told.
- Find out your institution's formal process. Ask early what the appeals route is, whether there is a student advocate, ombudsman, or students' union support service. Use them — that is what they are for, and they have seen this exact case before.
- Bring someone with you. Many institutions allow a supporter in academic integrity meetings. Ask whether yours does.
Know that institutions have already acted on this
You are not asking anyone to accept a fringe theory. A number of universities, including some large and well-known ones, have restricted or disabled AI detection tools in their learning systems, and the disproportionate false-positive rate for non-native English speakers has been an explicitly cited reason. Detector vendors themselves publish caveats about false positives and state that scores should not be used as sole evidence.
This means the position you are arguing from is the mainstream, cautious one. The person treating a percentage as proof is the one departing from both the research and the vendors' own guidance.
The wider fairness problem
Step back from your own case for a moment, because the shape of this matters.
A detection tool with an uneven error rate does not distribute harm randomly. It concentrates it on one group — here, international students, immigrants, and anyone writing academically in a language they learned rather than grew up with. These are frequently people already carrying more risk: visa conditions tied to academic standing, family funding, distance from support networks, and often less familiarity with how to challenge an institutional process.
The consequences compound in ways the tool's designers never see. Students start deliberately writing worse to avoid flags — adding errors, abandoning the disciplined habits their teachers taught them. That is a genuine harm to their education, produced by a tool meant to protect academic standards.
None of this is a reason to despair about your own case. It is a reason to push back with confidence rather than apology. You did the work. The tool has a known, published, institutionally-acknowledged failure mode that matches your situation exactly. That is a strong position, and you should occupy it calmly.
A last word
If you are dealing with this right now: gather your evidence before you do anything else, ask for a conversation rather than sending a defensive essay, and ask what evidence exists beyond the number. Most of these situations end quietly once a person who did the work sits down and talks about it.
And do not let this change how you write. Writing clearly, in vocabulary you can use precisely, with structures you control, is good writing. A tool that misreads that is the thing that is broken.
More from the blog
Using AI for a Literature Review Without Wrecking It
AI has a real role in a literature review, and it sits later than most people put it: afte...
Does Google Penalise AI Content? What the Policies Actually Say
No, Google does not penalise AI-written content. It penalises pages published at scale to...
AI Invented a Citation: How to Catch It Before Anyone Else Does
Two lawyers filed a brief citing six cases that did not exist, then asked ChatGPT whether...