Catching the wrong students: AI detection, international students and the fairness crisis in UK universities
This blog was kindly authored by Brendal Aformeziem, PhD candidate, University of Strathclyde.
Last summer, a student at an English university explained during a misconduct viva that they had used Google to find synonyms, because English was not their first language. The panel recorded this as an admission of using AI to paraphrase. The Office of the Independent Adjudicator (OIA), the official ombudsman for higher education in England and Wales, partly upheld the student’s complaint. Among its findings was that the university had never considered whether Turnitin’s AI detection might perform less reliably for non-native English speakers, despite a substantial body of research showing exactly that.
In July 2025, the OIA published four case summaries involving students accused of AI-assisted misconduct. Three involved international or second-language students, and three were either upheld or partly upheld against the university. These are not isolated failures but the predictable consequence of deploying unreliable tools against a student population that was never well-served by them.
The bias built into these tools has been documented for some time. A study published in the journal Patterns by Stanford University researchers tested seven widely-used AI detectors and found that over 61 per cent of essays written by non-native English speakers were misclassified as AI-generated, while the same tools showed near-perfect accuracy for native English writers. The structural explanation is that students writing in a second language tend to use shorter sentences, lower word variety and less idiomatic phrasing. These are precisely the patterns that detection algorithms associate with machine-generated text. The tool cannot tell the difference between artificial intelligence and a student doing their best in a language that is not their own.
The broader reliability picture offers no reassurance. A recent large-scale evaluation tested detection tools across 805 samples and found an average accuracy of just 39.5 per cent on unmodified AI-generated text, falling to 17.4 per cent when students applied simple techniques, with some tools misclassifying half of all human-written text as AI. Weber-Wulff et al’s study, in the most comprehensive independent evaluation to date, found that none of the tools tested met the standard required for reliable use in high-stakes decisions. Turnitin’s own Chief Product Officer has acknowledged that detection scores are probabilistic rather than definitive, and should not form the sole basis for misconduct proceedings. In UK universities, they frequently do.
The OIA cases illustrate how this plays out in practice. In a second case from July 2025, an international student flagged by Turnitin was not shown the evidence against them before their viva; was not given the opportunity to present mitigating evidence in writing; and had their use of Grammarly dismissed without explanation, despite the student’s belief that it was permitted as a language support tool. The OIA upheld the complaint in full. In a third case, a student with autism was accused, penalised, and ultimately cleared on reconsideration, having already had an earlier submission flagged by the same detection software and found, on that occasion too, to be entirely human-written. The tool had identified a writing style as suspicious. On both occasions, no AI had been used.
Research by Arslan et al. found that international students in the UK turn to AI tools primarily to overcome language barriers, navigate unfamiliar academic conventions, and meet the writing standards their programmes demand. They are not gaming the system but trying to meet its expectations, and the tools universities are using to catch misconduct are structurally more likely to flag them than any other group.
This matters financially as well as morally. International students represent 24 per cent of all UK higher education students and 51 per cent of all postgraduate students. At a time when 43 per cent of English universities are forecasting deficits, international tuition income is not a peripheral concern. Universities are financially dependent on these students while simultaneously deploying systems that treat them as the highest-risk group for academic dishonesty and that is neither a sustainable nor a defensible position.
Leading institutions elsewhere have already reached this conclusion. The University of Waterloo discontinued Turnitin’s AI detection in September 2025, citing unreliability and the risk of wrongful accusation. Curtin University followed in January 2026, and UCLA and UC San Diego made the same call in 2024. UK universities have not followed their lead, but the evidence that they should is now overwhelming.
The alternative to detection is not the abandonment of academic integrity but the pursuit of it through approaches that work. Francis et al. (2025) and Newton and Jones (2025) identify the same evidence-based alternatives, such as process-based assessment, staged submissions, oral defences and AI literacy embedded in curricula. MIT’s Sloan School of Management has concluded that AI detectors do not work and that the right institutional response is assessment redesign rather than better surveillance. These approaches verify learning rather than police outputs, and while they are harder to design, they are fairer, more educationally defensible, and far less likely to place an international student in front of a misconduct panel for the crime of writing in their second language.
The policy ask is specific. Universities should suspend AI detection as a primary evidence base in misconduct proceedings, pending independent validation of the tools they are using. The Quality Assurance Agency (QAA) and the OIA should issue joint guidance making clear that detection scores alone cannot ground disciplinary action. And universities should commit, with dedicated staff support, to assessment redesign programmes that reduce reliance on essay-based submissions where AI use cannot be reliably distinguished from legitimate writing assistance.
Three years into the generative AI era, UK universities are still treating academic integrity as a detection problem. The OIA cases published in July 2025 show what that choice costs. The question for 2026 is whether institutions will change course before more students pay the price.





Comments
Dr John Milliken says:
# Stop detecting and start designing: AI, assessment and the duty to make learning visible
Brendal Aformeziem is right to expose the potentially discriminatory consequences of AI-detection systems. When the language patterns of international students are more likely to be interpreted as machine-generated, a technical weakness quickly becomes a question of procedural fairness, institutional responsibility and natural justice.
But the problem goes deeper than unreliable software.
AI detection is being used to repair at the disciplinary stage what should have been addressed at the assessment-design stage. Universities are attempting to determine, after a piece of work has been submitted, whether genuine learning occurred during a process they have made largely invisible.
That is the fundamental design failure.
The conventional assessment sequence appears straightforward:
**Task → Student work → Submission → Marking → Outcome**
Yet this sequence tells the university very little about how the student reached the submitted answer. When doubts arise, an AI detector is inserted between submission and academic judgement:
**Task → Student work → Submission → Detection score → Suspicion → Investigation**
The detector does not observe the student learning. It does not know what the student understands, which sources were consulted, how ideas developed, what language support was used or why particular decisions were made. It merely identifies textual patterns that resemble other textual patterns.
A probability is then in danger of becoming an accusation.
For international students, the consequences are particularly troubling. Universities recruit students into an unfamiliar linguistic, cultural and academic environment. They expect them to master new conventions of argument, citation and formal written English, often within remarkably short periods. Students may be encouraged to use dictionaries, grammar checkers, writing centres, academic-skills tutors and other forms of language support.
Yet when the resulting prose appears controlled, simplified, formulaic or insufficiently idiomatic, those same linguistic characteristics may be interpreted as evidence of artificial authorship.
The student is first required to adapt and then risks being punished for the linguistic traces of that adaptation.
That is not simply a software problem. It is an institutional flow failure.
## From product inspection to learning verification
Universities have traditionally concentrated assessment on the final product. An essay, report or examination script is treated as the visible representation of learning. This worked imperfectly even before generative AI. Students could receive extensive assistance, imitate model answers, reproduce lecture material or construct plausible arguments without necessarily demonstrating deep understanding.
Generative AI has not created this weakness. It has exposed it.
The proper response is therefore not to build more elaborate systems for inspecting finished products. It is to redesign assessment so that learning becomes visible throughout the process.
A useful framework is:
**Recall → Comprehension → Application → Reflection**
Students should be required to demonstrate that they can:
* identify and recall relevant concepts;
* explain those concepts in their own terms;
* apply them to a specific problem, professional setting or personal context;
* reflect critically on the application, including what succeeded, what failed and what they would change.
This is not a theoretical aspiration. Versions of this approach have existed for decades.
Around 20 years ago, participants on the postgraduate teaching programme at Queen’s University Belfast were required to apply course concepts directly to their own teaching practice and then reflect upon the results. The assessment was not designed around the reproduction of generic educational theory. Participants had to show what they had attempted, why they had chosen a particular approach, what happened in their own classroom and what they had learned from the experience.
A contemporary AI system might help a participant improve the expression of that account. It might suggest terminology or assist with structure. But it could not easily manufacture a sustained and credible relationship between the participant’s teaching context, professional decisions, classroom experience, resulting evidence and subsequent reflection.
The authenticity came from contextualised application and reflective judgement, not from linguistic imperfection.
## Assessment as a flow process
A more robust assessment system would make the learning journey visible through a connected sequence:
**Assessment design → Student engagement → Evidence of development → Application → Reflection → Academic dialogue → Judgement → Feedback**
Each stage strengthens the next.
Assessment design should make clear what kind of learning is being demonstrated and what forms of AI or language assistance are permitted. Student engagement should generate evidence through notes, drafts, decisions, discussions or staged submissions. Application should require students to use knowledge in a sufficiently specific context. Reflection should reveal their reasoning and capacity for self-evaluation. Academic dialogue, including brief oral questioning where appropriate, should allow the assessor to test understanding without turning every student encounter into a misconduct interrogation.
Judgement then rests on a body of connected evidence rather than a numerical suspicion score.
This does not mean that every module must adopt multiple drafts, oral examinations and individual vivas. That would be impractical, particularly in large cohorts. Nor does it mean abandoning essays, which remain valuable forms of intellectual work.
It means designing proportionate points of verification.
A short commentary explaining how an argument developed may reveal more than another thousand words of polished prose. A five-minute discussion may establish understanding more reliably than an AI-detection percentage. A personalised application may be more educationally valuable than a generic question that thousands of students—and several AI systems—could answer in much the same way.
The objective is not to make assessment AI-proof. That is probably impossible. The objective is to make learning sufficiently visible that authorship is not reduced to guesswork.
## The danger of reversing the burden of proof
There is also a serious procedural issue.
Where an institution relies heavily on detection software, the student may effectively be required to prove that they did not use AI. This reverses the proper burden of evidence. A student confronted with an algorithmic score may struggle to explain why their writing resembles an unidentified statistical pattern, particularly when the detector itself cannot explain its reasoning in educationally meaningful terms.
For a student writing in a second language, a student with a disability or a student who has legitimately used grammar-support software, that burden may be especially difficult to discharge.
Universities should not begin with the assumption that unusual language indicates dishonesty. Nor should a student’s inability to explain the operation of a proprietary detection system be treated as evidence against them.
A detector may provide a reason to examine the assessment process more closely. It should not provide a substitute for evidence.
## What universities should do
Aformeziem is right to call for the suspension of AI-detection scores as the primary basis for misconduct action. But suspension alone will not solve the underlying problem. If universities simply remove the detector while retaining assessments that reveal almost nothing about the learning process, uncertainty will remain.
A more constructive response would combine procedural safeguards with systematic assessment redesign.
Universities should:
* prohibit disciplinary findings based solely or predominantly on AI-detection scores;
* require students to be shown the full evidence against them before any misconduct meeting;
* define clearly the permitted use of generative AI, grammar checkers and language-support tools;
* introduce proportionate opportunities for staged work, personal application and reflective explanation;
* use academic dialogue to verify understanding rather than to secure admissions;
* provide staff development in assessment design, international student learning and the limitations of detection technologies;
* evaluate misconduct procedures for differential effects on international, disabled and second-language students.
Above all, institutions should restore academic judgement to the centre of the process.
That judgement should not consist of an assessor deciding whether a sentence “sounds like AI”. It should involve examining whether the student can explain, apply, defend and reflect upon what has been submitted.
## A student-centred integrity system
Academic integrity should remain a central university responsibility. Students are entitled to know that their own work will not be devalued by widespread impersonation, purchased assignments or unacknowledged machine production.
But integrity cannot be protected through systems that are themselves procedurally unreliable.
A student-centred approach does not excuse misconduct. It creates assessments in which misconduct is more difficult to conceal, legitimate assistance is easier to identify, and genuine understanding is easier to demonstrate.
The distinction matters.
The choice is not between aggressive detection and institutional surrender. It is between policing textual outputs and verifying learning.
Three years into the generative AI era, universities are still asking how they can identify whether a machine contributed to a piece of writing. The more educationally important question is whether the student can demonstrate ownership of the ideas, decisions, applications and learning represented by that work.
Universities do not need better machinery for guessing who wrote a text. They need better assessment designs for establishing what a student knows, understands, can apply and has learned from the process.
When a university cannot distinguish linguistic difference from academic dishonesty, the student should not carry the burden of that institutional failure.
Reply
Stephen Langston says:
This is exactly the challenge we’ve been discussing through our work at the University of the West of Scotland and Cyberhare Solutions.
The problem isn’t simply whether AI was used. The problem is whether universities can make fair, evidence-based decisions when the evidence itself has known limitations. AI detection scores were never designed to be treated as proof of misconduct, particularly in high-stakes disciplinary processes.
We’ve consistently argued that academic integrity should move beyond “AI detection” and towards evidence of authorship, transparency and process. Rather than asking “Can we detect AI?”, perhaps the better question is “Can we give students fair opportunities to demonstrate how their work was created?”
That approach not only reduces the risk of false accusations, particularly for international students and those using legitimate language support, but also provides staff with richer, more defensible evidence on which to make human decisions.
Technology absolutely has a role to play, but it should support academic judgement, not replace it. The future of academic integrity is unlikely to be better AI detectors; it’s more likely to be better evidence, greater transparency and human-led decision making. Our solution will be available at http://www.cyberharesolutions.com from 1st October 2026 and it’s free for any one wishing to demonstrate their honest approach to creating any written documentation.
Reply
Charles Knight says:
This is not a pedagogical problem; it’s an economic one. We have a range of methods that are AI proof but takeaway assessment is economically the massproduced hamburger of assessment.
Replacing it is costly, which is why providers are trying to avoid it – not a single one of the proposed fixes for AI in terms of takeaway assessment, from workflows to notes to documentation, actually works.
Reply
Jonathan Alltimes says:
The production process is flawed.
It is a consequence of the scale of operation and the process of admission. University degrees are not a product you can buy off the shelf or a service you order through your phone, in the same way, studying is not a continuous flow it is a social relationship. I understand the financial constraints caused by the tuition fee freeze and the recruitment of international students. You do not know who is studying. In 1985, academics at the Department of Agricultural Botany, University of Reading, used a sheet of photo fit pictures for the first year undergraduates to match a face to the name, the cohort consisted of 10 students. You can redesign the AI detection software and alter the assessments, but these changes can not control fully for the effect of scale and admission: the AI is becoming more sophisticated, as it can copy your style and the answer each time is not simply a duplicate. Academic judgments rely on personal experience, as long as you maintain your social distance, the fudge production will continue.
Reply
Oluwaseun Temitayo Onwuka says:
Well detailed and informative.
Reply
Abdul says:
A fascinating read. I hope that educational institutions can be conscious of these biases when they start to draw their policies on AI uses to ensure fairness and prevent victimisation.
Reply
Add comment