Automated AI text detectors are not reliable enough to prove someone used AI, but online exams still catch AI misuse in other ways. OpenAI withdrew its own AI text classifier in 2023 for low accuracy, and detectors produce false positives. In a real exam, proctoring, exam design and the answers themselves are what flag AI use, not a single score.
Why AI text detectors are unreliable
AI writing detectors try to guess whether text was machine-generated. The problem is they are wrong often enough to be dangerous. OpenAI, the maker of ChatGPT, shut down its own AI text classifier in July 2023, citing a low rate of accuracy. Detectors flag genuine human writing as AI (false positives), miss AI text that has been lightly edited (false negatives), and can be evaded with simple rewording. That is why no responsible institution should accuse a candidate on a detector score alone.
How online exams actually catch AI use
Detection in a well-run exam rarely rests on a text classifier. It comes from several signals reviewed together:
- Proctoring signals: a live or recorded proctoring session can show eyes repeatedly moving off-screen, a second device, another voice in the room, or copy-and-paste behaviour.
- Lockdown controls: a locked-down browser stops a candidate switching tabs to an AI tool mid-exam.
- The answer itself: responses far above a candidate's demonstrated level, oddly generic phrasing, or near-identical answers across candidates prompt a human to look closer.
- Timing and process data: an essay that appears fully formed in seconds, with no drafting, looks nothing like one a person writes.
No single signal is proof. Together, and reviewed by a person against the exam's rules, they build a defensible case.
What this means for candidates
Generating answers with AI in a proctored, closed-book exam is a serious integrity breach that can void your result. Using AI to help you study beforehand is a completely different thing and is fine, as long as the exam itself is your own work. If you are unsure what your exam allows, read its rules before you sit rather than guessing.
What this means for trainers and institutions
Do not rely on an AI detector to make accusations; the false-positive risk is a real legal and reputational hazard. Instead, design assessments that are hard to outsource: proctored delivery, randomised question banks so no two candidates see the same paper, timed sections, and questions that test applied judgement rather than recall a model can generate in seconds. Robust process beats a magic detector. See our guides on how online proctoring works and preventing cheating in online exams.
Practise under real exam conditions
The best defence against AI shortcuts is an assessment candidates cannot easily game and are prepared for honestly. CandidatesPrep delivers timed, randomised, proctored-style mocks scored by domain, so candidates rehearse legitimately and trainers can spot anomalies in the data rather than trusting a detector.
The bottom line
AI text detectors are too unreliable to stand on their own, and OpenAI retiring its own is the clearest sign of that. Online exams still catch AI misuse, but through proctoring, exam design and human review, never a single number.
Build assessments that hold up: try the CandidatesPrep simulator, or book a demo to see proctored, randomised mocks in action.

