Education resources › Blog › AI, the Toupee Fallacy and the 88% Problem

A woman marking a potentially AI-generated coursework essay with green and red pens

AI, the Toupee Fallacy and the 88% Problem

3 min read
  • Phones, AI & technology

Is coursework dead? A new study just released could well be the final nail in the coffin. To understand why coursework now represents the worst form of assessment, we need to understand two different concepts: the toupee fallacy and the 88% problem.

The Toupee Fallacy

We all know a bad wig when we see it. But because we are easily able to spot really bad and obvious toupees, we tend to overestimate our ability to spot really good ones. Something similar happens with AI-generated work. Because we are good at spotting bad examples, we are less likely to spot good examples of AI-generated work.

There are some common features that people often associate with AI-generated work. These include predictable patterns and structure, overusing the rule of three, em-dashes and American spelling. For example, when looking at the following passage, most people would quickly and correctly realise they are reading an AI-generated text:

AI-generated quote about volcanoes and magma.

Because we find it easy to spot the above as AI, we become over-confident in spotting all AI-generated work. For example, consider this fascinating experiment run in the New York Times. They found that in a test where you don’t know if it is AI-generated vs some of the best historical human-written text, people’s preference was basically a coin toss between the two. And this was true across a range of genres: literary fiction, fantasy, science writing, historical fiction and poetry. You can try this test for yourself here. This is a perfect example of the Toupee Fallacy.

The 88% Problem

We have previously blogged about how AI detection software doesn’t really work and isn’t reliable (the exact quote from the research is about as conclusive as you tend to get in research papers: “Detection tools are neither accurate nor reliable”). And yet, schools and colleges are being increasingly sold software that promises to detect if students have used AI to generate entire essays.

Recent research provides further weight to this, with researchers concluding that “detectors appear more adept at identifying low-cognitive-load tasks but often fail when confronted with outputs of higher-order thinking”.

And to make matters worse, with a few simple tweaks, students can avoid it 88% of the time. When a student makes small, deliberate changes, the statistical fingerprint of the text can shift enough to avoid detection. This is known as The Laundering Effect and students are wise to it. In the last few months, we have heard of students doing the following to cover up their AI-generated work:

  • Using prompts to include a few spelling mistakes
  • Telling AI not to use em dashes
  • Using a few synonyms to change the AI output

Where does this leave us?

The evidence is pretty clear. We can’t really spot good AI-generated work using the old-fashioned eye test, and with minimal changes, technology really struggles to detect AI-generated work. This means that students who use AI are likely to get away with it, but it also leads to a culture which is ripe for false positives where genuine student work is incorrectly flagged as AI-generated.

As Daisy Christodoulou says in our recent Science of Learning podcast episode, the cost of this can be hugely damaging.

Final Thoughts

If we don’t know if students have done the work, then how much weight should we give it as part of a formal assessment? And what could replace unsupervised coursework? Oral assessments? These are very hard (maybe impossible?) to do at scale. Handwritten pen and paper exams? That seems more likely for now. But in an increasingly digital age, pressure will come from a range of sources to abandon this.

What this ultimately means for schools and colleges is difficult to say. The truth is there are no simple answers to complex challenges.  And AI is arguably the most significant and complex challenge currently facing us. We don’t have all the answers. No one does. But we firmly believe that research can illuminate. It can shine a light to help us not stumble completely in the darkness. Knowing about the Toupee Fallacy and the 88% Problem can hopefully offer a flicker of light.

Explore the good, the bad and the ugly sides of AI in education and develop strategies to integrate AI tools effectively in the classroom.


About the author

Bradley Busch

Bradley Busch

Bradley Busch is a Chartered Psychologist and a leading expert on illuminating Cognitive Science research in education. As Director at InnerDrive, his work focuses on translating complex psychological research in a way that is accessible and helpful. He has delivered thousands of workshops for educators and students, helping improve how they think, learn and perform. Bradley is also a prolific writer: he co-authored four books including Evidence-Informed Wisdom, Teaching & Learning Illuminated and The Science of Learning, as well as regularly featuring in publications such as The Guardian and The Telegraph.

Follow on XConnect on LinkedIn
Claire Badger

Dr Claire Badger

Dr Claire Badger is Head of Teacher Professional Development at InnerDrive. She specialises in helping schools and colleges develop their evidence-informed practice. Her main area of focus is on cognitive science, teacher development and maximising CPD within schools. She was an assistant head in charge of teaching and learning and holds a Masters in Teaching and Learning from the Institute of Education. She is the co-author of the book “Creativity for Teachers:  A Cognitive Science Approach“.

Follow on XConnect on LinkedIn