Testing Familiarity, Not Memory
Retrieval practice is one of the most reliably replicated findings in learning science. Asking learners to pull information out of memory produces better long-term retention than asking them to review the same material again. Roediger and Karpicke demonstrated this repeatedly in the mid-2000s, and the effect has held up across subject areas, age groups, and delays. So we all added quizzes. Knowledge checks after every module. A five-question gate before the next section. Reviews at the end of the course.
And most of those quizzes are multiple-choice, because multiple-choice is cheap to author, trivial to auto-grade, and works on a phone. That convenience carries a cost that rarely shows up in a completion report: a multiple-choice item does not necessarily test retrieval at all. It often tests recognition, and recognition is an easier, shallower, and much more forgiving process than recall. The result is a quiz that looks like retrieval practice, scores like success, and predicts performance poorly.
The Two Things A Question Can Ask For
Recall asks the learner to reconstruct information from memory with no support. "What are the three conditions that trigger the escalation path?" There is nothing on screen to lean on. The learner either produces it or does not.
Recognition asks the learner to decide whether something presented to them is familiar. "Which of these is a condition that triggers the escalation path? A / B / C / D." The answer is on the screen. The task is to pick it out.
These feel like the same question dressed differently. Cognitively they are not. Recognition can be supported by a vague sense of familiarity (the phrasing looks like something from slide 14) without the learner being able to produce, apply, or explain the content. Familiarity is a fast, low-effort signal, and it is exactly the signal a well-formatted answer option hands over for free.
Robert and Elizabeth Bjork's work on desirable difficulties gives the underlying principle: the effort of retrieval is not a side effect of learning, it is the mechanism. Make the retrieval easier and you get a more comfortable learner and a weaker memory trace. A four-option question with one plausible answer and three obviously wrong ones has removed most of the difficulty that was doing the work.
Where This Bites In Practice
- Scores inflate and stop being diagnostic.
With four options, a learner who knows nothing scores 25% by guessing. With 3 implausible distractors, a learner with partial familiarity scores far higher. When a cohort averages 88% on a knowledge check and then cannot perform the task on the floor, the quiz is usually the thing that was wrong, not the learners. - Learners' own judgments follow the score.
Karpicke, Butler, and Roediger found that learners systematically misjudge which study strategies are working, favouring the ones that feel fluent. An easy multiple-choice quiz supplies exactly that fluency. The learner walks away confident, which is the worst possible combination: no knowledge and no reason to seek any. - Wrong options can be learned.
Roediger and Marsh documented a negative suggestibility effect: after taking a multiple-choice quiz, learners sometimes later reproduce the incorrect alternatives they read as if they were facts. If your distractors are well written enough to be tempting, they are well written enough to be remembered. Butler and Roediger found that corrective feedback substantially reduces this, which makes feedback a requirement of the format rather than a nice-to-have.
The Case For The Defence
Multiple-choice quizzes are not irredeemable, and the research does not say they are. Little, Bjork, Bjork, and Angello showed that multiple-choice tests can produce retention benefits matching or exceeding short-answer tests, but only under a specific condition. The alternatives must be competitive: each one plausible enough that the learner has to actively retrieve information about it in order to rule it out. When that happens, a single four-option item triggers retrieval of several pieces of knowledge instead of one, and learners end up performing better on later tests of the incorrect alternatives too.
The mechanism cuts both ways, and the distractors decide which way. Lazy distractors produce a recognition task. Competitive distractors produce several retrieval attempts. Same format, opposite outcome, and the difference lives entirely in the authoring.
Five Changes That Recover The Difficulty
- Write distractors from real learner errors, not from the thesaurus.
Pull them from support tickets, open-response pilot data, help-desk logs, and the misconceptions your SMEs complain about. Distractors invented to fill slots are almost always implausible; distractors harvested from actual mistakes are automatically competitive, and each one ruled out is a genuine retrieval. - Kill the surface tells.
The longest option, the only grammatically agreeing option, the only one hedged with "usually", the only one written in the course's own phrasing: each of these lets a learner score without knowing anything. Read your item set with the stem covered and see how far you get. - Put a free-recall item first.
Before the multiple-choice block, ask one open question with no options: "Name what you can remember about X." It does not need to be auto-graded or even scored; the retrieval attempt is the point, and it happens before any options contaminate memory. A single unsupported attempt tends to be worth more than the block that follows it. - Give corrective feedback immediately, and explain why each distractor is wrong.
This addresses the suggestibility problem directly and converts a distractor into a teaching moment rather than a candidate false memory. - Space the items instead of massing them.
Ten questions at the end of a module produce a burst of retrieval from working memory that is still warm. The same ten spread across days force retrieval from long-term memory, which is the only kind that predicts later performance. Spacing costs nothing to implement and reliably outperforms the end-of-module block.
What To Measure Instead
If you take one operational change from this, make it delayed assessment. Immediate post-module scores are dominated by short-term availability and reward the wrong design choices. A short check one or two weeks later, on a subset of the same objectives, tells you what actually stuck.
Expect the delayed numbers to be worse, sometimes dramatically. That drop is not a regression. It is the measurement finally catching up with what was always true, and it is the only honest baseline against which a format change can be evaluated.
The uncomfortable version of all of this: a multiple-choice quiz that everyone passes has told you nothing. If your knowledge checks never surface a gap, they are not instruments. They are ceremony.
References:
- Roediger, H. L., and J. D. Karpicke. 2006. Test-Enhanced Learning: Taking Memory Tests Improves Long-Term Retention. Psychological Science 17 (3): 249-55.
- Karpicke, J. D., and H. L. Roediger. 2008. "The Critical Importance of Retrieval for Learning." Science 319 (5865): 966-68.
- Bjork, E. L., and R. A. Bjork. 2011. "Making Things Hard on Yourself, But in a Good Way: Creating Desirable Difficulties to Enhance Learning." In Psychology and the Real World.
- Karpicke, J. D., A. C. Butler, and H. L. Roediger. 2009. "Metacognitive Strategies in Student Learning: Do Students Practise Retrieval When They Study on Their Own?" Memory 17 (4): 471-79.
- Roediger, H. L., and E. J. Marsh. 2005. "The Positive and Negative Consequences of Multiple-Choice Testing." Journal of Experimental Psychology: Learning, Memory, and Cognition 31 (5): 1155-59.
- Butler, A. C., and H. L. Roediger. 2008. "Feedback Enhances the Positive Effects and Reduces the Negative Effects of Multiple-Choice Testing." Memory & Cognition 36 (3): 604-16.
- Little, J. L., E. L. Bjork, R. A. Bjork, and G. Angello. 2012. "Multiple-Choice Tests Exonerated, At Least of Some Charges." Psychological Science 23 (11): 1337-44.