Interleaved Recall Beats Re-Reading by 27% in Year-Long Tests
A year-long randomised controlled trial involving 1,240 university students found that interleaved recall practice improved long-term retention by 27% over an equivalent amount of re-reading time, as measured by a cumulative final examination administered 11 months after the initial study phase. The effect held across four distinct subject domains—statistics, art history, organic chemistry nomenclature, and introductory Portuguese vocabulary—with the smallest observed gain (19%) occurring in the chemistry group. These results challenge the long-standing assumption that repeated exposure, even when spaced, is the primary driver of durable memory formation.
Experimental Design and the 27% Figure
The study, conducted across three European universities between January 2023 and February 2024, assigned participants to one of four conditions: pure re-reading, interleaved recall, blocked recall, and a hybrid condition combining re-reading with interleaved recall. Each participant received 18 hours of total study time, distributed as three 30-minute sessions per week over 12 weeks. The interleaved recall condition involved presenting material in a shuffled sequence—for example, alternating between Portuguese verb conjugations, statistical formulas, and art historical periods within a single session—and requiring participants to generate answers from memory before checking their accuracy. The re-reading group received the same content in the same shuffled sequence but was instructed to read each item aloud twice without attempting recall.
The 27% figure refers to the difference in mean final exam scores between the interleaved recall group (78.4%) and the re-reading group (61.7%), a raw percentage-point gap of 16.7 points. When expressed as a relative improvement—the metric used in the study’s primary analysis—this constitutes a 27.1% advantage (Cohen’s d = 0.84, 95% CI [0.71, 0.97]). The effect was not a product of extra time on task: both groups engaged with the material for exactly 9,720 total seconds. The researchers verified compliance through a custom browser extension that logged keystroke patterns and screen focus, discarding data from 87 participants who failed to meet the 90% engagement threshold.
Why Interleaving and Recall Compound
The study’s theoretical contribution lies in isolating the additive effects of two separate mechanisms. Interleaving alone—without recall—produced a 9% improvement over re-reading, consistent with prior work on contextual interference. Recall alone—without interleaving—produced a 14% improvement, aligning with the testing effect literature. The combined condition yielded a 27% improvement, which is significantly greater than the sum of the individual effects (23%), suggesting a superadditive interaction (interaction term p = 0.03).
The Retrieval Effort Hypothesis
The researchers attribute the superadditivity to what they term “retrieval effort modulation.” When recall is attempted on interleaved material, the learner must not only retrieve the target information but also suppress competing responses from recently studied items in different domains. In the blocked recall condition, the retrieval set is homogeneous—all items come from the same category—so the discrimination process is trivial. In interleaved recall, each retrieval attempt requires category-level discrimination before item-level retrieval, a two-stage process that appears to create stronger episodic traces.
This hypothesis was supported by a secondary analysis of response latencies. In the interleaved recall condition, correct responses took an average of 2.8 seconds longer than in the blocked recall condition during the first week of the study. By week 12, that latency gap had narrowed to 0.4 seconds, indicating that the discrimination process had become automated. Critically, the magnitude of latency reduction predicted final exam performance within the interleaved group (r = 0.41, p < 0.001), but showed no predictive value within the blocked group (r = 0.06, ns).
Practical Parameters for Implementation
The study’s most actionable finding concerns the optimal schedule for interleaved recall. The researchers varied the interleaving granularity across four experimental sub-groups, ranging from fine-grained (switching topics every 2 minutes) to coarse-grained (switching every 15 minutes). The 2-minute condition produced the highest final exam scores (81.2%), but also the highest dropout rate (23% vs. 11% in the 15-minute condition). The 6-minute condition provided the best retention-to-attrition trade-off, with a mean score of 79.1% and a dropout rate of 14%.
The 6-Minute Rule
The 6-minute rule—switching topics every six minutes during a 30-minute study session—emerged as the practical sweet spot. Under this schedule, a typical session would involve five topic switches, each requiring the learner to reorient to a different domain. The researchers note that this aligns with the average attention span for deliberate practice tasks observed in prior laboratory studies, though they caution against overgeneralising this specific number to all learner populations. Older participants (over 40) showed a preference for the 10-minute interval, achieving 76.8% accuracy, while younger participants (18–25) performed best at 4-minute intervals (80.5%).
Limitations and the Replication Question
The study is not without methodological caveats. The participant pool was drawn exclusively from university students, 68% of whom were pursuing humanities or social science degrees. The organic chemistry nomenclature module—the most technical content—showed the smallest effect size, raising the possibility that interleaved recall benefits are domain-dependent. Additionally, the final examination was a multiple-choice test; the researchers did not assess free-recall or application-based outcomes, which might favour different study strategies.
There is also the question of ecological validity. The study’s browser-extension monitoring created an artificial accountability environment that likely increased motivation across all conditions. A follow-up survey administered one month after the study ended found that only 31% of participants in the interleaved recall condition had continued using the technique voluntarily, despite 89% reporting that they believed it was effective. The gap between perceived efficacy and actual adoption suggests that the technique’s cognitive demands—the very mechanism driving its effectiveness—may be a barrier to real-world use.
Implications for Assessment Design
If interleaved recall reliably produces 27% better retention, the most immediate implication may not be for learners but for assessment designers. Standardised tests that sample content in a blocked, topic-by-topic format may systematically underestimate what interleaved learners actually know, while overestimating the competence of re-reading learners who have developed strong within-topic familiarity. The study’s authors propose that high-stakes examinations should themselves be interleaved—randomising question order across topics—to produce a fairer measure of durable knowledge. This recommendation, if adopted, would represent a significant departure from current practice in most licensing exams and university finals.
Whether such a shift would close the gap between study-strategy efficacy and real-world adoption remains an open empirical question. What is clear is that the 27% advantage is not a statistical arteifact of a single cohort or a single content domain. The effect replicated across four subjects, two age bands, and three institutions, and it survived controls for baseline aptitude, working memory capacity, and prior domain knowledge. The next step is not to ask whether interleaved recall works, but rather to identify the conditions under which learners will actually tolerate the discomfort it requires.