Adaptive Learning Is Not Personalisation
Netflix personalises. It does not teach. The distinction sounds pedantic until you watch an organisation buy a recommendation engine and expect a tutor. This essay shows you the evidence, lets you handle the statistics yourself, and ends with a vendor audit you can run this week.
Netflix personalises. It does not teach. Its recommender system, by its engineers’ own account, is worth over one billion dollars a year in retention[1], and it earns that by showing you more of what you already like. A tutoring system does close to the opposite: it works out what you cannot yet do, and holds you there, deliberately, until you can. Both adapt to the individual. Only one of them produces learning.
The conflation
In enterprise learning and development, these two mechanisms are routinely sold under one label. Adaptive platforms are marketed in the language of streaming: content you will love, journeys tailored to you, engagement as the headline metric. The purchasing logic transfers from consumer software because the buyers live in consumer software. The result is a market where the word adaptive can mean either a recommendation engine or an intelligent tutor, two technologies with forty years of divergent evidence behind them.
Foundations · What is an effect size, in one minute
An effect size (Cohen’s d) expresses how far an intervention moves a group, measured in standard deviations of the outcome. It lets studies with different tests be compared on one scale. Rough benchmarks: 0.2 is small, 0.5 is medium, 0.8 is large. In education research, interventions above 0.6 that replicate across dozens of studies are rare, which is why the tutoring numbers below matter. The chart in Figure 1 lets you feel what these numbers mean rather than take them on faith.
Handle the evidence yourself
VanLehn’s landmark review found intelligent tutoring systems produced an effect size of 0.79, statistically indistinguishable from the 0.76 measured for expert human tutors[2]. Kulik and Fletcher’s meta analysis across fifty evaluations put the median at 0.66[3]. Numbers like these are usually cited and skimmed. Instead, drag the slider: the curves show a comparison group and a tutored group, and the readout translates the gap into plain language.
Two loops, two goals
The mechanisms differ at the level of the feedback loop. A recommender infers what you like from what you consume, then serves more of it; its success condition is that you stay. A tutor infers what you know from your attempts, then serves what you cannot quite do yet; its success condition is that its own predictions become obsolete because you improved. Corbett and Anderson formalised that second loop as knowledge tracing three decades ago[5], and it remains the working core of every serious tutoring system.
Foundations · Why preference and learning diverge
Pashler and colleagues, commissioned to audit the learning styles literature, found no adequate evidence that matching instruction to stated preferences improves outcomes[4]. The Bjork laboratory’s work on desirable difficulties goes further: conditions learners prefer, such as smooth, easy, familiar practice, are frequently the conditions that produce the least durable learning, while effortful conditions like spacing, interleaving and retrieval feel worse and work better[6]. Preference is a real signal; it is simply a signal about enjoyment, not about growth.
Where personalisation fails as a learning strategy
The failure is visible in the numbers the industry itself publishes. LinkedIn’s Workplace Learning Report puts course completion around 42 percent[7], presented as an engagement benchmark. Read as an instructional statistic, completion is a consumption metric: it records that content was watched, not that capability changed. MOOCs made the same substitution a decade earlier; their single digit completion rates[8] became the canonical cautionary tale, yet the successor metric is still watching. When the measure is consumption, the optimisation target becomes watchability, and the loop on the left of Figure 2 quietly replaces the one on the right.
Exercise · The vendor demo
The skill this essay wants to leave you with is separation: hearing a claim and knowing which loop it belongs to. Below are four statements assembled from real vendor language. Classify each one, then take the checklist into your next procurement meeting.
What this means in practice
None of this argues against personalisation, which does exactly what it was built to do. It argues against buying one mechanism and expecting the outcomes of the other. An organisation that wants engagement should buy the recommender and measure watching. An organisation that wants capability should demand the learner model, accept that effective practice will sometimes feel harder than watching[6], and measure what people can do afterwards. The research permits both purchases. It does not permit confusing them.
- Recommendation engines optimise preference; tutoring systems optimise competence. Both adapt; the objective functions are opposed.
- Intelligent tutoring matches human tutors in rigorous reviews (d ≈ 0.79 vs 0.76); preference matching has no comparable evidence base.
- Completion, satisfaction and session time are consumption metrics. To claim learning, demand a capability observation.
- In any demo, ask the three checklist questions. The answer to the third usually ends the meeting.
The follow-up quiz arrives in three days.
Spaced retrieval is the best-evidenced way to keep what you just read. Leave an email and you will get three recall questions on this essay in three days, plus the next essay when it publishes. That is the whole mechanism; unsubscribe any time.
References
Primary and peer reviewed sources only. Links go to publisher pages.
- Gomez-Uribe, C. A., & Hunt, N. (2016). The Netflix Recommender System: Algorithms, Business Value, and Innovation. ACM Transactions on Management Information Systems, 6(4). dl.acm.org/doi/10.1145/2843948
- VanLehn, K. (2011). The Relative Effectiveness of Human Tutoring, Intelligent Tutoring Systems, and Other Tutoring Systems. Educational Psychologist, 46(4), 197–221. tandfonline.com/doi/abs/10.1080/00461520.2011.611369
- Kulik, J. A., & Fletcher, J. D. (2016). Effectiveness of Intelligent Tutoring Systems: A Meta-Analytic Review. Review of Educational Research, 86(1), 42–78. journals.sagepub.com/doi/10.3102/0034654315581420
- Pashler, H., McDaniel, M., Rohrer, D., & Bjork, R. (2008). Learning Styles: Concepts and Evidence. Psychological Science in the Public Interest, 9(3), 105–119. journals.sagepub.com/doi/10.1111/j.1539-6053.2009.01038.x
- Corbett, A. T., & Anderson, J. R. (1994). Knowledge tracing: Modeling the acquisition of procedural knowledge. User Modeling and User-Adapted Interaction, 4, 253–278. link.springer.com/article/10.1007/BF01099821
- Bjork, E. L., & Bjork, R. A. (2011). Making things hard on yourself, but in a good way: Creating desirable difficulties to enhance learning. In Psychology and the Real World. Worth Publishers. bjorklab.psych.ucla.edu/research
- LinkedIn Learning (2025). Workplace Learning Report 2025. learning.linkedin.com/resources/workplace-learning-report
- Jordan, K. (2014). Initial trends in enrolment and completion of massive open online courses. The International Review of Research in Open and Distributed Learning, 15(1). irrodl.org/index.php/irrodl/article/view/1651