CoLearn: An Agentic Tutor that Learns its Learner in a Human--AI Co-Learning Loop

He, Wu, Zhang, Zhao, He and Li (King's College London) present CoLearn, an agentic tutor that keeps an evidence-grounded memory of each learner's mastery and misconceptions and uses it to choose the next question.
Ask this paper
Learner memory. Per-topic mastery is updated with a soft-evidence variant of Bayesian Knowledge Tracing in which the LLM serves as a continuous observation function.
Targeted questions. Question generation aims at the weakest topic and recurring misconceptions rather than drawing from a fixed item bank.
Evidence view. Live progress visualization and blind A/B comparison make the personalization visible and testable.
Results. Memory-conditioned questions are preferred over non-personalized ones 68 to 69% of the time in blind A/B tests, and in persona simulations the agent's belief converges toward hidden true mastery. Accepted to EMNLP 2026.
Abstract
Good tutoring adapts to the individual: it tracks what a learner knows, notices why they go wrong, and asks the next question that will help most. Most deployed tutoring tools instead serve fixed item banks and treat a wrong answer as a single bit of signal. We present CoLearn, an interactive, agentic tutor that supports an iterative tutoring loop: the learner practises, and the system builds an evidence-grounded memory of the learner's mastery and misconceptions. This memory is updated as evidence accumulates and is used to generate the next personalised question. CoLearn has three components: (i) a persistent learner-state memory that updates per-topic mastery with a soft-evidence variant of Bayesian Knowledge Tracing, where a large language model acts as a continuous observation function; (ii) adaptive question generation that targets the learner's weakest topic and recurring misconceptions; and (iii) an evidence view that makes personalisation visible and testable through live progress visualisation and blind A/B comparison. In blind A/B evaluation, questions conditioned on this memory are preferred over non-personalised ones 68-69% of the time, and in persona simulations with hidden ground-truth mastery the agent's belief converges toward the learner's true mastery.