Abstract
Many online learning platforms still rely on a rigid, one-size-fits-all approach where content is delivered linearly and learners receive essentially the same experience. An adaptive e-learning path recommender is introduced in this work that combines probabilistic reinforcement learning with generative AI explanations. Our method represents rich latent learner states using modular probabilistic encoders based on cognitive, noncognitive, and contextual factors, leverages small data by synthesizing episodes, and employs a conservative offline RL algorithm to learn safe and effective policies. Rationales are produced by a retrieval-augmented generative module to achieve pedagogically aligned, context-aware explanations. Synthetic expansion and rigorous offline evaluation experiments indicate significant improvements in expected reward, policy safety (low KL to behavior policy), and consistency of outcomes. The framework paves the way for practical, explainable, and safe personalization at scale.