ExploraTutor: A Dataset for Children’s Exploratory Dialogue by Integrating Multiple Educational Theories
-
Abstract
To address the lack of pedagogical intelligence in general-purpose Large Language Models (LLMs) for guiding children's deep exploration, this paper introduces a pathway of “Theory—Practice—Data—Model”. The study first analyzes authentic child-adult dialogues to distill the “implicit knowledge” according to the Scaffolding Theory, the Inquiry-Based Learning, and the Schema Theory. This knowledge is developed into a systematic annotation framework featuring pedagogical strategies, schema development goals, and cognitive alignment levels. Leveraging this framework, we constructed the ExploraTutor dataset through a dual-pathway approach that combines the augmentation of real data with theory-guided synthesis, resulting in 2,045 high-quality dialogues totaling 18,275 question-answer pairs. Experiments on mainstream models such as Qwen, Deepseek, and ChatGLM demonstrate that models fine-tuned on this dataset significantly outperform their baselines in heuristic guidance and cognitive adaptability.
-
-