高级检索

ExploraTutor:融合多元教育理论的儿童启发式对话数据集

ExploraTutor: A Dataset for Children’s Exploratory Dialogue by Integrating Multiple Educational Theories

  • 摘要: 为解决通用大语言模型缺乏引导儿童深度探索所需教育智慧的问题,该文提出了一条“理论—实践—数据—模型”的有效赋能路径。研究首先深入分析真实儿童-成人对话,从中提炼出支架理论、探究式学习与图式理论在实践中应用的“隐性知识”,并将其抽象为一个包含回应策略、图式发展目标与认知对齐水平的系统性标注框架。依托该框架,通过真实数据增强与理论指导生成相结合的双路径方法,构建了包含2 045条高质量对话,共18275个问答对的儿童启发式微调数据集ExploraTutor。在Qwen、DeepSeek、ChatGLM等主流模型上的实验证明,经该数据集微调的模型,其启发引导能力和认知适应性均显著超越基线模型,成功地将蕴含实践智慧的教育理论内化为模型的核心能力,使其从“知识问答机”转变为“认知启发者”。

     

    Abstract: To address the lack of pedagogical intelligence in general-purpose Large Language Models (LLMs) for guiding children's deep exploration, this paper introduces a pathway of “Theory—Practice—Data—Model”. The study first analyzes authentic child-adult dialogues to distill the “implicit knowledge” according to the Scaffolding Theory, the Inquiry-Based Learning, and the Schema Theory. This knowledge is developed into a systematic annotation framework featuring pedagogical strategies, schema development goals, and cognitive alignment levels. Leveraging this framework, we constructed the ExploraTutor dataset through a dual-pathway approach that combines the augmentation of real data with theory-guided synthesis, resulting in 2,045 high-quality dialogues totaling 18,275 question-answer pairs. Experiments on mainstream models such as Qwen, Deepseek, and ChatGLM demonstrate that models fine-tuned on this dataset significantly outperform their baselines in heuristic guidance and cognitive adaptability.

     

/

返回文章
返回