跳到正文
arXiv cs.CL· Clayton Cohn, Joyce Fonteles, Kirk Vanacore, Gianni Mazza, Candida Crawford, Tom Hooper, Gautam Biswas, Rene Kizilcec·· 2 天前AI 评分42

K-12 数学辅导中跨模型 LLM 共识不等于诊断学生失败模式的有效性

Agreement Is Not Validity: Cross-Model LLM Consensus in Diagnosing Student Failure Modes in K-12 Math Tutoring Dialogue

AI 导读

针对 K-12 数学辅导对话中五种学生失败模式的诊断,研究发现跨模型 LLM 共识(kappa = .755-.781)显著高于人机一致性(kappa = .524-.597)。这表明模型间的高一致性可能制造正确性的假象,不能替代对学习者解释有效性的独立证据。

来源:arXiv cs.CL · arxiv.org