ParaStudent: Closing the Sim2Real Gap in User Simulators for AI Tutor Evaluation
arXiv:2507.12674v3 Announce Type: replace-cross Abstract: Evaluating Artificial Intelligence (AI) tutor feedback before deployment requires anticipating student engagement, typically assessed through real interaction data.
We introduce ParaStudent, a fine-tuning framework for simulating novice programming revisions to support AI tutor evaluation.
Compared with prompted baselines, ParaStudent's revisions more closely match real student code distributions across functional, stylistic, and semantic metrics. Our best variant achieves AUCs of 0.80 for both feedback relevance and successful uptake when distinguishing streams with real engagement above versus at or below the median, while prompted baselines remain near chance on successful uptake.
These findings demonstrate the promise of simulated engagement for pre-deployment feedback triage.