
TutorMoments is a framework that assesses how well language models balance support and independent thinking in math tutoring. It uses real data from U.S. students in grades 2-7, analyzing key moments in tutoring sessions. Models often provide too much support, leaving students without opportunities for deeper thinking, and improved prompts help, but they still fall short of human-level performance. The dataset includes de-identified transcripts and code for reproducibility, and the framework evaluates scaffolding, rigor, and over-scaffolding. Preliminary results show that evaluation-aware prompts improve performance, but models vary in their interpretation of these prompts. The dataset focuses on U.S. elementary and middle-school math, and the findings may not generalize to other subjects or grade

