Challenge

Trace the Ace

Predict the learning gains from a tutoring session measured by quiz performance in this tutoring outcomes prediction challenge

Tasks
Tutoring
Data types
Human transcript
Tabular
$50,000 in prizes
4 weeks left model submissions
7 weeks left solution write-up
639 joined
Image generated using Gemini

About the sponsor

The National Tutoring Observatory (NTO) partners with tutoring organizations, school districts, and other education providers to develop evidence-based insights into effective teaching strategies. Their mission is to improve teaching and learning at scale by studying great tutors. As part of this effort, NTO is currently developing the “Million Tutor Moves” dataset, which will aggregate at least one million interactions between students and tutors.

About the data

The data for this competition come from two online tutoring providers: Eedi and Third Space Learning.

Eedi

Eedi is an online learning platform for students aged 9–16. Students complete problem sets assigned by their teachers and are optionally able to begin a chat with a human tutor about a problem at any time. Both student and tutor interact by typing in a chat box. This results in short student-tutor sessions discussing a specific pain point.

Student correctness is measured using the next question attempted after the tutoring session concludes, rather than the question discussed in the tutoring session itself.

Third Space Learning

Third Space Learning (TSL) is an online tutoring platform for grade school students. Human tutors engage in longer, voice-based lessons with students covering pre-determined lessons and slides. This results in longer transcripts of student-tutor conversations. These interactions were originally spoken and later transcribed into text.

After each tutoring session, students complete assessment questions aligned to one or more learning objectives covered during the session. The outcome is whether the student correctly answered a conceptual review question.

De-identification process

To protect user privacy, all transcripts were rigorously anonymized. Rather than replacing personally identifiable information (PII) with obvious placeholder tags like [REDACTED], the names are replaced with contextually relevant, synthetic surrogates. Names present in the transcripts do not reflect actual PII. Names in the competition data are fictitious surrogates carefully injected to preserve the natural conversational flow of the tutoring sessions while ensuring the real identities of students and tutors remain completely protected.

Transcripts were anonymized using a fully local AI cascade framework. For details of the process, see:

Zhang, H., Zhou, Z., Vanacore, K., Ahtisham, B., & Kizilcec, R. F. (2026). Redact or keep? A fully local AI cascade for educational dialogue de-identification. arXiv preprint. https://doi.org/10.48550/arXiv.2606.18372


Additional Resources:

For background on knowledge tracing and predicting student outcomes from tutoring dialogues, see:

  • Ikram, F., Scarlatos, A., & Lan, A. (2025). Exploring LLMs for predicting tutor strategy and student outcomes in dialogues. In Proceedings of the 20th Workshop on Innovative Use of NLP for Building Educational Applications (pp. 765–779). Association for Computational Linguistics. https://doi.org/10.48550/arXiv.2507.06910
  • Scarlatos, A., Baker, R. S., & Lan, A. (2025). Exploring knowledge tracing in tutor-student dialogues using LLMs. In Proceedings of the 15th International Learning Analytics and Knowledge Conference (LAK 2025). https://doi.org/10.1145/3706468.3706501
  • Zhou, X., Zhang, Z., Xie, X., et al. (2025). Deep learning based knowledge tracing in intelligent tutoring systems. Scientific Reports, 15, 21395. https://doi.org/10.1038/s41598-025-07422-7