At a glance
Authored by: Mohamed bin Zayed University of Artificial Intelligence
MRBench is a dataset of math tutoring dialogues with responses from seven LLMs. One subset contains conversations between human teachers and human students. The second subset contains conversations between human teachers and AI acting as student. Each LLM response is annotated across eight pedagogical dimensions: mistake identification, mistake location, revealing of the answer, providing guidance, actionability, coherence, tone, and human-likeness, providing a standard for evaluating LLMs as tutors.
This dataset is useful for studying LLMs as math tutors and evaluating the quality of their reponses.
Permissions and licensing
- License
- Creative Commons Attribution-ShareAlike 4.0 International
- Key restrictions
-
You may use this data for any purpose, even commercially. Any derivative work must give credit to the original source and be shared under the same license as the original.
How to access the data
This dataset is available to anyone. To access the data, head to the Data download tab and follow the instructions provided.