MRBench was compiled from two previous mistake remidiation datasets: Bridge and MathDial. Bridge dataset consists of dialogues between students at the elementary level and real human tutors. MathDial contains middle-school-level interactions between a real human tutor and LLM acting as student.
Both datasets were prepared to only maintain dialogues ending in a student mistake or confusion. Based on the conversation history and last utterance, a next response was generated using 7 LLMs, including: GPT-4, Gemini, Sonnet, Llama-3.1-8B, Llama-3.1-40.5B, and Phi3.
Then, trained annotators annotated each human tutor and LLM response using a validated taxonomy representing 8 pedegogical dimensions: mistake identification, mistake location, revealing of the answer, providing guidance, actionability, coherence, tone, and human-likeness. Each dimension was evaluated on a 3-item scale of yes, to some extent, or no.
In addition, two LLMs were used as evalutors: Prometheus2 and Llama-3.1-8B to evaluate the similarity of human and LLM-based response evaluation.
MRBench contains three versions. Version 1 and Version 2 were rated across all eight pedagogical dimensions. Version 3 was rated on only four dimensions.
For additional dataset info and data schema, see the dataset documentation on Github.
Appropriate uses and limitations
LLMs are evolving and improving rapidly. LLM responses on this dataset, and therefore the pedagogical ratings, may go out of date quickly. It may be interesting to update this dataset with new LLM responses to the same tutoring context.
This data should not be used to attempt to identify the children in this study.