Tutoring chat transcripts were acquired from an online tutoring provider. The tutor interacts primarily through a chat interface, while the student uses either the chat or a voice recording.
Opt-out was available
Unknown or not applicable
Data processing
The chat transcripts were deidentified by replacing the student's and tutor's name with [STUDENT]/[TUTOR] respectively. The original transcripts were subsetted to excerpts where a tutor responds to a mistake. Specifically, excerpts were identified by looking for a sequence of message patterns in the chat: 1) the tutor sends a message containing a "?", 2) the student responds, 3) the tutor uses with a signaling expression (e.g. "That is incorrect" or "good try").
Each tutoring segment was then evaluated by an expert teacher to identify the type of error the student made, the stragegy needed to correct that error, and the intention of the correction. The expert tutor additionally provided a response for each mistake.
'c_id': Conversation ID.
'lesson_topic': Lesson Topic.
'c_h': Conversation History. The last conversation turn is the student's, where they make a mistake.
'c_r': The original tutor's response to the student.
'c_r_': The expert math teacher's response.
'e': The error type identified by the expert math teacher.
'z_what': The strategy used by the expert math teacher.
'z_why': The intention used by the expert math teacher.
The data are taken from 1st - 5th grade math tutoring lessons and may not be generalizable beyond that demographic. Some chat transcripts do not contain the question the student is solving and the dataset still includes annotations for those excerpts.