Images were collected on the ASSISTments online learning platform, where students receive feedback form teachers on assigned work. The problems were drawn from Eureka Math, Open Up Resources, and Illustrative Math.
Opt-out was available
No
Data processing
The original data contained 60,000 images across 18 problems. A random sample of 15 images were selected for each problem for inclusion in the dataset. Images were cropped to ensure no personally identifying information was visible.
Three NYC-based expert math teachers were recruited to annotate the images. During round one, teachers described each image as thoroughly as possible. During a second phase, additional teachers augmented and revised the annotations as well as adding up to five question-answer pairs about the student's work (e.g. Q: "Did students label the number line correctly?", A: "The student labeled the number line correctly."), resulting in 11,661 human-created QA pairs. Synthetic question-answer pairs were also generated by prompting Claude-3.5 Sonnet and GPT-4o to decompose the human annotation and create close-ended question-answer pairs, resulting in 44,362 synthetic QA pairs.
Images are stored as a url in the dataset. This is useful for feeding to an LLM. If you want the images directly, you'll first need to query the urls in the "Image URL" column.
The dataset disk size shown does not include the images.
Appropriate uses and limitations
This dataset does not contain metadata on grade level or student age connected to each problem. Students' handwriting may change as they age, creating a challenge for model interpretation.
DrawEduMath: Evaluating Vision Language Models with Expert-Annotated Students’ Hand-Drawn Math Images
Sami Baral1
,
Li Lucy2
,
Ryan Knight3
,
Alice Ng4
,
Luca Soldaini5
,
Neil Heffernan1
,
Kyle Lo5
1Heffernan-WPI-Lab / Worcester Polytechnic Institute
2University of California Berkeley
3Insource Services
4Teaching LAB
5Allen Institute for Artificial Intelligence
Organization
Insource Services, Teaching Lab, University of California Berkeley, Allen Institute for Artificial Intelligence, Heffernan-WPI-Lab / Worcester Polytechnic Institute
Authors
Li Lucy, Ryan Knight, Alice Ng, Sami Baral, Luca Soldaini, Kyle Lo, Neil Heffernan