Benchmark

SAGE: Science Answer Grading & Evaluation

Compare models to automatically grade written responses to science questions

Tasks
Grading science/math
Grading writing
Data types
Tabular
Eleni Koureas on Unsplash, generated using Gemini

Overview

Recent advances in generative AI are opening up new possibilities for education, particularly when it comes to reducing the burden of time-consuming tasks. One area of growing interest is the assessment of student writing, which can be both labor-intensive and critical for understanding student learning. By helping teachers evaluate written responses more efficiently, AI-powered assessment tools could free up valuable time for individualized instruction, allowing educators to focus more attention on students who need it most and ultimately supporting better learning outcomes.

At the same time, assessing student writing remains a challenging task for AI. Student responses often express the same ideas in different ways, contain incomplete reasoning, or mix correct concepts with misconceptions. Accurately evaluating these responses requires models to reason about scientific concepts and relationships rather than simply matching keywords or phrases.

This benchmark evaluates how well different pre-trained models and modeling approaches can assess students' short-answer science responses. Given a question, a reference answer, and a students' response, models must determine whether the response is correct, partially correct, contradictory, irrelevant, or non-domain. By measuring performance on this task, the benchmark provides insight into current capabilities of AI systems for educational assessment.

How to use this benchmark

The benchmark leaderboard provides a snapshot of how well current AI models can evaluate student-written science responses. By comparing performance across a range of modeling approaches, it highlights both the progress and remaining challenges in automated assessment.

Educators, researchers, and developers can use the leaderboard to understand which models are most effective at reasoning about student explanations and distinguishing between different types of responses. Benchmark performance can serve as a useful reference point when deciding which pre-trained model to use:

  • As a baseline to build on in research to push performance forwards
  • In a real-world use cases to help inform human decision makers

This benchmark is not yet open for users to add new model submissions.

Benchmark task

This benchmark is based on the SemEval dataset, which contains short-answer science questions, reference answers, and graded student responses collected from an online tutoring system. Questions span a variety of science domains, including electricity, electronics, geology, chemistry, and experimental design.

The goal is to predict one of five assessment categories: correct, partially correct, contradictory, irrelevant, or non-domain.

An example question, reference answer, and graded student responses.

An example question, reference answer, and graded student responses.

Success on this task requires both understanding the scientific concepts being discussed, and accurately interpreting a student response.

References

Any work produced from this data or benchmark should site the original works: