Graded essays were collected from a K-12 Education Organization.
Opt-out was available
No
Data processing
Images, questions, and essays were retained from the original dataset. Essays were removed that were incomplete, low quality, or that didn't meet the topic criteria for reliability and diversity.
Then, two experienced experts in English education assessed each essay for all 10 traits. If the difference between the two scorers was less than 1, the average of the two scores was taken as the ground-truth score. When the difference score exceeded 1, another team of three senior annotators reviewed the essays and came to a consensus on the final ground-truth score.
For additional methods and data processing descriptions, see the dataset publication.
Tips for using the data
Technical tips
Scripts to automate LLM evaluation are provided in the GitHub code/ directory.
Appropriate uses and limitations
This dataset only contains scores for lexical quality and does not include the overall grade on the response. Essays were written by english language learners.
1The Hong Kong University of Science and Technology
2Tsinghua University
3Guangxi Zhuang Autonomous Region Big Data Research Institute
Organization
Beijing Future Brain Education and Technology, Guangxi Zhuang Autonomous Region Big Data Research Institute, Tsinghua University, The Hong Kong University of Science and Technology
Authors
Jiamin Su, Yibo Yan, Fangteng Fu, Han Zhang, Jingheng Ye, Xiang Liu, Jiahao Huo, Xuming Hu, Huiyu Zhou