The narrative texts used in this dataset are classic fairytales gathered from the Project Gutenberg website. A subset of the most popular fairytales were selected for inclusion.
Opt-out was available
Unknown or not applicable
Data processing
Text were cleaned to replace outdated vocabulary. Each story was evaluated for its reading dificulting using the textstat Python package. Stories with levels above10th grade were excluded.
Stories were chunked by human annotators into 100-300 word sections containing meaningful content and natural story breaks. Often this resulted in a single paragraph of the original story, though sometimes more than one paragraph was included in a chunk.
Five human annotators developed QA pairs for each chunk of the story. They were instructed to only generate open-ended questions and avoid any yes/no questions. Additionally, they were asked to create diverse questions covering seven narrative elements with both explicit and implicit answers.
9The Hong Kong University of Science and Technology
10Georgia Institute of Technology
11UCLA
12Columbia University
Organization
University of Washington, IBM Research, Tencent/WeChat AI, University of Notre Dame, Syracuse University, IBM Research Ireland, HKUST, University of California, Irvine, UCLA, Rensselaer Polytechnic Institute, Georgia Institute of Technology, Columbia University
Authors
Mark Warschauer, Ying Xu, Dakuo Wang, Mo Yu, Daniel Ritchie, Tongshuang Wu, Zheng Zhang, Toby Jia-Jun Li, Nora Bradford, Branda Sun, Tran Bao Hoang, Bingsheng Yao, Yisi Sang, Yufang Hou, Xiaojuan Ma, Diyi Yang, Nanyun Peng, Zhou Yu