We are pleased to present to the Arab research community a dataset that is the largest and most important of its kind in terms of size, tagging, and diversity. This dataset supports the automated essay scoring of high school (and university) students writing in Arabic. The new dataset, named “LAILA”, comprises eight persuasive and expository writing tasks, encompassing 7,859 essays collected from 24 schools in Qatar. These essays were written by 4,372 high school students and tagged by nine Arabic language teachers based on seven different writing attributes: relevance, development of ideas and content, style and structural cohesion, grammar, mechanics, vocabulary, and organization. Holistic evaluation was also included. This entire process took nearly a year, from preparation and ethical approvals to compilation, tagging, and documentation.
What we are proud of and delighted about is that “LAILA” rivals, in size, tagging, and quality, the standard corpora widely used in research in this field.
We pray and hope that making “LAILA” available will contribute to supporting research and researchers (especially Arabs) in developing Arabic systems in this research area, which is unfortunately rare due to the lack of readily available corpora, despite research on them in other languages having begun more than 55 years ago, and their use in international English language tests for over twenty years.
For more details about “Layla,” you can read a (preliminary version) of the research paper that was accepted today (praise be to God) at the EACL 2026 conference, one of the largest conferences in the field of Natural Language Processing (NLP): https://aclanthology.org/2026.eacl-long.142.
The corpora is fully available to researchers via the GitLab link: https://gitlab.com/bigirqu/laila.
This work This project was completed—through a remarkable year of effort—by a large team of researchers from Qatar University and Carnegie Mellon University in Qatar, led by Ms. May Bashendy and Dr. Walid Massoud, with Sohaila Eltanbouly, Salam Albatarni, Marwan Sayed, Abrar Abir, and Dr. Houda Bouamor.
Special thanks to the research funding provided by the Qatar Research, Development and Innovation Council (QRDI), and to the wonderful collaboration of the Qatari Ministry of Education and Higher Education.
*The dataset is named “LAILA” because we have a previous blog (albeit much smaller) called “QAES” 🙂

Leave a Reply