AmQA: Amharic Question Answering Dataset

Tilahun Abedissa; Ricardo Usbeck; Yaregal Assabie

doi:10.48550/arXiv.2303.03290

AmQA: Amharic Question Answering Dataset

Research output: Contributions to collected editions/works › Article in conference proceedings › Research

Standard

AmQA: Amharic Question Answering Dataset. / Abedissa, Tilahun; Usbeck, Ricardo; Assabie, Yaregal.
Conference XXX. 2023.

Research output: Contributions to collected editions/works › Article in conference proceedings › Research

Harvard

Abedissa, T, Usbeck, R & Assabie, Y 2023, AmQA: Amharic Question Answering Dataset. in Conference XXX. https://doi.org/10.48550/arXiv.2303.03290

APA

Abedissa, T., Usbeck, R., & Assabie, Y. (2023). AmQA: Amharic Question Answering Dataset. Manuscript in preparation. In Conference XXX https://doi.org/10.48550/arXiv.2303.03290

Vancouver

Abedissa T, Usbeck R, Assabie Y. AmQA: Amharic Question Answering Dataset. In Conference XXX. 2023 doi: 10.48550/arXiv.2303.03290

Bibtex

@inbook{5abb1ef212ab4e4ba203ca5560e52c5e,

title = "AmQA: Amharic Question Answering Dataset",

abstract = " Question Answering (QA) returns concise answers or answer lists from natural language text given a context document. Many resources go into curating QA datasets to advance robust models' development. There is a surge of QA datasets for languages like English, however, this is not true for Amharic. Amharic, the official language of Ethiopia, is the second most spoken Semitic language in the world. There is no published or publicly available Amharic QA dataset. Hence, to foster the research in Amharic QA, we present the first Amharic QA (AmQA) dataset. We crowdsourced 2628 question-answer pairs over 378 Wikipedia articles. Additionally, we run an XLMR Large-based baseline model to spark open-domain QA research interest. The best-performing baseline achieves an F-score of 69.58 and 71.74 in reader-retriever QA and reading comprehension settings respectively. ",

keywords = "cs.CL, cs.AI, cs.IR, Informatics",

author = "Tilahun Abedissa and Ricardo Usbeck and Yaregal Assabie",

year = "2023",

month = mar,

day = "6",

doi = "10.48550/arXiv.2303.03290",

language = "English",

booktitle = "Conference XXX",

}

RIS

TY - CHAP

T1 - AmQA

T2 - Amharic Question Answering Dataset

AU - Abedissa, Tilahun

AU - Usbeck, Ricardo

AU - Assabie, Yaregal

PY - 2023/3/6

Y1 - 2023/3/6

N2 - Question Answering (QA) returns concise answers or answer lists from natural language text given a context document. Many resources go into curating QA datasets to advance robust models' development. There is a surge of QA datasets for languages like English, however, this is not true for Amharic. Amharic, the official language of Ethiopia, is the second most spoken Semitic language in the world. There is no published or publicly available Amharic QA dataset. Hence, to foster the research in Amharic QA, we present the first Amharic QA (AmQA) dataset. We crowdsourced 2628 question-answer pairs over 378 Wikipedia articles. Additionally, we run an XLMR Large-based baseline model to spark open-domain QA research interest. The best-performing baseline achieves an F-score of 69.58 and 71.74 in reader-retriever QA and reading comprehension settings respectively.

AB - Question Answering (QA) returns concise answers or answer lists from natural language text given a context document. Many resources go into curating QA datasets to advance robust models' development. There is a surge of QA datasets for languages like English, however, this is not true for Amharic. Amharic, the official language of Ethiopia, is the second most spoken Semitic language in the world. There is no published or publicly available Amharic QA dataset. Hence, to foster the research in Amharic QA, we present the first Amharic QA (AmQA) dataset. We crowdsourced 2628 question-answer pairs over 378 Wikipedia articles. Additionally, we run an XLMR Large-based baseline model to spark open-domain QA research interest. The best-performing baseline achieves an F-score of 69.58 and 71.74 in reader-retriever QA and reading comprehension settings respectively.

KW - cs.CL

KW - cs.AI

KW - cs.IR

KW - Informatics

U2 - 10.48550/arXiv.2303.03290

DO - 10.48550/arXiv.2303.03290

M3 - Article in conference proceedings

BT - Conference XXX

ER -

Other publications by the same author(s)

ShortPathQA: A Dataset for Controllable Fusion of Large Language Models with Knowledge Graphs

Salnikov, M., Sakhovskiy, A., Nikishina, I., Usmanova, A., Kraft, A., Möller, C., Banerjee, D., Huang, J., Jiang, L., Abdullah, R., Yan, X., Tutubalina, E., Usbeck, R. & Panchenko, A., 2026, Natural Language Processing and Information Systems: 30th International Conference on Applications of Natural Language to Information Systems, NLDB 2025, Proceedings. Ichise, R. (ed.). Springer Science and Business Media Deutschland, p. 95-110 16 p. (Lecture Notes in Computer Science; vol. 15836 LNCS).

Research output: Contributions to collected editions/works › Article in conference proceedings › Research › peer-review

Analyzing the Influence of Knowledge Graph Information on Relation Extraction.

Möller, C. & Usbeck, R., 2025

Research output: other publications › Other › Research

Analyzing the Influence of Knowledge Graph Information on Relation Extraction

Möller, C. & Usbeck, R., 2025, The Semantic Web: 22nd European Semantic Web Conference, ESWC 2025 Portoroz, Slovenia, June 1–5, 2025 Proceedings, Part I. Curry, E., Acosta, M., Poveda-Villalón, M., van Erp, M., Ojo, A., Hose, K., Shimizu, C. & Lisena, P. (eds.). Cham: Springer Nature Switzerland AG, Vol. 1. p. 460-480 21 p. (Lecture Notes in Computer Science ; vol. 15718).

Research output: Contributions to collected editions/works › Article in conference proceedings › Research › peer-review

ASK-DBLP: Answering Questions over DBLP

Taffa, T., Neises, P., Ollinger, S., Westphal, P., Ackermann, M. R., Banerjee, D. & Usbeck, R., 02.11.2025, ISWC-C 2025, Industry, Doctoral Consortium, Posters and Demos at ISWC 2025: Joint Proceedings of Industry, Doctoral Consortium, Posters and Demos of the 24th International Semantic Web Conference (ISWC-C 2025), ISWC 2025 Companion Volume. Celino, I., Hassanzadeh, O., Bernstein, A., Noy, N., Cheng, G., Wang, S., Ferrada, S., Soulard, T., Kozaki, K., Takeda, H. & Gentile, A. L. (eds.). Aachen: Sun Site Central Europe (RWTH Aachen University), p. 435-440 6 p. D13. (CEUR Workshop Proceedings; vol. 4085).

Research output: Contributions to collected editions/works › Article in conference proceedings › Research › peer-review

Automating SPARQL Query Translations between DBpedia and Wikidata

Bartels, M. C., Banerjee, D. & Usbeck, R., 14.07.2025, Linking Meaning: Semantic Technologies Shaping the Future of AI: Cover 74617 Proceedings of the 21st International Conference on Semantic Systems, 3-5 September 2025, Vienna, Austria. Spahiu, B., Vahdati, S., Salatino, A., Pellegrini, T. & Havur, G. (eds.). IOS Press BV, p. 176-193 18 p. (Studies on the Semantic Web; vol. 62).

Research output: Contributions to collected editions/works › Article in conference proceedings › Research

DOI

https://doi.org/10.48550/arXiv.2303.03290
Submitted manuscript