IDRISI-D: Arabic and English Datasets and Benchmarks for Location Mention Disambiguation over Disaster Microblogs

Suwaileh, Reem; Elsayed, Tamer; Imran, Muhammad

View/Open

2023.arabicnlp-1.14.pdf (895.6Kb)

Date

2023

Author

Suwaileh, Reem
Elsayed, Tamer
Imran, Muhammad

Metadata

Show full item record

Abstract

Extracting and disambiguating geolocation information from social media data enables effective disaster management, as it helps response authorities; for example, locating incidents for planning rescue activities and affected people for evacuation. Nevertheless, the dearth of resources and tools hinders the development and evaluation of Location Mention Disambiguation (LMD) models in the disaster management domain. Consequently, the LMD task is greatly understudied, especially for the low resource languages such as Arabic. To fill this gap, we introduce IDRISI-D, the largest to date English and the first Arabic public LMD datasets. Additionally, we introduce a modified hierarchical evaluation framework that offers a lenient and nuanced evaluation of LMD systems. We further benchmark IDRISI-D datasets using representative baselines and show the competitiveness of BERT-based models.

DOI/handle

http://dx.doi.org/10.18653/v1/2023.arabicnlp-1.14
http://hdl.handle.net/10576/60869

Collections

Computer Science & Engineering [‎2484‎ items ]