Seham Basabain S.Basabain@napier.ac.uk
Research Student
Arabic Short-text Dataset for Sentiment Analysis of Tourism and Leisure Events
Basabain, Seham; Al‐Dubai, Ahmed; Cambria, Erik; Alomar, Khalid; Hussain, Amir
Authors
Prof Ahmed Al-Dubai A.Al-Dubai@napier.ac.uk
Professor
Erik Cambria
Khalid Alomar
Prof Amir Hussain A.Hussain@napier.ac.uk
Professor
Abstract
The focus of this study is to present the detailed process of collecting a dataset of Arabic short-text in the tourism context and annotating this dataset for the task of sentiment analysis using an automatic zero-shot labelling technique utilising transformer-based models. This is benchmarked against a baseline manual annotation approach utilising native Arab human annotators. This study also introduces an approach exploiting both manual/handcrafted and automatically generated annotations of the dataset tweets for the task of sentiment analysis as part of a cross-domain approach using a model trained on sarcasm labels and vice versa. The total collected corpus size is 2293 tweets; after annotation, these tweets were labelled in a three-way classification approach as either positive, negative or neutral. We run different experiments to provide benchmark results of Arabic sentiment classification. Comparative results on our dataset show that the highest performing baseline model when utilising manual labels was MARBERT, with an accuracy of up to 87%, which was pre-trained for Arabic on a massive amount of data. It should be noted that this model enhanced its performance additionally after pre-training on a dialectical Arabic and modern standard Arabic corpus. On the other hand, zero-shot automatically generated labels achieved an 84% accuracy rate in predicting sarcasm classes from sentiment labels.
Citation
Basabain, S., Al‐Dubai, A., Cambria, E., Alomar, K., & Hussain, A. (2025). Arabic Short-text Dataset for Sentiment Analysis of Tourism and Leisure Events. Expert Systems, 42(5), Article e70030. https://doi.org/10.1111/exsy.70030
Journal Article Type | Article |
---|---|
Acceptance Date | Feb 10, 2025 |
Online Publication Date | Mar 22, 2025 |
Publication Date | 2025-05 |
Deposit Date | Feb 24, 2025 |
Publicly Available Date | Mar 25, 2025 |
Journal | Expert Systems |
Print ISSN | 0266-4720 |
Electronic ISSN | 1468-0394 |
Publisher | Wiley |
Peer Reviewed | Peer Reviewed |
Volume | 42 |
Issue | 5 |
Article Number | e70030 |
DOI | https://doi.org/10.1111/exsy.70030 |
Keywords | Arabic sentiment analysis, automatic labelling, Saudi tourism, twitter, zero-shot learning |
Public URL | http://researchrepository.napier.ac.uk/Output/4129503 |
Files
Arabic Short-text Dataset for Sentiment Analysis of Tourism and Leisure Events
(664 Kb)
PDF
Publisher Licence URL
http://creativecommons.org/licenses/by/4.0/