To read this content please select one of the options below:

A lexicon based approach for classifying Arabic multi-labeled text

Ismail Hmeidi (Jordan University of Science and Technology, Irbid, Jordan)
Mahmoud Al-Ayyoub (Jordan University of Science and Technology, Irbid, Jordan)
Nizar A. Mahyoub (Jordan University of Science and Technology, Irbid, Jordan)
Mohammed A. Shehab (Jordan University of Science and Technology, Irbid, Jordan)

International Journal of Web Information Systems

ISSN: 1744-0084

Article publication date: 7 November 2016

349

Abstract

Purpose

Multi-label Text Classification (MTC) is one of the most recent research trends in data mining and information retrieval domains because of many reasons such as the rapid growth of online data and the increasing tendency of internet users to be more comfortable with assigning multiple labels/tags to describe documents, emails, posts, etc. The dimensionality of labels makes MTC more difficult and challenging compared with traditional single-labeled text classification (TC). Because it is a natural extension of TC, several ways are proposed to benefit from the rich literature of TC through what is called problem transformation (PT) methods. Basically, PT methods transform the multi-label data into a single-label one that is suitable for traditional single-label classification algorithms. Another approach is to design novel classification algorithms customized for MTC. Over the past decade, several works have appeared on both approaches focusing mainly on the English language. This work aims to present an elaborate study of MTC of Arabic articles.

Design/methodology/approach

This paper presents a novel lexicon-based method for MTC, where the keywords that are most associated with each label are extracted from the training data along with a threshold that can later be used to determine whether each test document belongs to a certain label.

Findings

The experiments show that the presented approach outperforms the currently available approaches. Specifically, the results of our experiments show that the best accuracy obtained from existing approaches is only 18 per cent, whereas the accuracy of the presented lexicon-based approach can reach an accuracy level of 31 per cent.

Originality/value

Although there exist some tools that can be customized to address the MTC problem for Arabic text, their accuracies are very low when applied to Arabic articles. This paper presents a novel method for MTC. The experiments show that the presented approach outperforms the currently available approaches.

Keywords

Citation

Hmeidi, I., Al-Ayyoub, M., Mahyoub, N.A. and Shehab, M.A. (2016), "A lexicon based approach for classifying Arabic multi-labeled text", International Journal of Web Information Systems, Vol. 12 No. 4, pp. 504-532. https://doi.org/10.1108/IJWIS-01-2016-0002

Publisher

:

Emerald Group Publishing Limited

Copyright © 2016, Emerald Group Publishing Limited

Related articles