The Library
English-Arabic cross-language plagiarism detection
Tools
Alotaibi, Naif and Joy, Mike (2021) English-Arabic cross-language plagiarism detection. In: International Conference on Recent Advances in Natural Language Processing (RANLP 2021), Online, 1–3 Sep 2021. Published in: Proceedings of the International Conference on Recent Advances in Natural Language Processing (RANLP 2021) pp. 44-52. ISBN 9789544520724. doi:10.26615/978-954-452-072-4_006
An open access version can be found in:
Official URL: http://dx.doi.org/10.26615/978-954-452-072-4_006
Abstract
The advancement of the web and information technology has contributed to the rapid growth of digital libraries and automatic machine translation tools which easily translate texts from one language into another. These have increased the content accessible in different languages, which results in easily performing translated plagiarism, which are referred to as “cross-language plagiarism”. Recognition of plagiarism among texts in different languages is more challenging than identifying plagiarism within a corpus written in the same language. This paper proposes a new technique for enhancing English-Arabic cross-language plagiarism detection at the sentence level. This technique is based on semantic and syntactic feature extraction using word order, word embedding and word alignment with multilingual encoders. Those features, and their combination with different machine learning (ML) algorithms, are then used in order to aid the task of classifying sentences as either plagiarized or non-plagiarized. The proposed approach has been deployed and assessed using datasets presented at SemEval-2017. Analysis of experimental data demonstrates that utilizing extracted features and their combinations with various ML classifiers achieves promising results.
Item Type: | Conference Item (Paper) | ||||||
---|---|---|---|---|---|---|---|
Divisions: | Faculty of Science, Engineering and Medicine > Science > Computer Science | ||||||
Journal or Publication Title: | Proceedings of the International Conference on Recent Advances in Natural Language Processing (RANLP 2021) | ||||||
Publisher: | INCOMA Ltd | ||||||
ISBN: | 9789544520724 | ||||||
Book Title: | Proceedings of the Conference Recent Advances in Natural Language Processing - Deep Learning for Natural Language Processing Methods and Applications | ||||||
Official Date: | September 2021 | ||||||
Dates: |
|
||||||
Page Range: | pp. 44-52 | ||||||
DOI: | 10.26615/978-954-452-072-4_006 | ||||||
Status: | Peer Reviewed | ||||||
Publication Status: | Published | ||||||
Access rights to Published version: | Open Access (Creative Commons) | ||||||
Conference Paper Type: | Paper | ||||||
Title of Event: | International Conference on Recent Advances in Natural Language Processing (RANLP 2021) | ||||||
Type of Event: | Conference | ||||||
Location of Event: | Online | ||||||
Date(s) of Event: | 1–3 Sep 2021 | ||||||
Related URLs: | |||||||
Open Access Version: |
Request changes or add full text files to a record
Repository staff actions (login required)
View Item |