Deep Investigation of Cross-Language Plagiarism Detection Methods

Jeremy Ferrero,Laurent Besacier,Didier Schwab,Frederic Agnes
DOI: https://doi.org/10.48550/arXiv.1705.08828
2017-05-24
Abstract:This paper is a deep investigation of cross-language plagiarism detection methods on a new recently introduced open dataset, which contains parallel and comparable collections of documents with multiple characteristics (different genres, languages and sizes of texts). We investigate cross-language plagiarism detection methods for 6 language pairs on 2 granularities of text units in order to draw robust conclusions on the best methods while deeply analyzing correlations across document styles and languages.
Computation and Language
What problem does this paper attempt to address?