We found a match
Your institution may have rights to this item. Sign in to continue.
- Title
Locating similar names through locality sensitive hashing and graph theory.
- Authors
Turrado García, Fernando; García Villalba, Luis Javier; Sandoval Orozco, Ana Lucila; Aranda Ruiz, Francisco Damián; Aguirre Juárez, Andrés; Kim, Tai-Hoon
- Abstract
Locality Sensitive Hashing is a known technique applied for finding similar texts and it has been applied to plagiarism detection, mirror pages identification or to identify the original source of a news article. In this paper we will show how can Locality Sensitive Hashing be applied to identify misspelled people names (name, middle name and last name) or near duplicates. In our case, and due to the short length of the texts, using two similarity functions (the Jaccard Similarity and the Full Damerau-Levenshtein Distance) for measuring the similarity of the names allowed us to obtain better results than using a single one. All the experimental work was made using the statistical software R and the libraries: textreuse and stringdist.
- Subjects
GRAPH theory; HASHING; STATISTICAL software; LIBRARY software; ATTRIBUTION of news
- Publication
Multimedia Tools & Applications, 2019, Vol 78, Issue 21, p29853
- ISSN
1380-7501
- Publication type
Article
- DOI
10.1007/s11042-018-6375-9