Yayın: Semantic Similarity Based Filtering for Turkish Paraphrase Dataset Creation
Yükleniyor...
Tarih
Danışman
item.page.editor
Editör
Bölüm / Program
Dergi Başlığı
Dergi ISSN
Cilt Başlığı
Yayıncı
ASSOC COMPUTATIONAL LINGUISTICS-ACL
DOI
Özet
In this study, we introduce a new method for creating paraphrase datasets from parallel bilingual corpora. We also introduce large paraphrase datasets created using this method. We utilize machine translation to create paraphrase datasets by translating the English phrases in Turkish-English parallel datasets to Turkish. Detailed pre-processing steps are applied to the text pairs. A sample from our translated datasets was annotated by native speakers for semantic similarity, and a model with the same task was chosen based on the correlation with human annotations. We then filtered the preprocessed and translated text pairs by semantic similarity calculated by the chosen model. Two pre-trained encoder-decoder architectures were fine-tuned on the datasets that we created. We present results asserting our data collection and filtering method's effectiveness.
Tanım
Dergi veya Seri
PROCEEDINGS OF THE 5TH INTERNATIONAL CONFERENCE ON NATURAL LANGUAGE AND SPEECH PROCESSING, ICNLSP 2022
ISSN
ISBN
978-1-959429-36-4