Yayın:
Semantic Similarity Based Filtering for Turkish Paraphrase Dataset Creation

dc.contributor.authorAlkurdi, Besher
dc.contributor.authorSarioglu, Hasan Yunus
dc.contributor.authorAmasyali, Mehmet Fatih
dc.date.accessioned2026-06-27T14:43:12Z
dc.date.issued2022
dc.description.abstractIn this study, we introduce a new method for creating paraphrase datasets from parallel bilingual corpora. We also introduce large paraphrase datasets created using this method. We utilize machine translation to create paraphrase datasets by translating the English phrases in Turkish-English parallel datasets to Turkish. Detailed pre-processing steps are applied to the text pairs. A sample from our translated datasets was annotated by native speakers for semantic similarity, and a model with the same task was chosen based on the correlation with human annotations. We then filtered the preprocessed and translated text pairs by semantic similarity calculated by the chosen model. Two pre-trained encoder-decoder architectures were fine-tuned on the datasets that we created. We present results asserting our data collection and filtering method's effectiveness.en
dc.description.sponsorshipScientific and Technological Research Council of Turkey (TUBITAK) [120E100]
dc.identifier.endpage127
dc.identifier.isbn978-1-959429-36-4
dc.identifier.startpage119
dc.identifier.urihttps://hdl.handle.net/20.500.14981/63910
dc.identifier.wos001511025100014
dc.language.isoeng
dc.publisherASSOC COMPUTATIONAL LINGUISTICS-ACL
dc.relation.conference5th International Conference on Natural Language and Speech Processing
dc.relation.ispartofPROCEEDINGS OF THE 5TH INTERNATIONAL CONFERENCE ON NATURAL LANGUAGE AND SPEECH PROCESSING, ICNLSP 2022
dc.subjectComputer Science
dc.titleSemantic Similarity Based Filtering for Turkish Paraphrase Dataset Creation
dc.typeProceedings Paper
dspace.entity.typePublication
local.import.sourceWOS

Dosyalar

Koleksiyonlar