Yayın:
Speaker Diarization using Embedding Vectors

dc.contributor.authorToruk, Mesut
dc.contributor.authorBilgin, Gokhan
dc.contributor.authorSerbes, Ahmet
dc.date.accessioned2026-06-27T14:28:27Z
dc.date.issued2020
dc.description.abstractIn recent years, with the rapid increase of voice data, solutions are being sought extensively for examination and indexing in the field of speech processing. One of these solutions is the speaker diarization, which is used to examine speech records that include multi-speaker. The speaker diarization system basically splits the speech file into segments using the speech file's silence fields and examines the similarity between the segments. In this study, deep learning based embedding vectors are used for speaker representation. The vector embedding proposed for speaker representation is performed with x-vectors extracted using time-delayed deep neural network and d-vectors extracted using LSTM. Then the system is tested with xd-vectors consisting of the combination of these two vectors. As a result, the effect of representative embedding vectors on the performance of the diarization system is examined in the scope of this paper.en
dc.description.urihttps://doi.org/10.1109/siu49456.2020.9302162
dc.identifier.doi10.1109/siu49456.2020.9302162
dc.identifier.isbn978-1-7281-7206-4
dc.identifier.issn2165-0608
dc.identifier.urihttps://hdl.handle.net/20.500.14981/61013
dc.identifier.wos000653136100136
dc.language.isotur
dc.publisherIEEE
dc.relation.conference28th Signal Processing and Communications Applications Conference (SIU)
dc.relation.ispartof2020 28TH SIGNAL PROCESSING AND COMMUNICATIONS APPLICATIONS CONFERENCE (SIU)
dc.subjectSpeaker diarization
dc.subjectx-vector
dc.subjecttime-delay neural network
dc.subjectd-vector
dc.subjectLSTM
dc.subjectEngineering
dc.subjectTelecommunications
dc.titleSpeaker Diarization using Embedding Vectors
dc.typeProceedings Paper
dspace.entity.typePublication
local.import.sourceWOS

Dosyalar

Koleksiyonlar