Yayın:
Audio fingerprinting using wavelet transform

dc.contributor.authorKanalıcı, Evren
dc.date.accessioned2022-08-09T11:30:15Z
dc.date.accessioned2026-06-21T05:18:31Z
dc.date.available2022-08-09T11:30:15Z
dc.date.issued2019
dc.descriptionTez (Yüksek Lisans) - Yıldız Teknik Üniversitesi, Fen Bilimleri Enstitüsü, 2019en_US
dc.description.abstractAudio fingerprinting systems have many real-world use-cases such as digital rights management/copyright detection, duplicated audio detection, untagged audio labelling or identify/query-by-example recognition systems. Nowadays, there are popular online platforms that offer identify/query-by-example music recognition services where users can query by snippets of recorded audio to retrieve the matched song metadata. The compact, robust and fast retrieving fingerprint design is the cornerstone of these systems. Although short-term Fourier transform and Mel-spectral representations are common tools that come to mind, these feature extraction methods suffer from being unstable and having somehow limited resolution. In order to overcome these challenges, scattering wavelet transform (SWT) provides an alternative solution to these limitations by recovering information loss, while ensuring translation invariance and stability. In this study, a two-stage audio fingerprint characteristic/feature extraction framework is introduced using SWT integrated with Siamese neural network hashing model for musical audio identification. Similarity-preserving hashes provided by the Siamese neural network model correspond to sound fingerprints and can be defined by a similarity distance metric in the embedded hashing space. The Siamese neural network hashing model was trained by two-layer scattering wavelet transform coefficients using relatively aligned segments of the same music files and segments of different music files. The proposed system achieves successful performance scores under environmental noise, modeling the challenges of detecting music and audio data that may be encountered in everyday life. Using very compact storage, it has been shown to achieve high ROC-AUC scores both by one-to-one comparison and by using locality-sensitive hashing (LSH) for content storage.en_US
dc.identifier.urihttps://hdl.handle.net/20.500.14981/12955
dc.language.isoenen_US
dc.subjectAudio fingerprintingen_US
dc.subjectMusic information retrievalen_US
dc.subjectWavelet transformen_US
dc.subjectScattering wavelet transformen_US
dc.subjectSiamese neural networksen_US
dc.titleAudio fingerprinting using wavelet transformen_US
dc.typemasterThesisen_US
dspace.entity.typePublication

Dosyalar

Orijinal paket

Şimdi gösterimde1 - 1 of 1
Yükleniyor...
Küçük Resim
Adı:
0098.pdf
Boyut:
4,16 MB
Format:
Adobe Portable Document Format

Koleksiyonlar