Yayın:
KFD: Selective Token Filtering and Adaptive Weighting for Efficient Knowledge Distillation

dc.contributor.authorYuce, Muzaffer Kaan
dc.contributor.authorAmasyali, Mehmet Fatih
dc.date.accessioned2026-06-27T15:36:41Z
dc.date.issued2026
dc.description.abstractKnowledge distillation (KD) transfers knowledge from large language models (LLMs) to smaller or similarly sized models in order to obtain efficient yet capable systems. However, performing distillation over all tokens is computationally expensive and may weaken the transfer signal. To address this limitation, Knowledge-Filtered Distillation (KFD) is introduced as a selective distillation approach in which tokens are filtered according to the divergence KL(M2 divided by M0) between a teacher model (M2) and a base model (M0), while the student model (M1) is also derived from the same base model. Only tokens whose divergence exceeds a predefined threshold are distilled. For the selected tokens, the teacher distribution is normalized over the Top-5 predictions, whereas tokens outside this case receive a label-ranking bonus. The proposed conditional Top-5/bonus target design is shown theoretically to yield a lower label-focused target error than using only Top-5 normalization or only the bonus across all tokens. In addition, the KL and cross-entropy (CE) losses are balanced through a dynamically computed batch-level coefficient alpha. Experiments on multiple Turkish text datasets show that KFD consistently outperforms CE-only training, achieving higher accuracy with less data and shorter training time. KFD also outperforms entropy-based token selection methods and highlights the role of student initialization in effective knowledge transfer, thereby providing an efficient and scalable distillation framework for teacher-student models of equal size.en
dc.description.sponsorshipScientific and Technological Research Council of Turkey (TUBITAK) [124E055]
dc.description.sponsorshipYildiz Technical University Scientific Research Projects Coordination Unit [FDK-2024-6421]
dc.description.urihttps://doi.org/10.3390/sym18040667
dc.identifier.doi10.3390/sym18040667
dc.identifier.eissn2073-8994
dc.identifier.issue4
dc.identifier.urihttps://hdl.handle.net/20.500.14981/71982
dc.identifier.volume18
dc.identifier.wos001750384300001
dc.language.isoeng
dc.publisherMDPI
dc.relation.ispartofSYMMETRY-BASEL
dc.rightsopenAccess
dc.subjectknowledge distillation
dc.subjecttoken-level filtering
dc.subjectselective knowledge transfer
dc.subjectTop-k filtration
dc.subjectKL divergence
dc.subjectdynamic weighting
dc.subjectbonus-based regularization
dc.subjectScience & Technology - Other Topics
dc.titleKFD: Selective Token Filtering and Adaptive Weighting for Efficient Knowledge Distillation
dc.typeArticle
dspace.entity.typePublication
local.import.sourceWOS

Dosyalar

Koleksiyonlar