Yayın: KFPT: Reliability and uncertainty filtered self-distillation for language model training
| dc.contributor.author | Yuce, Muzaffer Kaan | |
| dc.contributor.author | Amasyali, Mehmet Fatih | |
| dc.date.accessioned | 2026-06-27T15:30:24Z | |
| dc.date.issued | 2026 | |
| dc.description.abstract | Transformer-based large language models are typically trained on large text corpora using the next-token cross-entropy (CE) objective. Although CE is scalable and stable, in practice it can exhibit limitations such as overconfidence, weak learning signals on hard/rare tokens, and a mismatch between the training objective and generation-time behavior. In this work, we propose Knowledge-Filtered Phase Training (KFPT), a two-phase scheme that strengthens the training signal without requiring an additional teacher/model. In the first phase, KFPT augments CE with a selective regularization term (RU) and, at fixed intervals, performs a second forward pass on the same text by masking small blocks in the attention mask, averaging the CE losses to make updates more stable. In the second phase, KFPT adds a one-way KL-consistency term by taking the distribution from a span-drop-induced second view as the target; this term is selectively weighted and strengthened only at useful positions based on the reference view's reliability (gold-margin and correctness) and the student's uncertainty (entropy). We also analyze why the additional terms used in Phase 1 and Phase 2 can be effective through mathematical theorems and proofs. In comprehensive experiments, we compare KFPT across multiple model architectures and training regimes against a strong CE baseline and prior teacher-free objective-improvement methods. The results show that KFPT generally improves accuracy and reduces perplexity, outperforming teacher-free alternatives in the literature. | en |
| dc.description.sponsorship | Scientific and Technological Research Council of Turkey (TUBITAK) [124E055] | |
| dc.description.sponsorship | Yildiz Technical University Scientific Research Projects Coordination Unit [FDK-2024-6421] | |
| dc.description.uri | https://doi.org/10.1016/j.knosys.2026.115880 | |
| dc.identifier.doi | 10.1016/j.knosys.2026.115880 | |
| dc.identifier.eissn | 1872-7409 | |
| dc.identifier.issn | 0950-7051 | |
| dc.identifier.uri | https://hdl.handle.net/20.500.14981/71302 | |
| dc.identifier.volume | 343 | |
| dc.identifier.wos | 001745893800001 | |
| dc.language.iso | eng | |
| dc.publisher | ELSEVIER | |
| dc.relation.ispartof | KNOWLEDGE-BASED SYSTEMS | |
| dc.subject | Large language models | |
| dc.subject | KFPT | |
| dc.subject | Self-distillation | |
| dc.subject | Two-phase training | |
| dc.subject | Reliability weighting | |
| dc.subject | Uncertainty (entropy)-based filtering | |
| dc.subject | Continual pretraining | |
| dc.subject | Computer Science | |
| dc.title | KFPT: Reliability and uncertainty filtered self-distillation for language model training | |
| dc.type | Article | |
| dspace.entity.type | Publication | |
| local.import.source | WOS |