Yayın: A cyclical loss-based optimization algorithm for pretraining LLMs on noisy data
| dc.contributor.author | Kesgin, H. Toprak | |
| dc.contributor.author | Amasyali, M. Fatih | |
| dc.date.accessioned | 2026-06-27T15:19:33Z | |
| dc.date.issued | 2025 | |
| dc.description.abstract | Large language models (LLMs) depend on vast web-scale datasets, which frequently include noisy or low-quality samples that degrade performance and fairness-despite conventional data cleaning. This paper introduces an in-training filtering approach that selectively ignores noisy data points based on real-time loss statistics during training. The approach combines deterministic and probabilistic selection mechanisms using robust loss-based metrics and cyclically adjusted thresholds to balance stability and diversity. Evaluations on Turkish-language datasets demonstrate that this strategy reduces validation loss and improves downstream task accuracy without any preprocessing. By integrating filtering directly into the training loop, the method maintains data diversity, requires minimal overhead, and improves learning efficiency-offering a scalable alternative for robust LLM pre-training in noisy or low-resource environments. | en |
| dc.description.sponsorship | Scientific and Technological Research Council of Turkey (TUBITAK) Grant [124E055] | |
| dc.description.sponsorship | Yildiz Technical University Scientific Research Projects Coordination Unit [FDK-2024-6422] | |
| dc.description.uri | https://doi.org/10.1016/j.knosys.2025.114189 | |
| dc.identifier.doi | 10.1016/j.knosys.2025.114189 | |
| dc.identifier.eissn | 1872-7409 | |
| dc.identifier.issn | 0950-7051 | |
| dc.identifier.uri | https://hdl.handle.net/20.500.14981/69746 | |
| dc.identifier.volume | 328 | |
| dc.identifier.wos | 001554098100001 | |
| dc.language.iso | eng | |
| dc.publisher | ELSEVIER | |
| dc.relation.ispartof | KNOWLEDGE-BASED SYSTEMS | |
| dc.subject | Large language models | |
| dc.subject | LLM pretraining | |
| dc.subject | Noisy data | |
| dc.subject | Loss-based filtering | |
| dc.subject | Cyclical scheduling | |
| dc.subject | Adaptive optimization | |
| dc.subject | In-training filtering | |
| dc.subject | Data selection | |
| dc.subject | Web-scale datasets | |
| dc.subject | Low-resource languages | |
| dc.subject | Computer Science | |
| dc.title | A cyclical loss-based optimization algorithm for pretraining LLMs on noisy data | |
| dc.type | Article | |
| dspace.entity.type | Publication | |
| local.import.source | WOS |