Yayın:
Imbalanced generative sampling of training data for improving quality of machine learning model

dc.contributor.authorCoskun, Umut Can
dc.contributor.authorDogan, Kemal Mert
dc.contributor.authorGunpinar, Erkan
dc.date.accessioned2026-06-27T15:10:25Z
dc.date.issued2024
dc.description.abstractDesign exploration in engineering applications often requires a meticulous experimental or numerical study to evaluate performance ( Y) of each design, which may require great effort, time or resources. Reducing the number of these tests for finding a good design is of paramount importance in all engineering fields. This study aims at computing a machine learning (ML) model using less number of designs as training data. Uniform sampling (US) in the design space (based on predefined design parameters) to obtain a training data is a promising approach. We further extend this sampling concept to obtain designs in the design space by also employing the ML model. The designs are selected via two non -uniform (imbalanced) sampling methods (namely, height -based sampling - HBS and gradient -based sampling - GBS) while considering their Y and gradient, dY, values. These values are divided into uniform intervals, and we aim at equalizing the number of designs in the training data at each interval as much as possible. This can force designs to have minimum or maximum Y or dY values, which, in fact, lie on small portion of the design space, in general. Therefore, capturing designs from all design space portions can be enabled. Results of the proposed methods are compared against US along with two well studied non -uniform sampling strategies, Stratified Over Sampling (SOS) and Gaussian -Process Based Sampling (GPBS). To reliably investigate quality of ML models obtained using designs sampled via US, SOS, GPBS, HBS and GBS, we utilize standard test (known) functions (such as Easom and Beale ) as substitutes for engineering problems. According to the results presented, ML models using HBS and GBS have either better prediction accuracy or wider applicability compared to all other tested sampling methods.en
dc.description.sponsorshipScientific and Technological Research Council of Turkiye (TUBITAK) [2230037]
dc.description.urihttps://doi.org/10.1016/j.aei.2024.102631
dc.identifier.doi10.1016/j.aei.2024.102631
dc.identifier.eissn1873-5320
dc.identifier.issn1474-0346
dc.identifier.urihttps://hdl.handle.net/20.500.14981/68565
dc.identifier.volume62
dc.identifier.wos001252910400001
dc.language.isoeng
dc.publisherELSEVIER SCI LTD
dc.relation.ispartofADVANCED ENGINEERING INFORMATICS
dc.subjectImbalanced sampling
dc.subjectMachine learning
dc.subjectComputer-aided design
dc.subjectDesign exploration
dc.subjectTraining data
dc.subjectComputational fluid dynamics
dc.subjectDESIGN
dc.subjectOPTIMIZATION
dc.subjectPERFORMANCE
dc.subjectUNCERTAINTY
dc.subjectALGORITHM
dc.subjectSYSTEM
dc.subjectComputer Science
dc.subjectEngineering
dc.titleImbalanced generative sampling of training data for improving quality of machine learning model
dc.typeArticle
dspace.entity.typePublication
local.import.sourceWOS

Dosyalar

Koleksiyonlar