Yayın:
A New Reward System Based on Human Demonstrations for Hard Exploration Games

dc.contributor.authorTareq, Wadhah Zeyad
dc.contributor.authorAmasyali, Mehmet Fatih
dc.date.accessioned2026-06-27T14:35:48Z
dc.date.issued2022
dc.description.abstractThe main idea of reinforcement learning is evaluating the chosen action depending on the current reward. According to this concept, many algorithms achieved proper performance on classic Atari 2600 games. The main challenge is when the reward is sparse or missing. Such environments are complex exploration environments like Montezuma's Revenge, Pitfall, and Private Eye games. Approaches built to deal with such challenges were very demanding. This work introduced a different reward system that enables the simple classical algorithm to learn fast and achieve high performance in hard exploration environments. Moreover, we added some simple enhancements to several hyperparameters, such as the number of actions and the sam-pling ratio that helped improve performance. We include the extra reward within the human demonstrations. After that, we used Prioritized Double Deep Q-Networks (Prioritized DDQN) to learning from these demonstra-tions. Our approach enabled the Prioritized DDQN with a short learning time to finish the first level of Montezuma's Revenge game and to perform well in both Pitfall and Private Eye. We used the same games to compare our results with several baselines, such as the Rainbow and Deep Q-learning from demonstrations (DQfD) algorithm. The results showed that the new rewards system enabled Prioritized DDQN to out-perform the baselines in the hard exploration games with short learning time.en
dc.description.urihttps://doi.org/10.32604/cmc.2022.020036
dc.identifier.doi10.32604/cmc.2022.020036
dc.identifier.eissn1546-2226
dc.identifier.endpage2414
dc.identifier.issn1546-2218
dc.identifier.issue2
dc.identifier.startpage2401
dc.identifier.urihttps://hdl.handle.net/20.500.14981/62483
dc.identifier.volume70
dc.identifier.wos000705060700017
dc.language.isoeng
dc.publisherTECH SCIENCE PRESS
dc.relation.ispartofCMC-COMPUTERS MATERIALS & CONTINUA
dc.rightsopenAccess
dc.subjectDeep reinforcement learning
dc.subjecthuman demonstrations
dc.subjectprioritized double deep q-networks
dc.subjectatari
dc.subjectComputer Science
dc.subjectMaterials Science
dc.titleA New Reward System Based on Human Demonstrations for Hard Exploration Games
dc.typeArticle
dspace.entity.typePublication
local.import.sourceWOS

Dosyalar

Koleksiyonlar