Yayın:
Q-Learning with probability based action policy

dc.contributor.authorUgurlu, Ekin Su
dc.contributor.authorBiricik, Goksel
dc.date.accessioned2026-06-27T13:00:56Z
dc.date.issued2006
dc.description.abstractIn Q-learning, the aim is to reach the goal by using state and action pairs. When the goal is set as a big reward, the optimal path is found as soon as the reward accumulated reaches its highest value. Upon modification of the start and goal points, the information concerning how to reach the goal becomes useless even if the environment does not change. In this study, Q-learning is improved by making the usage of the past data possible. To achieve this, action probabilities for certain start and goal points are found and a neural network is trained with those values to estimate the action probabilities for other start and goal points. A radial basis function network is used as neural network for it can support local representation and can learn fast when there is a few number of inputs. When Q-learning is run with the found action probabilities, an increase in speed is observed in reaching the goal.en
dc.identifier.endpage+
dc.identifier.isbn978-1-4244-0238-0
dc.identifier.startpage210
dc.identifier.urihttps://hdl.handle.net/20.500.14981/49002
dc.identifier.wos000245347800054
dc.language.isotur
dc.publisherIEEE
dc.relation.conferenceIEEE 14th Signal Processing and Communications Applications
dc.relation.ispartof2006 IEEE 14TH SIGNAL PROCESSING AND COMMUNICATIONS APPLICATIONS, VOLS 1 AND 2
dc.subjectComputer Science
dc.subjectEngineering
dc.subjectImaging Science & Photographic Technology
dc.titleQ-Learning with probability based action policy
dc.typeProceedings Paper
dspace.entity.typePublication
local.import.sourceWOS

Dosyalar

Koleksiyonlar