DSpace at KOASAS: Final Iteration Convergence Bound of Q-Learning: Switching System Approach

DSpace at KOASAS

College of Engineering(공과대학)School of Electrical Engineering(전기및전자공학부)EE-Journal Papers(저널논문)

Final Iteration Convergence Bound of Q-Learning: Switching System Approach

Cited 0 time in webofscience

Cited 0 time in

Hit : 13
Download : 0

Export

DC Field	Value	Language
dc.contributor.author	Lee, Donghwan	ko
dc.date.accessioned	2024-08-23T10:00:09Z	-
dc.date.available	2024-08-23T10:00:09Z	-
dc.date.created	2024-08-23	-
dc.date.created	2024-08-23	-
dc.date.created	2024-08-23	-
dc.date.issued	2024-07	-
dc.identifier.citation	IEEE TRANSACTIONS ON AUTOMATIC CONTROL, v.69, no.7, pp.4765 - 4772	-
dc.identifier.issn	0018-9286	-
dc.identifier.uri	http://hdl.handle.net/10203/322400	-
dc.description.abstract	Q-learning is known as one of the fundamental reinforcement learning (RL) algorithms. Its convergence has been the focus of extensive research over the past several decades. Recently, a new finite-time error bound and analysis for Q-learning was introduced using a switching system framework. This approach views the dynamics of Q-learning as a discrete-time stochastic switching system. The prior study established a finite-time error bound on the averaged iterates using Lyapunov functions, offering further insights into Q-learning. While valuable, the analysis focuses on error bounds of the averaged iterate, which comes with the inherent disadvantages: It necessitates extra averaging steps, which can decelerate the convergence rate. Moreover, the final iterate, being the original format of Q-learning, is more commonly used and is often regarded as a more intuitive and natural form in the majority of iterative algorithms. In this article, we present a finite-time error bound on the final iterate of Q-learning based on the switching system framework. The proposed error bounds have different features compared to the previous works, and cover different scenarios. Finally, we expect that the proposed results provide additional insights on Q-learning via connections with discrete-time switching systems, and can potentially present a new template for finite-time analysis of more general RL algorithms.	-
dc.language	English	-
dc.publisher	IEEE-INST ELECTRICAL ELECTRONICS ENGINEERS INC	-
dc.title	Final Iteration Convergence Bound of Q-Learning: Switching System Approach	-
dc.type	Article	-
dc.identifier.wosid	001259639500022	-
dc.identifier.scopusid	2-s2.0-85182936662	-
dc.type.rims	ART	-
dc.citation.volume	69	-
dc.citation.issue	7	-
dc.citation.beginningpage	4765	-
dc.citation.endingpage	4772	-
dc.citation.publicationname	IEEE TRANSACTIONS ON AUTOMATIC CONTROL	-
dc.identifier.doi	10.1109/TAC.2024.3355326	-
dc.contributor.localauthor	Lee, Donghwan	-
dc.description.isOpenAccess	N	-
dc.type.journalArticle	Article	-
dc.subject.keywordAuthor	Q-learning	-
dc.subject.keywordAuthor	Convergence	-
dc.subject.keywordAuthor	Switching systems	-
dc.subject.keywordAuthor	Switches	-
dc.subject.keywordAuthor	Symmetric matrices	-
dc.subject.keywordAuthor	Markov processes	-
dc.subject.keywordAuthor	Behavioral sciences	-
dc.subject.keywordAuthor	finite-time analysis	-
dc.subject.keywordAuthor	reinforcement learning (RL)	-
dc.subject.keywordAuthor	switching system	-
dc.subject.keywordPlus	STOCHASTIC-APPROXIMATION	-
dc.subject.keywordPlus	RATES	-

Appears in Collection: EE-Journal Papers(저널논문)

Files in This Item: There are no files associated with this item.

Display Simple Item Record

qr_code

트윗하기

KOASAS

Knowledge Service Development Team, KAIST 291 Daehak-ro, Yuseong-gu, Daejeon 34141, Republic of Korea. T. 82-42-350-4493 Email. koasas@kaist.ac.kr
Copyright © 2016. Korea Advanced Institute of Science and Technology. All Rights Reserved.

KOASAS

KOASAS

Browse

Final Iteration Convergence Bound of Q-Learning: Switching System Approach

KOASAS

Communities & Collections