| 講演抄録/キーワード |
| 講演名 |
2024-12-20 09:40
事前知識を用いたQ学習 ○今川孝久・榎田修一(九工大) IBISML2024-34 |
| 抄録 |
(和) |
強化学習は報酬が高い行動を学習する方法であり,その代表的アルゴリズムの一つがQ-学習である.
Q-学習においてregretの解析が進んできたが,多くの分析は一から学習するということを前提としている.
もし,学習済みの知識を活かすことができれば,転移学習などの研究で示されているように改善が可能である.
そこで,本研究では事前に学習対象に対する知識を得られるという仮定の下でのQ-学習アルゴリズムBiased Exploration Q-learning (BEQ)を提案する.
そしてBEQのregretの解析を行い,regretの上界が$mathcal{O}(hat{w}_{max}sqrt{H^2S'A'Tiota} + HS'A'Delta_Q)$であり,知識を使っているという違いがあるものの,既存手法より小さいことを示す.
また,実験を通じて,事前知識を導入する既存手法Potential Based Reward Shapingに対して,BEQは優れていることを示す. |
| (英) |
Reinforcement learning (RL) is a method for learning actions which are highly rewarding and one of the representative methods of RL is Q-learning.
Regret of Q-learning have been analyzed, however most of them assume that learning takes place from scratch.
If the agent has acquired prior knowledge about the learning domain, its learning will be more efficient as suggested in transfer learning research.
Therefore, we propose a Q-learning method, Biased Exploration Q-learning (BEQ), which assumes that the agent can acquire domain knowledge in advance.
We analyze regret of BEQ and show that its upper bound is $mathcal{O}(hat{w}_{max}sqrt{H^2S'A'Tiota} + HS'A'Delta_Q)$.
BEQ and existing methods differ in terms of whether the agent can obtain the domain knowledge or not, but this bound is smaller than those of existing methods.
Also experiments show that BEQ outperforms Potential Based Reward Shaping (PBRS), which is a common method for introducing domain knowledge. |
| キーワード |
(和) |
強化学習 / Q-学習 / 転移学習 / 事前知識 / / / / |
| (英) |
Reinforcement Learning / Q-Learning / Transfer Learning / Prior Knowledge / / / / |
| 文献情報 |
信学技報, vol. 124, no. 321, IBISML2024-34, pp. 14-27, 2024年12月. |
| 資料番号 |
IBISML2024-34 |
| 発行日 |
2024-12-13 (IBISML) |
| ISSN |
Online edition: ISSN 2432-6380 |
著作権に ついて |
技術研究報告に掲載された論文の著作権は電子情報通信学会に帰属します.(許諾番号:10GA0019/12GB0052/13GB0056/17GB0034/18GB0034) |
| PDFダウンロード |
IBISML2024-34 |