| 講演抄録/キーワード |
| 講演名 |
2011-11-09 15:45
状態遷移の推定に基づく能動的価値関数推定法 ○幸島匡宏(東工大) IBISML2011-51 |
| 抄録 |
(和) |
強化学習において価値関数の精度の良い推定は重要である. 本研究では, 状態遷移確率の推定の後に, 価値関数の推定を行う手法における能動学習法を提案する.提案手法では, データ採取点は推定価値関数と真の価値関数の漸近平均二乗誤差を最小にする最適データ比率を基に決定される. 数値実験により, 導出した漸近平均二乗誤差値の検証と提案手法の有効性の確認を行う. |
| (英) |
It is considered to be a great importance in reinforcement learning to estimate value function precisely. In this study, the author proposes active learning algorithm for model based value function estimation, which computes value function using estimated transition probability. Its data sampling scheme is based on optimal ratio of the number of data to minimize asymptotic squared error of value function.
Experimental results show the effectiveness of proposed algorithm. |
| キーワード |
(和) |
強化学習 / マルコフ決定過程 / 状態遷移確率推定 / モデルベースアルゴリズム / 能動学習 / 漸近理論 / / |
| (英) |
reinforcement learning / markov decision processes / transition probability estimation / model based algorithm / active learning / asymptotic theory / / |
| 文献情報 |
信学技報, vol. 111, no. 275, IBISML2011-51, pp. 61-66, 2011年11月. |
| 資料番号 |
IBISML2011-51 |
| 発行日 |
2011-11-02 (IBISML) |
| ISSN |
Print edition: ISSN 0913-5685 Online edition: ISSN 2432-6380 |
著作権に ついて |
技術研究報告に掲載された論文の著作権は電子情報通信学会に帰属します.(許諾番号:10GA0019/12GB0052/13GB0056/17GB0034/18GB0034) |
| PDFダウンロード |
IBISML2011-51 |