| 講演抄録/キーワード |
| 講演名 |
2007-12-22 15:45
Adaptive Importance Sampling with Automatic Model Selection in Value Function Approximation ○Hirotaka Hachiya・Takayuki Akiyama・Masashi Sugiyama(Tokyo Inst. of Tech.) NC2007-84 |
| 抄録 |
(和) |
(まだ登録されていません) |
| (英) |
Off-policy reinforcement learning is aimed at efficiently reusing data samples gathered in the past. A common approach is to use importance sampling techniques for compensating for the bias caused by the difference between data-collecting policies and the target policy. However, existing off-policy methods do not often take the variance of value function estimators explicitly into account and therefore their performance tends to be unstable. To cope with this problem, we propose using an adaptive importance sampling technique which allows us to actively control the trade-off between bias and variance. We further provide a method for optimally determining the trade-off parameter based on a statistical machine learning theory. |
| キーワード |
(和) |
/ / / / / / / |
| (英) |
Off-policy / Reinforcement learning / Value function approximation / Importance sampling / Importance weighted cross validation / / / |
| 文献情報 |
信学技報, vol. 107, no. 410, NC2007-84, pp. 75-80, 2007年12月. |
| 資料番号 |
NC2007-84 |
| 発行日 |
2007-12-15 (NC) |
| ISSN |
Print edition: ISSN 0913-5685 Online edition: ISSN 2432-6380 |
著作権に ついて |
技術研究報告に掲載された論文の著作権は電子情報通信学会に帰属します.(許諾番号:10GA0019/12GB0052/13GB0056/17GB0034/18GB0034) |
| PDFダウンロード |
NC2007-84 |