| 講演抄録/キーワード |
| 講演名 |
2024-03-15 12:20
Decision Transformerモデルの拡張による将来報酬和と動作制御の研究 ○稲田泰生・久保田 繁(山形大) SIS2023-58 |
| 抄録 |
(和) |
強化学習は制御の分野において、自動運転の速度調や、対戦ゲームの敵AIの強さなど動的な調整が必要になることがある。Decision Transformerというオフライン強化学習モデルでは、入力した将来報酬和に対応する報酬が得られるように出力される特性がある。本研究では、そのDecision Transformerの特性を応用して、将来報酬和以外の出力の調整を行えるように、入力層の次元数を変更することで、将来報酬和に加えて、動き方についても出力の調整を可能にした。この結果はDecision Transformerの制御の分野への応用の可能性があることを示している。 |
| (英) |
In the field of control, reinforcement learning is sometimes required to dynamically adjust the speed of automatic driving, the strength of enemy AI in competitive games, etc. An off-line reinforcement learning model called Decision Transformer has the property of outputting the reward corresponding to the input sum of future rewards. In this study, we applied the characteristics of the Decision Transformer to change the dimensionality of the input layer so that the output can be adjusted in addition to the sum of future rewards, and thus the output can be adjusted not only for the sum of future rewards but also for the movement. The results indicate that Decision Transformer has potential applications in the field of control. |
| キーワード |
(和) |
制御 / オフライン強化学習 / Transformer / / / / / |
| (英) |
Control / Off-line reinforcement learning / Transformer / / / / / |
| 文献情報 |
信学技報, vol. 123, no. 440, SIS2023-58, pp. 73-76, 2024年3月. |
| 資料番号 |
SIS2023-58 |
| 発行日 |
2024-03-07 (SIS) |
| ISSN |
Online edition: ISSN 2432-6380 |
著作権に ついて |
技術研究報告に掲載された論文の著作権は電子情報通信学会に帰属します.(許諾番号:10GA0019/12GB0052/13GB0056/17GB0034/18GB0034) |
| PDFダウンロード |
SIS2023-58 |