| 講演抄録/キーワード |
| 講演名 |
2026-01-08 15:10
Certified Actor Critic法における勾配の簡易計算法 ○佐藤達志・潮 俊光(南山大) MSS2025-43 |
| 抄録 |
(和) |
システムの振る舞いが状態空間上で指定された安全領域の中に留まり続けるという制御仕様を安全仕様という.安全仕様を満たしつつシステムを安定化する制御法を学習する方法にcertified actor critic(CAC)がある.この方法では,最初に安全仕様を満たすように強化学習を行い,その後,安全仕様と安定化に対して個別にcriticネットワークを使って評価し,その評価をもとにactorの更新を行う.このとき,各criticネットワークに対応する勾配からactorの更新に用いる勾配を最適化問題に帰着して決定している.本報告では,この勾配の簡易計算法を3つ提案し,シミュレーションによりその学習能力を検討する. |
| (英) |
A control specification that requires the system's behavior to remain within a designated safe region in the state space is referred to as a safety specification. One method for learning a control policy that stabilizes the system while satisfying the safety specification is the Certified Actor-Critic (CAC) approach. In this method, reinforcement learning is first conducted to satisfy the safety specification. Then, separate critic networks are used to evaluate the satisfaction of the safety specification and the stabilization performance, respectively. Based on these evaluations, the actor is updated. At this stage, the gradient used to update the actor is determined by formulating an optimization problem that combines the gradients from each critic network. In this report, we propose three simplified methods for calculating this gradient and investigate their learning performance by simulation. |
| キーワード |
(和) |
深層強化学習 / SAC / 安全 / 制御バリア関数 / / / / |
| (英) |
deep reinforcement learning / SAC / safety / control barrier function / / / / |
| 文献情報 |
信学技報, vol. 125, no. 310, MSS2025-43, pp. 55-59, 2026年1月. |
| 資料番号 |
MSS2025-43 |
| 発行日 |
2026-01-01 (MSS) |
| ISSN |
Online edition: ISSN 2432-6380 |
著作権に ついて |
技術研究報告に掲載された論文の著作権は電子情報通信学会に帰属します.(許諾番号:10GA0019/12GB0052/13GB0056/17GB0034/18GB0034) |
| PDFダウンロード |
MSS2025-43 |