| 講演抄録/キーワード |
| 講演名 |
2022-12-16 10:15
姿勢変換のための姿勢知覚トランスフォーマーネットワーク ○柴崎 圭・池原雅章(慶大) PRMU2022-44 |
| 抄録 |
(和) |
姿勢変換は,ソースの画像とその姿勢の情報,ターゲットの姿勢情報から人物画像の姿勢変換を行うタスクである.従来手法の多くは,追加のパース情報・タスクの必要性があり,実用性が制限される.また,CNNを用いるため画像全体の整合性を考慮できない.本論文では画像の整合性の問題に対応した実用的な姿勢変換ネットワークを提案する.提案手法では,姿勢変換というタスクを,「大まかな姿勢の変換」と「詳細なテクスチャの生成」という2つのタスクに分離する.前者のタスクでは低解像度の特徴マップに対して, Axial Transformer を含むブロックで変換を行う.後者のタスクはCNNネットワークを用いている.提案ネットワークは非常に軽量であるが優れた性能を獲得している. |
| (英) |
Pose Guided Person Image Generation (PGPIG) is the task that transforms the pose of a person image from the source image, its pose information and the target pose information. Most existing PGPIG methods require additional pose information or tasks, limiting their application. In addition, all input information is combined and fed into the network, and CNNs are used as the feature extractor. However, CNNs can only extract features from neighboring pixels and cannot consider the consistency of the entire image. Furthermore, they combine the input information before extracting enough features, making it unclear which task the network should learn, which degrades the network performance. This paper proposes a PGPIG network that addresses the image consistency problem and clarifies which task the network should learn. The proposed method disentangles the PGPIG task into two sub tasks: “rough pose transformation” and “detailed texture generation”. In the former task, low-resolution feature maps are transformed by blocks containing Axial Transformer with a large receptive field. These blocks employ an Encoder-Decoder structure, which allows the network to use the pose information well and improves the stability and performance of the training. The latter task uses a CNN network with Adaptive Instance Normalization. Experiments show that the proposed method has competitive performance with other state-of-the-art methods. Furthermore, despite achieving excellent performance, the proposed network has a significantly fewer parameters than existing methods. |
| キーワード |
(和) |
深層学習 / 画像処理 / 姿勢変換 / トランスフォーマー / マルチスケール / / / |
| (英) |
Deep learning / Image Processing / Pose Guided Person Image Generation / Transformer / Multi-scale Network / / / |
| 文献情報 |
信学技報, vol. 122, no. 314, PRMU2022-44, pp. 63-69, 2022年12月. |
| 資料番号 |
PRMU2022-44 |
| 発行日 |
2022-12-08 (PRMU) |
| ISSN |
Online edition: ISSN 2432-6380 |
著作権に ついて |
技術研究報告に掲載された論文の著作権は電子情報通信学会に帰属します.(許諾番号:10GA0019/12GB0052/13GB0056/17GB0034/18GB0034) |
| PDFダウンロード |
PRMU2022-44 |