| 講演抄録/キーワード |
| 講演名 |
2020-03-06 10:10
時空間的特徴を考慮したDNNによる手話翻訳手法の比較検討 ○渡邊滉大・亀山 渉(早大) IMQ2019-68 IE2019-150 MVE2019-89 |
| 抄録 |
(和) |
動画からの手話翻訳において、AlexNetと呼ばれる2DCNN(2次元畳み込みニューラルネットワーク)とSeq2Seqと呼ばれる機械翻訳モデルを組み合わせた手話翻訳モデルが提案されている。これは、2DCNNによって空間的な情報を失った特徴量からGRU(Gated Recurrent Unit)によって時系列的な特徴量を抽出している手法と考えられる。しかし、手話の動作は手及び指の位置や形とその動きによって形成されるため、空間的な情報を保ったまま時系列的な情報を考慮できる手法がより適していると考えられる。そこで、本稿では、動画の各フレームから特徴量を抽出する段階で、時系列的な情報を考慮する様々な手法を提案し、比較検討を行った。時空間的特徴量抽出器の比較実験の結果、本実験で使用したデータセットでは、最適化されるパラメータの数と手話翻訳性能が反比例することが示唆された。そのため、パラメータ数が最も少ないOptical Flowのみを入力としたモデルが高い手話翻訳性能を示したと考えられる。 |
| (英) |
In Neural Sign Language Translation, a model based on 2DCNN (2 Dimensional Convolutional Neural Network) called AlexNet and a neural machine translation model called Seq2Seq has been proposed. In this model, temporal information is extracted by GRU (Gated Recurrent Unit) from the features in which the spatial information is lost by 2DCNN. However, since sign language uses position, shape and motion of hands and fingers, a model that can extract temporal information from the features that contain spatial information seems to be more suitable. Therefore, in this paper, we propose various methods and compare them that extract temporal information at the stage of extracting spatial features from each frame of video. As the result of the comparison experiment of the various spatio-temporal feature extractors, it is suggested that the number of to-be-optimized parameters and the performance of sign language translation are inversely proportional on the dataset used in this experiment. That seems the reason why the model using only Optical Flow shows the highest performance in sign language translation because it has the least number of parameters to be trained. |
| キーワード |
(和) |
手話翻訳 / 時空間的特徴 / DNN / Optical Flow / / / / |
| (英) |
Neural Sign Language Translation / Spatio-temporal Features / DNN / Optical Flow / / / / |
| 文献情報 |
信学技報, vol. 119, no. 456, IE2019-150, pp. 273-278, 2020年3月. |
| 資料番号 |
IE2019-150 |
| 発行日 |
2020-02-27 (IMQ, IE, MVE) |
| ISSN |
Print edition: ISSN 0913-5685 Online edition: ISSN 2432-6380 |
著作権に ついて |
技術研究報告に掲載された論文の著作権は電子情報通信学会に帰属します.(許諾番号:10GA0019/12GB0052/13GB0056/17GB0034/18GB0034) |
| PDFダウンロード |
IMQ2019-68 IE2019-150 MVE2019-89 |
| 研究会情報 |
| 研究会 |
IE IMQ MVE CQ |
| 開催期間 |
2020-03-05 - 2020-03-06 |
| 開催地(和) |
九州工業大学 戸畑キャンパス |
| 開催地(英) |
Kyushu Institute of Technology |
| テーマ(和) |
五感メディア,マルチメディア,メディアエクスペリエンス, 映像符号化,イメージメディアの品質,ネットワークの品質 および信頼性,一般 (魅力工学(AC)研究会協賛) |
| テーマ(英) |
|
| 講演論文情報の詳細 |
| 申込み研究会 |
IE |
| 会議コード |
2020-03-IE-IMQ-MVE-CQ |
| 本文の言語 |
日本語 |
| タイトル(和) |
時空間的特徴を考慮したDNNによる手話翻訳手法の比較検討 |
| サブタイトル(和) |
|
| タイトル(英) |
A Comparison Study of Neural Sign Language Translation Methods with Spatio-Temporal Features |
| サブタイトル(英) |
|
| キーワード(1)(和/英) |
手話翻訳 / Neural Sign Language Translation |
| キーワード(2)(和/英) |
時空間的特徴 / Spatio-temporal Features |
| キーワード(3)(和/英) |
DNN / DNN |
| キーワード(4)(和/英) |
Optical Flow / Optical Flow |
| キーワード(5)(和/英) |
/ |
| キーワード(6)(和/英) |
/ |
| キーワード(7)(和/英) |
/ |
| キーワード(8)(和/英) |
/ |
| 第1著者 氏名(和/英/ヨミ) |
渡邊 滉大 / Kodai Watanabe / ワタナベ コウダイ |
| 第1著者 所属(和/英) |
早稲田大学 (略称: 早大)
Waseda University (略称: Waseda Univ.) |
| 第2著者 氏名(和/英/ヨミ) |
亀山 渉 / Wataru Kameyama / カメヤマ ワタル |
| 第2著者 所属(和/英) |
早稲田大学 (略称: 早大)
Waseda University (略称: Waseda Univ.) |
| 第3著者 氏名(和/英/ヨミ) |
/ / |
| 第3著者 所属(和/英) |
(略称: )
(略称: ) |
| 第4著者 氏名(和/英/ヨミ) |
/ / |
| 第4著者 所属(和/英) |
(略称: )
(略称: ) |
| 第5著者 氏名(和/英/ヨミ) |
/ / |
| 第5著者 所属(和/英) |
(略称: )
(略称: ) |
| 第6著者 氏名(和/英/ヨミ) |
/ / |
| 第6著者 所属(和/英) |
(略称: )
(略称: ) |
| 第7著者 氏名(和/英/ヨミ) |
/ / |
| 第7著者 所属(和/英) |
(略称: )
(略称: ) |
| 第8著者 氏名(和/英/ヨミ) |
/ / |
| 第8著者 所属(和/英) |
(略称: )
(略称: ) |
| 第9著者 氏名(和/英/ヨミ) |
/ / |
| 第9著者 所属(和/英) |
(略称: )
(略称: ) |
| 第10著者 氏名(和/英/ヨミ) |
/ / |
| 第10著者 所属(和/英) |
(略称: )
(略称: ) |
| 第11著者 氏名(和/英/ヨミ) |
/ / |
| 第11著者 所属(和/英) |
(略称: )
(略称: ) |
| 第12著者 氏名(和/英/ヨミ) |
/ / |
| 第12著者 所属(和/英) |
(略称: )
(略称: ) |
| 第13著者 氏名(和/英/ヨミ) |
/ / |
| 第13著者 所属(和/英) |
(略称: )
(略称: ) |
| 第14著者 氏名(和/英/ヨミ) |
/ / |
| 第14著者 所属(和/英) |
(略称: )
(略称: ) |
| 第15著者 氏名(和/英/ヨミ) |
/ / |
| 第15著者 所属(和/英) |
(略称: )
(略称: ) |
| 第16著者 氏名(和/英/ヨミ) |
/ / |
| 第16著者 所属(和/英) |
(略称: )
(略称: ) |
| 第17著者 氏名(和/英/ヨミ) |
/ / |
| 第17著者 所属(和/英) |
(略称: )
(略称: ) |
| 第18著者 氏名(和/英/ヨミ) |
/ / |
| 第18著者 所属(和/英) |
(略称: )
(略称: ) |
| 第19著者 氏名(和/英/ヨミ) |
/ / |
| 第19著者 所属(和/英) |
(略称: )
(略称: ) |
| 第20著者 氏名(和/英/ヨミ) |
/ / |
| 第20著者 所属(和/英) |
(略称: )
(略称: ) |
| 第21著者 氏名(和/英/ヨミ) |
/ / |
| 第21著者 所属(和/英) |
(略称: )
(略称: ) |
| 第22著者 氏名(和/英/ヨミ) |
/ / |
| 第22著者 所属(和/英) |
(略称: )
(略称: ) |
| 第23著者 氏名(和/英/ヨミ) |
/ / |
| 第23著者 所属(和/英) |
(略称: )
(略称: ) |
| 第24著者 氏名(和/英/ヨミ) |
/ / |
| 第24著者 所属(和/英) |
(略称: )
(略称: ) |
| 第25著者 氏名(和/英/ヨミ) |
/ / |
| 第25著者 所属(和/英) |
(略称: )
(略称: ) |
| 第26著者 氏名(和/英/ヨミ) |
/ / |
| 第26著者 所属(和/英) |
(略称: )
(略称: ) |
| 第27著者 氏名(和/英/ヨミ) |
/ / |
| 第27著者 所属(和/英) |
(略称: )
(略称: ) |
| 第28著者 氏名(和/英/ヨミ) |
/ / |
| 第28著者 所属(和/英) |
(略称: )
(略称: ) |
| 第29著者 氏名(和/英/ヨミ) |
/ / |
| 第29著者 所属(和/英) |
(略称: )
(略称: ) |
| 第30著者 氏名(和/英/ヨミ) |
/ / |
| 第30著者 所属(和/英) |
(略称: )
(略称: ) |
| 第31著者 氏名(和/英/ヨミ) |
/ / |
| 第31著者 所属(和/英) |
(略称: )
(略称: ) |
| 第32著者 氏名(和/英/ヨミ) |
/ / |
| 第32著者 所属(和/英) |
(略称: )
(略称: ) |
| 第33著者 氏名(和/英/ヨミ) |
/ / |
| 第33著者 所属(和/英) |
(略称: )
(略称: ) |
| 第34著者 氏名(和/英/ヨミ) |
/ / |
| 第34著者 所属(和/英) |
(略称: )
(略称: ) |
| 第35著者 氏名(和/英/ヨミ) |
/ / |
| 第35著者 所属(和/英) |
(略称: )
(略称: ) |
| 第36著者 氏名(和/英/ヨミ) |
/ / |
| 第36著者 所属(和/英) |
(略称: )
(略称: ) |
| 講演者 |
第1著者 |
| 発表日時 |
2020-03-06 10:10:00 |
| 発表時間 |
25分 |
| 申込先研究会 |
IE |
| 資料番号 |
IMQ2019-68, IE2019-150, MVE2019-89 |
| 巻番号(vol) |
vol.119 |
| 号番号(no) |
no.454(IMQ), no.456(IE), no.457(MVE) |
| ページ範囲 |
pp.273-278 |
| ページ数 |
6 |
| 発行日 |
2020-02-27 (IMQ, IE, MVE) |
|