| 講演抄録/キーワード |
| 講演名 |
2023-09-22 10:30
視覚言語モデルに関する順序数の的確な把握と活用能力の調査 ○増田琉斗・宮森 恒(京都産大) DE2023-20 |
| 抄録 |
(和) |
本稿では,視覚言語モデルが,順序数の概念を的確に把握し活用する能力をどの程度有するのかについて調査する.Transformerベースの大規模事前学習モデルは,四則演算といった単純な算術問題等のタスクにおいて高い正答率を示しているが,モデルが数の概念をどのように捉え,活用しているのかについては不明な点も多い.本研究では,数の概念の一つとして,順序数に焦点をあて,Transformerベースの視覚言語モデルが,順序数の概念をどの程度把握し活用する能力をもつのかについて調査する.具体的には,順序数の数え上げに焦点を当てた参照表現理解タスクのためのデータセットを新たに構築する.画像中に複数物体を配置したCG画像を生成し,物体間関係や数え上げが必須となるような参照表現を付与する.実験では,構築したデータセットを用いて,代表的な視覚言語モデルに対する参照表現理解タスクの性能評価を実施し,順序数に対する的確な把握と活用能力について分析する. |
| (英) |
In this paper, we investigate the extent to which visual language models have the ability to accurately grasp and utilize the concept of ordinal numbers. Although the Transformer-based large-scale pre-training models show high correct response rates for tasks such as simple arithmetic operations, it is still unclear how these models capture and utilize the concept of numbers.In this study, we focus on ordinal numbers as one of the concepts of numbers and investigate to what extent Transformer-based visual language models have the ability to grasp and utilize the concept of ordinal numbers.Specifically, we construct a new dataset for referring expression comprehension focusing on counting via ordinal numbers. CG images are generated with multiple objects placed in the image, and the objects are annotated with referring expressions which require understanding inter-object relations and counting them up.In the experiments, we evaluate the performance of referring expression comprehension tasks by typical visual language models using the constructed dataset and analyze the ability to accurately grasp and utilize the ordinal numbers. |
| キーワード |
(和) |
順序数 / 概念理解 / 視覚言語モデル / 数え上げ / 推論 / / / |
| (英) |
ordinal numbers / concept understanding / visual language model / counting operation / reasoning / / / |
| 文献情報 |
信学技報, vol. 123, no. 192, DE2023-20, pp. 54-59, 2023年9月. |
| 資料番号 |
DE2023-20 |
| 発行日 |
2023-09-14 (DE) |
| ISSN |
Online edition: ISSN 2432-6380 |
著作権に ついて |
技術研究報告に掲載された論文の著作権は電子情報通信学会に帰属します.(許諾番号:10GA0019/12GB0052/13GB0056/17GB0034/18GB0034) |
| PDFダウンロード |
DE2023-20 |
| 研究会情報 |
| 研究会 |
DE IPSJ-DBS IPSJ-IFAT |
| 開催期間 |
2023-09-21 - 2023-09-22 |
| 開催地(和) |
北九州国際会議場 |
| 開催地(英) |
Kitakyushu International Conference Center |
| テーマ(和) |
ビッグデータを対象とした管理・情報検索・知識獲得および一般 |
| テーマ(英) |
Bigdata management, information retrieval, knowledge discovery, etc. |
| 講演論文情報の詳細 |
| 申込み研究会 |
DE |
| 会議コード |
2023-09-DE-DBS-IFAT |
| 本文の言語 |
日本語 |
| タイトル(和) |
視覚言語モデルに関する順序数の的確な把握と活用能力の調査 |
| サブタイトル(和) |
|
| タイトル(英) |
Probing the ability to accurately understand and utilize the ordinal numbers by visual language models |
| サブタイトル(英) |
|
| キーワード(1)(和/英) |
順序数 / ordinal numbers |
| キーワード(2)(和/英) |
概念理解 / concept understanding |
| キーワード(3)(和/英) |
視覚言語モデル / visual language model |
| キーワード(4)(和/英) |
数え上げ / counting operation |
| キーワード(5)(和/英) |
推論 / reasoning |
| キーワード(6)(和/英) |
/ |
| キーワード(7)(和/英) |
/ |
| キーワード(8)(和/英) |
/ |
| 第1著者 氏名(和/英/ヨミ) |
増田 琉斗 / Ryuto Masuda / マスダ リュウト |
| 第1著者 所属(和/英) |
京都産業大学 (略称: 京都産大)
Kyoto Sangyo University (略称: Kyoto Sangyo Univ.) |
| 第2著者 氏名(和/英/ヨミ) |
宮森 恒 / Hisashi Miyamori / ミヤモリ ヒサシ |
| 第2著者 所属(和/英) |
京都産業大学 (略称: 京都産大)
Kyoto Sangyo University (略称: Kyoto Sangyo Univ.) |
| 第3著者 氏名(和/英/ヨミ) |
/ / |
| 第3著者 所属(和/英) |
(略称: )
(略称: ) |
| 第4著者 氏名(和/英/ヨミ) |
/ / |
| 第4著者 所属(和/英) |
(略称: )
(略称: ) |
| 第5著者 氏名(和/英/ヨミ) |
/ / |
| 第5著者 所属(和/英) |
(略称: )
(略称: ) |
| 第6著者 氏名(和/英/ヨミ) |
/ / |
| 第6著者 所属(和/英) |
(略称: )
(略称: ) |
| 第7著者 氏名(和/英/ヨミ) |
/ / |
| 第7著者 所属(和/英) |
(略称: )
(略称: ) |
| 第8著者 氏名(和/英/ヨミ) |
/ / |
| 第8著者 所属(和/英) |
(略称: )
(略称: ) |
| 第9著者 氏名(和/英/ヨミ) |
/ / |
| 第9著者 所属(和/英) |
(略称: )
(略称: ) |
| 第10著者 氏名(和/英/ヨミ) |
/ / |
| 第10著者 所属(和/英) |
(略称: )
(略称: ) |
| 第11著者 氏名(和/英/ヨミ) |
/ / |
| 第11著者 所属(和/英) |
(略称: )
(略称: ) |
| 第12著者 氏名(和/英/ヨミ) |
/ / |
| 第12著者 所属(和/英) |
(略称: )
(略称: ) |
| 第13著者 氏名(和/英/ヨミ) |
/ / |
| 第13著者 所属(和/英) |
(略称: )
(略称: ) |
| 第14著者 氏名(和/英/ヨミ) |
/ / |
| 第14著者 所属(和/英) |
(略称: )
(略称: ) |
| 第15著者 氏名(和/英/ヨミ) |
/ / |
| 第15著者 所属(和/英) |
(略称: )
(略称: ) |
| 第16著者 氏名(和/英/ヨミ) |
/ / |
| 第16著者 所属(和/英) |
(略称: )
(略称: ) |
| 第17著者 氏名(和/英/ヨミ) |
/ / |
| 第17著者 所属(和/英) |
(略称: )
(略称: ) |
| 第18著者 氏名(和/英/ヨミ) |
/ / |
| 第18著者 所属(和/英) |
(略称: )
(略称: ) |
| 第19著者 氏名(和/英/ヨミ) |
/ / |
| 第19著者 所属(和/英) |
(略称: )
(略称: ) |
| 第20著者 氏名(和/英/ヨミ) |
/ / |
| 第20著者 所属(和/英) |
(略称: )
(略称: ) |
| 第21著者 氏名(和/英/ヨミ) |
/ / |
| 第21著者 所属(和/英) |
(略称: )
(略称: ) |
| 第22著者 氏名(和/英/ヨミ) |
/ / |
| 第22著者 所属(和/英) |
(略称: )
(略称: ) |
| 第23著者 氏名(和/英/ヨミ) |
/ / |
| 第23著者 所属(和/英) |
(略称: )
(略称: ) |
| 第24著者 氏名(和/英/ヨミ) |
/ / |
| 第24著者 所属(和/英) |
(略称: )
(略称: ) |
| 第25著者 氏名(和/英/ヨミ) |
/ / |
| 第25著者 所属(和/英) |
(略称: )
(略称: ) |
| 第26著者 氏名(和/英/ヨミ) |
/ / |
| 第26著者 所属(和/英) |
(略称: )
(略称: ) |
| 第27著者 氏名(和/英/ヨミ) |
/ / |
| 第27著者 所属(和/英) |
(略称: )
(略称: ) |
| 第28著者 氏名(和/英/ヨミ) |
/ / |
| 第28著者 所属(和/英) |
(略称: )
(略称: ) |
| 第29著者 氏名(和/英/ヨミ) |
/ / |
| 第29著者 所属(和/英) |
(略称: )
(略称: ) |
| 第30著者 氏名(和/英/ヨミ) |
/ / |
| 第30著者 所属(和/英) |
(略称: )
(略称: ) |
| 第31著者 氏名(和/英/ヨミ) |
/ / |
| 第31著者 所属(和/英) |
(略称: )
(略称: ) |
| 第32著者 氏名(和/英/ヨミ) |
/ / |
| 第32著者 所属(和/英) |
(略称: )
(略称: ) |
| 第33著者 氏名(和/英/ヨミ) |
/ / |
| 第33著者 所属(和/英) |
(略称: )
(略称: ) |
| 第34著者 氏名(和/英/ヨミ) |
/ / |
| 第34著者 所属(和/英) |
(略称: )
(略称: ) |
| 第35著者 氏名(和/英/ヨミ) |
/ / |
| 第35著者 所属(和/英) |
(略称: )
(略称: ) |
| 第36著者 氏名(和/英/ヨミ) |
/ / |
| 第36著者 所属(和/英) |
(略称: )
(略称: ) |
| 講演者 |
第1著者 |
| 発表日時 |
2023-09-22 10:30:00 |
| 発表時間 |
25分 |
| 申込先研究会 |
DE |
| 資料番号 |
DE2023-20 |
| 巻番号(vol) |
vol.123 |
| 号番号(no) |
no.192 |
| ページ範囲 |
pp.54-59 |
| ページ数 |
6 |
| 発行日 |
2023-09-14 (DE) |