| 講演抄録/キーワード |
| 講演名 |
2008-02-21 11:30
文字仮説の多重生成による帳票画像からの単語抽出方式 ○武部浩明・藤本克仁(富士通研) PRMU2007-217 |
| 抄録 |
(和) |
帳票画像の論理構造を認識するとき,帳票の見出しとデータの文字列である単語を抽出することが必要であるが,レイアウト認識と文字認識の失敗により,単語を構成する文字が正しく認識されずに単語を抽出できないことがある.特に,データの文字列は類似文字が多い数字や記号で構成されていることが多く,文字認識誤りが原因で抽出に失敗することが多い.さらに,データの文字列は文字数が可変の正規表現で指定される特徴を持つ.そこで,文字認識誤りに対応するため,文字認識対象カテゴリを制御することにより通常の文字認識では得られない文字仮説を生成し,切り出し候補の多重性と単語の正規表現の可変性を同時に考慮した、局所最適性に基づくラティス間のマッチングを行い,単語を抽出する方式を提案する.実帳票を用いた見出しとデータの抽出精度評価を行い,本方式により精度が向上することを確認した. |
| (英) |
It is necessary to extract words which are headers and data for recognizing logical structure of form images. However, word extraction fails if the part of characters is not extracted because of layout analysis error or character recognition error. Words of data usually consist of numbers or symbols which are similar to many characters, so they are often mis-extracted due to character recognition error. Moreover, words of data are represented by regular expression which includes variable numbers of wild-card character. We propose the word extraction method which generates multiple character hypotheses by controlling recognition target, and searches the combination of recognition result with character of regular expression by lattice matching based on the local optimum. It was confirmed that the proposed method improved extraction accuracy of header and data by the experiment for the real form images. |
| キーワード |
(和) |
帳票画像 / 単語 / 見出し / データ / 認識誤り / 正規表現 / ラティス / |
| (英) |
Form Image / Word / Header / Data / Recognition Error / Regular Expression / Lattice / |
| 文献情報 |
信学技報, vol. 107, no. 491, PRMU2007-217, pp. 19-24, 2008年2月. |
| 資料番号 |
PRMU2007-217 |
| 発行日 |
2008-02-14 (PRMU) |
| ISSN |
Print edition: ISSN 0913-5685 Online edition: ISSN 2432-6380 |
著作権に ついて |
技術研究報告に掲載された論文の著作権は電子情報通信学会に帰属します.(許諾番号:10GA0019/12GB0052/13GB0056/17GB0034/18GB0034) |
| PDFダウンロード |
PRMU2007-217 |