| 講演抄録/キーワード |
| 講演名 |
2006-07-28 09:30
インド系言語レガシデータの符号化方式自動識別 ~ クメール語フォントにおける非標準レガシ符号の調査 ~ ○鈴木俊哉(広島大)・佐藤 大(東北大) |
| 抄録 |
(和) |
国際標準準拠の観点からUnicodeを用いてインド系言語を符号化するとテキストのレンダリングが困難であるために、WWWなどでは標準とは別にフォント個別に定義した符号化方式が用いられている。また複雑なレンダリング機構を用いない場合の符号化方式について、準拠すべき標準がなく、無名の符号化方式がフォント製品ごとに用いられている。これらのフォントを利用したドキュメントが多数流通しているが、ファイル形式はデジタルドキュメントであるが画像と同様のものになっており、情報交換性を阻害している。
典型的な例として、カンボジアの公用語であるクメール語のレガシ符号を調査すると、符号化方式が乱立しており事実上標準も存在していないことが分かった。本稿では、クメール文字フォントのレガシ符号を網羅的に調査した結果を整理し、符号化方式の自動識別可能性について述べる。 |
| (英) |
Although Unicode text layout systems are introduced into modern text processing softwares, still legacy character encodings are widely used for
Indic scripts in South and South East Asia to work with systems missing intelligent text layout functionalities.
Some de-jure or de-facto legacy standards are used for some scripts, but there are scripts whose encodings was not standardized before ISO 10646. If the document uses the fonts with unstandardized encodings, usually the text data extraction from the document is difficult.
In this paper, we take a concrete example of such script: Khmer script. It is suppored by Unicode standard, not fully supported by applications, and no legacy encodings were standardized.
We investigate the encodings in the freely distributed Khmer TrueType fonts and propose algorithm to identify which encoding is used in the fonts. By the algorithm, the text extraction of the document including legacy TrueType fonts can be automated. |
| キーワード |
(和) |
インド系言語 / クメール文字 / TrueType / フォント / レガシ符号 / 文字コード / 符号化方式 / 自動識別 |
| (英) |
Indic script / Khmer script / TrueType / font / legacy encoding / Unicode / character encoding / auto detection |
| 文献情報 |
信学技報, vol. 106, pp. 1-8, 2006年7月. |
| 資料番号 |
|
| 発行日 |
2006-07-21 (OIS) |
| ISSN |
Print edition: ISSN 0913-5685 |
| PDFダウンロード |
|