Information: Join today and make your research activities more affordable! Technical workshop participation fees and annual registration fees are available at member rates.
Notice: [Important] Announcement of Changes to Registration Fee Payment and Manuscript Upload Procedures for IEICE Technical Meetings
IEICE Technical Committee Submission System
Conference Paper's Information
Online Proceedings
[Sign in]
Tech. Rep. Archives
 Go Top Page Go Previous   [Japanese] / [English] 

Paper Abstract and Keywords
Presentation 2018-07-28 10:00
Improve accuracy of predicting confidential words in juditial precedents
Masakazu Kanazawa, Atsushi Ito (Utsunomiya Univ.), Yuya Kiryu (KDDI), Kazuyuki Yamaswa (TKC), Takehiko Kasahara (Toin Yokohama Univ.) TL2018-12
Abstract (in Japanese) (See Japanese page) 
(in English) The judicial system in which IT technology was introduced is called Cyber Court. In judicial field in Japan, confidential words such as personal names and personal places are converted to meaningless words manually because Japanese values do not allow the disclosure of the individual information. It is not easy to construct a comprehensive dictionary for detecting confidential words. We have already proposed two models that predict confidential words automatically by using neural networks. We used long short-term memory (LSTM) and continuous bag-of-words (CBOW) as our language models. Firstly, we explained the possibility of detecting the words surrounding a confidential word by using CBOW. Then, we proposed two models to predict the confidential words from the neighboring words by applying LSTM. The first model imitates the anonymization work by a human being, and the second model was based on CBOW. The results show that the first model is more effective for predicting confidential words than the simple LSTM model. We expected the second model to have paraphrasing ability to increase the possibility of finding other paraphraseable One is Bi-directional LSTM LR model and the other is Sum-LSTM based on the CBOW model. The two proposed models were effective for predicting all the words. However, only the Bi-directional LSTM LR model was effective for predicting confidential words. This could have happened because Sum-LSTM was based on CBOW. CBOW is an effective model for paraphrasing words; therefore, Sum-LSTM also has that mechanism. Therefore, when Sum-LSTM predicted a word whose answer was “confidential,” the CW_PPL (Confidential word perplexity) became worse because there was a possibility of paraphrasing words of paraphrasing words such as “plaintiff,” “defendant,” “doctor,” and “teacher.” Knowing the paraphrased words of the confidential words meant that the embedding vectors of the confidential words could be successfully generated. This meant that the model could recognize the meaning of “confidential.” However, the prediction accuracy did not improve; therefore, there was a problem in calculating the probability of the prediction task. To solve this problem, we could exclude these paraphraseable words from the choices when calculating the probability. It is also important to examine scores other than PPL. Then we consider to improve accuracy of predicting confidential words. At first we focus the parameter of neural network. We would change the window size to large because we think the surrounding words of the target words (that is window size) is larger, accuracy will be high. Then, we are experimenting but we can’t get the results by now to spent more memory and more calculating time. Next, we would use the proper noun dictionary combined neural network. But we don’t discuss in the this paper.
Keyword (in Japanese) (See Japanese page) 
(in English) Bi-directional LSTM / CBOW / ppl / cw-ppl / window size / / /  
Reference Info. IEICE Tech. Rep., vol. 118, no. 163, TL2018-12, pp. 1-6, July 2018.
Paper # TL2018-12 
Date of Issue 2018-07-21 (TL) 
ISSN Print edition: ISSN 0913-5685    Online edition: ISSN 2432-6380
Copyright
and
reproduction
All rights are reserved and no part of this publication may be reproduced or transmitted in any form or by any means, electronic or mechanical, including photocopy, recording, or any information storage and retrieval system, without permission in writing from the publisher. Notwithstanding, instructors are permitted to photocopy isolated articles for noncommercial classroom use without fee. (License No.: 10GA0019/12GB0052/13GB0056/17GB0034/18GB0034)
Download PDF TL2018-12

Conference Information
Committee TL  
Conference Date 2018-07-28 - 2018-07-29 
Place (in Japanese) (See Japanese page) 
Place (in English) Keio University 
Topics (in Japanese) (See Japanese page) 
Topics (in English) Human Language Processing and Learning 
Paper Information
Registration To TL 
Conference Code 2018-07-TL 
Language English 
Title (in Japanese) (See Japanese page) 
Sub Title (in Japanese) (See Japanese page) 
Title (in English) Improve accuracy of predicting confidential words in juditial precedents 
Sub Title (in English)  
Keyword(1) Bi-directional LSTM  
Keyword(2) CBOW  
Keyword(3) ppl  
Keyword(4) cw-ppl  
Keyword(5) window size  
Keyword(6)  
Keyword(7)  
Keyword(8)  
1st Author's Name Masakazu Kanazawa  
1st Author's Affiliation Utsunomiya University (Utsunomiya Univ.)
2nd Author's Name Atsushi Ito  
2nd Author's Affiliation Utsunomiya University (Utsunomiya Univ.)
3rd Author's Name Yuya Kiryu  
3rd Author's Affiliation KDDI Corporation (KDDI)
4th Author's Name Kazuyuki Yamaswa  
4th Author's Affiliation TKC Corporation (TKC)
5th Author's Name Takehiko Kasahara  
5th Author's Affiliation Toin Yokohama Univeecity (Toin Yokohama Univ.)
6th Author's Name  
6th Author's Affiliation ()
7th Author's Name  
7th Author's Affiliation ()
8th Author's Name  
8th Author's Affiliation ()
9th Author's Name  
9th Author's Affiliation ()
10th Author's Name  
10th Author's Affiliation ()
11th Author's Name  
11th Author's Affiliation ()
12th Author's Name  
12th Author's Affiliation ()
13th Author's Name  
13th Author's Affiliation ()
14th Author's Name  
14th Author's Affiliation ()
15th Author's Name  
15th Author's Affiliation ()
16th Author's Name  
16th Author's Affiliation ()
17th Author's Name  
17th Author's Affiliation ()
18th Author's Name  
18th Author's Affiliation ()
19th Author's Name  
19th Author's Affiliation ()
20th Author's Name  
20th Author's Affiliation ()
21st Author's Name  
21st Author's Affiliation ()
22nd Author's Name  
22nd Author's Affiliation ()
23rd Author's Name  
23rd Author's Affiliation ()
24th Author's Name  
24th Author's Affiliation ()
25th Author's Name  
25th Author's Affiliation ()
26th Author's Name / /
26th Author's Affiliation ()
()
27th Author's Name / /
27th Author's Affiliation ()
()
28th Author's Name / /
28th Author's Affiliation ()
()
29th Author's Name / /
29th Author's Affiliation ()
()
30th Author's Name / /
30th Author's Affiliation ()
()
31st Author's Name / /
31st Author's Affiliation ()
()
32nd Author's Name / /
32nd Author's Affiliation ()
()
33rd Author's Name / /
33rd Author's Affiliation ()
()
34th Author's Name / /
34th Author's Affiliation ()
()
35th Author's Name / /
35th Author's Affiliation ()
()
36th Author's Name / /
36th Author's Affiliation ()
()
Speaker Author-1 
Date Time 2018-07-28 10:00:00 
Presentation Time 30 minutes 
Registration for TL 
Paper # TL2018-12 
Volume (vol) vol.118 
Number (no) no.163 
Page pp.1-6 
#Pages
Date of Issue 2018-07-21 (TL) 


[Return to Top Page]

[Return to IEICE Web Page]


The Institute of Electronics, Information and Communication Engineers (IEICE), Japan