| Paper Abstract and Keywords |
| Presentation |
2018-07-28 10:00
Improve accuracy of predicting confidential words in juditial precedents Masakazu Kanazawa, Atsushi Ito (Utsunomiya Univ.), Yuya Kiryu (KDDI), Kazuyuki Yamaswa (TKC), Takehiko Kasahara (Toin Yokohama Univ.) TL2018-12 |
| Abstract |
(in Japanese) |
(See Japanese page) |
| (in English) |
The judicial system in which IT technology was introduced is called Cyber Court. In judicial field in Japan, confidential words such as personal names and personal places are converted to meaningless words manually because Japanese values do not allow the disclosure of the individual information. It is not easy to construct a comprehensive dictionary for detecting confidential words. We have already proposed two models that predict confidential words automatically by using neural networks. We used long short-term memory (LSTM) and continuous bag-of-words (CBOW) as our language models. Firstly, we explained the possibility of detecting the words surrounding a confidential word by using CBOW. Then, we proposed two models to predict the confidential words from the neighboring words by applying LSTM. The first model imitates the anonymization work by a human being, and the second model was based on CBOW. The results show that the first model is more effective for predicting confidential words than the simple LSTM model. We expected the second model to have paraphrasing ability to increase the possibility of finding other paraphraseable One is Bi-directional LSTM LR model and the other is Sum-LSTM based on the CBOW model. The two proposed models were effective for predicting all the words. However, only the Bi-directional LSTM LR model was effective for predicting confidential words. This could have happened because Sum-LSTM was based on CBOW. CBOW is an effective model for paraphrasing words; therefore, Sum-LSTM also has that mechanism. Therefore, when Sum-LSTM predicted a word whose answer was “confidential,” the CW_PPL (Confidential word perplexity) became worse because there was a possibility of paraphrasing words of paraphrasing words such as “plaintiff,” “defendant,” “doctor,” and “teacher.” Knowing the paraphrased words of the confidential words meant that the embedding vectors of the confidential words could be successfully generated. This meant that the model could recognize the meaning of “confidential.” However, the prediction accuracy did not improve; therefore, there was a problem in calculating the probability of the prediction task. To solve this problem, we could exclude these paraphraseable words from the choices when calculating the probability. It is also important to examine scores other than PPL. Then we consider to improve accuracy of predicting confidential words. At first we focus the parameter of neural network. We would change the window size to large because we think the surrounding words of the target words (that is window size) is larger, accuracy will be high. Then, we are experimenting but we can’t get the results by now to spent more memory and more calculating time. Next, we would use the proper noun dictionary combined neural network. But we don’t discuss in the this paper. |
| Keyword |
(in Japanese) |
(See Japanese page) |
| (in English) |
Bi-directional LSTM / CBOW / ppl / cw-ppl / window size / / / |
| Reference Info. |
IEICE Tech. Rep., vol. 118, no. 163, TL2018-12, pp. 1-6, July 2018. |
| Paper # |
TL2018-12 |
| Date of Issue |
2018-07-21 (TL) |
| ISSN |
Print edition: ISSN 0913-5685 Online edition: ISSN 2432-6380 |
Copyright and reproduction |
All rights are reserved and no part of this publication may be reproduced or transmitted in any form or by any means, electronic or mechanical, including photocopy, recording, or any information storage and retrieval system, without permission in writing from the publisher. Notwithstanding, instructors are permitted to photocopy isolated articles for noncommercial classroom use without fee. (License No.: 10GA0019/12GB0052/13GB0056/17GB0034/18GB0034) |
| Download PDF |
TL2018-12 |
| Conference Information |
| Committee |
TL |
| Conference Date |
2018-07-28 - 2018-07-29 |
| Place (in Japanese) |
(See Japanese page) |
| Place (in English) |
Keio University |
| Topics (in Japanese) |
(See Japanese page) |
| Topics (in English) |
Human Language Processing and Learning |
| Paper Information |
| Registration To |
TL |
| Conference Code |
2018-07-TL |
| Language |
English |
| Title (in Japanese) |
(See Japanese page) |
| Sub Title (in Japanese) |
(See Japanese page) |
| Title (in English) |
Improve accuracy of predicting confidential words in juditial precedents |
| Sub Title (in English) |
|
| Keyword(1) |
Bi-directional LSTM |
| Keyword(2) |
CBOW |
| Keyword(3) |
ppl |
| Keyword(4) |
cw-ppl |
| Keyword(5) |
window size |
| Keyword(6) |
|
| Keyword(7) |
|
| Keyword(8) |
|
| 1st Author's Name |
Masakazu Kanazawa |
| 1st Author's Affiliation |
Utsunomiya University (Utsunomiya Univ.) |
| 2nd Author's Name |
Atsushi Ito |
| 2nd Author's Affiliation |
Utsunomiya University (Utsunomiya Univ.) |
| 3rd Author's Name |
Yuya Kiryu |
| 3rd Author's Affiliation |
KDDI Corporation (KDDI) |
| 4th Author's Name |
Kazuyuki Yamaswa |
| 4th Author's Affiliation |
TKC Corporation (TKC) |
| 5th Author's Name |
Takehiko Kasahara |
| 5th Author's Affiliation |
Toin Yokohama Univeecity (Toin Yokohama Univ.) |
| 6th Author's Name |
|
| 6th Author's Affiliation |
() |
| 7th Author's Name |
|
| 7th Author's Affiliation |
() |
| 8th Author's Name |
|
| 8th Author's Affiliation |
() |
| 9th Author's Name |
|
| 9th Author's Affiliation |
() |
| 10th Author's Name |
|
| 10th Author's Affiliation |
() |
| 11th Author's Name |
|
| 11th Author's Affiliation |
() |
| 12th Author's Name |
|
| 12th Author's Affiliation |
() |
| 13th Author's Name |
|
| 13th Author's Affiliation |
() |
| 14th Author's Name |
|
| 14th Author's Affiliation |
() |
| 15th Author's Name |
|
| 15th Author's Affiliation |
() |
| 16th Author's Name |
|
| 16th Author's Affiliation |
() |
| 17th Author's Name |
|
| 17th Author's Affiliation |
() |
| 18th Author's Name |
|
| 18th Author's Affiliation |
() |
| 19th Author's Name |
|
| 19th Author's Affiliation |
() |
| 20th Author's Name |
|
| 20th Author's Affiliation |
() |
| 21st Author's Name |
|
| 21st Author's Affiliation |
() |
| 22nd Author's Name |
|
| 22nd Author's Affiliation |
() |
| 23rd Author's Name |
|
| 23rd Author's Affiliation |
() |
| 24th Author's Name |
|
| 24th Author's Affiliation |
() |
| 25th Author's Name |
|
| 25th Author's Affiliation |
() |
| 26th Author's Name |
/ / |
| 26th Author's Affiliation |
()
() |
| 27th Author's Name |
/ / |
| 27th Author's Affiliation |
()
() |
| 28th Author's Name |
/ / |
| 28th Author's Affiliation |
()
() |
| 29th Author's Name |
/ / |
| 29th Author's Affiliation |
()
() |
| 30th Author's Name |
/ / |
| 30th Author's Affiliation |
()
() |
| 31st Author's Name |
/ / |
| 31st Author's Affiliation |
()
() |
| 32nd Author's Name |
/ / |
| 32nd Author's Affiliation |
()
() |
| 33rd Author's Name |
/ / |
| 33rd Author's Affiliation |
()
() |
| 34th Author's Name |
/ / |
| 34th Author's Affiliation |
()
() |
| 35th Author's Name |
/ / |
| 35th Author's Affiliation |
()
() |
| 36th Author's Name |
/ / |
| 36th Author's Affiliation |
()
() |
| Speaker |
Author-1 |
| Date Time |
2018-07-28 10:00:00 |
| Presentation Time |
30 minutes |
| Registration for |
TL |
| Paper # |
TL2018-12 |
| Volume (vol) |
vol.118 |
| Number (no) |
no.163 |
| Page |
pp.1-6 |
| #Pages |
6 |
| Date of Issue |
2018-07-21 (TL) |
|