| 講演抄録/キーワード |
| 講演名 |
2026-03-03 16:20
LLMを用いた医療面接の総合評価に関する研究 ○森山絢太(長崎大)・遠山修平・村井浩一(システック井上)・川尻真也・山梨啓友(長崎大)・小林 透(駒澤大)・今井哲郎(長崎大) LOIS2025-76 |
| 抄録 |
(和) |
医療面接教育では模擬患者を用いた訓練が普及している一方,総合評価は教員の負担が大きく,
評価者間のばらつきが課題である.本研究では,医学生--模擬患者アバターの会話ログ全文から
教員による12段階総合評価を推定する,LLMを用いた医療面接自動評価手法を検討した.
総合評価に必要な根拠情報を段階的に与えるため,
OSCEを参考とした質問項目リスト(39項目)達成情報および
RIASに基づく発話タグ付けを会話入力に付加し,その効果を比較した.
さらに,単一教員評価値の推定,複数教員評価値の平均推定,同時推定を比較した.
RIASタグ付き会話ログ71件(3症例)を用いた交差検証の結果,
質問項目達成情報とRIASタグを併用し複数教員評価値を同時学習した条件において,
許容誤差付き一致率84.5%,絶対誤差平均0.887を達成した.
教員間一致率39.4%とばらつきの大きい条件下でも,
構造化情報の付加と学習問題設計により実用的な総合評価推定が可能であることを示した. |
| (英) |
Medical interview education widely employs training with simulated patients; however,
assigning overall performance scores places a substantial burden on faculty and is
affected by considerable inter-rater variability.
This study investigates an LLM-based automatic evaluation method that estimates
12-point global ratings assigned by faculty directly from full conversation logs
between medical students and simulated patient avatars.
To provide evidence required for global assessment in a stepwise manner,
achievement information derived from an OSCE-referenced question item list (39 items)
and utterance-level annotations based on the Roter Interaction Analysis System (RIAS)
were incorporated into the conversation input, and their effects were compared.
In addition, we evaluated different learning problem settings, including estimation of
single-rater scores, estimation of averaged multi-rater scores, and simultaneous
estimation of multiple rater scores.
Cross-validation experiments using 71 RIAS-annotated interview logs from three cases
demonstrated that the joint incorporation of question item achievement information
and RIAS tags with simultaneous multi-rater learning achieved the best performance,
yielding a tolerance-based accuracy of 84.5% and a mean absolute error of 0.887.
These results indicate that, even under conditions of substantial inter-rater
variability (inter-rater agreement of 39.4%), practical estimation of global
interview performance is feasible through the integration of structured intermediate
evaluation information and appropriate learning problem design. |
| キーワード |
(和) |
医療面接 / 評価の予測 / 大規模言語モデル / RIAS / OSCE / / / |
| (英) |
medical interview / global rating prediction / large language model / RIAS / OSCE / / / |
| 文献情報 |
信学技報, vol. 125, no. 375, LOIS2025-76, pp. 162-167, 2026年3月. |
| 資料番号 |
LOIS2025-76 |
| 発行日 |
2026-02-23 (LOIS) |
| ISSN |
Online edition: ISSN 2432-6380 |
著作権に ついて |
技術研究報告に掲載された論文の著作権は電子情報通信学会に帰属します.(許諾番号:10GA0019/12GB0052/13GB0056/17GB0034/18GB0034) |
| PDFダウンロード |
LOIS2025-76 |