| 講演抄録/キーワード |
| 講演名 |
2026-03-04 13:50
APNを用いた分散KVキャッシュ共有による大規模言語モデル推論のリソース効率向上 ○近藤汰一・小野翔多・中尾彰宏(東大) NS2025-243 |
| 抄録 |
(和) |
大規模言語モデルを用いたサービスでは,推論時に保持される Key-Value(KV)キャッシュが大容量であ
るため,計算資源の容量を圧迫し,同時に処理可能なユーザ数が制約されるという課題がある.本研究の目的は,KVキャッシュ容量に起因する GPU メモリ制約を緩和し,応答品質を維持したまま多数のユーザを収容可能な推論基盤を実現することである.本研究では,All-Photonics Network(APN)を用いた分散 KV キャッシュ共有機構を提案する.実験の結果,動的な KV キャッシュ移送により負荷偏在時における収容可能ユーザ数を 1.38 倍に,GPU メモリ使用率を 1.32 倍に拡大できることを示す.さらに,KV キャッシュを再生成する場合と比較して Time-To-First-Token(TTFT)を 53.9% 低減できることを確認し,計算資源の制約下におけるサービス継続性向上に寄与することを示す. |
| (英) |
In services powered by large language models, the Key–Value (KV) cache maintained during inference consumes a large amount of memory, placing significant pressure on computational resources and limiting the number of users that can be served concurrently. The objective of this study is to mitigate GPU memory constraints caused by the KV cache and to realize an inference infrastructure capable of accommodating many users while maintaining response quality. To achieve this goal, we propose a distributed KV cache sharing mechanism over the All-Photonics Network (APN). Experimental results show that dynamic KV cache migration increases the number
of supported users by 1.38 × and improves GPU memory utilization by 1.32 × under load imbalance. Furthermore, compared with regenerating the KV cache, the proposed method reduces the Time-To-First-Token (TTFT) by 53.9%, thereby contributing to improved service continuity under limited computational resources. |
| キーワード |
(和) |
LLM / KVキャッシュ / APN / / / / / |
| (英) |
LLM / KV Cache / All-Photonics Network / / / / / |
| 文献情報 |
信学技報, vol. 125, no. 385, NS2025-243, pp. 131-136, 2026年3月. |
| 資料番号 |
NS2025-243 |
| 発行日 |
2026-02-25 (NS) |
| ISSN |
Online edition: ISSN 2432-6380 |
著作権に ついて |
技術研究報告に掲載された論文の著作権は電子情報通信学会に帰属します.(許諾番号:10GA0019/12GB0052/13GB0056/17GB0034/18GB0034) |
| PDFダウンロード |
NS2025-243 |