ご案内 入会して研究会活動をもっとお得に!研究会参加費・年間登録費が会員価格になります。
お知らせ 【重要】研究会参加費の支払いおよび原稿アップロード手続きの変更に関するご案内
電子情報通信学会 研究会発表申込システム
講演論文 詳細
技報閲覧サービス
[ログイン]
技報アーカイブ
 トップに戻る 前のページに戻る   [Japanese] / [English] 

講演抄録/キーワード
講演名 2012-08-02 17:00
An Optimal Parallel Prefix-sums Algorithm on the Memory Machine Models for GPUs
Koji NakanoHiroshima Univ.CPSY2012-15
抄録 (和) 本論文では,GPU向けの理論計算モデルDMMとUMM上の最適な並列接頭部和アルゴリズムを示す.
これらのモデルは,3つのパラメタ,スレッド数$p$,メモリ幅$w$,メモリアクセスレイテンシ$l$を持つ.
まず,$n$個の数の合計が$O({n\over w}+{nl\over p}+l\log n)$時間で求められることを示す.
そして,合計を求める計算が少なくとも$\Omega({n\over w}+{nl\over p}+l\log n)$時間
必要であることを示す.
最後に,接頭部和が最適な$O({n\over w}+{nl\over p}+l\log n)$ 時間で求められることを示す. 
(英) The main contribution of this paper is to show optimal algorithms
computing the sum and the prefix-sums on two memory machine models,
the Discrete Memory Machine (DMM) and the Unified Memory Machine (UMM).
The DMM and the UMM are theoretical parallel computing models
that capture the essence of the shared memory and the global memory of GPUs.
These models have three parameters, the number $p$ of threads, the width $w$ of
the memory, and the memory access latency $l$.
We first show that the sum of $n$ numbers can be computed
in $O({n\over w}+{nl\over p}+l\log n)$ time units on the DMM and the UMM.
We then go on to show that $\Omega({n\over w}+{nl\over p}+l\log n)$ time units
are necessary to compute the sum.
Finally, we show an optimal parallel algorithm that computes the prefix-sums of
$n$ numbers in $O({n\over w}+{nl\over p}+l\log n)$ time units on the DMM and the UMM.
キーワード (和) メモリマシンモデル / 接頭部和 / 並列アルゴリズム / GPU / CUDA / / /  
(英) Memory machine models / Prefix-sums computation / Parallel algorithms / GPU / CUDA / / /  
文献情報 信学技報, vol. 112, no. 173, CPSY2012-15, pp. 37-42, 2012年8月.
資料番号 CPSY2012-15 
発行日 2012-07-26 (CPSY) 
ISSN Print edition: ISSN 0913-5685    Online edition: ISSN 2432-6380
著作権に
ついて
技術研究報告に掲載された論文の著作権は電子情報通信学会に帰属します.(許諾番号:10GA0019/12GB0052/13GB0056/17GB0034/18GB0034)
PDFダウンロード CPSY2012-15

研究会情報
研究会 DC CPSY  
開催期間 2012-08-02 - 2012-08-03 
開催地(和) とりぎん文化会館 
開催地(英) Torigin Bunka Kaikan 
テーマ(和) 2012年並列/分散/協調処理に関する『鳥取』サマー・ワークショップ(SWoPP鳥取2012) 
テーマ(英) Summer United Workshops on Parallel, Distributed and Cooperative Processing "Tottori" (SWoPP 2012) 
講演論文情報の詳細
申込み研究会 CPSY 
会議コード 2012-08-DC-CPSY 
本文の言語 英語 
タイトル(和)  
サブタイトル(和)  
タイトル(英) An Optimal Parallel Prefix-sums Algorithm on the Memory Machine Models for GPUs 
サブタイトル(英)  
キーワード(1)(和/英) メモリマシンモデル / Memory machine models  
キーワード(2)(和/英) 接頭部和 / Prefix-sums computation  
キーワード(3)(和/英) 並列アルゴリズム / Parallel algorithms  
キーワード(4)(和/英) GPU / GPU  
キーワード(5)(和/英) CUDA / CUDA  
キーワード(6)(和/英) /  
キーワード(7)(和/英) /  
キーワード(8)(和/英) /  
第1著者 氏名(和/英/ヨミ) 中野 浩嗣 / Koji Nakano / ナカノ ヒロツグ
第1著者 所属(和/英) 広島大学 (略称: 広島大)
Hiroshima Univ. (略称: Hiroshima Univ.)
第2著者 氏名(和/英/ヨミ) / /
第2著者 所属(和/英) (略称: )
(略称: )
第3著者 氏名(和/英/ヨミ) / /
第3著者 所属(和/英) (略称: )
(略称: )
第4著者 氏名(和/英/ヨミ) / /
第4著者 所属(和/英) (略称: )
(略称: )
第5著者 氏名(和/英/ヨミ) / /
第5著者 所属(和/英) (略称: )
(略称: )
第6著者 氏名(和/英/ヨミ) / /
第6著者 所属(和/英) (略称: )
(略称: )
第7著者 氏名(和/英/ヨミ) / /
第7著者 所属(和/英) (略称: )
(略称: )
第8著者 氏名(和/英/ヨミ) / /
第8著者 所属(和/英) (略称: )
(略称: )
第9著者 氏名(和/英/ヨミ) / /
第9著者 所属(和/英) (略称: )
(略称: )
第10著者 氏名(和/英/ヨミ) / /
第10著者 所属(和/英) (略称: )
(略称: )
第11著者 氏名(和/英/ヨミ) / /
第11著者 所属(和/英) (略称: )
(略称: )
第12著者 氏名(和/英/ヨミ) / /
第12著者 所属(和/英) (略称: )
(略称: )
第13著者 氏名(和/英/ヨミ) / /
第13著者 所属(和/英) (略称: )
(略称: )
第14著者 氏名(和/英/ヨミ) / /
第14著者 所属(和/英) (略称: )
(略称: )
第15著者 氏名(和/英/ヨミ) / /
第15著者 所属(和/英) (略称: )
(略称: )
第16著者 氏名(和/英/ヨミ) / /
第16著者 所属(和/英) (略称: )
(略称: )
第17著者 氏名(和/英/ヨミ) / /
第17著者 所属(和/英) (略称: )
(略称: )
第18著者 氏名(和/英/ヨミ) / /
第18著者 所属(和/英) (略称: )
(略称: )
第19著者 氏名(和/英/ヨミ) / /
第19著者 所属(和/英) (略称: )
(略称: )
第20著者 氏名(和/英/ヨミ) / /
第20著者 所属(和/英) (略称: )
(略称: )
第21著者 氏名(和/英/ヨミ) / /
第21著者 所属(和/英) (略称: )
(略称: )
第22著者 氏名(和/英/ヨミ) / /
第22著者 所属(和/英) (略称: )
(略称: )
第23著者 氏名(和/英/ヨミ) / /
第23著者 所属(和/英) (略称: )
(略称: )
第24著者 氏名(和/英/ヨミ) / /
第24著者 所属(和/英) (略称: )
(略称: )
第25著者 氏名(和/英/ヨミ) / /
第25著者 所属(和/英) (略称: )
(略称: )
第26著者 氏名(和/英/ヨミ) / /
第26著者 所属(和/英) (略称: )
(略称: )
第27著者 氏名(和/英/ヨミ) / /
第27著者 所属(和/英) (略称: )
(略称: )
第28著者 氏名(和/英/ヨミ) / /
第28著者 所属(和/英) (略称: )
(略称: )
第29著者 氏名(和/英/ヨミ) / /
第29著者 所属(和/英) (略称: )
(略称: )
第30著者 氏名(和/英/ヨミ) / /
第30著者 所属(和/英) (略称: )
(略称: )
第31著者 氏名(和/英/ヨミ) / /
第31著者 所属(和/英) (略称: )
(略称: )
第32著者 氏名(和/英/ヨミ) / /
第32著者 所属(和/英) (略称: )
(略称: )
第33著者 氏名(和/英/ヨミ) / /
第33著者 所属(和/英) (略称: )
(略称: )
第34著者 氏名(和/英/ヨミ) / /
第34著者 所属(和/英) (略称: )
(略称: )
第35著者 氏名(和/英/ヨミ) / /
第35著者 所属(和/英) (略称: )
(略称: )
第36著者 氏名(和/英/ヨミ) / /
第36著者 所属(和/英) (略称: )
(略称: )
講演者 第1著者 
発表日時 2012-08-02 17:00:00 
発表時間 30分 
申込先研究会 CPSY 
資料番号 CPSY2012-15 
巻番号(vol) vol.112 
号番号(no) no.173 
ページ範囲 pp.37-42 
ページ数
発行日 2012-07-26 (CPSY) 


[研究会発表申込システムのトップページに戻る]

[電子情報通信学会ホームページ]


IEICE / 電子情報通信学会