ExTensor: An Accelerator for Sparse Tensor Algebra

Generalized tensor algebra is a prime candidate for acceleration via customized ASICs. Modern tensors feature a wide range of data sparsity, with the density of non-zero elements ranging from 10^-6% to 50%. This paper proposes a novel approach to accelerate tensor kernels based on the principle of hierarchical elimination of computation in the presence of sparsity. This approach relies on rapidly finding intersections -- situations where both operands of a multiplication are non-zero -- enabling new data fetching mechanisms and avoiding memory latency overheads associated with sparse kernels implemented in software. We propose the ExTensor accelerator, which builds these novel ideas on handling sparsity into hardware to enable better bandwidth utilization and compute throughput. We evaluate ExTensor on several kernels relative to industry libraries (Intel MKL) and state-of-the-art tensor algebra compilers (TACO). When bandwidth normalized, we demonstrate an average speedup of 3.4x, 1.3x, 2.8x, 24.9x, and 2.7x on SpMSpM, SpMM, TTV, TTM, and SDDMM kernels respectively over a server class CPU.

Authors

Kartik Hegde (University of Illinois at Urbana-Champaign)

Hadi Asghari-Moghaddam (University of Illinois at Urbana-Champaign)

Michael Pellauer

Neal Crago

Aamer Jaleel

Edgar Solomonik (University of Illinois at Urbana-Champaign)

Joel Emer

Christopher W. Fletcher (University of Illinois Urbana-Champaign)

Publication Date

Saturday, October 12, 2019

Published in

International Symposium on Microarchitecture (MICRO)

Research Area

Computer Architecture

External Links

ACM Digital Library

Uploaded Files

Published manuscript738.64 KB

Award

IEEE Micro Top Picks in Computer Architecture (Honorable Mention)

Copyright

Copyright by the Association for Computing Machinery, Inc. Permission to make digital or hard copies of part or all of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for components of this work owned by others than ACM must be honored. Abstracting with credit is permitted. To copy otherwise, to republish, to post on servers, or to redistribute to lists, requires prior specific permission and/or a fee. Request permissions from Publications Dept, ACM Inc., fax +1 (212) 869-0481, or permissions@acm.org. The definitive version of this paper can be found at ACM's Digital Library http://www.acm.org/dl/.