Tailors: Accelerating Sparse Tensor Algebra by Overbooking Buffer Occupancy @ MICRO 2023
October 28, 2023
562
10
1.78%
Search the Record
IndexedEvery word spoken in this episode is indexed. Type any phrase to jump straight to the moment it was said.
Type any word or phrase that may have been spoken. Click a result to seek the player to that exact moment.
Try a name, a topic, or a quoted line
Mit Eems Group Pi Vivienne Sze Episodes Around October 28, 2023
See what was published immediately before and after this episode.
46:08Systematic Modeling and Design of Sparse Tensor Accelerators [Nellie Wu]
4:08Now PlayingTailors: Accelerating Sparse Tensor Algebra by Overbooking Buffer Occupancy @ MICRO 2023
YouTube Description
as posted by the channelZ. Y. Xue, Y. N. Wu, J. S. Emer, V. Sze, "Tailors: Accelerating Sparse Tensor Algebra by Overbooking Buffer Occupancy," ACM/IEEE International Symposium on Microarchitecture (MICRO), October 2023
Project Website
Abstract: Sparse tensor algebra is a challenging class of workloads to accelerate due to low arithmetic intensity and varying sparsity patterns. Prior sparse tensor algebra accelerators have explored tiling sparse data to increase exploitable data reuse and improve throughput, but typically allocate tile size in a given buffer for the worst-case data occupancy. This severely limits the utilization of available memory resources and reduces data reuse. Other accelerators employ complex tiling during preprocessing or at runtime to determine the exact tile size based on its occupancy.
This paper proposes a speculative tensor tiling approach, called overbooking, to improve buffer utilization by taking advantage of the distribution of nonzero elements in sparse tensors to construct larger tiles with greater data reuse. To ensure correctness, we propose a low-overhead hardware mechanism, Tailors, that can tolerate data overflow by design while ensuring reasonable data reuse. We demonstrate that Tailors can be easily integrated into the memory hierarchy of an existing sparse tensor algebra accelerator. To ensure high buffer utilization with minimal tiling overhead, we introduce a statistical approach, Swiftiles, to pick a tile size so that tiles usually fit within the buffer's capacity, but can potentially overflow, i.e., it overbooks the buffers. Across a suite of 22 sparse tensor algebra workloads, we show that our proposed overbooking strategy introduces an average speedup of 52.7× and 2.3× and an average energy reduction of 22.5× and 2.5× over ExTensor without and with optimized tiling, respectively.
Information about accessibility can be found at
Links & Promotions
Guests & Subjects Covered
Sentinel Indexing in Progress
Metadata and chapters are available. Claim extraction for this episode is pending.
All video content is delivered via YouTube embedded players in accordance with the YouTube Terms of Service. Sentinel provides research tools that promote discovery and accountability across political media.







