【書報討論】9月9日(三)葉宗泰 教授 (國立陽明交通大學資工系)

2026-09-04 16:25:57

-演講時間: 115年9月9日(三) 14:00~16:00

-演講地點: E6-A207教室

-演講者: 葉宗泰 教授 (國立陽明交通大學資工系)

-演講主題: Energy-Efficient LUT-based GEMM Accelerator with Hardware-Aware KV Cache Quantization

-演講摘要:

The energy efficiency of Large Language Models (LLMs) model inference is becoming critical for edge AI applications. Recent quantization produces low-bit LLMs (1.58-bit weights) while minimizing the accuracy loss during model inference. However, the KV cache significantly increases model capacity, and the LLM decoding stage begins to dominate overall inference time as the context length grows. This talk will present our research on the energy-efficient GEMM accelerator for low-bit and mixed-precision LLM models. I will introduce our proposed Omni-LUT GEMM hardware accelerator architecture and a novel hardware-aware KV-cache quantization method. Our proposed hardware-software co-design architectural support enables efficient mixed-precision computations in the prefill and decode stages. Evaluations show that our proposed Omni-LUT accelerator significantly reduces model inference latency and energy consumption compared to state-of-the-art AI accelerators.

-講者簡歷:

Dr. Tsung Tai Yeh received his Ph.D degree from the Department of Electrical and Computer Engineering at Purdue University in 2020. He is currently an Associate Professor in the Computer Science (CS) Department at National Yang Ming Chiao Tung University (NYCU), Taiwan. Dr. Yeh has also been recognized with several awards. He received a Fellowship of the HEA from Advance HE in 2023, as well as the Excellent Mentor Award from NYCU in 2021 and 2025. He won the 2025 Qualcomm Innovation Fellowship. He is also a recipient of the 2030 Cross-Generation Young Scholars Program (NSTC), Taiwan, 2025. His research interests span computer architecture, computer systems, and programming languages, with a primary focus on the GPU, Domain-Specific Accelerators, and AI compiler systems. His compiler research was nominated for the Best Paper Award at the PPoPP conference, and his work was published in multiple top-ranking conference proceedings (ISCA, ASPLOS, HPCA, PPoPP, ICRA, NeurIPS).