Jingwei Cai
I received my Ph.D. in Computer Systems Architecture from the Institute for Interdisciplinary Information Sciences (IIIS), Tsinghua University in 2026, under the supervision of Prof. Kaisheng Ma. During my Ph.D., I worked on DNN accelerator architecture and compiler design for 2.5D/3D/wafer-scale chiplets, wafer-scale silicon photonic computing, recommendation-system acceleration, and cloud LLM inference optimization.
I am currently with the Heterogeneous Computing team at ByteDance SEED, where I study cloud AI workloads and optimize online LLM inference systems. I also develop early-stage architecture simulators and contribute to hardware-software co-design for emerging 3D DRAM and AI accelerator platforms, turning production requirements into practical system and architecture improvements.
My work connects academic research with real-world systems. I have published 7 papers, including 5 first-author papers in CCF-A computer architecture venues, and received the HPCA 2024 Distinguished Artifact Award. I have also led architecture and compiler development for multiple generations of Qiming DNN accelerators at Arctic Xiongxin, bringing research ideas into production compiler stacks and chip designs.
For more details, please see my CV. I am currently working at ByteDance SEED and remain open to exceptional opportunities where I can tackle ambitious problems and make a broader impact.
news
selected publications
- HPCACharacterizing Cloud-Native LLM Inference at ByteDance and Exposing Optimization Challenges and Opportunities for Future AI AcceleratorsIn IEEE International Symposium on High-Performance Computer Architecture, Industry Track , 2026
- HPCAIdentifying, Exploring, and Understanding the DRAM Communication Scheduling Space for DNN AcceleratorsIn IEEE International Symposium on High-Performance Computer Architecture , 2025
- DACDiscovering and Exploiting Untapped Buffer Resources in Many-Core DNN AcceleratorsIn Design Automation Conference , 2025
- HPCAGemini: Mapping and Architecture Co-exploration for Large-scale DNN Chiplet AcceleratorsIn IEEE International Symposium on High-Performance Computer Architecture , 2024
- ISCAInter-layer Scheduling Space Definition and Exploration for Tiled AcceleratorsIn Proceedings of the 50th Annual International Symposium on Computer Architecture , 2023