Jingwei Cai

photo_jingwei.jpg

I received my Ph.D. in Computer Systems Architecture from the Institute for Interdisciplinary Information Sciences (IIIS), Tsinghua University in 2026, under the supervision of Prof. Kaisheng Ma. During my Ph.D., I worked on DNN accelerator architecture and compiler design for 2.5D/3D/wafer-scale chiplets, wafer-scale silicon photonic computing, recommendation-system acceleration, and cloud LLM inference optimization.

I am currently with the Heterogeneous Computing team at ByteDance SEED, where I study cloud AI workloads and optimize online LLM inference systems. I also develop early-stage architecture simulators and contribute to hardware-software co-design for emerging 3D DRAM and AI accelerator platforms, turning production requirements into practical system and architecture improvements.

My work connects academic research with real-world systems. I have published 7 papers, including 5 first-author papers in CCF-A computer architecture venues, and received the HPCA 2024 Distinguished Artifact Award. I have also led architecture and compiler development for multiple generations of Qiming DNN accelerators at Arctic Xiongxin, bringing research ideas into production compiler stacks and chip designs.

For more details, please see my CV. I am currently working at ByteDance SEED and remain open to exceptional opportunities where I can tackle ambitious problems and make a broader impact.

news

I received the Wang Dazhong Scholarship and was nominated by IIIS for Tsinghua University’s Top-Grade Scholarship. :trophy:
I was named a 2025 ByteDance Scholar, one of 20 doctoral recipients from China and Singapore. :tada:
I joined the Heterogeneous Computing team at ByteDance SEED, working on workload characterization, chip co-design, and LLM inference optimization. :rocket:
I was selected for the inaugural CAST Young Talent Support Program for Doctoral Students as IIIS’s sole recipient. :sparkles:
Our work SoMa was accepted to HPCA 2025. :tada:

selected publications

  1. HPCA
    Characterizing Cloud-Native LLM Inference at ByteDance and Exposing Optimization Challenges and Opportunities for Future AI Accelerators
    Jingwei Cai, Dehao Kong , Hantao Huang , and 8 more authors
    In IEEE International Symposium on High-Performance Computer Architecture, Industry Track , 2026
  2. HPCA
    Identifying, Exploring, and Understanding the DRAM Communication Scheduling Space for DNN Accelerators
    Jingwei Cai, Xuan Wang , Mingyu Gao , and 5 more authors
    In IEEE International Symposium on High-Performance Computer Architecture , 2025
  3. DAC
    Discovering and Exploiting Untapped Buffer Resources in Many-Core DNN Accelerators
    Yuchen Wei , Jingwei Cai, Mingyu Gao , and 4 more authors
    In Design Automation Conference , 2025
  4. HPCA
    Gemini: Mapping and Architecture Co-exploration for Large-scale DNN Chiplet Accelerators
    Jingwei Cai, Zuotong Wu , Sen Peng , and 5 more authors
    In IEEE International Symposium on High-Performance Computer Architecture , 2024
  5. ISCA
    Inter-layer Scheduling Space Definition and Exploration for Tiled Accelerators
    Jingwei Cai, Yuchen Wei , Zuotong Wu , and 2 more authors
    In Proceedings of the 50th Annual International Symposium on Computer Architecture , 2023