Boxun Xu
Ph.D. in Electrical & Computer Engineering, UC Santa Barbara
I received my Ph.D. in Electrical and Computer Engineering at UC Santa Barbara, advised by Prof. Peng Li (IEEE Fellow). My research focuses on efficient generative models, multimodal content generation, and ML systems & hardware co-design, building toward scalable, real-time multimodal and world models. I received consecutive William J. McCalla Best Paper Award nominations at ICCAD 2024 and ICCAD 2025.
I interned at Meta (2024) and Meta Superintelligence Labs (2025-2026), where I integrated Video Sparse Attention into MovieGen-30B, delivering 1.55× tuning-free end-to-end speedup, and extending it from inference to large-scale sparse distillation across 256 H100s.
Prior to UCSB, I received my M.S. in Electrical and Computer Engineering from the University of Michigan, Ann Arbor, advised by Prof. David Blaauw (IEEE Fellow) and Prof. Dennis Sylvester (IEEE Fellow), and my B.S. in Electronic Engineering from the University of Electronic Science and Technology of China.
Research Focus
- Efficient Generative Modeling and Multimodal & Interactive World Modeling
- Hardware / Algorithm Co-design & ML Systems & Electronic Design Automation
News
| 2026.09 | VAR-Q will appear at NeurIPS 2026! Our project page is now live, featuring tuning-free KV cache quantization for autoregressive image and video generation, with results across nine models. Read the paper. |
|---|---|
| 2026.07 | Successfully defended my Ph.D. dissertation, “Co-Architecting Models and Systems for Efficient Perception and Generation,” at UC Santa Barbara. |
| 2026.07 | Received Second Place at the DAC Ph.D. Forum and a travel grant from ACM SIGDA / IEEE CEDA. |
| 2026.05 | Recognized as an ICML 2026 Gold Reviewer. |
| 2026.04 | Our Sparse Forcing preprint is available: native trainable sparse attention for real-time autoregressive video generation, in collaboration with Meta Superintelligence Labs. Paper. |
| 2026.04 | Awarded the Radhakrishnan Nagarajan Family Fellowship at UCSB. |
| 2026.02 | Paper on VLM hallucination mitigation (VEGAS) accepted at CVPR 2026 Findings. |
| 2025.11 | Papers on adaptive KV caching for visual autoregressive models and KAN-based graph contrastive learning accepted at AAAI 2026. |
| 2025.10 | 🏆 Paper on 3D MoE spiking transformers nominated for the William J. McCalla Best Paper Award at ICCAD 2025 — second consecutive year. |
| 2025.06 | Paper on 3D MoE spiking transformer acceleration accepted at ICCAD 2025. |
| 2025.05 | Paper on transfer learning for Vmin prediction in advanced nodes accepted at ITC 2025. |
| 2025.04 | Paper on heterogeneous quantization for spiking vision transformers accepted at ASAP 2025. |
| 2025.03 | Paper on heterogeneous-core acceleration of spiking transformers with error-constrained pruning accepted at ISCA 2025. |
| 2025.01 | Paper on network-hardware co-optimization for sparse SNN accelerators accepted at TCAD as a long paper. |
| 2025.01 | Joining Meta Superintelligence Labs this summer in Seattle, working on efficient movie generation. |
| 2024.10 | 🏆 Paper on 3D spiking transformer accelerators nominated for the William J. McCalla Best Paper Award at ICCAD 2024. |
| 2024.07 | Papers on 3D spiking transformer accelerators and LLM-guided analog design accepted at ICCAD 2024. |
| 2024.06 | Started summer internship at Meta AI, working on multi-teacher knowledge distillation of multi-modal multi-task foundation models. |
| 2024.05 | Paper on a multi-modal IoT SoC with on-chip MRAM accepted at JSSC. |
Selected Publications
I have published papers in top conferences in machine learning / computer architecture / design automation, including NeurIPS, ISCA, AAAI, CVPR, ICCV, ICCAD, TCAD and JSSC.
Efficient Generative Modeling
- AAAI’26
★ AMS-KV: Adaptive KV Caching in Multi-Scale Visual Autoregressive TransformersIn AAAI Conference on Artificial Intelligence (main track)(Acceptance Rate: 17.6%) , 2026First efficient KV-caching design tailored for multi-scale visual AR transformers. - Preprint
★ Sparse Forcing: Native Trainable Sparse Attention for Real-time Autoregressive Diffusion Video Generation2026First native trainable sparse-attention framework enabling real-time autoregressive video generation.Work done during internship at Meta Superintelligence Labs. - NeurIPS’26
★ VAR-Q: Tuning-free KV Cache Quantization for Visual Autoregressive Image and Video GenerationIn Advances in Neural Information Processing Systems (NeurIPS), 2026Extended image and video paper. Earlier version appeared at the ICCV 2025 Binary and Extreme Quantization workshop.Tuning-free channel-block KV quantization across nine image and video autoregressive models. - CVPR’26
VEGAS: Mitigating Hallucinations in Large Vision-Language Models via Vision-Encoder Attention Guided Adaptive SteeringIn IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Findings, 2026
Hardware/Algorithm Co-design and EDA
- ISCA’25
★ Bishop: Sparsified Bundling Spiking Transformers on Heterogeneous Cores with Error-Constrained PruningIn International Symposium on Computer Architecture (ISCA)(Acceptance Rate: 22.2%) , 2025First SW/HW co-design framework for neuromorphic transformers. - ICCAD’25
🏆 Nominated as William J. McCalla Best Paper Award in 2025★ 3D Acceleration for Mixture-of-Experts and Multi-Head Attention Spiking Transformers with Dynamic Head PruningIn ACM/IEEE International Conference on Computer-Aided Design (ICCAD)(Acceptance Rate: 24.7%) , 2025First 3D-integrated accelerator for Mixture-of-Experts spiking transformers with dynamic head pruning. - ICCAD’24
🏆 Nominated as William J. McCalla Best Paper Award in 2024★ Spiking Transformer Hardware Accelerators in 3D IntegrationIn ACM/IEEE International Conference on Computer-Aided Design (ICCAD)(Acceptance Rate: 24%) , 2024First 3D-integrated hardware accelerator for spiking transformers. - TCAD’25
SpikeX: Exploring Accelerator Architecture and Network-Hardware Co-Optimization for Sparse Spiking Neural NetworksIn IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems(TCAD), 2025 - ASAP’25
Trimming Down Large Spiking Vision Transformers via Heterogeneous Quantization SearchIn IEEE International Conference on Application-specific Systems, Architectures and Processors (ASAP), 2025 - TMLR
DS2TA: Denoising Spiking Transformer with Attenuated Spatiotemporal AttentionIn Transactions on Machine Learning Research (TMLR, under review), 2024 - ICCAD’24
ADO-LLM: Analog Design Bayesian Optimization with In-Context Learning of Large Language ModelsIn ACM/IEEE International Conference on Computer-Aided Design (ICCAD), 2024First work to bring LLMs into analog circuit design, pairing in-context priors with Bayesian optimization for sample-efficient sizing. - Preprint
LASER: Language Model Regression for Semi-Structured Workflow Resource and Runtime Estimation2026Under review at Transactions on Machine Learning Research (TMLR) - ITC’25
Transfer Learning for Minimum Operating Voltage Prediction in Advanced Technology Nodes: Leveraging Legacy Data and Silicon Odometer SensingIn ACM/IEEE International Test Conference (ITC), 2025 - JSSC’24
AIMMI: Audio and Image Multi-Modal Intelligence via a Low-Power SoC With 2-MByte On-Chip MRAM for IoT DevicesIn IEEE Journal of Solid-State Circuits(JSSC), 2024 - VLSI’22
Audio and Image Cross-Modal Intelligence via a 10TOPS/W 22nm SoC with Back-Propagation and Dynamic Power GatingIn 2022 IEEE Symposium on VLSI Technology and Circuits (VLSI-Symposium), 2022
Other Publications
- AAAI’26
Khan-GCL: Kolmogorov-Arnold Network Based Graph Contrastive Learning with Hard NegativesIn AAAI Conference on Artificial Intelligence (main track)(Acceptance Rate: 17.6%) , 2026