Memory-Efficient Visual Autoregressive Modeling with Scale-Aware KV Cache
NeurIPS 2025
Scale-aware cache management reduces KV memory by 90% without quality loss, making visual autoregressive generation faster and practical at resolutions up to 4K.
I am a first-year PhD student at UC Berkeley, interested in building efficient and scalable AI systems.
Previously, I graduated with honors from the National University of Singapore and spent a term at the University of Washington as an exchange student.
NeurIPS 2025
Scale-aware cache management reduces KV memory by 90% without quality loss, making visual autoregressive generation faster and practical at resolutions up to 4K.
CVPR 2025 Highlight (Top 3%)
An end-to-end learnable depth-pruning framework that halves Diffusion Transformer depth and parameters, delivering 2× faster inference at only 7% of the original training cost.
ACM/IEEE IPSN 2024 Best Demo Runner-Up
A low-power camera system that combines monochrome images with sensor signals and diffusion models to visualize information beyond the visible spectrum.

Ph.D. in Computer Science
Berkeley, California
B.S. in Computer Engineering
Singapore
Exchange Student
Seattle, WashingtonI feel most at home outdoors: hiking, exploring, and photographing the small moments along the way. I also follow basketball and tennis, and I am always happy to talk about research, sports, or the next trail to try.
Let’s connect