About Me

I am an undergraduate at MIT studying Physics and Artificial Intelligence & Decision Making, expected to graduate in 2027. I am an undergraduate researcher in Professor Kaiming He’s group, where I have worked on computer vision, generative modeling and (most recently) multi-agent systems.

My research experience spans diffusion models, one-step image generation, and multimodal training. Alongside this work, I build training infrastructure, including JAX/TPU workflows, distributed optimization, and tools for managing research experiments. I am especially interested in large-scale model training and applying my research experience to build real-world AI products that benefit people broadly.

I am now seeking full-time Research Engineer or Machine Learning Engineer roles, with particular interest in teams working on agents, multimodal models, or training infrastructure. Please feel free to reach out about relevant opportunities or collaborations.

Beyond academics, I also enjoy engaging with people who share similar interests and chatting about anything from research ideas to life experiences. Feel free to reach out if you’d like to connect!

My resume is linked here.

Publications & Projects


ELF: Embedded Language Flows

K. Hu*, L. Qiu*, Y. Lu, H. Zhao, T. Li, Y. Kim, J. Andreas, and K. He


[Paper] [Code]


One-step Latent-free Image Generation with Pixel Mean Flows

Y. Lu*, S. Lu*, Q. Sun*, H. Zhao*, Z. Jiang, X. Wang, T. Li, Z. Geng, and K. He

(ICML 2026)


[Paper] [Code]


Bidirectional Normalizing Flow: From Data to Noise and Back

Y. Lu*, Q. Sun*, X. Wang*, Z. Jiang, H. Zhao, and K. He

(CVPR 2026)


[Paper] [Code]


Is Noise Conditioning Necessary for Denoising Generative Models?

Q. Sun*, Z. Jiang*, H. Zhao*, and K. He

(ICML 2025)


[Paper]

Other Projects

  • Speeding Up Diffusion Models with One-step Generators

    This is the final project for the seminar course 6.S978: Deep Generative Models at MIT. In the project, we proposed a new method to speed up the training of diffusion models by using one-step generators. On toy experiments, this reduces NFE by half while maintaining the sample quality. We also wrote a blog post, explaining the motivation of the experiment from a higher perspective.