Bowen Zhang

Research Scientist/Engineer at Luma AI
bowen.zh@outlook.com

I am Bowen Zhang (张博文), a Research Scientist/Engineer on the Omni team at Luma AI, where I contribute to Uni-1, Luma's unified understanding-and-generation foundation model. I did my Ph.D. at the University of Science and Technology of China (USTC) (2021-2026), enrolled in the USTC-Microsoft Research Asia (MSRA) Joint-PhD Program and co-advised by Dr. Baining Guo and Prof. Feng Zhao. Before that, I received my Bachelor's Degree in the School of the Gifted Young, USTC, in 2021.

My research focuses on multimodal foundation models that unify language, vision, reasoning, and generation, building on my earlier work on 2D/3D/4D generative modeling. Before joining Luma AI, I was a research intern on the SAM 3D team at Meta Superintelligence Labs, mentored by Dr. Hao Tang. Prior to that, I spent over three years as a research intern in the Visual Computing Group at Microsoft Research Asia, where I had the opportunity to learn from and collaborate with many incredible researchers, including Bo Zhang, Jianmin Bao, Dong Chen and Jiaolong Yang.

For a detailed overview of my work, please see my CV. If my background aligns with your interests, please don't hesitate to reach out.


Education

University of Science and Technology of China

School of Information Science and Technology
Ph.D., USTC & Microsoft Research Asia Joint Ph.D. Program
September 2021 - June 2026

University of Science and Technology of China

School of the Gifted Young
Major in Computer Science & Member of the USTC Honor Class of Artificial Intelligence
September 2017 - June 2021

Linyi NO.1 High School of Shandong Province


September 2015 - June 2017

Experience

Luma AI

Research Scientist/Engineer, Omni Team
Contributor to Uni-1, Luma's unified understanding-and-generation foundation model.
February 2026 - Present

Meta Superintelligence Labs

Research Intern, SAM 3D Team
Researched generalizable 3D generative modeling, mentored by Dr. Hao Tang.
June 2025 - November 2025

Microsoft Research Asia

Research Intern, Visual Computing Group
Researched high-quality visual content creation and 2D/3D/4D generative modeling, mentored by Dr. Dong Chen and Dr. Baining Guo.
March 2022 - June 2025

Publications

SAM 3D: 3Dfy Anything in Images

SAM 3D Team, including Bowen Zhang

We present a generalizable model that reconstructs 3D geometry, texture and layout of objects from a single natural image, trained with a human- and model-in-the-loop annotation pipeline that produces visually grounded 3D data at unprecedented scale.

CVPR 2026
Best Paper Honorable Mention

Gaussian Variation Field Diffusion for High-fidelity Video-to-4D Synthesis

Bowen Zhang, Sicheng Xu, Chuxin Wang, Jiaolong Yang, Feng Zhao, Dong Chen, Baining Guo

We present a novel video-to-4D framework that performs diffusion modeling on a compressed representation of Gaussian Splats animations, enabling high-quality generation from a single-view video input.

ICCV 2025

Structured 3D Latents for Scalable and Versatile 3D Generation

Jianfeng Xiang, Zelong Lv, Sicheng Xu, Yu Deng, Ruicheng Wang, Bowen Zhang, Dong Chen, Xin Tong, Jiaolong Yang

We introduce a unified Structured LATent (SLAT) representation for versatile and high-quality 3D asset creation, which is achieved by integrating a sparsely-populated 3D grid with dense multiview visual features extracted from a powerful vision foundation model, comprehensively capturing both geometry and appearance information while maintaining flexibility during decoding.

CVPR 2025 Highlight

GaussianCube: A Structured and Explicit Radiance Representation for 3D Generative Modeling

Bowen Zhang, Yiji Cheng, Jiaolong Yang, Chunyu Wang, Feng Zhao, Yansong Tang, Dong Chen, Baining Guo.

We present a structured and explicit representation for 3D generative modeling by structuring Gaussian Splatting using Optimal Transport, achieving state-of-the-art generation quality.

NeurIPS 2024

RodinHD: High-Fidelity 3D Avatar Generation with Diffusion Models

Bowen Zhang*, Yiji Cheng*, Chunyu Wang, Ting Zhang, Jiaolong Yang, Yansong Tang, Feng Zhao, Dong Chen, Baining Guo.

We introduce RodinHD, which can generate high-fidelity 3D avatars from a frontal view portrait image.

ECCV 2024

MetaPortrait: Identity-Preserving Talking Head Generation with Fast Personalized Adaptation

Bowen Zhang*, Chenyang Qi*, Pan Zhang, Bo Zhang, HsiangTao Wu, Dong Chen, Qifeng Chen, Yong Wang, Fang Wen.

We propose an ID-preserving talking head generation framework and achieves fast personalized adaptation through the proposed meta-learning approach.

CVPR 2023

StyleSwin: Transformer-based GAN for High-resolution Image Generation

Bowen Zhang, Shuyang Gu, Bo Zhang, Jianmin Bao, Dong Chen, Fang Wen, Yong Wang, Baining Guo.

A transformer-based GAN crafted for high-resolution image synthesis, which achieves SOTA generation quality on the resolution of 1024, which for the first time, approaches the best performed ConvNets.

CVPR 2022

Awards

  • 2025 National Scholarship (Top 0.2%)
  • 2021 USTC Outstanding Graduate Student (Top 5%)
  • 2020 USTC Outstanding Student Scholarship, Gold Award
  • 1st Place - ISC2020 SCC (International Super Computing Student Cluster Competition)
  • 2019 8700, Special Class for Gifted Young Innovation Scholarship
  • 2019 USTC Outstanding Student Scholarship, Silver Award
  • 3rd Place - 2019 PAC (Parallel Application Competition)
  • 2018 USTC Outstanding Student Scholarship, Silver Award
  • 2018 iGEM USTC Software, Silver Award
  • 2018 USTC 60th Anniversary Outstanding Volunteer
  • 2017 USTC Outstanding Student Scholarship, Bronze Award

Interests

My research focuses on multimodal foundation models that unify language, vision, reasoning, and generation. I am particularly interested in scalable training and post-training data pipelines, instruction fine-tuning and alignment, and agentic systems that can plan, act, and self-correct across understanding, generation, and editing.

I Love playing Rubik Cubes in my spare time, and I have taken part in world level competitions many times. Making decisions about what to do next in just seconds is really exciting, and you can check out my personal best record at the World Cube Association website.

I also love music and singing. I have made a home studio and have covered some of my favorite songs, you can check out these covers at my radio in NetEase Cloud Music. Besides, I'm now trying to write my songs.


Skills

Programming Languages & Tools
  • C/C++
  • Python
  • Latex
  • Linux
Language
  • Chinese
  • English
    • TOFEL: Total 101 (Reading 28 Listening 23 Speaking 23 Writing 27)
Others
  • Basic music theory
  • Basic music recording and mixing techniques

Services

Conference Reviewer
  • IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
  • International Conference on Computer Vision (ICCV)
  • European Conference on Computer Vision (ECCV)
  • AAAI Conference on Artificial Intelligence (AAAI)
  • Annual Conference on Neural Information Processing Systems (NeurIPS)
Journal Reviewer
  • International Journal of Computer Vision (IJCV)