I am Bowen Zhang (张博文), a Research Scientist/Engineer on the Omni team at Luma AI, where I contribute to Uni-1, Luma's unified understanding-and-generation foundation model. I did my Ph.D. at the University of Science and Technology of China (USTC) (2021-2026), enrolled in the USTC-Microsoft Research Asia (MSRA) Joint-PhD Program and co-advised by Dr. Baining Guo and Prof. Feng Zhao. Before that, I received my Bachelor's Degree in the School of the Gifted Young, USTC, in 2021.
My research focuses on multimodal foundation models that unify language, vision, reasoning, and generation, building on my earlier work on 2D/3D/4D generative modeling. Before joining Luma AI, I was a research intern on the SAM 3D team at Meta Superintelligence Labs, mentored by Dr. Hao Tang. Prior to that, I spent over three years as a research intern in the Visual Computing Group at Microsoft Research Asia, where I had the opportunity to learn from and collaborate with many incredible researchers, including Bo Zhang, Jianmin Bao, Dong Chen and Jiaolong Yang.
For a detailed overview of my work, please see my CV. If my background aligns with your interests, please don't hesitate to reach out.
We present a generalizable model that reconstructs 3D geometry, texture and layout of objects from a single natural image, trained with a human- and model-in-the-loop annotation pipeline that produces visually grounded 3D data at unprecedented scale.
We present a novel video-to-4D framework that performs diffusion modeling on a compressed representation of Gaussian Splats animations, enabling high-quality generation from a single-view video input.
We introduce a unified Structured LATent (SLAT) representation for versatile and high-quality 3D asset creation, which is achieved by integrating a sparsely-populated 3D grid with dense multiview visual features extracted from a powerful vision foundation model, comprehensively capturing both geometry and appearance information while maintaining flexibility during decoding.
We present a structured and explicit representation for 3D generative modeling by structuring Gaussian Splatting using Optimal Transport, achieving state-of-the-art generation quality.
We introduce RodinHD, which can generate high-fidelity 3D avatars from a frontal view portrait image.
We propose an ID-preserving talking head generation framework and achieves fast personalized adaptation through the proposed meta-learning approach.
A transformer-based GAN crafted for high-resolution image synthesis, which achieves SOTA generation quality on the resolution of 1024, which for the first time, approaches the best performed ConvNets.
My research focuses on multimodal foundation models that unify language, vision, reasoning, and generation. I am particularly interested in scalable training and post-training data pipelines, instruction fine-tuning and alignment, and agentic systems that can plan, act, and self-correct across understanding, generation, and editing.
I Love playing Rubik Cubes in my spare time, and I have taken part in world level competitions many times. Making decisions about what to do next in just seconds is really exciting, and you can check out my personal best record at the World Cube Association website.
I also love music and singing. I have made a home studio and have covered some of my favorite songs, you can check out these covers at my radio in NetEase Cloud Music. Besides, I'm now trying to write my songs.