Ji Xie

I am a first-year PhD student at CMU LTI, advised by Prof. Louis-Philippe Morency. Previously, I was a research intern at ByteDance Seed and a visiting student at the Berkeley AI Research (BAIR) Lab, UC Berkeley, advised by Prof. XuDong Wang and Prof. Trevor Darrell.

My research focuses on scalable, unified, controllable multimodal models, with a long-term goal of developing world models for embodied intelligence.

Email  /  Google Scholar  /  CV  /  GitHub  /  Twitter

profile avatar

Industry Experience

Seedream 5.0 series (Lite / Pro)
Core contributor to interactive precision editing

Led precise color control and palette-guided image generation and editing, supporting Hex color codes and external color swatches.

Core contributor to layer separation, text edit, and spatial grounding for precise local editing.


tech blog / project page

Favorite Works

Foundation Models

reca Reconstruction Alignment Improves Unified Multimodal Models
Ji Xie, Trevor Darrell, Luke Zettlemoyer, Xudong Wang
ICLR 2026
paper / code / model

Unlocking the massive zero-shot potential in unified multimodal models through self-supervised learning.

Application

paint-anything Paint-Anything: Toward Any-Color Controllable Image Generation and Editing
Ji Xie, Dewei Zhou, Xinyu Huang, Xun Wang
Seed Technical Report 2026

Paint-Anything enables unified, prompt-native 24-bit hex-color control for both image generation and editing in a single model.

VideoCoF: Unified Video Editing with Temporal Reasoner
Xiangpeng Yang, Ji Xie, Yiyuan Yang, Yue Ma, Yan Huang, Min Xu, Qiang Wu
CVPR 2026 (Highlight)
paper / code / model / project page

VideoCoF uses a see-reason-edit pipeline for mask-free, precise video editing and strong long-video extrapolation.

icedit In-Context Edit: Enabling Instructional Image Editing with In-Context Generation in Large-Scale Diffusion Transformer
Zechuan Zhang, Ji Xie, Yu Lu, Zongxin Yang, Yi Yang
NeurIPS 2025
paper / code (2K Stars🌟) / model

Image editing is worth a single LoRA! With only 0.1% training data, ICEdit delivers fantastic instructional image editing.

3dis 3DIS: Depth-Driven Decoupled Instance Synthesis for Text-to-Image Generation
Dewei Zhou*, Ji Xie*, Zongxin Yang, Yi Yang
(* denotes equal contribution)
ICLR 2025 (Spotlight)
paper / code / model

3DIS uses depth-driven decoupled instance synthesis for controllable text-to-image generation.

Experience

Carnegie Mellon University seal
First-year PhD Student, Language Technologies Institute, Carnegie Mellon University
Advisor: Louis-Philippe Morency
August 2026 ~ Present
ByteDance icon
Research Intern, ByteDance Seed
Mentor: Xun Wang
October 2025 ~ August 2026
UC Berkeley seal Visiting Student, Berkeley AI Research (BAIR), UC Berkeley
March 2025 ~ December 2025

Invited Talks

"Reconstruction Alignment Improves Unified Multimodal Models"
Apple Research · Invited Talk · Hosted by Chen Chen and Yinfei Yang
October 2025

Selected Honors & Awards

SenseTime Scholarship
Top 30 recipients annually in China
June 2025
Zhejiang Provincial Higher Mathematics Competition, First Prize
June 2024
Zhejiang Provincial Collegiate Programming Contest, Gold Medal
April 2024, April 2023
International Collegiate Programming Contest (ICPC), Shenyang Site Gold Medal
October 2022
China Collegiate Programming Contest (CCPC), Guangzhou Site Gold Medal
October 2022

Miscellaneous

I have competitive-programming experience in ACM/ICPC and achieved a rating of 2478 on Codeforces. You can find my old blog here — it contains my competitive-programming notes :)


Website template from Jon Barron