I am actively seeking for 27 Summer Research Internship, with interest in Multiodal and Embodied AI. Feel free to reach out if you have any opportunities!
We propose Fast Spatial Memory (FSM), which leverages scalable elastic test-time training to enable efficient and high-quality 4D reconstruction from dynamic videos.
We propose Mirage, interleaving latent visual tokens, which represent compact imagery visual features, with explicit text tokens to solve diverse multimodal reasoning tasks, boosting the reasoning performance without the full pixel-level image generation.
We introduce VCA, a curiosity-driven video agent with self-exploration capability, which autonomously navigates video segments and efficiently builds a comprehensive understanding of complex video sequences.
We propose a simple yet effective MLLMs for language-instructed video segmentation. It emphasizes global-local video understanding and achieves SOTA performance on multiple benchmarks.
We show that synthetic videos and natural images can replace real videos for pre-training, achieving comparable or better performance while offering a controllable, transparent alternative.