Posts by Collection

portfolio

publications

Secure On-Device Video OOD Detection Without Backpropagation

Published in International Conference on Computer Vision, 2025

SecDOOD: a secure cloud-device collaboration framework for efficient on-device OOD detection without requiring device-side backpropagation

Recommended citation: @article{li2025secure, title={Secure on-device video ood detection without backpropagation}, author={Li, Shawn and Cai, Peilin and Zhou, Yuxiao and Ni, Zhiyu and Liang, Renjie and Qin, You and Nian, Yi and Tu, Zhengzhong and Hu, Xiyang and Zhao, Yue}, journal={arXiv preprint arXiv:2503.06166}, year={2025} }
Download Paper

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations

Published in arxiv preprint, 2025

PersonaConvBench: a large-scale benchmark for evaluating personalized reasoning and generation in multi-turn conversations with LLMs

Recommended citation: @article{li2025personalized, title={A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations}, author={Li, Li and Cai, Peilin and Rossi, Ryan A and Dernoncourt, Franck and Kveton, Branislav and Wu, Junda and Yu, Tong and Song, Linxin and Yang, Tiankai and Qin, Yuehan and others}, journal={arXiv preprint arXiv:2505.14106}, year={2025} }
Download Paper

The Earth Simulator: Street View World Modeling with 3D Gaussian Memory and Camera Control

Published in Submission, 2025

The Earth Simulator is a generative street-view world model that combines a 3D Gaussian spatial memory with camera-controlled video generation to synthesize long-horizon exploration videos from sparse pose-free images.

Recommended citation: @misc{cai2025earth, title={The Earth Simulator: Street View World Modeling with 3D Gaussian Memory and Camera Control}, author={Cai, Peilin and Yuan, Weiduo and He, Sicheng and Wu, Cho-Ying and Paz, David and Zhang, Hengyuan and Guo, Yuliang and Huang, Xinyu and Ren, Liu and Mao, Jiageng and Wang, Yue}, note={In submission}, year={2025} }

LAM: Language Articulated Object Modelers

Published in IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2026

LAM generates articulated 3D objects from text by coordinating LLM/VLM agents that design part hierarchies, write geometry and articulation code, debug failures, and refine results through visual feedback.

Recommended citation: @inproceedings{gao2026lam, title = {LAM: Language Articulated Object Modelers}, author = {Gao, Yipeng and Ge, Yunhao and Cai, Peilin and Seita, Daniel and Itti, Laurent}, booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, year = {2026} }
Download Paper

talks

teaching