About Me
I am a master's student in Computational Science and Engineering at Harvard University, working with Prof. Mengyu Wang (Harvard AI and Robotics Lab) and Prof. Paul Pu Liang (MIT Media Lab). My research sits at the intersection of robotics, embodied AI, and 3D vision.
I received my B.S.E. in Computer Engineering from the University of Michigan, where I worked with Prof. Ram Vasudevan in the ROAHM Lab on differentiable models of deformable linear objects and multi-modal Bayesian perception.
My work builds physically grounded models that robots can actually act on β from real-time differentiable simulators for deformable objects, to perception that fuses vision and touch, to geometry-aware video world models for manipulation. I'm always open to research discussions and collaboration. Feel free to reach out via yuzhenchen@g.harvard.edu.
Research Interests
My goal is to give robots a physically grounded understanding of the world, the ability to reason about how it will change, and the means to act on it alongside people. My research spans three connected threads:
- Predictive World Models: Building models that can imagine plausible futures β anticipating how a scene will change and what people intend β so robots can reason and act reliably over long horizons.
- Human-Interactive, Agentic Robotics: Moving robots beyond one-shot instruction following toward agents that interact with people as fluently as language models do in software β able to explain their decisions, and to be redirected, corrected, and taught through natural interaction.
- Active Perception & Embodied Understanding: Robots that do not just perceive passively, but actively decide where to look and how to move to understand a dynamic 3D world β its geometry, motion, and interactions β and act robustly within it.
News
- [Aug. 2026] π€ Serving as a Junior Organizer of the 1st Workshop on Physical World AI (PhysWorldAI) at NeurIPS 2026.
- [Jul. 2026] π Learning Ophthalmologist Clinical Reasoning for Glaucoma Diagnosis is out on medRxiv.
- [Jul. 2026] π Act2See is accepted by the RSS 2026 FM4RoboPlan Workshop.
- [Jul. 2026] π World Cognition Model (WCM) is accepted by the RSS 2026 Workshop on Human-centric Mobile Manipulation.
- [Jun. 2026] π Awarded the OMPP Summer Research Fellowship at Harvard University.
- [Jun. 2026] π GEM-4D is accepted by ECCV 2026.
- [Jun. 2026] π GEM-4D β Best Paper Award, Runner-Up at the CVPR 2026 WMAS Workshop.
- [Jun. 2026] β¨ GEM-4D β Spotlight at the CVPR 2026 3D-LLM/VLA Workshop.
- [Mar. 2026] π SLIM-VDB is accepted by IEEE RA-L.
- [May. 2025] π MatPredict β a dataset and benchmark for learning material properties of indoor objects β is released.
- [Feb. 2025] π€ DEFT is on arXiv β differentiable branched discrete elastic rods for furcated DLOs.
- [Sep. 2024] π€ DEER is accepted by CoRL 2024.
- [Apr. 2024] β You've Got to Feel It to Believe It is accepted by RSS 2024.
- [Oct. 2023] π Multi-Modal Semantic Perception Using Bayesian Inference β Best Poster at the IROS 2023 Workshop.
Publications
* denotes equal contribution.
-
Learning Ophthalmologist Clinical Reasoning for Glaucoma Diagnosis from Fundus Images
The first clinically annotated fundus reasoning dataset β 1,077 photographs paired with expert-authored six-step diagnostic reports β plus a reasoning-driven visionβlanguage framework that mirrors an ophthalmologist's diagnostic workflow, generating structured clinical reasoning before reaching a glaucoma diagnosis.
-
WCM: World-Cognition Model for Generalizable Human-Robot Interaction
A human-centered embodied agent built on the SLAK architecture with an asynchronous runtime β it stays transparent about why it acts and can be interactively taught long-horizon tasks, reaching 73.8% success across nine real-world humanβrobot interaction tasks. Accepted to the RSS 2026 Workshop on Human-centric Mobile Manipulation.
-
Act to See: Structured Articulated Representations through Robot Interaction
A robot autonomously inspects an unknown articulated object and builds a structured articulated representation (an explicit URDF) online, frame-by-frame, through active interaction. Accepted to the RSS 2026 FM4RoboPlan Workshop.
-
GEM-4D: Geometry-Enhanced Video World Models for Robot Manipulation
Distills dense 4D correspondence supervision from a geometry foundation model into video world models, lifting real-world manipulation success from 61% to 81%. Best Paper Runner-Up (CVPR 2026 WMAS Workshop); Spotlight (CVPR 2026 3D-LLM/VLA Workshop).
-
SLIM-VDB: A Real-Time 3D Probabilistic Semantic Mapping Framework
A real-time probabilistic semantic mapping framework built on sparse VDB volumes for large-scale 3D scenes.
-
MatPredict: A Dataset and Benchmark for Learning Material Properties of Diverse Indoor Objects
A dataset and benchmark pairing 18 indoor objects with 14 materials, for inferring physical material properties from camera images.
-
DEFT: Differentiable Branched Discrete Elastic Rods for Modeling Furcated DLOs in Real-Time
Extends differentiable discrete elastic rods to branched topologies, enabling real-time modeling of furcated deformable linear objects.
-
Differentiable Discrete Elastic Rods for Real-Time Modeling of Deformable Linear Objects
A differentiable discrete elastic rod formulation that identifies material parameters and predicts DLO dynamics in real time.
-
You've Got to Feel It to Believe It: Multi-Modal Bayesian Inference for Semantic and Property Prediction
Fuses vision and touch in a Bayesian framework to jointly infer object semantics and physical properties.
-
Multi-Modal Semantic Perception Using Bayesian Inference
Bayesian multi-modal fusion for semantic perception in robot manipulation. Best Poster, IROS 2023 Workshop.
No publications match this filter.
Education

Harvard University, Cambridge, MA
M.S. in Computational Science and Engineering
Coursework: Video Generation, Multi-modal World Models & VLA.

University of Michigan, Ann Arbor, MI
B.S.E. in Computer Engineering Β· GPA 3.9 / 4.0
Machine Learning, Computer Vision, SLAM, Robot Perception, GPU Parallel Computation, Robot 3D Reconstruction.
Experience

Harvard AI and Robotics Lab
Research Assistant Β· advised by Prof. Mengyu Wang
Geometry-enhanced video world models for robot manipulation (GEM-4D); VLM-guided active URDF reconstruction (Act2See); interpretable medical VLMs.

Multisensory Intelligence Lab, MIT Media Lab
Research Assistant Β· advised by Prof. Paul Pu Liang
Geometry-enhanced video world models for robot manipulation (GEM-4D).

ROAHM Lab
Research Assistant Β· advised by Prof. Ram Vasudevan
Real-time differentiable modeling of deformable linear objects (DEER); multi-modal Bayesian semantic perception; conformalized safe motion planning (CROWS); material-property learning from images (MatPredict, with Prof. Arpan Kusari).
Honors & Awards
- [2026] π OMPP Summer Research Fellowship, Harvard University.
- [2026] π Best Paper Award, Runner-Up β CVPR 2026 WMAS Workshop.
- [2026] β¨ Spotlight β CVPR 2026 3D-LLM/VLA Workshop.
- [2023] π₯ Best Poster Award β IROS 2023.
- [2023β24] ποΈ James B. Angell Scholar, University of Michigan.
- [2022β25] π Dean's List, University of Michigan.
Services & Teaching
- [2026β] Robotics Community Event Organizer, MIT CHIEF.
- [2022β25] Teaching Assistant, EECS 442 (Computer Vision) & Physics 140X, University of Michigan.
- [2023β24] Departmental Ambassador, ENGR 110 & UMich Robotics Department.
- [2023] Student Tutor, UMich Engineering Center for Academic Success.