About Me

I am a master's student in Computational Science and Engineering at Harvard University, working with Prof. Mengyu Wang (Harvard AI and Robotics Lab) and Prof. Paul Pu Liang (MIT Media Lab). My research sits at the intersection of robotics, embodied AI, and 3D vision.

I received my B.S.E. in Computer Engineering from the University of Michigan, where I worked with Prof. Ram Vasudevan in the ROAHM Lab on differentiable models of deformable linear objects and multi-modal Bayesian perception.

My work builds physically grounded models that robots can actually act on β€” from real-time differentiable simulators for deformable objects, to perception that fuses vision and touch, to geometry-aware video world models for manipulation. I'm always open to research discussions and collaboration. Feel free to reach out via yuzhenchen@g.harvard.edu.

Research Interests

My goal is to give robots a physically grounded understanding of the world, the ability to reason about how it will change, and the means to act on it alongside people. My research spans three connected threads:

  • Predictive World Models: Building models that can imagine plausible futures β€” anticipating how a scene will change and what people intend β€” so robots can reason and act reliably over long horizons.
  • Human-Interactive, Agentic Robotics: Moving robots beyond one-shot instruction following toward agents that interact with people as fluently as language models do in software β€” able to explain their decisions, and to be redirected, corrected, and taught through natural interaction.
  • Active Perception & Embodied Understanding: Robots that do not just perceive passively, but actively decide where to look and how to move to understand a dynamic 3D world β€” its geometry, motion, and interactions β€” and act robustly within it.

News

  • [Aug. 2026] 🀝 Serving as a Junior Organizer of the 1st Workshop on Physical World AI (PhysWorldAI) at NeurIPS 2026.
  • [Jul. 2026] πŸ“„ Learning Ophthalmologist Clinical Reasoning for Glaucoma Diagnosis is out on medRxiv.
  • [Jul. 2026] πŸŽ‰ Act2See is accepted by the RSS 2026 FM4RoboPlan Workshop.
  • [Jul. 2026] πŸŽ‰ World Cognition Model (WCM) is accepted by the RSS 2026 Workshop on Human-centric Mobile Manipulation.
  • [Jun. 2026] πŸŽ“ Awarded the OMPP Summer Research Fellowship at Harvard University.
  • [Jun. 2026] πŸŽ‰ GEM-4D is accepted by ECCV 2026.
  • [Jun. 2026] πŸ† GEM-4D β€” Best Paper Award, Runner-Up at the CVPR 2026 WMAS Workshop.
  • [Jun. 2026] ✨ GEM-4D β€” Spotlight at the CVPR 2026 3D-LLM/VLA Workshop.
  • [Mar. 2026] πŸŽ‰ SLIM-VDB is accepted by IEEE RA-L.
  • [May. 2025] πŸ“Š MatPredict β€” a dataset and benchmark for learning material properties of indoor objects β€” is released.
  • [Feb. 2025] πŸ€– DEFT is on arXiv β€” differentiable branched discrete elastic rods for furcated DLOs.
  • [Sep. 2024] πŸ€– DEER is accepted by CoRL 2024.
  • [Apr. 2024] βœ‹ You've Got to Feel It to Believe It is accepted by RSS 2024.
  • [Oct. 2023] πŸ† Multi-Modal Semantic Perception Using Bayesian Inference β€” Best Poster at the IROS 2023 Workshop.

Publications

* denotes equal contribution.

  1. Teaser for Glaucoma Reasoning

    Learning Ophthalmologist Clinical Reasoning for Glaucoma Diagnosis from Fundus Images

    Kaichen Zhou*, Yuzhen Chen*, Elif Yildiz*, Min Shi, David Dai, Grace Chen, Jiale Zheng, He Wang, Fangneng Zhan, Chhavi Saini, Lucy Q. Shen, Yike Guo, Paul Pu Liang, Mengyu Wang

    The first clinically annotated fundus reasoning dataset β€” 1,077 photographs paired with expert-authored six-step diagnostic reports β€” plus a reasoning-driven vision–language framework that mirrors an ophthalmologist's diagnostic workflow, generating structured clinical reasoning before reaching a glaucoma diagnosis.

    medRxiv preprint Β· 2026 Medical AI
  2. Teaser for WCM

    WCM: World-Cognition Model for Generalizable Human-Robot Interaction

    Yuzhen Chen, Kaichen Zhou

    A human-centered embodied agent built on the SLAK architecture with an asynchronous runtime β€” it stays transparent about why it acts and can be interactively taught long-horizon tasks, reaching 73.8% success across nine real-world human–robot interaction tasks. Accepted to the RSS 2026 Workshop on Human-centric Mobile Manipulation.

    RSS 2026 Workshop on Human-centric Mobile Manipulation Β· 2026 Robotics
  3. Teaser for GEM-4D

    GEM-4D: Geometry-Enhanced Video World Models for Robot Manipulation

    Kaichen Zhou*, Yuzhen Chen*, Fangneng Zhan, Hang Hua, Grace Chen, Xinhai Chang, Ao Qu, Yilun Du, Zhuang Liu, Paul Pu Liang, Mengyu Wang

    Distills dense 4D correspondence supervision from a geometry foundation model into video world models, lifting real-world manipulation success from 61% to 81%. Best Paper Runner-Up (CVPR 2026 WMAS Workshop); Spotlight (CVPR 2026 3D-LLM/VLA Workshop).

    ECCV Β· 2026 Robotics

Education

Harvard University

Harvard University, Cambridge, MA

Sept. 2025 – May 2027 (Expected)

M.S. in Computational Science and Engineering

Coursework: Video Generation, Multi-modal World Models & VLA.

University of Michigan

University of Michigan, Ann Arbor, MI

Aug. 2021 – May 2025

B.S.E. in Computer Engineering Β· GPA 3.9 / 4.0

Machine Learning, Computer Vision, SLAM, Robot Perception, GPU Parallel Computation, Robot 3D Reconstruction.

Experience

Harvard AI and Robotics Lab

Harvard AI and Robotics Lab

Sept. 2025 – Present

Research Assistant Β· advised by Prof. Mengyu Wang

Geometry-enhanced video world models for robot manipulation (GEM-4D); VLM-guided active URDF reconstruction (Act2See); interpretable medical VLMs.

MIT Media Lab

Multisensory Intelligence Lab, MIT Media Lab

Sept. 2025 – Present

Research Assistant Β· advised by Prof. Paul Pu Liang

Geometry-enhanced video world models for robot manipulation (GEM-4D).

ROAHM Lab

ROAHM Lab

Apr. 2023 – May 2025

Research Assistant Β· advised by Prof. Ram Vasudevan

Real-time differentiable modeling of deformable linear objects (DEER); multi-modal Bayesian semantic perception; conformalized safe motion planning (CROWS); material-property learning from images (MatPredict, with Prof. Arpan Kusari).

Honors & Awards

  • [2026] πŸŽ“ OMPP Summer Research Fellowship, Harvard University.
  • [2026] πŸ† Best Paper Award, Runner-Up β€” CVPR 2026 WMAS Workshop.
  • [2026] ✨ Spotlight β€” CVPR 2026 3D-LLM/VLA Workshop.
  • [2023] πŸ₯‡ Best Poster Award β€” IROS 2023.
  • [2023–24] πŸŽ–οΈ James B. Angell Scholar, University of Michigan.
  • [2022–25] πŸ“œ Dean's List, University of Michigan.

Services & Teaching

  • [2026–] Robotics Community Event Organizer, MIT CHIEF.
  • [2022–25] Teaching Assistant, EECS 442 (Computer Vision) & Physics 140X, University of Michigan.
  • [2023–24] Departmental Ambassador, ENGR 110 & UMich Robotics Department.
  • [2023] Student Tutor, UMich Engineering Center for Academic Success.