About Me

Open to Collaboration

Hi, I am Zhenyu Yi (伊振宇). I am now a first-year M.S. student at Shanghai Jiao Tong University, advised by Lichi Zhang and Qian Wang. I also worked as a Research Intern at DAMO Academy, Alibaba Group. Before that, I received my B.S. degree from Huazhong University of Science and Technology (2021-2025), where I collaborated with Qiang Hu.

My research focuses on computer vision and medical AI, with particular interest in endoscopic image&video analysis and brain imaging diagnosis. I am also interested in data-limited learning in medical scenarios, including learning with noisy labels, semi-supervised learning, and weakly supervised learning, to improve model robustness under realistic clinical annotation settings. In parallel, I explore multimodal representation learning through vision-language alignment and pretraining, and I further work on multimodal large language models, especially efficient MLLMs for clinically grounded reasoning and deployment.

Research Interests

Endoscopic Image and Video Analysis Noisy, Semi-, and Weak Supervision Representation Learning Multimodal Learning Efficient MLLMs

Highlights

  • 🏆
    MICCAI Oral
    Lead published work: SALI
  • 📄
    8+ Papers
    Published / Under review

🔥 News

  • 2026.08 ⚡ Our paper PhenoMIL was accepted by Medical Image Analysis (MedIA).
  • 2026.08 ⚡ One paper CARVE was posted on arXiv.
  • 2026.06EndoVLM and Brain-Adapter were accepted to MICCAI 2026.
  • 2026.06E-MRL and CerviThink were accepted to MICCAI 2026.
  • 2026.02 ⚡ One paper SAMIX was accepted by CVPR 2026. Congratulations to Qiang Hu!
  • 2024.12 ⚡ One paper MonoBox was accepted by AAAI 2025.
  • 2024.09 ⚡ Our paper SALI was invited as an Oral presentation (<3%) in MICCAI 2024.

📝 Selected Publications

Note: * indicates co-first author.

MICCAI 2026
EndoVLM teaser
EndoVLM: An Endoscopy Vision-Language Pre-training Model via Anatomy-Guided Sparsity and Progressive Alignment
Zhenyu Yi*, Jianwei Xu*, Yue Hu, Zhongwei Qiu, Sijing Li, Liang Huang, Bin Lv, Ling Zhang, Yingda Xia.
  • An endoscopy VLM pretraining framework with anatomy-guided sparsity and progressive alignment.
MICCAI 2026
Brain-Adapter teaser
Brain-Adapter: A Dual-Stream Vision-Language MIL Framework for Comprehensive 3D CT Diagnosis of Acute Intracranial Pathologies
Zhenyu Yi, Zhiyun Song, Yusong Sun, Zelin Liu, Manman Fei, Zhenhao Li, Jiaxuan Zhao, Xu Han, Lichi Zhang.
  • A dual-stream vision-language MIL framework for 3D CT intracranial pathology diagnosis.
MedIA
PhenoMIL teaser
Learning from Limited Phenotype-Level Annotations for Promoting Multiple Instance Learning in Endoscopic Helicobacter pylori Infection Diagnosis
Zhenyu Yi*, Jianwei Xu*, Yue Hu*, Liang Huang, Zhilin Zheng, Haifeng Jin, Panpan Ma, Tanzhou Chen, Jie Pan, Xiaoyun Ding, Fangfang Zhang, Jiang Liu, Xiaoteng Wang, Yingda Xia, Bin Lv, Ling Zhang.
  • A phenotype-aware MIL framework for endoscopic H. pylori infection diagnosis with limited fine-grained annotations.
MICCAI 2024 Oral
SALI teaser
SALI: Short-Term Alignment and Long-Term Interaction Network for Colonoscopy Video Polyp Segmentation
Qiang Hu*, Zhenyu Yi*, Ying Zhou, Fang Peng, Mei Liu, Qiang Li, Zhiwei Wang.
  • A hybrid temporal interaction network for colonoscopy video polyp segmentation.
CVPR 2026
SAMIX teaser
SAMIX: Reinforcing SAM2 with Semantic Adapter and Reference Selecting Policy for Mix-Supervised Segmentation
Qiang Hu, Jiajie Wei, Zhenyu Yi, Zhifen Yan, Yingjie Guo, Hongkuan Shi, Ge-Peng Ji, Qiang Li, and Zhiwei Wang.
  • A mix-supervised segmentation framework with semantic adaptation and RL-based reference selection.
AAAI 2025
MonoBox teaser
MonoBox: Tightness-free Box-supervised Polyp Segmentation using Monotonicity Constraint
Qiang Hu, Zhenyu Yi, Ying Zhou, Fan Huang, Mei Liu, Qiang Li, Zhiwei Wang.
  • A box-supervised polyp segmentation method with a tightness-free monotonicity constraint.
MICCAI 2026
E-MRL teaser
E-MRL: Cross-view Aligned Evidence-driven Multimodal Reinforcement Learning for Reliable 3D Tumor Analysis
Sijing Li, Zhongwei Qiu, Zhuoya Wang, Boxiang Yun, Zhenyu Yi, Jianwei Xu, Wenqiao Zhang, Yingda Xia, Ling Zhang.
  • A cross-view evidence-driven multimodal RL framework for reliable 3D tumor analysis.
MICCAI 2026
CerviThink teaser
CerviThink: A Reinforced Visual Reasoning Framework for Cervical Cancer Cell Classification
Manman Fei, Haotian Jiang, Zhenyu Yi, Qian Wang, Lichi Zhang.
  • An RL-based visual reasoning framework for cervical cancer cell classification.
arXiv 2026
CARVE teaser
CARVE: Cross-Slice Anisotropic Reallocation of Visual Evidence for Efficient 3D Medical Volume Understanding
Zhenyu Yi, Qiang Hu, Zhenhao Li, Jiaxuan Zhao, Yusong Sun, Lichi Zhang.
  • A training-free token allocation framework for efficient 3D medical volume understanding with cross-slice evidence reallocation.

🎖 Honors and Awards

  • 2025.06   Outstanding Graduate, Huazhong University of Science and Technology.
  • 2025.05   Future Technology Taihu Scholarship.
  • 2024.10   First Class Scholarship, Huazhong University of Science and Technology.
  • 2024.05   Future Technology Taihu Scholarship.
  • 2023.10   First Class Scholarship, Huazhong University of Science and Technology.
  • 2022.10   First Class Scholarship, Huazhong University of Science and Technology.

🎓 Education

  • 2025-present   M.S. Student, Shanghai Jiao Tong University.
  • 2021-2025   B.S. Student, Huazhong University of Science and Technology.