I am an Associate Research Fellow at Zhejiang Lab. My research
focuses on multimodal foundation vision-language models, with an emphasis on their applications in Earth
science. My work spans visual perception and representation learning, aiming to build visual and
embodied intelligence systems with strong generalization capability.
I received my Ph.D. from KAIST in 2023,
during which I was awarded the Robert Bosch Ph.D. Fellowship and the
Qualcomm Innovation
Fellowship. After my Ph.D., I conducted postdoctoral research at the
University of Michigan, working with Prof.
Stella X. Yu on large-scale
vision-language models.
目前于之江实验室任副研究员,主要研究多模态基础视觉语言模型,并探索其在地球科学领域的应用。
研究方向包括视觉感知与表征学习,致力于构建强泛化能力的视觉及具身智能系统。
2023 年于韩国科学技术院(KAIST)取得博士学位,博士在读期间曾获
罗伯特‑博世博士奖学金(Robert Bosch PhD Fellowship)与
高通创新奖学金(Qualcomm Innovation Fellowship)。
博士毕业后,于密歇根大学(University of Michigan)从事博士后研究,与
Stella X. Yu 教授合作开展大规模视觉语言模型相关的工作。
Multi-modal Vision-Language Models and their applications to Earth science.
多模态视觉语言模型及其在地球科学中的应用。
Embodied AI and robotics — 3D perception, robotic grasping, VLA model fine-tuning, and on-device AI model training and deployment.
具身智能与机器人方向 — 3D 视觉感知、机械臂抓取、VLA 模型微调、端侧AI模型训练与部署。
Large-scale vision-language models, diffusion models, and ViTs, with Prof. Stella X. Yu.
与 Stella X. Yu 教授合作,研究大规模视觉语言模型、扩散模型与 ViTs。
YOLO-based multi-camera vehicle detection with GAN-based synthetic data augmentation.
基于 YOLO 的多路摄像头车辆检测,结合 GAN 合成数据增强。
Domain-adaptive and self-supervised semantic segmentation using GANs, Cut & Mixing, and geometry-guided motion cues.
基于 GAN、Cut & Mixing 、3D视觉运动信息的域自适应与自监督语义分割研究。
-
[1]Unsupervised Whole Object Discovery by Contextual Grouping with RepulsionNeurIPS 2025 Workshop.Existing self-supervised ViT features (e.g., DINO) mainly segment local object parts. This work introduces attraction-and-repulsion cues into a graph-cut framework to refine these features, enabling unsupervised segmentation of complete objects across image and video datasets without any prior knowledge of object categories. 现有的自监督 ViT 特征(如 DINO)主要针对对象局部区域进行分割;本工作在图割框架中引入吸引与排斥机制, 进一步优化 ViT 特征,使其能够在图像和视频数据集中实现对完整对象的无监督分割,而无需任何关于对象的先验知识。 Read More展开摘要
-
[2]DHR: Dual Features-Driven Hierarchical Rebalancing in Inter- and Intra-Class Regions for Weakly-Supervised Semantic SegmentationECCV, 2024.
-
[3]Fine-grained Background Representation for Weakly Supervised Semantic SegmentationIEEE Trans. on Circuits and Systems for Video Technology, 2024.IEEE 电路与系统视频技术汇刊,2024。
-
Best Paper Award最佳论文奖
[4]MoDA: Leveraging Motion Priors from Videos for Advancing Unsupervised Domain Adaptive SegmentationCVPR Workshop (Learning with Limited Labelled Data for Image and Video Understanding)(有限标注数据学习在图像与视频理解中的应用研讨会), 2024.We consider a practical UDA setting where the target domain contains sequential video frames. MoDA is a motion-guided domain adaptive segmentation framework that uses self-supervised object motion to learn effective target-domain representations, handling foreground and background domain gaps with different strategies — foreground object discovery and semantic mining guided by instance-level motion, and background adversarial training with a category-specific discriminator. MoDA outperforms prior methods on multiple domain-adaptive image and video segmentation benchmarks and can be combined with existing SOTA approaches for further gains. 我们考虑目标域包含连续视频帧的实际 UDA 场景,设计了运动引导的域自适应分割框架 MoDA, 利用自监督物体运动学习有效的目标域表征,分别处理前景与背景的域差异 —— 前景通过实例级运动引导进行对象发现与语义挖掘, 背景通过类别特定判别器进行对抗训练。MoDA 在多个域自适应图像与视频分割基准上超越已有方法,并可与现有 SOTA 方法结合以进一步提升性能。 Read More展开摘要
-
Highlight (Acceptance Rate < 2.0%)Highlight(录用率 < 2.0%)
[5]ImageNet-D: Benchmarking Neural Network Robustness on Diffusion Synthetic ObjectCVPR, 2024.We use diffusion models as a data source to synthesize hard images that benchmark robustness, providing more diversified backgrounds, textures, and materials — with higher synthetic quality — than prior benchmarks such as ImageNet-C/-9 and Stylized ImageNet. ImageNet-D causes a significant accuracy drop across vision models, from standard ResNet classifiers to foundation models like CLIP and MiniGPT-4, reducing accuracy by up to 60%. 我们利用扩散模型作为数据源合成困难图像用于鲁棒性基准测试,相比 ImageNet-C、ImageNet-9、 Stylized ImageNet 等以往基准,在背景、纹理和材料上提供了更高的多样性与合成质量。实验表明,ImageNet-D 导致从 ResNet 到 CLIP、MiniGPT-4 等大模型的视觉模型准确率显著下降,最高降幅达 60%。 Read More展开摘要
-
[6]Zero-shot Building Attribute Extraction from Large-Scale Vision and Language ModelsWACV, 2024.We propose a zero-shot workflow for building attribute extraction that uses large-scale vision-language models to reduce reliance on human-annotated data, combining image-level and segment-level captioning with domain vocabularies from structural and civil engineering to semantically match visual and textual representations. 提出一种利用大规模视觉语言模型进行建筑属性提取的零样本工作流,通过图像级与分割级描述结合结构与土木工程领域词汇, 实现视觉与文本表征的语义匹配,从而减少对人工标注数据的依赖。 Read More展开摘要
-
Best Paper Nominated最佳论文提名
[7]Masking-augmented Collaborative Domain Congregation for Multi-target Domain Adaptation in Semantic SegmentationIEEE Intelligent Vehicles Symposium (IV), 2024.MacDC handles both style and contextual gaps among multiple target domains through collaborative domain congregation (image- and region-level data mixing) and multi-context masking consistency, learning a single model for multi-target adaptation without multiple networks or distillation. MacDC 通过协同域聚合(图像级与区域级数据混合)与多上下文掩码一致性,同时处理多个目标域间的风格差异与上下文差异, 无需多网络训练或蒸馏即可学习单一模型完成多目标域自适应。 Read More展开摘要
-
[8]CCTV-Calib: A Toolbox to Calibrate Surveillance Cameras Around the GlobeMachine Vision and Applications, 2023.Machine Vision and Applications 期刊,2023。CCTV-Calib is a user-friendly toolbox to calibrate perspective and fisheye surveillance cameras using satellite views, estimating intrinsic/extrinsic parameters and GPS location with an automated keypoint-matching stage, without restrictive structural or semantic assumptions. CCTV-Calib 是一个基于卫星视图标定透视与鱼眼监控摄像头的易用工具箱,通过自动化关键点匹配估计相机内外参与 GPS 位置, 无需依赖限制性的结构或语义假设。 Read More展开摘要
-
[9]ML-BPM: Multi-teacher Learning with Bidirectional Photometric Mixing for Open Compound Domain Adaptation in Semantic SegmentationECCV, 2022.We introduce a multi-teacher framework with automatic domain separation and bidirectional photometric mixing to separately adapt to every target subdomain, followed by adaptive distillation and consistency regularization to train a generalizable student model, achieving SOTA on compound and open-domain segmentation benchmarks. 提出多教师框架,通过自动域划分与双向光度混合分别适配每个目标子域,并结合自适应蒸馏与一致性正则化训练出泛化能力强的学生模型, 在复合域与开放域分割基准上取得当前最优性能。 Read More展开摘要
-
[10]Attentive and Contrastive Learning for Joint Depth and Motion Field EstimationICCV, 2021.A self-supervised framework for 3D object motion field estimation from monocular videos that disentangles ego-motion and object motion via a two-stage projection pipeline with a dynamics attention module, plus contrastive sample consensus for object motion estimation using weak semantic priors and geometric constraints. 提出自监督单目视频 3D 物体运动场估计框架,通过带动态注意力模块的两阶段投影管线解耦自车运动与物体运动, 并结合弱语义先验与几何约束提出对比样本一致性方法估计物体运动。 Read More展开摘要
-
[11]Two-phase Pseudo Label Densification for Self-training based Domain AdaptationECCV, 2020.TPLD densifies sparse pseudo labels in self-training UDA via sliding-window vote propagation and confidence-based easy/hard classification, using full pseudo labels for easy samples and adversarial hard-to-easy feature alignment for hard ones, achieving new SOTA when combined with CRST. TPLD 通过滑动窗口投票传播与基于置信度的难易样本划分来加密自训练 UDA 中稀疏的伪标签, 对简单样本使用完整伪标签,对困难样本采用对抗式难易特征对齐,与 CRST 结合后取得新的 SOTA 结果。 Read More展开摘要
-
Oral Presentation (Acceptance Rate < 0.7%)口头报告(录用率 < 0.7%)
[12]Unsupervised Intra-domain Adaptation for Semantic Segmentation through Self-SupervisionCVPR, 2020.A two-step self-supervised domain adaptation approach that minimizes both inter-domain and intra-domain gaps: first adapting source to target, splitting the target into easy/hard subsets by entropy-based ranking, then self-supervised adaptation from easy to hard split. 提出两阶段自监督域自适应方法,同时缩小域间与域内差异:先进行源域到目标域的自适应, 基于熵排序将目标域划分为简单/困难两部分,再通过自监督方式由简单向困难部分适应。 Read More展开摘要
-
[13]Variational Prototyping-Encoder: One-Shot Learning with Prototypical ImagesCVPR, 2019.VPE tackles open-set graphic-symbol recognition via one-shot classification with prototypical images, learning an image-translation task from real images to prototypes as a meta-task that induces a generalizable embedding space. VPE 通过以原型图像为单一训练样本的单样本分类,解决开放集图形符号识别问题, 以从真实图像到原型图像的翻译任务作为元任务,学习出可泛化的嵌入空间。 Read More展开摘要
-
[14]Driver Drowsiness Detection System Based on Feature Representation Learning Using Various Deep NetworksACCV Workshops, 2016.DDD combines three deep networks for background/environment robustness and local facial-movement and head-gesture recognition, achieving 73.06% detection accuracy on the NTHU drowsy driver benchmark. DDD 网络融合三个深度网络分别学习环境鲁棒性表征以及面部局部动作与头部姿态特征, 在 NTHU 疲劳驾驶检测基准数据集上取得 73.06% 的检测准确率。 Read More展开摘要
CVPR Workshop on Learning with Limited Labelled Data for Image and Video Understanding. CVPR「有限标注数据学习在图像与视频理解中的应用」研讨会。
Awarded for outstanding research achievements during graduate study in Korea. 因在韩国攻读研究生期间的杰出研究成果而获得。
Supported research on intelligent transportation systems. 支持在智能交通系统方面的创新研究。
Outstanding graduate, Xidian University. 西安电子科技大学本科优秀毕业生。
For outstanding academic performance, Xidian University. 因本科期间的优异学术表现而获得。
- Journal Review:期刊审稿: TPAMI, CVIU, Neurocomputing, Pattern Recognition Letters
- Conference Review:会议审稿: CVPR, ICCV, ECCV, NeurIPS, AAAI