I am a Young Researcher at Shanghai Artificial Intelligence Laboratory, with research interests in Vision-Language-Action (VLA) and Robotic Manipulation.
I received my Ph.D. from Beihang University in Nov. 2025, advised by
Di Huang.
I'm interested in robotic learning, computer vision and embodied AI. Most
of my research is about inferring the grasp poses from
images and point-clouds.
EruDiff refactors world knowledge in text-to-image diffusion models to improve synthesis from implicit prompts across scientific and commonsense knowledge benchmarks.
A unified VLA framework that combines understanding, visual foresight, and action generation for robust robotic manipulation in dynamic and static scenarios.
GraspLDP injects grasp priors into latent diffusion policy learning to improve grasp precision and generalization in both simulation and real-world manipulation.
Focus on the problem of feature learning in the presence of scale imbalance for 6-DoF grasp detection and propose a novel approach to especially address the difficulty in dealing with small-scale samples.