Multimodal Perception
Representing the physical world
We study 3D geometry, multimodal fusion and scene representations, combining visual, geometric and other sensory information to understand objects and their surroundings.
Research & publicationsOMEGA LAB · HKUST(GZ)
Omnimodal Multi-Embodiment
Generalist Agents
Our research connects multimodal perception, intent and motion modeling, and autonomous decision-making in physical systems and digital workflows.
Explore our researchModalities & representations
Land, water, air & robotic systems










World models, reasoning & decision-making










Omnimodal Multi-Embodiment Generalist Agents
Select any number of modalities to highlight their matching orbits. Select again to deselect. Drag along the orbit to change its position; release to continue orbiting. Arrow keys also adjust position. Escape returns to the overview.
The three O, ME and GA sections fade, and their letters join to spell OMEGA, then condense into the lab’s Greek Ω logo. A nucleus and orbital paths then appear. After orbiting, the graphic unfolds back into the three sections and repeats. Each modality has its own tilted orbital plane around the central nucleus, with vision, language, geometry and tactile emphasized. The nine modalities and representations are: vision, language, geometry, tactile, audio, radio, proprioception, intent and trajectory. Inside the nucleus, ten kinds of embodiments form a ring around the central Ω: quadrupeds, UGVs, autonomous cars, USVs and UAVs; then dexterous hands, robot arms, bimanual systems, wheeled dual-arm robots and humanoids. Labels and tracks become smaller, softer and occluded as they pass behind the translucent nucleus. Multiple modalities can be selected independently; dragging pauses only the dragged modality until release. The Generalist Agents section forms a six-stage loop: Perceive, Model, Predict, Reason, Plan and Act. At the center, digital tasks and physical tasks connect through an agent and tool use. Software tools and control interfaces let agents act in digital workflows and physical environments, including autonomous navigation. Outward calls and returning results or observations form a shared feedback loop.
Energy-Aware Wind-Resilient Routing for Truck-Assisted Multi-UAV Delivery under Wind Uncertainty, by Tianshun Li, Yanggang Sheng, Hongliang Lu, Zhongzhen Wang, Haoang Li and Xinhu Zheng.
Led by Xinhu Zheng, the HKUST(GZ) team joined Annu to compete in the inaugural Shenzhi Cup AI Innovation Competition, also taking first place in material handling.
The award recognizes Intelligent Multi-Modal Sensing-Communication Integration: Synesthesia of Machines, co-authored by Xinhu Zheng and published in IEEE Communications Surveys & Tutorials.
OUR RESEARCH
We study how agents perceive, model, reason and act, from robot policies and world models to tool use, multi-agent workflows and the safety of embodied systems.
Explore researchMultimodal Perception
We study 3D geometry, multimodal fusion and scene representations, combining visual, geometric and other sensory information to understand objects and their surroundings.
Research & publicationsRobot Learning & Manipulation
We develop vision-language-action models that connect reasoning with manipulation, studying motion generation, policy improvement and efficient inference.
Research & publicationsWorld Models & Prediction
We investigate predictive world models and future motion, with a focus on how behavior, driving style and physical dynamics inform planning.
Research & publicationsVision-Language Navigation
We study how agents follow language instructions in physical environments, connecting spatial perception, motion prediction and planning for ground and aerial navigation.
Research & publicationsIntent & Decision-Making
We model intent and human driving behavior, and examine how visual evidence and context shape autonomous decisions under real-world constraints.
Research & publicationsMulti-Agent Systems
We study how agents use specialized models and tools, share information and coordinate tasks, spanning digital workflows for analysis and generation as well as collaborative aerial and ground systems.
Research & publicationsEmbodied AI Safety
We probe how perception, planning and detection models fail under backdoor, adversarial and out-of-distribution threats, designing stealthy yet effective attacks to expose vulnerabilities and developing mitigation strategies for robust autonomous driving and multimodal systems.
Research & publicationsRESEARCH IN FOCUS
Recent work by Xinhu Zheng and collaborators in 3D perception, multimodal learning and embodied intelligence.
View all publications
Multimodal Perception
UniT brings online perception and offline 3D reconstruction into one geometry model. Group autoregression accommodates different views and auxiliary inputs, while supporting metric-scale prediction and bounded memory over long sequences.

Robot Learning & Manipulation
ForesightFlow learns from successful and failed robot experience. A single flow model generates action chunks and estimates their success potential, guiding action selection without a separate critic. The policy is evaluated on simulated and real-world bimanual tasks.

World Models & Prediction
PLAN-S decodes style-conditioned semantic cost maps from latent world models, making planning preferences inspectable before trajectory selection. The same bridge supports regression and anchor-score planners, evaluated on nuScenes and NAVSIM.

Vision-Language Navigation
P³Nav brings perception, prediction and planning into one vision-and-language navigation model. Object and map cues describe the current scene; predicted waypoints and future scene features help the agent plan toward a language-specified goal.

Intent & Decision-Making
A structured visual-perturbation study examines how VLA driving decisions depend on camera input. Evaluations compare open-loop trajectory prediction with closed-loop driving behavior, revealing that visual grounding varies across perturbations and evaluation settings.

Multi-Agent Systems
EWR plans truck-assisted UAV delivery routes with an energy graph updated by wind estimates and payload state. It continually checks return feasibility, aiming to preserve enough battery energy to reach the truck or depot as conditions change.
PEOPLE
Meet the principal investigator and explore our member directory.
View all people
Assistant Professor
The Hong Kong University of Science and Technology (Guangzhou)
Intelligent Transportation · Internet of Things
Xinhu Zheng studies multimodal perception, intent and trajectory prediction, and autonomous systems.
Platforms, instruments, and experimental setups.
PROSPECTIVE STUDENTS & RESEARCHERS
Bring your questions about multimodal intelligence and agents. Explore PhD, MPhil and postdoctoral pathways, university support, and how to get in touch.
Opportunities & applicationsGET IN TOUCH
For research collaboration, prospective student enquiries and academic exchange, contact Xinhu Zheng.
Office W2 L5 509
The Hong Kong University of Science and Technology (Guangzhou)