Event
Ph.D. Dissertation Defense: Mohamed Elnoor
Monday, August 10, 2026
11:00 a.m.
AVW1146
Emily Irwin
301 405 0680
eirwin@umd.edu
ANNOUNCEMENT: Ph.D. Dissertation Defense
Name: Mohamed Elnoor
Committee:
Professor Dinesh Manocha, Chair/Advisor
Professor Pratap Tokekar
Professor Kaiqing Zhang
Professor Sanghamitra Dutta
Professor Michael Otte, Dean's Representative
Date/time: Monday, August 10 at 11:00 AM
Location: AVW 1146
Date/time: Monday, August 10 at 11:00 AM
Location: AVW 1146
Zoom Link: umd.zoom.us/my/dmanocha
Title: Proprioception-based Perception and Physically Grounded Vision-Language Reasoning for Autonomous Navigation
Abstract:
Autonomous mobile robots increasingly operate in complex outdoor environments where navigation requires reasoning about terrain geometry, physical properties, environmental conditions, and semantic context. Existing perception systems rely primarily on vision-based sensors, which often fail when terrain appearance does not accurately reflect traversability or when sensing conditions deteriorate.
In this dissertation, we present several novel algorithms built on two complementary ideas. The first, proprioception-based perception, uses measurements of the robot’s own body, including joint positions, forces, motor currents, and platform state, to estimate terrain traversability, stability, and other physical properties that may not be reliably inferred from visual appearance alone. The second, physically grounded vision-language reasoning, uses Vision-Language Models (VLMs) for high-level scene understanding while incorporating measured physical information, such as the robot’s interaction with the environment or its current state, into the reasoning and navigation process. Building on these ideas, our algorithms estimate terrain traversability from proprioceptive signals, adaptively combine vision and proprioception under variable sensing conditions, ground VLM predictions in physical feedback, distill VLM reasoning into lightweight on-board models, and apply state-conditioned vision-language reasoning to safety-critical two-wheeled platforms. We implement and evaluate the proposed navigation methods on real legged and wheeled robots and on a simulated motorcycle, demonstrating improvements in navigation success rate, energy consumption, and stability over state-of-the-art baselines.
The first part develops proprioception-based perception for legged robots. We propose ProNav, which uses proprioceptive signals from joint encoders, force, and current sensors for traversability estimation, gait selection, and preemptive crash prediction, showing up to 40% improvement in success rate and up to 15.1% reduction in energy consumption over exteroceptive-based methods. Building on this, AMCO adaptively couples vision and proprioception through three cost maps weighted by the visual stream’s estimated reliability, observing 10.8%–34.9% reduction in stability metrics and up to 50% improvement in success rate over prior methods.
In the second part, we integrate physically grounded vision-language reasoning through VLMs along two complementary directions. We present VLM-GroNav, which integrates VLMs with physical grounding using in-context learning to dynamically update traversability estimates from the robot’s real-time physical interactions and inform both local and global planners, observing up to 50% increase in success rate. Furthermore, ViLAM distills vision-language reasoning from large VLMs into spatial attention maps used as planning cost maps by a sampling-based local planner, demonstrating 14.2%–50% improvements in success rate over existing methods.
The final part extends the vision-language reasoning framework to safety-critical two-wheeled platforms. We present an Advanced Rider Assistance System (ARAS) that leverages VLM chain-of-thought reasoning conditioned on the vehicle’s speed and lean angle to produce a dense per-pixel risk map, and drives a local planner adapted to motorcycle dynamics, improving success rate and hazard-clearance distance over baselines in the CARLA simulator.
Together, these contributions establish a unified approach for proprioception-based perception and physically grounded vision-language reasoning that combines physical interaction, adaptive sensor fusion, and semantic reasoning to enable reliable autonomous navigation across diverse robotic platforms.
Title: Proprioception-based Perception and Physically Grounded Vision-Language Reasoning for Autonomous Navigation
Abstract:
Autonomous mobile robots increasingly operate in complex outdoor environments where navigation requires reasoning about terrain geometry, physical properties, environmental conditions, and semantic context. Existing perception systems rely primarily on vision-based sensors, which often fail when terrain appearance does not accurately reflect traversability or when sensing conditions deteriorate.
In this dissertation, we present several novel algorithms built on two complementary ideas. The first, proprioception-based perception, uses measurements of the robot’s own body, including joint positions, forces, motor currents, and platform state, to estimate terrain traversability, stability, and other physical properties that may not be reliably inferred from visual appearance alone. The second, physically grounded vision-language reasoning, uses Vision-Language Models (VLMs) for high-level scene understanding while incorporating measured physical information, such as the robot’s interaction with the environment or its current state, into the reasoning and navigation process. Building on these ideas, our algorithms estimate terrain traversability from proprioceptive signals, adaptively combine vision and proprioception under variable sensing conditions, ground VLM predictions in physical feedback, distill VLM reasoning into lightweight on-board models, and apply state-conditioned vision-language reasoning to safety-critical two-wheeled platforms. We implement and evaluate the proposed navigation methods on real legged and wheeled robots and on a simulated motorcycle, demonstrating improvements in navigation success rate, energy consumption, and stability over state-of-the-art baselines.
The first part develops proprioception-based perception for legged robots. We propose ProNav, which uses proprioceptive signals from joint encoders, force, and current sensors for traversability estimation, gait selection, and preemptive crash prediction, showing up to 40% improvement in success rate and up to 15.1% reduction in energy consumption over exteroceptive-based methods. Building on this, AMCO adaptively couples vision and proprioception through three cost maps weighted by the visual stream’s estimated reliability, observing 10.8%–34.9% reduction in stability metrics and up to 50% improvement in success rate over prior methods.
In the second part, we integrate physically grounded vision-language reasoning through VLMs along two complementary directions. We present VLM-GroNav, which integrates VLMs with physical grounding using in-context learning to dynamically update traversability estimates from the robot’s real-time physical interactions and inform both local and global planners, observing up to 50% increase in success rate. Furthermore, ViLAM distills vision-language reasoning from large VLMs into spatial attention maps used as planning cost maps by a sampling-based local planner, demonstrating 14.2%–50% improvements in success rate over existing methods.
The final part extends the vision-language reasoning framework to safety-critical two-wheeled platforms. We present an Advanced Rider Assistance System (ARAS) that leverages VLM chain-of-thought reasoning conditioned on the vehicle’s speed and lean angle to produce a dense per-pixel risk map, and drives a local planner adapted to motorcycle dynamics, improving success rate and hazard-clearance distance over baselines in the CARLA simulator.
Together, these contributions establish a unified approach for proprioception-based perception and physically grounded vision-language reasoning that combines physical interaction, adaptive sensor fusion, and semantic reasoning to enable reliable autonomous navigation across diverse robotic platforms.
