Back to feed
News Story
Meta AI (FAIR)
1 sources

Reimagining Independence: How Meta's AI Models Are Helping the University of Pittsburgh Transform Assistive Robotics

Meta's AI models are being used by the University of Pittsburgh to advance assistive robotics, aiming to enhance independence for individuals with disabilities. The collaboration leverages Meta's AI research to improve robotic assistance in daily tasks.

SynthePulse Insight · AI deep reading

Edge Intelligence and Open Models: How Meta AI Powers Real-Time Perception for University of Pittsburgh Assistive Robots

Version 1 · 1 source

With up to $41.5 million in ARPA-H funding, the University of Pittsburgh's HERL lab partners with Meta to deploy open-source vision models like DINOv3 and SAM on edge devices, building the next-generation assistive mobility platform RAMMP, aiming for reliable, low-latency object recognition and navigation assistance in real dynamic environments.

  • The RAMMP project receives up to $41.5 million from ARPA-H, aiming to improve independence and safety for approximately 5.5 million wheelchair users in the U.S. through AI and robotics.
  • The project integrates Meta's open-source vision models DINOv3 and SAM, enabling real-time object detection and segmentation on battery-powered edge devices without relying on network connectivity.
  • The perception system is based on the RF-DETR detection model, fine-tuned with DINOv2 embeddings, and uses SAM to automatically annotate training data covering various angles, heights, backgrounds, and lighting conditions.
  • Engineering trade-offs: To meet real-time requirements on edge devices, slight compromises in boundary precision or feature detail are sometimes made in exchange for speed and stability.
  • The prototype already integrates DINO-based query capabilities to detect automatic door buttons, cups, curbs, etc., with voice and touch input planned as next steps.
Open section navigationProject Background and Funding Scale

Project Background and Funding Scale

The RAMMP project, led by the University of Pittsburgh's Human Engineering Research Laboratories (HERL), receives up to $41.5 million from the Advanced Research Projects Agency for Health (ARPA-H), a U.S. research funding agency focused on transformative biomedical breakthroughs. Project partners include assistive technology company ATDev.

Approximately 5.5 million wheelchair users in the U.S. experience over 100,000 wheelchair-related injuries treated in emergency rooms annually, many caused by trips and falls. The project aims to address shortcomings in current assistive mobility platform designs by integrating advanced robotics, novel operating systems, and digital twin technology.

Core Role of Meta's Open-Source Models

RAMMP integrates multiple open-source AI vision models from Meta, including DINO (self-supervised vision Transformer) and Segment Anything Model (SAM). DINO excels at learning visual representations from unlabeled data, suitable for scenarios with scarce labeled data; SAM can identify and outline any object in images or videos with minimal prompts.

DINOv3 serves as a compact and efficient 'visual brain' that can be topped with lightweight task modules (e.g., object detection, motion tracking), enabling reuse of visual data and power savings. SAM is used to automatically annotate training data, allowing the team to quickly generate high-quality annotations covering various angles, heights, backgrounds, and lighting conditions.

The perception system is based on the RF-DETR lightweight detection model, fine-tuned with DINOv2 embeddings, and uses SAM for automatic data annotation. Combined with data augmentation and multi-view strategies, the system achieves real-time 360-degree environmental perception and adaptive object detection.

Engineering Challenges and Trade-offs in Edge Deployment

Deploying DINOv3 and SAM on battery-powered, size- and weight-constrained hardware presents significant engineering challenges, including battery life, heat dissipation, and unreliable network connectivity. The project adopts edge computing, processing camera images and sensor data directly on the device to ensure immediate response.

Engineers optimize models by reducing memory footprint, using low-precision inference, and deploying formats suitable for real-world conditions. At practical resolutions and efficient batch processing, models remain fast and reliable on compact hardware, sometimes sacrificing slight boundary precision or feature details for speed and stability.

RAMMP project chief of staff Sivashankar Sivakanthan states: 'For assistive robots, performance is not measured by benchmark accuracy, but by whether the system can operate reliably amid the unpredictability of daily life. Running DINOv3 and SAM on-device enables real-time perception that users can trust, without relying on connectivity or compromising safety.'

Prototype Capabilities and User Interaction

The prototype already integrates DINO-based tools that query the robot's image sensor to detect automatic door buttons, cups, and curbs/ground for navigation assistance. Engineers are focusing on voice and touch input to allow users to select and interact with surrounding objects.

Combining natural language with image data allows users to query the environment and issue commands in a more natural way, directly reducing cognitive load and context-switching costs. For example, a user can simply say 'pick up the cup on the table' to operate.

Current challenges include ensuring accuracy and temporal consistency of model outputs, as well as robustness and predictability across user prompts and inputs.

Collaboration Model and Future Outlook

HERL leads the project, setting the vision with deep expertise in biomedical engineering and user-centered research; ATDev provides an engineering perspective to turn translational research into real-world devices. The combination leverages strengths in academic leadership and technological innovation.

The project emphasizes user-centered design, aiming to make assistive technology reflect the complexity of the real world, enhancing user confidence, independence, and physical safety.

Credibility boundary

The information in this article primarily comes from the Meta AI official blog, a first-party source. Project funding, model capabilities, engineering details, etc., are clearly attributed. Some descriptions of user interaction effects (e.g., reducing cognitive load) are the project team's views and have not been independently verified.

Insight takeaway

Meta's open-source vision models DINOv3 and SAM, deployed on edge devices, are bringing advanced AI perception capabilities to assistive robotics, but engineering trade-offs between real-time performance and accuracy, as well as robustness across user inputs, remain key challenges in practical applications.

Primary report

Meta AI (FAIR)

Primary source