FindingDory evaluates long-range memory and reasoning in embodied agents.
Self-Grounded Verification mitigates agreement bias in MLLM verifiers for web navigation, computer use, and robotic manipulation.
Memo enables efficient memory formation and usage for long-horizon embodied RL tasks, improving generalization and efficiency.
We investigate representations from pre-trained text-to-image diffusion models for control tasks and showcase competitive performance across a wide range of tasks.
We present a new embodied question answering (EQA) dataset with open vocabulary questions.
We conduct a study on using pre-trained visual representations (PVRs) to train robots for real-world tasks.
We present the largest and most comprehensive empirical study of visual foundation models for Embodied AI (EAI).
We propose a combined simulation and real-world benchmark on the problem of Open-Vocabulary Mobile Manipulation (OVMM).
We present a modular system that can perform well on the Instance ImageNav task in both simulation and the real world.
We present Habitat-Matterport 3D Semantics (HM3DSEM), the largest dataset of 3D real-world spaces with densely annotated semantics.