Mohamed El Banani

personal_pic_v2.jpg

I am exploring spatial intelligence at World Labs.

I am broadly interested in computer vision, machine learning, and cognitive science. My goal is to understand how visual agents learn to represent their world with minimal supervision and easily generalize to novel objects and scenes.

I received my PhD from the University of Michigan where I was advised by Justin Johnson. During my PhD, I was fortunate to work with David Fouhey and John Laird at UM, Benjamin Graham at FAIR, and Varun Jampani and Leo Guibas at Google Research. I did my undergraduate studies at Georgia Tech where I got the chance to work with Maithilee Kunda and Jim Rehg on cognitive modeling, as well as Omer Inan and Todd Sulchek on biomedical devices.

news

Jun 2026 Our work on generative pixel-aligned geometry beyond the visible is on arXiv.
Nov 2025 We released Marble, a multimodal world model.
Oct 2025 We released RTFM, a real-time frame model.
Feb 2024 Our work on the 3D awareness of visual foundation models was accepted at CVPR 2024.
Jan 2024 I successfully defended my thesis and joined World Labs as a founding member.

selected publications

  1. marble_teaser.jpg
    Marble: A Multimodal World Model
    World Labs Team
    World Labs Blog, 2025
    TL;DR: Announcing Marble, our multimodal world model, available to all users.
  2. teaser_probe3d.png
    Probing the 3D Awareness of Visual Foundation Models
    In CVPR, 2024
    TL;DR: Visual foundation models can classify, delineate, and localize objects in 2D. We study how well these models represent the 3D world that images depict?
  3. lgssl_teaser.png
    Learning Visual Representations via Language-Guided Sampling
    Mohamed El BananiKaran Desai, and Justin Johnson
    In CVPR, 2023
    TL;DR: A picture is worth a thousand words, but a caption can describe a thousand images. We use language models to find image pairs with similar captions, and use them for stronger contrastive learning.
  4. byoc_teaser.png
    Bootstrap your own correspondences
    Mohamed El Banani, and Justin Johnson
    In ICCV, 2021 (Oral)
    TL;DR: Good features get us accurate correspondence, accurate correspondence is good for feature learning. We leverage this to learn {visual, geometric} features via self-supervised point cloud registration.