Mohamed El Banani
I am exploring spatial intelligence at World Labs.
I am broadly interested in computer vision, machine learning, and cognitive science. My goal is to understand how visual agents learn to represent their world with minimal supervision and easily generalize to novel objects and scenes.
I received my PhD from the University of Michigan where I was advised by Justin Johnson. During my PhD, I was fortunate to work with David Fouhey and John Laird at UM, Benjamin Graham at FAIR, and Varun Jampani and Leo Guibas at Google Research. I did my undergraduate studies at Georgia Tech where I got the chance to work with Maithilee Kunda and Jim Rehg on cognitive modeling, as well as Omer Inan and Todd Sulchek on biomedical devices.
news
| Jun 2026 | Our work on generative pixel-aligned geometry beyond the visible is on arXiv. |
|---|---|
| Nov 2025 | We released Marble, a multimodal world model. |
| Oct 2025 | We released RTFM, a real-time frame model. |
| Feb 2024 | Our work on the 3D awareness of visual foundation models was accepted at CVPR 2024. |
| Jan 2024 | I successfully defended my thesis and joined World Labs as a founding member. |
selected publications
-
Marble: A Multimodal World ModelWorld Labs Blog, 2025TL;DR: Announcing Marble, our multimodal world model, available to all users. -
Learning Visual Representations via Language-Guided SamplingIn CVPR, 2023TL;DR: A picture is worth a thousand words, but a caption can describe a thousand images. We use language models to find image pairs with similar captions, and use them for stronger contrastive learning.