Robostral Navigate from Mistral: Robot Sees Through One Camera and Navigates Itself
Mistral introduced Robostral Navigate — an 8B model for autonomous robot navigation that works with video from a single RGB camera. Other systems use LiDAR, depth sensors or multiple cameras simultaneously. Robostral achieves 76.6% success on R2R-CE benchmarks (Room-to-Room) and outperforms multi-sensor approaches by 4.5 percentage points. Works on wheeled, walking and flying robots. Trained entirely in simulation on 400k trajectories.
AI-processed from Mistral AI News; edited by Hamidun News
Mistral AI presented Robostral Navigate — its first model for autonomous robot navigation, which orients itself in space using video from a single standard RGB camera and demonstrates 76.6% successful attempts on the R2R-CE test for environments not encountered during training.
What the model can do
Robostral Navigate is an 8-billion-parameter model that, based on a camera frame and a text instruction, guides a robot through complex environments: offices, residential or commercial buildings, outdoor spaces. The instruction might sound like: "Exit the lobby, walk down the corridor, enter the storage closet and stop across from the second shelf" — and the model guides the robot through the entire route autonomously, without human intervention.
- Success rate of 79.4% on R2R-CE validation in familiar scenarios and 76.6% in unseen ones
- Superiority over the best single-camera approach — 9.7 points, over the best system with lidar or multiple cameras — 4.5 points
- Works on wheeled, legged, and flying robots, resistant to differences in camera parameters
- Training dataset — approximately 400,000 trajectories collected in 6,000 simulated scenarios
How the model finds the way
Robostral Navigate predicts not metric displacement, but a point in the image: target coordinates in the current camera frame and the required orientation upon arrival. This "pointing" method of navigation is resistant to lens changes and scene scale variation — unlike commands based on metric distances. When the target moves out of view, the model switches to local coordinate displacements relative to the robot, for example: "Advance 2 meters forward, 1.5 meters left and turn 25 degrees left."
How Robostral Navigate was trained
The model is built entirely within Mistral AI and does not rely on open-source vision-language models — it is initialized based on the company's own model, specialized in pointing, counting, and object localization tasks. A key training element is the prefix-caching method: it compresses an entire episode into one sequence and allows training on all time steps in a single pass, reducing training tokens by 22 times compared to training on individual steps. After the supervised learning stage, the model is further improved through online reinforcement learning using the CISPO algorithm, which allows it to learn from its own mistakes and recover from failures.
"Today we present
Robostral Navigate — our first model for embodied navigation," according to the official Mistral AI blog.
What this means
Navigation using only a single camera without lidar and depth sensors reduces the cost and complexity of robots for warehouses, delivery, hotels, and manufacturing — Mistral AI directly names these industries as the first beneficiaries of the technology.
Frequently asked questions
What robots can Robostral Navigate work on?
The model works on wheeled, legged, and flying robots and generalizes to different platform sizes — Mistral AI confirms this in the model description.
How much data was required for training?
The company collected approximately 400,000 trajectories in 6,000 simulated scenarios — the entire data generation pipeline is built entirely in simulation, without using real robots at the data collection stage.
How much more accurate is
Robostral Navigate than competitors?
On the R2R-CE test for unseen environments, the model surpasses the best single-camera approach by 9.7 points and the best system with multiple cameras or depth sensor — by 4.5 points, with final 76.6% successful attempts.
How does the R2R-CE metric differ from standard navigation tests?
R2R-CE (Room-to-Room in Continuous Environments) tests how well the model follows text instructions in environments not included in the training sample — on this scenario Robostral Navigate achieves 76.6% result, not on familiar scenarios where the result is even higher.
Want to stop reading about AI and start using it?
AI News is a curated feed of AI/tech news. Hamidun Academy teaches you to use AI systematically in your work.
The AI world, distilled — once a week
Seven stories that actually mattered, hand-picked. No noise, no reposts, no press releases.
Done! Check your inbox for a confirmation.