MarkTechPost→ original

Ant Group Presented LingBot-VLA 2.0 — Open Model for Controlling Robots of Any Design

Robbyant, part of Ant Group, released LingBot-VLA 2.0, an open vision-language-action model for cross-platform robot control. The 6-billion parameter model trained on 60 thousand hours of data: 50 thousand hours of video from 20 robot configurations and 10 thousand hours of first-person video. It maps any robot to a unified 55-dimensional action space for arms, deltoid grippers, torsos and mobile bases. On GM-100 benchmark LingBot-VLA 2.0 surpasses π0.5 and version 1.0.

AI-processed from MarkTechPost; edited by Hamidun News
Ant Group Presented LingBot-VLA 2.0 — Open Model for Controlling Robots of Any Design
Source: MarkTechPost. Collage: Hamidun News.
◐ Listen to article

Robbyant, the robotics division of Ant Group, released LingBot-VLA 2.0 on July 8, 2026 — an open vision-language-action model with 6 billion parameters for controlling robots of different designs under the Apache 2.0 license.

What data the model was trained on

LingBot-VLA 2.0 was pretrained on approximately 60,000 hours of data, according to MarkTechPost. Of these, 50,000 hours are robot trajectories collected across 20 different robot configurations, and another 10,000 hours are egocentric video of people shot from a first-person perspective while performing household and manipulation actions.

  • Developer — Robbyant, Ant Group's robotics division; release occurred on July 8, 2026
  • License — Apache 2.0, model checkpoint is open for download
  • Model size — 6 billion parameters (6B)
  • Volume of pretraining data — approximately 60,000 hours (50,000 hours of robot trajectories across 20 robot configurations + 10,000 hours of video with people)
  • Unified action space — 55 dimensions for all robot types

How the model controls different robots

LingBot-VLA 2.0 converts control of any robot into a unified 55-dimensional canonical action space that encompasses arms, dexterous grippers, waist, head, and mobile platforms, explains MarkTechPost. This approach solves the main problem of cross-embodiment models: previously, for each robot construction — with different numbers of joints, gripper type, or presence of a mobile base — you had to train a separate control model, and a unified action space allows the same checkpoint to work across 20 configurations.

Token-level action module based on Mixture-of-Experts helps scale the model without quality loss, bypassing the need for an auxiliary loss function to balance load between experts. This allows increasing model capacity without adding an extra term to the loss function, which typically slows down and complicates training of such architectures.

Where the model gets geometry and time

Dual-query distillation from LingBot-Depth and DINO-Video models adds geometric and temporal supervision, which helps LingBot-VLA 2.0 account for future scene state when choosing an action, according to MarkTechPost. In other words, the model "understands" in advance how scene geometry and object positions will change after performing the current step, rather than only reacting to a static frame.

How much LingBot-VLA 2.0 outperforms competitors

On the GM-100 test set for universal robot agents, LingBot-VLA 2.0 surpassed both the π0.5 model and the previous LingBot-VLA-1.0 version on both tested robot platforms, reports MarkTechPost. Superior performance over both the predecessor and the third-party model π0.5 on the same benchmark confirms that improvements — unified action space, MoE module, and dual-query distillation — deliver measurable quality gains, not just increased parameter count.

What this means

An open model with 6 billion parameters that works equally well across 20 different robot configurations lowers the entry barrier for teams wanting to train universal robot agents without collecting separate datasets for each specific robot design.

Frequently Asked Questions

What does the VLA abbreviation in the model name stand for?

VLA stands for vision-language-action — the model processes an image and text instruction and generates a robot action in a single pass.

Can LingBot-VLA 2.0 be used for free?

Yes, the model is released under the Apache 2.0 license, which permits free use, modification, and commercial application of the checkpoint, including integration into proprietary robotics products.

ZK
Hamidun News
AI news without noise. Daily editorial selection from 50+ sources. A product by Zhemal Khamidun, Head of AI at Alpina Digital.

Need AI working inside your business — not just in your newsfeed?

I build production AI for companies — custom CRM, internal tools, autonomous agents, workflow automation. Owned by you, shaped to your process, no per-seat tax. Built by Zhemal Khamidun, CPO of AlpinaGPT (AI platform, 6,000+ users).

What do you think?
Loading comments…