Как МТС учит гуманоида с VLA-моделью брать движущиеся по конвейеру предметы
МТС в RnD-направлении Physical AI разрабатывает VLA-политику (Vision-Language-Action) для полноростового гуманоида. Главная проблема: робот уверенно берёт предмет со стола в статичной сцене, но начинает промахиваться, когда объект едет по конвейеру. Полноростовой робот при этом должен ещё и делать шаги, удерживая баланс тела, — а это многократно усложняет расчёт движений.
AI-processed from Habr AI; edited by Hamidun News
22 July 2026 ML engineer at MTS Dania published a corporate blog analysis on Habr about how the company's Physical AI R&D team is teaching a full-sized humanoid to pick up objects moving along a conveyor belt using a VLA policy (Vision-Language-Action).
What the main problem is
The humanoid confidently picks up an object from a table, but starts missing when the same object moves along a conveyor belt — this is exactly the task the MTS team is focused on. In a static scene the object is motionless, lighting is stable, and the camera is pointed at the work area, so modern VLA models show high control quality. As soon as the object starts moving, the usual grasp accuracy disappears: a model trained on static scenes cannot keep up with the changing geometry of the task.
According to the engineer, this is not a one-off failure but a systemic limitation of an approach that works great in lab demonstrations but transfers poorly to a real conveyor belt.
The robot confidently picks up an object from a table, but starts
missing when the object moves along a conveyor belt.
— Dania, ML engineer, MTC Web Services
Why a dynamic scene is harder than a static one
A dynamic environment adds two levels of complexity at once: the object itself moves, and the robot that has to adjust to it moves too. Most impressive robot demos are still filmed in static conditions — objects are neatly laid out on a table, and the scene barely changes. Real-world tasks in everyday life and manufacturing don't look like that: at the very least, objects move along a conveyor belt, and the grasp has to be calculated taking their speed and trajectory into account.
Key project facts:
- Team — Physical AI R&D unit at MTC Web Services
- Technology — VLA policy (Vision-Language-Action) for a humanoid
- Author of the analysis — Dania, ML engineer at MTS
- Focus — grasping objects moving along a conveyor belt, not lying on a table
- Publication date of the analysis — 22 July 2026
Why it's harder for a humanoid than for a manipulator arm
A full-sized robot has to not only reach for the object but also keep the balance of its entire body. A fixed manipulator arm solves a simpler task: its base is fixed, and the whole model only has to control the arm's movement. A humanoid, for the same grasp, has to take steps and compensate for the shift in its center of gravity with additional body movements.
Every such movement changes the position of the camera and the arm relative to the object, which is also moving at that moment. As a result, calculating the movements becomes far more complex: the model has to simultaneously predict the object's trajectory, plan the grasp, and maintain the robot's stability. This is exactly the combination the MTS team is now trying to put together within its Physical AI direction.
What this means
Moving from static demonstrations to working with moving objects is one of the main barriers on humanoids' path from the lab to real production. The MTS analysis shows that even the basic operation of "picking up an object" stops being trivial as soon as the scene comes alive, and that industrial deployment of VLA robots is limited not by grip strength but by the dynamics of the environment.
Frequently asked questions
What is a VLA policy?
VLA (Vision-Language-Action) is a robot control model that links vision, a language instruction, and action: the robot "sees" the scene, understands the task, and plans movement. In static scenes, according to the MTS analysis, such models show high grasp quality.
Why is a dynamic environment harder for a humanoid?
Because both the object and the robot itself are moving. To grasp the object, a full-sized humanoid has to take steps and maintain body balance with additional movements, so calculating the movements becomes far more complex compared to a fixed manipulator arm.
Need AI working inside your business — not just in your newsfeed?
I build production AI for companies — custom CRM, internal tools, autonomous agents, workflow automation. Owned by you, shaped to your process, no per-seat tax. Built by Zhemal Khamidun, CPO of AlpinaGPT (AI platform, 6,000+ users).
The AI world, distilled — once a week
Seven stories that actually mattered, hand-picked. No noise, no reposts, no press releases.
Done! Check your inbox for a confirmation.