36Kr (36氪)→ original

Из Tencent Hunyuan ушёл глава мультимодального ИИ Ху Хань — команда переходит к world models

Ху Хань, руководитель направления мультимодального понимания в Tencent Hunyuan, подал заявление об уходе — он основывает собственный стартап. Об этом сообщило издание «Интеллектуальное возникновение» (36Kr). Ранее Ху Хань был principal researcher в Microsoft Research Asia, а в Tencent пришёл в начале 2025 года. Его бывшую группу глава департамента больших языковых моделей Яо Шуньюй, скорее всего, переориентирует на передовые исследования world models.

AI-processed from 36Kr (36氪); edited by Hamidun News
Из Tencent Hunyuan ушёл глава мультимодального ИИ Ху Хань — команда переходит к world models
Source: 36Kr (36氪). Collage: Hamidun News.
◐ Listen to article

Hu Han, head of multimodal understanding at Tencent Hunyuan, has filed his resignation to found his own startup — this was reported on July 23, 2026, by the Chinese outlet "Intelligent Emergence" (36Kr, 智能涌现). According to the outlet, his former research team at Tencent is planned to be redirected toward the study of world models.

Who is Hu Han

Before Tencent, Hu Han was a principal researcher in the visual computing group at Microsoft Research Asia. He joined Tencent in early 2025 and led research on visual large models, and after an internal reorganization moved to the "Frontier" advanced research group within the large language model department, reporting to Shunyu Yao.

According to 36Kr, Hu Han was responsible, among other things, for the development of world models — systems that simulate physical reality rather than simply describing an image. It is precisely this direction, the outlet believes, that could become the new focus of his former team.

Why Tencent is rethinking multimodality

Tencent is reallocating resources toward foundational language models, and research into multimodal understanding is losing priority. Recognition accuracy for text, images, and video in this field has already exceeded 85%, so further investment yields diminishing returns — as a researcher cited by 36Kr points out.

A key role in the pivot is played by Shunyu Yao, a former OpenAI researcher who now heads Tencent's large language model department. After his arrival, the company disbanded AI Lab and merged its core developers into the LLM department, while also stepping up hiring — in particular from ByteDance's Seed Infra and post-training teams.

  • Hu Han has filed his resignation and is founding a startup (per a 36Kr report from July 23, 2026)
  • Previously a principal researcher at Microsoft Research Asia, at Tencent since early 2025
  • His replacement for the VLM direction is Tian Yonglong, a former OpenAI researcher and MIT PhD, who joined Tencent in early July 2026
  • Tencent's capital expenditure for 2025 was 79.2 billion yuan
  • For comparison: Alibaba spent 126 billion yuan in the fiscal year ending March 2026; ByteDance plans to spend up to $70 billion on AI infrastructure in 2026 (according to Bloomberg)

What matters more — recognition or monetization

Multimodal understanding does not monetize well directly, and this is one reason for the shift in priorities. Users are not willing to pay for scenarios like image recognition or "image captioning," whereas they do pay for document processing, presentation creation, and reports — because these rely on the model's reasoning, coding, and agentic capabilities.

"No user is willing to pay for image recognition — the market is full of free alternatives," a product manager for the

Yuanbao app explained the logic of the rethink to 36Kr.

Shunyu Yao enforces strict discipline in the new team. A month before the release of the Hy3 model, when problems were found with part of the data, he reportedly warned the team, according to the outlet "LatePost":

"Data is extremely important.

If this happens again, it's an immediate firing," — Shunyu Yao, head of Tencent's large language model department.

What this means

Tencent is deliberately slowing down the mature but weakly monetizable direction of multimodal understanding and pulling talent and compute toward foundational models and cutting-edge topics like world models. The first result of this strategy is the Hy3 model from July 6, 2026, which, according to 36Kr's assessment, already competes with GLM-5.2 and DeepSeek V4 Pro in the mid-size class.

Frequently asked questions

What are world models and why is Tencent betting on them

World models are models that learn to simulate the physical world and its dynamics, rather than simply recognizing the contents of an image. Tencent considers this direction more promising than mature multimodal understanding, and it is precisely toward this, according to 36Kr, that Hu Han's former team may pivot.

Who will replace Hu Han at Tencent

The visual-language model (VLM) direction will be led by Tian Yonglong — a former OpenAI researcher and MIT PhD who joined Tencent's large language model department in early July 2026 and reports to Shunyu Yao.

What is the Hy3 model

Hy3 is a Tencent large language model released on July 6, 2026, the first major result of Shunyu Yao's tenure at the company. According to 36Kr, in the mid-size class it is comparable to GLM-5.2 and DeepSeek V4 Pro.

ZK
Hamidun News
AI news without noise. Daily editorial selection from 50+ sources. A product by Zhemal Khamidun, Head of AI at Alpina Digital.

Want to stop reading about AI and start using it?

AI News is a curated feed of AI/tech news. Hamidun Academy teaches you to use AI systematically in your work.

What do you think?
Loading comments…