Together AI brings nine research papers to ICML 2026 conference in Seoul
Together AI presented nine research papers at ICML 2026 in Seoul — from agentic frameworks to GPU kernels. ThunderAgent accelerates agent inference by 3.6x, Aurora (adaptive speculative decoding) already runs in production as ATLAS-speculator, and Untied Ulysses maintains 5 million token context on a single node. The company emphasizes: gains at one stack level are useless without the rest.
AI-processed from Together AI Blog; edited by Hamidun News
Together AI announced on June 30, 2026, that nine research papers from the company and its partners were accepted to the ICML 2026 conference in Seoul — presentations will be held at booth B714 and cover the entire AI infrastructure stack, from agents to GPU cores.
How the research stack is organized
Together AI emphasizes: frontier AI is not built at one level. A win in one layer is useless if neighboring layers cannot keep up with it. The company laid out nine papers across five levels of the stack — from agents at the very top to GPU cores at the foundation — and showed how each level feeds the next, with Together's production workloads indicating what research challenge to solve next.
ICML (International Conference on Machine Learning) is one of the world's leading machine learning conferences alongside NeurIPS and ICLR, and getting a paper into its program is traditionally considered a strong signal of research quality in the industry. In 2026, the conference is held in Seoul, where Together AI will have its own booth B714 for meeting with participants.
What's new in agents and model training
Three papers are devoted to agents that do real work and are evaluated on tasks that cannot be "faked" as solved. DSGym is a framework for evaluating and training agents on data: over 1,000 tasks in 10+ domains under a single API. ThunderAgent accelerates agent inference up to 3.6 times. TTT-Discover is an open model that, according to the authors, exceeds the best human results in its category of tasks.
Two more papers address how to turn a base model into a "reasoner" even where there is no ready reference answer to check against. RARO trains a model without a verifier and gives 25% wins compared to baseline. V1 adds up to 10% correct answers. Both papers solve one and the same practical problem: how to improve a model on tasks where answer correctness cannot be verified by simple comparison to a "dictionary of correct answers."
How inference and GPU cores are accelerated
At the algorithmic optimization level, Aurora implements adaptive speculative decoding — a technique in which a lighter model pre-guesses several tokens while the main model checks them, speeding up text generation. Aurora adjusts to live traffic and provides up to 1.25x speedup. According to Together AI, this same line of research is already working in production on their platform under the name ATLAS-speculator — that is, the path from a paper at ICML to a real service took not years but one development cycle.
Key facts on system and hardware levels:
- Untied Ulysses holds a context of 5 million tokens on a single compute node
- Opportunistic Expert Activation (OEA) accelerates MoE model decoding up to 39%
- ParallelKernelBench is a set of 87 tasks on multi-GPU kernels, the foundation of the stack where raw hardware becomes useful speed
- A total of 9 papers from Together AI and partners are accepted to ICML 2026
- The company presents them at booth B714 in Seoul
Why this matters to the company
Together AI directly links research to production: Aurora advances are already used in its platform as ATLAS-speculator. The logic of the cycle the company describes is this — research results become part of the Together platform, and production workloads on that platform show what research challenge to solve next.
What this means
ICML 2026 shows that the race for faster and cheaper AI inference is not just at the level of the models themselves, but also at the level of agent frameworks, decoding algorithms, and GPU cores — and companies like Together AI are building research across this entire stack at once, rather than at a single point.
Want to stop reading about AI and start using it?
AI News is a curated feed of AI/tech news. Hamidun Academy teaches you to use AI systematically in your work.
The AI world, distilled — once a week
Seven stories that actually mattered, hand-picked. No noise, no reposts, no press releases.
Done! Check your inbox for a confirmation.