Air Quality Arena: бенчмарк для прогноза загрязнения воздуха на foundation-моделях
Исследователи представили Air Quality Arena (AQA) — крупный открытый датасет и бенчмарк для прогноза загрязнения воздуха. В него вошли 6 основных загрязнителей за 3 года, свыше 14 000 рядов «станция-загрязнитель» из 7 стран на 4 континентах. На нём сравнили 11 time-series foundation-моделей: они превзошли классические методы в zero-shot режиме, а лучший результат показала кросс-модальная архитектура на базе vision-модели.
AI-processed from arXiv cs.LG; edited by Hamidun News
Why a new benchmark was needed
In July 2026, an international team of researchers introduced AirQualityArena (AQA) — a large open dataset and benchmark for air pollution forecasting that covers 6 major pollutants, 7 countries across 4 continents, and more than 14,000 station-pollutant time series. The work was published as a preprint on arXiv, and the data itself has been released in open access.
Why a new benchmark was needed
Previous air quality forecasting benchmarks are too narrow — both in geography and in the set of pollutants covered — and barely test modern time-series foundation models (TSFM) on real large-scale data. This is exactly the gap AirQualityArena fills. Machine learning is increasingly used to forecast pollution levels, but without a single broad standard it is hard to compare models.
The problem is not academic. According to the authors' estimate, air pollution causes around 7.9 million premature deaths every year, which makes accurate forecasting a public health priority.
"Air pollution causes an estimated 7.9 million premature deaths
annually, making accurate forecasting a critical public health priority," the arXiv preprint states.
What went into the dataset
AirQualityArena consists of two parts: the data (AQA-Data) and the benchmark (AQA-Bench). The dataset collects ground-based measurements of 6 major pollutants over a three-year period across 7 different countries on 4 continents — more than 14,000 station-pollutant series in total. This is noticeably broader than most previous datasets, which were limited to a single region or one or two substances.
*6 major air pollutants
*Observation period — 3 years
*7 countries across 4 continents
*More than 14,000 station-pollutant series
*11 tested foundation models plus classical baseline methods
The AQA-Bench benchmark evaluates models on the task of short-term forecasting — predicting pollutant levels for the near future. The authors emphasize the diversity of countries in the sample: this makes it possible to check how universally a model performs across different conditions.
Why foundation models won
Time-series foundation models outperformed classical methods because they deliver accurate forecasts in zero-shot mode — without fine-tuning for a specific station or region. The authors ran 11 leading TSFMs and classical baseline models through AQA-Bench on the task of short-term pollution forecasting. According to the study, foundation models consistently proved more accurate than classical methods across all regions.
Zero-shot capability matters in practice: a monitoring system can be deployed in a new city that does not yet have a long history of observations for training.
The best result was shown by a cross-modal architecture: to forecast time series, it employs a vision foundation model — a model originally trained on images. In other words, a "visual" neural network transfers its image-recognition experience to numerical measurement series.
"Foundation models for time series are effective zero-shot forecasters
and consistently outperform classical baseline methods," the authors note in the arXiv preprint.
What this means
Air quality forecasting is shifting toward universal foundation models: a single pretrained model can be applied across different countries at once without local tuning. The open dataset of 14,000+ series gives researchers a common reference point and speeds up the development of early warning systems for smog and pollution.
Frequently Asked Questions
What is AirQualityArena?
It is an open dataset (AQA-Data) and benchmark (AQA-Bench) for air pollution forecasting: 6 pollutants, 3 years of observations, 7 countries across 4 continents, more than 14,000 station-pollutant series.
Which model showed the best result?
The best result came from a cross-modal architecture that uses a vision foundation model for time-series forecasting. Overall, all tested time-series foundation models outperformed classical methods in zero-shot mode.
How many models were tested in the benchmark?
The authors compared 11 leading time-series foundation models and several classical baseline methods on the task of short-term air quality forecasting.
Want to stop reading about AI and start using it?
AI News is a curated feed of AI/tech news. Hamidun Academy teaches you to use AI systematically in your work.
The AI world, distilled — once a week
Seven stories that actually mattered, hand-picked. No noise, no reposts, no press releases.
Done! Check your inbox for a confirmation.