#LLM
Curated coverage on «LLM»: releases, research, benchmarks, open-source.

IBM and Artificial Analysis create benchmark: AI agents fail at IT tasks
Large language models scored less than 50% on the new ITBench-AA benchmark for assessing AI agents' ability to solve corporate IT tasks. Thi

DeepSeek Makes 75% Discount on V4 Pro Permanent — Prices Fall 4x
DeepSeek announced that the 75% discount on its flagship V4 Pro model is becoming permanent, significantly complicating competitors' fight i

World Models: How AI Learns to Understand Reality Instead of Text

Anthropic blacklisted by Pentagon, but NSA cleared to use Claude

Claude Code Leak: Anthropic Accidentally Publishes Entire Source Code to Public npm

Spotify Added AI-Generated Briefs and Q&A for Podcasts

AI Tools Flood Linux Maintainers with Duplicate Vulnerability Reports
Linux kernel maintainers are overwhelmed by the stream of automatic vulnerability reports generated by AI tools designed to find bugs.

Why RAG Chatbots Work Great in Demos but Produce Nonsense in Production
RAG bots shine on internal documentation demos but confidently generate nonsense in real work. A story about the gap between five prepared q













