Publisher · verified by editors

MarkTechPost

AI news source. Articles are auto-selected and adapted by Hamidun News editors.

290 articles in Hamidun·Latest: July 17· Active·marktechpost.com ↗

Latest publications

StepFun Releases StepAudio 2.5 Realtime Voice Model with Roleplay Support
LLMMarkTechPost

StepFun Releases StepAudio 2.5 Realtime Voice Model with Roleplay Support

Chinese AI lab StepFun introduced the real-time voice model StepAudio 2.5 Realtime, which outperforms competitors in speech naturalness and can adapt voice to user scenarios.

May 25, 2026·2 min
Langfuse for LLM Engineers: Complete Tracing and Experimentation Pipeline
LLMMarkTechPost

Langfuse for LLM Engineers: Complete Tracing and Experimentation Pipeline

Langfuse is a tool for debugging and optimizing LLM applications. Learn how to set up a complete monitoring pipeline, prompt management, and experiments without paid models.

May 25, 2026·2 min
WorkOS Introduces auth.md — An Open Protocol for AI Agent Registration
LLMMarkTechPost

WorkOS Introduces auth.md — An Open Protocol for AI Agent Registration

WorkOS has released auth.md — an open standard that enables AI agents to register in applications via Markdown file without human intervention.

May 25, 2026·3 min
ByteDance Presented Lance: One Model for Video Understanding, Generation, and Editing
LLMMarkTechPost

ByteDance Presented Lance: One Model for Video Understanding, Generation, and Editing

ByteDance released Lance — an open-source model that operates in a single framework with images and video: understands, generates, and edits content using only 3B active parameters.

May 25, 2026·2 min
Cohere releases Command A+: 218 billion parameters for agents on two GPUs
LLMMarkTechPost

Cohere releases Command A+: 218 billion parameters for agents on two GPUs

Cohere has unveiled the open Command A+ model with 218 billion parameters and multimodal capabilities, running on two H100 GPUs and supporting 48 languages.

May 25, 2026·2 min
Perplexity Opens Bumblebee Scanner to Protect Developer Systems
LLMMarkTechPost

Perplexity Opens Bumblebee Scanner to Protect Developer Systems

Perplexity has published the source code for Bumblebee, a tool for scanning vulnerabilities in developer system dependencies without running any code.

May 25, 2026·2 min
Alibaba Introduces Qwen3.7-Max: An Agent with Million-Token Context
LLMMarkTechPost

Alibaba Introduces Qwen3.7-Max: An Agent with Million-Token Context

Alibaba introduced Qwen3.7-Max — the most advanced agent model from Qwen with a 1M-token context window and reasoning mode for complex multi-step tasks.

May 25, 2026·3 min
CopilotKit Redefines Architecture for AI Agents in 2026
LLMMarkTechPost

CopilotKit Redefines Architecture for AI Agents in 2026

CopilotKit released a new stack for agentic AI developers: AG-UI protocol, AIMock testing platform, and Pathfinder server—a complete production solution.

May 25, 2026·3 min
OpenMythos: Building Advanced Transformers with MLA and GQA in Colab
LLMMarkTechPost

OpenMythos: Building Advanced Transformers with MLA and GQA in Colab

OpenMythos enables building recurrent transformers in Google Colab, comparing MLA and GQA architectures. A new video tutorial verifies model stability through spectral radius analysis of injection matrices.

May 25, 2026·2 min
Nous Research Introduces CNA: Controlling LLM Behavior Without Retraining
LLMMarkTechPost

Nous Research Introduces CNA: Controlling LLM Behavior Without Retraining

Nous Research introduced the Contrastive Neuron Attribution (CNA) method, which enables controlling the behavior of large language models by identifying and disabling individual neural circuits without retraining and wit

May 25, 2026·3 min
Eight best authentication platforms for AI-agents and MCP in 2026
LLMMarkTechPost

Eight best authentication platforms for AI-agents and MCP in 2026

MCP reached 97 million SDK downloads per month. AI agents are massively transitioning from experiments to production environments, and choosing the right authentication platform has become one of the…

May 25, 2026·2 min
SuperClaude Framework helps structure workflows for Claude API
LLMMarkTechPost

SuperClaude Framework helps structure workflows for Claude API

SuperClaude Framework provides developers with built-in components for creating advanced AI workflows: commands, agents, execution modes, and session memory — all in one system.

May 25, 2026·2 min
Tencent Released a Local Memory System for AI Agents TencentDB
LLMMarkTechPost

Tencent Released a Local Memory System for AI Agents TencentDB

Tencent open-sourced TencentDB Agent Memory — a local memory system for AI agents that reduces token consumption by 61% and improves accuracy by 28%.

May 25, 2026·3 min
NVIDIA Introduces Gated DeltaNet-2: Linear Attention with Separate Memory Gates
LLMMarkTechPost

NVIDIA Introduces Gated DeltaNet-2: Linear Attention with Separate Memory Gates

NVIDIA has created a new linear attention mechanism, Gated DeltaNet-2, which improves memory management in large language models through separate erase and write gates instead of a single gate.

May 25, 2026·3 min
Google Introduces Gemini 3.5 Flash: Fast and Affordable Model for Coding and AI Agents
LLMMarkTechPost

Google Introduces Gemini 3.5 Flash: Fast and Affordable Model for Coding and AI Agents

At I/O 2026, Google introduced Gemini 3.5 Flash — a model that is 75% cheaper than the flagship version, runs 4 times faster, and outperforms it on coding and automation tasks.

May 21, 2026·3 min
Alibaba releases a translator with 2.8-second latency across 60 languages
LLMMarkTechPost

Alibaba releases a translator with 2.8-second latency across 60 languages

Alibaba introduced a model for real-time translation of video and speech simultaneously across 60 languages — with minimal latency and preservation of the speaker’s voice.

May 21, 2026·2 min
NVIDIA introduced Nemotron-Labs-Diffusion: a model with triple decoding
LLMMarkTechPost

NVIDIA introduced Nemotron-Labs-Diffusion: a model with triple decoding

NVIDIA has released the Nemotron-Labs-Diffusion language model, which combines three decoding modes and processes tokens 6 times faster than Qwen3-8B.

May 21, 2026·2 min
Generating knowledge graphs from text: a practical guide with kg-gen and NetworkX
LLMMarkTechPost

Generating knowledge graphs from text: a practical guide with kg-gen and NetworkX

A tutorial on automatically extracting entities and relations from text with kg-gen, building interactive knowledge graphs, and analyzing them with NetworkX.

May 21, 2026·3 min
Turbovec: a Rust vector index with Google Research's TurboQuant algorithm
LLMMarkTechPost

Turbovec: a Rust vector index with Google Research's TurboQuant algorithm

Turbovec uses Google's TurboQuant algorithm to compress vectors 16x without pre-training, simplifying the deployment of RAG applications.

May 21, 2026·2 min
Best platforms for agentic AI in 2026: ranking of Salesforce, Microsoft, and others
LLMMarkTechPost

Best platforms for agentic AI in 2026: ranking of Salesforce, Microsoft, and others

Companies are moving from pilots to production. MarkTechPost compiled a top-10 ranking of agentic AI platforms: Salesforce Agentforce, Microsoft Copilot Studio, ServiceNow, and others. Verified pricing and real-world dep

May 19, 2026·3 min
NVIDIA developed a method for training neural networks at 4-bit precision
LLMMarkTechPost

NVIDIA developed a method for training neural networks at 4-bit precision

NVIDIA introduced NVFP4, a methodology for training large models at 4-bit precision instead of the standard 8-bit, halving memory use without loss of quality.

May 19, 2026·3 min
OpenAI unveils the MRC protocol for supercomputer networks with millions of GPUs
LLMMarkTechPost

OpenAI unveils the MRC protocol for supercomputer networks with millions of GPUs

OpenAI has created a new open network protocol, MRC, for large AI clusters. It distributes data across hundreds of paths and recovers from failures in microseconds, enabling the construction of supercomputers with 100,00

May 17, 2026·3 min
Meta AI introduces NeuralBench — a framework for testing brain activity models
LLMMarkTechPost

Meta AI introduces NeuralBench — a framework for testing brain activity models

Meta released NeuralBench, an open framework for standardized testing of EEG-based AI models, bringing 36 tasks, 94 datasets, and 13,603 hours of brain recordings into a single interface.

May 17, 2026·2 min
How to compress a language model 3x: a guide to FP8, GPTQ, and SmoothQuant
LLMMarkTechPost

How to compress a language model 3x: a guide to FP8, GPTQ, and SmoothQuant

Developers received a step-by-step guide to compressing large language models with llmcompressor, comparing the effectiveness of FP8, GPTQ, and SmoothQuant quantization to reduce hardware load.

May 17, 2026·3 min
OpenAI released three audio models: translation, transcription, and real-time reasoning
LLMMarkTechPost

OpenAI released three audio models: translation, transcription, and real-time reasoning

OpenAI expanded the Realtime API with three new audio models for voice processing: reasoning agents, multilingual translation, and streaming transcription.

May 17, 2026·2 min
Anthropic created a tool to translate Claude's thoughts into human language
LLMMarkTechPost

Anthropic created a tool to translate Claude's thoughts into human language

Anthropic developed Natural Language Autoencoders, a technology that translates Claude's internal activations into textual explanations, revealing how the neural network works.

May 17, 2026·2 min
NVIDIA packed 3 models into one file and made training 360× more efficient
LLMMarkTechPost

NVIDIA packed 3 models into one file and made training 360× more efficient

NVIDIA introduced Star Elastic, a method that packs three models of different sizes into a single checkpoint and enables training that is 360× more efficient.

May 17, 2026·3 min
NVIDIA released cuda-oxide: a compiler for Rust code on GPUs
LLMMarkTechPost

NVIDIA released cuda-oxide: a compiler for Rust code on GPUs

NVIDIA introduced cuda-oxide, a tool for compiling Rust functions directly into PTX code for GPUs. This will simplify the development of CUDA applications in Rust and make parallel computing more accessible.

May 17, 2026·1 min
NadirClaw: saving on LLM requests with smart prompt routing
LLMMarkTechPost

NadirClaw: saving on LLM requests with smart prompt routing

NadirClaw is a tool for intelligent prompt routing that classifies requests as simple or complex, sending them to the appropriate model to reduce costs.

May 17, 2026·2 min
Hermes Agent by Nous Research took the lead in token usage on OpenRouter
LLMMarkTechPost

Hermes Agent by Nous Research took the lead in token usage on OpenRouter

The open-source AI agent Hermes Agent by Nous Research surpassed the closed-source platform OpenClaw and took first place on OpenRouter, generating 224 billion tokens a day. This happened in just three months and shows t

May 17, 2026·3 min