Skip to content
AI News HubLIVE
Public articles 9Collected articles 9Trust 88Refresh 120 min
Health HealthySource type OfficialFull-text rights Official full textLast ingested 2026-07-28ID kimi-blogStatus Enabled

Official Kimi/Moonshot blog listing; verify terms before displaying full body.

Latest public articles

PerceptionBench: Evaluating Atomic Visual Perception in MLLMs

PerceptionBench is a new benchmark that isolates and evaluates atomic visual perception capabilities in multimodal large language models. Derived from model failures across 42 benchmarks, it defines 10 atomic perceptual categories and includes 3,000 verified questions. No frontier model reaches 60% accuracy, and perception-related hallucination is the weakest capability. The benchmark aims to expose where perception breaks and drive progress toward faithful multimodal AI.

Kimi BlogIn-site articlePerceptionBench: Evaluating Atomic Visual Perception in MLLMs

Kimi K3 Tech Blog: Open Frontier Intelligence

Kimi introduces its most capable model, K3, with 2.8 trillion parameters, built on Kimi Delta Attention and Attention Residuals, featuring native vision and a 1-million-token context window. It is the world's first open 3T-class model, designed for frontier intelligence across long-horizon coding, knowledge work, and reasoning. While still trailing the most powerful proprietary models, it outperforms other tested models in evaluations.

Kimi BlogIn-site articleKimi K3 Tech Blog: Open Frontier Intelligence

Kimi K2: Open Agentic Intelligence

Kimi K2 is an open agentic intelligence platform featuring tools for Excel formulas, document conversion, AI agent deployment, code assistance, and research advancements like Kimi K2.6 and Agent Swarm.

Kimi BlogIn-site articleKimi K2: Open Agentic Intelligence

Kimi K2 Thinking

Kimi K2 is an open-source thinking model with tools for Excel, docs, code, and web browsing, supporting agent swarms and deep research.

Kimi BlogIn-site articleKimi K2 Thinking

Kimi Vendor Verifier

Kimi open-sources the Vendor Verifier (KVV) to help users verify the accuracy of inference implementations of open-source models. It includes six critical benchmarks for detecting common deployment issues and encourages infrastructure providers to fix root causes.

Kimi BlogIn-site articleKimi Vendor Verifier

Kimi K2.5 Tech Blog: Visual Agentic Intelligence

Kimi K2.5 is an open-source multimodal model that delivers state-of-the-art performance in coding and vision tasks. It features a self-directed agent swarm capable of orchestrating up to 100 sub-agents for parallel execution, reducing task completion time by up to 4.5x. The model also excels in office productivity, handling complex documents, spreadsheets, and presentations. Available on multiple platforms, Kimi K2.5 represents a significant step toward AGI for the open-source community.

Kimi BlogIn-site articleKimi K2.5 Tech Blog: Visual Agentic Intelligence

WorldVQA: Measuring Atomic World Knowledge in MLLMs

WorldVQA is a new benchmark to evaluate factual correctness of MLLMs on visual world knowledge. It includes 3,500 high-quality image-question pairs across 9 categories, with a focus on head vs tail distribution. Frontier models achieve below 50% accuracy, revealing overconfidence and gaps in visual knowledge.

Kimi BlogIn-site articleWorldVQA: Measuring Atomic World Knowledge in MLLMs

Kimi Agent Swarm: 100 Sub-Agents at Scale

Kimi launches Agent Swarm, a multi-agent architecture enabling up to 100 parallel sub-agents for horizontal scaling. The system self-organizes into roles like CEO, researcher, and analyst, autonomously decomposing tasks, assigning agents, and synthesizing results. It is up to 4.5x faster than sequential execution and excels in broad research, batch processing, and multi-perspective analysis. Now available in preview for top-tier subscribers.

Kimi BlogIn-site articleKimi Agent Swarm: 100 Sub-Agents at Scale

Kimi K2.6 Tech Blog: Advancing Open-Source Coding

Kimi K2.6 is a new open-source model with state-of-the-art coding, long-horizon execution, and agent swarm capabilities. This blog details its features, benchmarks, and community feedback.

Kimi BlogIn-site articleKimi K2.6 Tech Blog: Advancing Open-Source Coding

All sources