GLM-5.5
❖ The GLM-5 series is our flagship model family for coding and long-horizon tasks. It marks a substantial leap in long-horizon task capability over its predecessor and delivers that capability on a solid 1M-token context. It is pure open with an MIT open-source license — no regional limits, technical access without borders.
Chatbot Arena +
❖ This leaderboard is based on the following benchmarks. Arena + — an agent-driven battle platform for large language models (LLMs). We use LLM-as-a-judge to compute Elo ratings. AAII — Artificial Analysis Intelligence Index aggregating 10 challenging evaluations. ARC-AGI — Artificial General Intelligence benchmark v2 to measure fluid intelligence.
Qwen3.8
❖ We are officially releasing Qwen 3.8-Max, the most capable model in the Qwen family to date. This also marks the first time we will open-source the weights of a Qwen-Max-class model. Built upon the architectural foundation of Qwen 3.5, Qwen 3.8-Max scales to 2.4 trillion parameters, delivering comprehensive improvements across coding, work, research, and long-horizon tasks.
Kimi K3
❖ We introduce Kimi K3 — our most capable model. Kimi K3 is a 2.8T-parameter model built on our Kimi Delta Attention and Attention Residuals, with native vision capabilities and a 1-million-token context window. It is the world’s first open 3T-class model, designed for frontier intelligence across long-horizon coding, knowledge work, and reasoning.
LongCat-2.0
❖ We introduce LongCat-2.0, a large-scale MoE language model with 1.6 trillion total parameters and ~48 billion activated per token — a substantial step up from previous LongCat models, accompanied by several architectural improvements. Both the full training run and the large-scale deployment are built entirely on AI ASIC superpods.
Agents-A1
❖ Agents-A1 is a 35B Mixture-of-Experts Agentic Model that reaches trillion-parameter-level performance by scaling the agent horizon. We investigate agent-horizon scaling from two perspectives: scaling long-horizon trajectories and scaling heterogeneous agent abilities.
Agent Harnesses
❖ Curated, ranked list of AI agent harnesses — the runtimes that close the loop between a stateless model and the outside world. 150+ harnesses · 12 categories · weekly-rescored · MCP-ready : claude mcp add agent-harnesses -- uvx agent-harnesses-mcp
Kimi K2.7
❖ Kimi K2.7 Code is a coding-focused agentic model built upon Kimi K2.6. With substantial improvements on real-world long-horizon coding tasks, it strengthens end-to-end task completion across complex software engineering workflows while improving token efficiency, reducing thinking-token usage by approximately 30% compared with Kimi K2.6.