Claude Sonnet 4.5 deprecated
Anthropic announced deprecation of claude-sonnet-4-5-20250929, with Claude API retirement scheduled for 2026-11-30 and Sonnet 5.5 as the recommended replacement.
Model launches, retirements, API and SDK releases, MCP spec changes and price moves that affect code you have already shipped. Newest first, each entry linked to the vendor's own notes.
Anthropic announced deprecation of claude-sonnet-4-5-20250929, with Claude API retirement scheduled for 2026-11-30 and Sonnet 5.5 as the recommended replacement.
OpenAI released gpt-6.1-sol for coding and professional work at $2 input / $0.10 cached / $10 output (<=272K prompts), with multi-agent subagent delegation in beta in the Responses API.
At DevDay 2026 OpenAI announced Dots persistent agents, GPT-6.1 Sol, the Ultrafast speed tier, a Decisions API limited preview, Amazon Bedrock Managed Agents powered by OpenAI, and ChatGPT support for the draft MCP Events spec.
Agents built on the Agents API can now complete tasks in an OpenAI-hosted browser, while the developer's application handles website access approvals and sign-in.
gpt-6-astra can be called with service_tier "ultrafast" in the Responses API at $60 input / $300 output per 1M tokens (short context), with global and US data-residency processing only.
claude-sonnet-5-5 launched at $2/$10 per MTok with 1M context and 128K output, with five breaking changes vs Sonnet 5, including forced tool_choice returning 400 and computer_20251124 no longer accepted on the Claude API.
Refusals that arrive before any output are billed again when stop_details.category is bio, frontier_llm or reasoning_extraction; early refusals in other categories remain free.
OpenClaw 2026.9.6 added managed-update recovery, 30-day usage reporting, remote-workspace Files/Memory/Skills, and chat-model support for Claude Opus 5.5, GPT-6 Sol/Luna and Grok 4.7.
OpenAI released the reasoning models gpt-6-sol ($2/$0.20/$10) and gpt-6-luna ($0.10/$0.01/$0.50) on the Responses and Chat Completions APIs.
claude-opus-5-5 launched at $4/$20 per MTok (cheaper than Opus 5) with 1M context, 128K output and always-on adaptive thinking; thinking cannot be disabled and computer use requires computer_toolset_20260801.
With the inline-tools-2026-09-15 beta header, tools (including MCP toolsets via mcp-client-2026-09-15) can be added or changed in a mid-conversation system message without invalidating the prompt cache.
grok-4.7 launched with a 500K context window at $2 input / $0.50 cached / $6 output (doubling for prompts >=200K), function calling and structured outputs; the Batch API is not supported.
Google limited Gemini 2.5 model access to users who had already used them (not deprecated) and told new projects to use Gemini 3.5 Flash-Lite or 3.8 Flash.
OpenClaw 2026.9.5 introduced Atomic Updates, plugins that install without restarting the Gateway, shareable conversations, GPT Live in meetings and calls, and guided setup for teams of specialist agents.
Google released managed agent antigravity-preview-09-2026 with renamed PascalCase built-in tools and line-range file edits; antigravity-preview-05-2026 shuts down on 2026-10-05.
The Messages API can compact a conversation on request (compact-2026-09-04 beta header) and returns a signed compaction block that replaces earlier messages in later requests.
OpenAI launched the Agents API in public beta, exposing the open-source Codex harness as a managed service with durable sessions, MCP/tools, subagents and OpenAI-hosted, self-hosted or partner sandboxes, with no fee beyond tokens and tools.
gpt-live-1 full-duplex voice sessions became GA at $0.05 per minute (billed per second), with backend model or agent reasoning delegated via Responses or the client.
DeepSeek released the multimodal V4.1-Flash under the model name deepseek-flash, retired V4-Flash, and cut prices (peak $0.30 in / $1.20 out, off-peak half).
Claude Managed Agents gained an auto permission policy in which the server evaluates each agent or MCP tool call and runs it, denies it, or pauses for approval.
OpenAI released gpt-6-astra ($10/$1/$50) as its most capable model, adding async tool calling, mid-turn steering over WebSockets and mid-conversation reasoning-effort changes; tool calling requires the Responses API.
Google released gemini-3.8-flash, aimed at long-horizon software engineering and autonomous agents, at an introductory $0.75/$3.75 per 1M tokens through 2026-12-31.
Meta released Muse Spark 1.3 with max reasoning on Muse Code and the Meta Model API at $1.25/$4.25 per 1M tokens (contributor tier $0.10/$0.20).
claude-fable-5-1 launched at $10/$50 per MTok with cache reads cut to $0.25 (0.025x), 1M context and 128K output, alongside Claude Mythos 5.1 for Project Glasswing participants; forced tool_choice is unsupported.
Mutual TLS and X.509 workload identity federation became generally available for the OpenAI API, configurable in the Platform console.
The Assistants API shut down; OpenAI also announced that whisper-1 and the gpt-4o transcribe models will shut down on 2027-02-26.
GPT-5.6 Sol dropped to $4 input / $20 output per 1M tokens (20% lower input, 33% lower output), a promotional price available at least through 2026-11-21.
Computer use left beta as computer_toolset_20260801 (batch actions, zoom on by default), a browser_toolset_20260801 client toolset launched, and the Files API and Agent Skills/Skills API dropped their beta headers.
openai-agents-python v0.22.0 shipped runtime hardening, guardrail-output redaction, provider-option validation and expanded handoff graph visualization, with breaking changes listed in the release.
Z.ai released GLM-5.3 (same base as GLM-5.2, improved via post-training) at $1.40/$4.40 per 1M tokens; thinking can no longer be disabled.
DeepSeek shipped V4-Pro GA with stronger agent benchmarks, native OpenAI Responses API format (adapted for Codex) and low/high/max effort, and announced peak/off-peak pricing effective 2026-08-16.
Google released gemini-3.7-flash for coding and agentic workflows at an introductory price through 2026-12-31.
claude-agent-sdk v0.2.137 added ConversationResetMessage, message origin fields and resume_session_at/resume_drops_turn options for safely truncating resumed sessions.
Anthropic cancelled the scheduled 2026-09-01 increase to $3/$15, keeping Claude Sonnet 5 at $2/$10 per MTok.
Claude Managed Agents sessions gained hard spend budgets (pausing with budget_reached), an advisor model the agent can consult mid-turn, per-agent inference_geo, and auto-loading of skills from mounted GitHub repos.
Claude Enterprise organizations can route governed prompts to their own AI security server for allow/deny verdicts before inference, and claude-opus-4-1-20250805 was retired.
Alibaba launched Qwen3.8-Max, a 2.4T-parameter MoE with 1M context, on Model Studio, promising open weights the following week.
Priority processing was renamed Fast mode (2x price, up to 2.5x speed for GPT-5.6 Sol), and GPT-5.6 Luna became 80% cheaper and GPT-5.6 Terra 20% cheaper.
The new MCP revision removes the initialize handshake and protocol sessions, adds a mandatory server/discover RPC and the Multi Round-Trip Request pattern, moves tasks into an extension, and deprecates Roots, Sampling and Logging.
claude-opus-5 launched at $5/$25 per MTok with 1M context and 128K output, along with mid-conversation tool changes (beta) and a default server-side fallback mode.
The legacy model names deepseek-chat and deepseek-reasoner were scheduled for discontinuation on this date (announced 2026-04-24 with the V4 launch).
Google made gemini-3.6-flash and gemini-3.5-flash-lite generally available and deprecated the temperature, top_p and top_k sampling parameters.
Moonshot launched Kimi K3, a 2.8T-parameter open-weight MoE with 1M context, on the Kimi API at $3/$15 per 1M tokens, with weights released by 2026-07-27.
claude-sonnet-5 launched with 1M context and a new tokenizer (~30% more tokens); manual extended thinking and non-default sampling parameters now return 400.
Google launched a public preview of the Computer Use tool in Gemini 3.5 Flash, with browser, mobile and desktop environments, configurable safety policies and prompt-injection detection.
claude-sonnet-4-20250514 and claude-opus-4-20250514 were retired; requests now return errors.
gemini-2.0-flash, gemini-2.0-flash-001, gemini-2.0-flash-lite and gemini-2.0-flash-lite-001 were shut down, with gemini-3.5-flash or gemini-3.1-flash-lite as replacements.