Claude Opus 4.8
Claude Opus 4.8 is available today at the same price as Opus 4.7: $5 per million input tokens and $25 per million output tokens. Fast mode runs at 2.5x speed and is now three times cheaper than previous models.
The Deploy Log | Month by month
43 rows dated May 2026 are in The Deploy Log: 11 leads, 14 in Models and APIs, 16 in Tools and products, 1 in Robotics, hardware and chips and 1 in Breakthroughs, from 10 publishers. They include Claude Opus 4.8 from Anthropic, Rosalind Biodefense from OpenAI, KPMG Claude alliance from Anthropic, Stable Audio 3.0 from Stability AI, Gemini 3.5 Flash from Google DeepMind and 38 others. Every item sourced to the publisher's own page.
Claude Opus 4.8 is available today at the same price as Opus 4.7: $5 per million input tokens and $25 per million output tokens. Fast mode runs at 2.5x speed and is now three times cheaper than previous models.
OpenAI is launching Rosalind Biodefense to sponsor access to GPT-Rosalind for trusted developers, and extending trusted access to select U.S. government and allied partners.
KPMG embeds Claude inside Digital Gateway, its client work platform, and gives all 276,000+ employees access to Claude, with Claude Cowork and Managed Agents for building AI capabilities in minutes.
Stable Audio 3.0 releases four models: Small SFX for on-device sound effects, Small for full music composition on-device, Medium for up to 6:20 tracks, and Large via API for low-latency high-volume generation.
Gemini 3.5 Flash is generally available via Google Antigravity, the Gemini API in Google AI Studio and Android Studio, Gemini Enterprise Agent Platform and Gemini Enterprise, and in the Gemini app and AI Mode in Search. It scores 76.2% on Terminal-Bench 2.1, 1656 Elo on GDPval-AA, and 83.6% on MCP Atlas, outperforming Gemini 3.1 Pro.
Anthropic doubled Claude Code's five-hour rate limits for Pro, Max, Team, and seat-based Enterprise plans, removed the peak hours limit reduction on Claude Code for Pro and Max accounts, and raised API rate limits considerably for Claude Opus models, all effective May 6, 2026.
GPT-5.5 Instant replaces GPT-5.3 Instant as the default model for all ChatGPT users and in the API as chat-latest. It produced 52.5% fewer hallucinated claims than GPT-5.3 Instant on high-stakes prompts covering medicine, law, and finance, and reduced inaccurate claims by 37.3% on conversations users flagged for factual errors.
GPT-Realtime-2, GPT-Realtime-Translate, and GPT-Realtime-Whisper are available in the Realtime API. GPT-Realtime-2 scores 15.2% higher on Big Bench Audio than GPT-Realtime-1.5 and 13.8% higher on Audio MultiChallenge for instruction following. Context window increases from 32K to 128K.
ChatGPT Enterprise and API Platform achieved FedRAMP 20x Moderate authorization. Agencies can access GPT-5.5 in the FedRAMP environment, and Codex Cloud access via FedRAMP ChatGPT Enterprise workspace is coming soon.
Workflows is released in public preview as the orchestration layer for enterprise AI, built on Temporal's durable execution engine. It includes durable execution, observability in Studio, human-in-the-loop approval via a single wait_for_input() call, and deployment where workers run in your environment.
OpenAI models including GPT-5.5, Codex, and Amazon Bedrock Managed Agents powered by OpenAI all launched in limited preview on AWS. Codex usage can apply toward AWS cloud commitments, and all customer data is processed by Amazon Bedrock.
Project Genie now grounds generated worlds in Google Street View imagery for places in the U.S., and access is rolling out to all eligible Google AI Ultra $200 subscribers globally.
OlmoEarth v1.1 is a new family of Earth observation models that cuts compute costs by up to 3x while maintaining OlmoEarth v1's performance on research benchmarks.
OpenAI and Dell Technologies partner to bring Codex to hybrid and on-premises environments, with more than 4 million developers now using Codex every week.
PaddleOCR 3.5 brings OCR and document parsing tasks closer to the Hugging Face ecosystem, with supported models able to run with Transformers as an inference backend.
Databricks is making GPT-5.5 available through AI Unity Gateway after the model became the first to surpass 50% accuracy on OfficeQA Pro and reduced errors by 46% compared to GPT-5.4.
Anthropic and the Gates Foundation committed $200 million in grant funding, Claude usage credits, and technical support over four years for programs in global health, life sciences, education, and economic mobility.
IBM released two Apache 2.0 multilingual embedding models built on ModernBERT: a 97M-parameter compact model scoring 60.3 on MTEB Multilingual Retrieval and a 311M full-size model scoring 65.2, both covering 200+ languages with 32K-token context.
OpenAI is launching the OpenAI Deployment Company with more than $4 billion of initial investment and has agreed to acquire Tomoro, bringing approximately 150 experienced Forward Deployed Engineers and Deployment Specialists from day one.
PwC will roll out Claude Code and Cowork starting with U.S. teams and expanding toward a global workforce of hundreds of thousands of professionals, with a program to train and certify 30,000 PwC professionals on Claude.
ChatGPT personal finance preview lets Pro users in the U.S. connect financial accounts via Plaid, with support for 12,000+ institutions. It provides a dashboard and grounded answers, with Intuit support coming soon.
Codex is now in the ChatGPT mobile app in preview, and more than 4 million people use Codex every week.
GPT-5.5-Cyber is rolling out in limited preview to defenders responsible for securing critical infrastructure, with more permissive behavior for authorized red teaming, penetration testing, and controlled validation, while GPT-5.5 with Trusted Access for Cyber remains the recommended starting point for most security workflows.
OpenAI expands ChatGPT ads with beta self-serve Ads Manager and cost-per-click bidding, adding partners like Dentsu, Omnicom, Publicis, and WPP.
NVIDIA released Nemotron 3 Nano Omni, a 30B-A3B omni-modal model for documents, audio, video, and agentic computer use, with OCRBenchV2-En 65.8, MMLongBench-Doc 57.5, Video-MME 72.2, and VoiceBench 89.4, and checkpoints in BF16, FP8, and NVFP4 on Hugging Face.
Vercel Function invocations move from package-based to per-unit pricing at $0.0000006 per invocation for Pro customers, previously $0.60 per 1M invocations.
Vercel Sandbox now supports installing and running Docker inside a sandbox, with the example using dnf install docker and starting dockerd.
Mistral released Search Toolkit in public preview, a composable framework for building production search pipelines with ingestion, retrieval, and evaluation in one interface.
Notion now lets you merge cells in simple tables, with the option to select multiple cells and open the cell menu to choose Merge cells.
The Vercel AI Gateway plugin gives any WordPress site access to hundreds of models from 40+ providers through a single API key, requiring WordPress 7.0.
Fast mode for Claude Opus 4.7 is available on AI Gateway in research preview, delivering about 2.5x faster output token generation at 6x standard Opus rates, with input at $30 per 1M tokens and output at $150 per 1M tokens.
Vercel Sandbox now supports Node.js version 26 by upgrading @vercel/sandbox to 1.10.2 or later and setting the runtime property to node26.
Claude for Small Business ships with 15 ready-to-run agentic workflows and 15 skills, integrating with QuickBooks, PayPal, HubSpot, Canva, Docusign, Google Workspace, and Microsoft 365.
Notion launched a Developer Platform with Workers, External Agents API, database sync, and a CLI, letting developers and agents build on Notion infrastructure.
Vercel's Chat SDK now supports cross-platform conversation history through transcripts and identity options, letting the same user keep up to 200 messages per user for 30 days across every platform adapter.
Notion added a Custom Agent Directory in your Library where you can browse all your workspace's agents, pin favorites, and create new ones to automate your team's busywork 24/7.
Vercel open sourced deepsec, a security harness powered by coding agents that runs on your own infrastructure, uses Opus 4.7 at max effort and GPT 5.5 at xhigh reasoning, and scales up to 1,000+ concurrent sandboxes on Vercel's codebases.
Vercel's Chat SDK adds a Messenger adapter, extending support to Slack, Discord, GitHub, Teams, and Telegram, with messages, reactions, attachments, and postback buttons.
Notion introduces Plan Mode, where agents ask clarifying questions and build a detailed plan before making significant changes like rewriting pages or bulk database updates.
Starting April 29th, the maximum retention policy for Vercel Hobby plans is capped at 30 days, with deployments outside the retention window automatically removed except the 10 most recent production deployments and any aliased deployments.
Custom Agents can now read and reply in private Slack channels, with access controlled by inviting agents to specific channels.
Waymo is welcoming its first public riders to the Ojai, a rider-first vehicle built on the Waymo Driver that has served over 20 million fully autonomous trips across 11+ cities.
Ai2 releases EMO, a 1B-active, 14B-total-parameter MoE trained on 1 trillion tokens, where modular structure emerges from data, allowing selective use of just 12.5% of experts with near full-model performance.
The edition every row came from arrives in your inbox, free.
Join free