GPT-6 Astra Ultrafast
GPT-6 Astra Ultrafast runs on NVIDIA Blackwell GPUs and is available now in the OpenAI API and to eligible ChatGPT Work and Codex users. It offers up to 8x faster token generation than the Astra Standard mode.
Every deployment deployedbyai has carried, in an edition or on the wire, 1110 rows so far. Every item sourced to the publisher's own page.
A row enters The Deploy Log when an edition carries it or the wire posts it. Each one names the publisher and the thing that shipped, with the number when there is one, and links to the page it came from, so you can open the source and check it yourself.
What shipped in AI, week by week: every row grouped by the week it shipped, then by lane.
What shipped in AI, month by month: every row grouped by the month it shipped, then by lane.
AI models, platform by platform: every model named by rows from 2 or more companies, with each of those rows.
What shipped in AI, platform by platform: every platform that has a page, most rows first.
GPT-6 Astra Ultrafast runs on NVIDIA Blackwell GPUs and is available now in the OpenAI API and to eligible ChatGPT Work and Codex users. It offers up to 8x faster token generation than the Astra Standard mode.
Cloudflare released two trained decision models, Clef and Clef-flash, hosted on Workers AI. Clef is currently the leader when evaluated against the Jev Decision Index, and the models are open-sourced on Hugging Face under an Apache 2.0 license.
Gemini 3.8 Live with Live Avatar adds near real-time visual presence to live dialogue, with lip-syncing, expressions, and turn-taking across 97 languages. It is available in Gemini Enterprise starting today.
Gemini 3.8 Flash TTS and Flash-Lite TTS generate speech in over 100 languages, with voice creation from natural language prompts, 2,000+ production-ready voices, and voice replication from a 30-second sample. Both are rolling out today in the Gemini API and Google AI Studio.
Google Vids now lets anyone with a Google or Google Workspace account generate 1080p HD videos at no cost using Gemini Omni 1.1 Flash, with scene extension, exact clip durations, and upscaling of existing AI clips.
Moonshot AI's Kimi K3 is available on Amazon Bedrock, the first open model to reach 2.8 trillion parameters, with a 1-million-token context window and a 2.5x improvement in scaling efficiency over Kimi K2.
Amazon SageMaker HyperPod Inference Gateway is a Kubernetes-native, GPU-aware routing system that deploys as a single EKS managed addon and reduces first-token latency by up to 82%.
OpenAI introduced Astra for Law, combining GPT-6 Astra with a legal search index covering more than 230 million URLs and 26 partner-built plugins, passing the Vals AI Legal Research Bench overall correctness check on 54.0% of questions versus 38.7% for GPT-6 Astra with web search alone.
GPT-6 Astra, rolling out to the OpenAI API and through Microsoft Azure and Amazon Bedrock at $10 per million input tokens and $50 per million output, the same headline price as Anthropic's Claude Fable 5.1, which shipped the same week. Cache reads are priced separately, and Fast mode costs 2x.
Waymo began welcoming first public riders in Denver, San Diego, and Tampa on September 1, 2026, marking 14 cities with fully autonomous trips. Tens of thousands in each city have signed up.
Mac Studio with M5 Ultra scales to a 36-core CPU, up to an 80-core GPU, and 512GB of unified memory with 1.2TB/s bandwidth, delivering up to 4.3x the peak AI compute performance of M3 Ultra and 9.8x more than M1 Ultra.
ChatGPT for Teens places users estimated under 18 or stating age 13 to 17 into a protected experience with Study Mode, responsible homework reminders, quizzes, learning visualizations, Study Hours, break reminders, and parent controls including Quiet Hours and safety notifications.
Replit introduces Free Mode powered by GPT-5.6 Luna, letting users get answers, suggestions, feedback, and analysis in seconds without consuming usage, with routing to GPT-5.6 Sol for tasks requiring more advanced reasoning.
Premium seats are now available on ChatGPT Business at $125 per user per month, or $100 per user per month billed annually, with 5x more usage than Standard seats and no five-hour usage limit.
GPT-5.6 Sol on Ultrafast mode runs up to 14x faster than Standard processing and generates up to 750 output tokens per second, launching first in the OpenAI API and powered by Cerebras.
Gemini 3.7 Flash is available at $0.75 per 1M input tokens and $3.75 per 1M output tokens through the end of the year, half the original 3.6 Flash cost per million tokens.
GPT-5.6 Sol in ChatGPT now gives more focused answers and makes factual errors 68% less common than GPT-5.5 Instant on financial, medical, and legal prompts. A new slider controls how much thought each response gets.
Grok Imagine Image 2.0 Preview from xAI is available on AI Gateway. It follows detailed instructions closely and plans typography and layout together, so dense visuals like infographics, posters, and title screens hold their structure and small text stays legible.
Shieldstral is a 3B open-weights multimodal safety classifier under Apache 2.0 that matches or outperforms open guard models up to 7x its size on text safety, refusal detection, policy adaptability, and multimodal benchmarks. It runs on a single 16GB NVIDIA GPU.
OpenAI is giving 100,000 researchers at selected academic institutions free access to frontier models including GPT-5.6 Sol Pro, starting with 10,000 researchers this summer and expanding through 2027, with access already available at the Institute for Advanced Study and École normale supérieure.
Lakeflow pipelines can now set run-as identity to an account group, in addition to user or service principal.
Version 1.0.92-4 adds copilot config subcommands and improves MCP server startup responsiveness.
Prime Intellect released Qwen3.5-0.8B-Reverse-Text-RL, a short RL fine-tune for the reverse-text environment.
Prime Intellect released Qwen3.5-0.8B-Reverse-Text-SFT, a short SFT fine-tune of Qwen3.5-0.8B.
Ultralytics 8.4.173 adds AMD Xilinx Vitis AI export for Versal AI Edge Gen 2 NPUs.
openai-java v4.76.0 adds custom voices and agent session webhook events.
openai-node v7.28.0 adds custom voice creation and agent session events.
supervision 0.30.7 fixes video processing hangs, dataset export, mAP, and Transformers loading.
Cloudflare's AI Search is generally available, combining Workers AI, Vectorize, R2, and Browser Run into a fully managed index and retrieval pipeline, with billing starting November 1, 2026.
Cloudflare Basin is generally available as a serverless data analytics platform built on Apache Iceberg and R2 Object Storage, including Basin Pipelines, Basin Catalog, and Basin SQL.
Workers KV Instant, powered by Quicksilver, offers 100 times faster p99 reads than classic mode with reads resolving in under two milliseconds at p99, and writes replicating in around 250ms for 99% of writes.
NVIDIA DGX Spark will be available with 64GB of unified memory from Acer, ASUS, Dell, Gigabyte, HP and MSI starting Friday, Oct. 23 at $4,999, supporting up to 100-billion-parameter models fully on device.
What shipped in AI, company by company: every company with 3 or more rows, most rows first.
The edition every row came from arrives in your inbox, free.
Join free