WebGPU kernels
Hugging Face released @huggingface/kernels with 207 WebGPU kernels, 2.57x faster than ORT WebGPU by geometric mean across 809 matching test cases on an Apple M4 GPU.
Every deployment from Hugging Face in The Deploy Log, 74 rows, newest first.
Hugging Face released @huggingface/kernels with 207 WebGPU kernels, 2.57x faster than ORT WebGPU by geometric mean across 809 matching test cases on an Apple M4 GPU.
Meta released Muse Glimmer, a 30B multimodal model under Apache 2.0, with day-0 support in transformers, llama.cpp, vLLM, and Inference Endpoints.
Baseten is now a supported Inference Provider on the Hugging Face Hub, launching support for conversational and text-generation tasks with open-weight LLMs such as Kimi K3, DeepSeek V4 Flash, and GLM-5.2.
Nunchaku Lite brings 4-bit diffusion inference to Diffusers, cutting peak VRAM by up to 50% and improving latency by 30%.
LeRobot v0.6.0 ships world model policies VLA-JEPA, FastWAM, and LingBot-VA, new VLAs including GR00T N1.7 and MolmoAct2, reward models Robometer and TOPReward, and six new simulation benchmarks under lerobot-eval.
The transformers vLLM backend now meets or beats native throughput on Qwen3 4B dense, 32B dense, and 235B-parameter FP8 MoE models, using torch.fx static analysis and ast source rewriting to apply inference-specific layer fusions at runtime.
Hugging Face and Cerebras demonstrate a real-time speech-to-speech pipeline pairing Gemma 4 31B on Cerebras with Nvidia's Parakeet for speech recognition and Alibaba's Qwen3TTS for text-to-speech, already powering more than 9,000 Reachy Mini robots.
PP-OCRv6 scales from 1.5M to 34.5M parameters across tiny, small, and medium tiers, with the medium and small tiers supporting 50 languages and the medium tier reaching 86.2% detection Hmean and 83.2% recognition accuracy.
GLM-5.2 is an MIT-licensed open-source model with a 1M-token context. It scores 81.0 on Terminal-Bench 2.1 and 62.1 on SWE-bench Pro, and is the highest-ranked open-source model on FrontierSWE, PostTrainBench, and SWE-Marathon.
Strands Robots is an Apache 2.0 SDK from AWS that exposes robot abstractions, simulation, and the LeRobot stack as AgentTools, with sim and hardware datasets sharing the same on-disk format.
JetBrains released Mellum2, a 12B-parameter Mixture-of-Experts model that activates only 2.5B parameters per token and delivers more than 2x faster inference than similar-sized models under the Apache 2.0 license.
NVIDIA released Nemotron 3.5 Content Safety, a 4B-parameter model that unifies multimodal input, 12-language explicit coverage, custom policy enforcement, and auditable reasoning traces in one inference call.
OlmoEarth v1.1 is a new family of Earth observation models that cuts compute costs by up to 3x while maintaining OlmoEarth v1's performance on research benchmarks.
PaddleOCR 3.5 brings OCR and document parsing tasks closer to the Hugging Face ecosystem, with supported models able to run with Transformers as an inference backend.
IBM released two Apache 2.0 multilingual embedding models built on ModernBERT: a 97M-parameter compact model scoring 60.3 on MTEB Multilingual Retrieval and a 311M full-size model scoring 65.2, both covering 200+ languages with 32K-token context.
NVIDIA released Nemotron 3 Nano Omni, a 30B-A3B omni-modal model for documents, audio, video, and agentic computer use, with OCRBenchV2-En 65.8, MMLongBench-Doc 57.5, Video-MME 72.2, and VoiceBench 89.4, and checkpoints in BF16, FP8, and NVFP4 on Hugging Face.
DeepInfra is now a supported Inference Provider on Hugging Face Hub, offering over 100 models including DeepSeek V4, Kimi-K2.6, GLM-5.1, with PRO users getting $2 monthly Inference credits.
IBM released Granite 4.1, a family of dense LLMs (3B, 8B, 30B) trained on ~15T tokens with 512K context, Apache 2.0, with the 8B matching the previous 32B MoE.
DeepSeek released V4 today with two MoE checkpoints on the Hub: DeepSeek-V4-Pro at 1.6T total parameters with 49B active, and DeepSeek-V4-Flash at 284B total with 13B active, both with a 1M-token context window.
Hugging Face published a guide on using Transformers.js in a Chrome extension, with a demo powered by Gemma 4 E2B for local AI features.
Sentence Transformers v5.4 adds multimodal embedding and reranker models that encode and compare text, images, audio, and video through the same API, with Qwen3-VL-Embedding-2B and Qwen3-VL-Reranker-2B supported.
Overworld released Waypoint-1.5, a real-time video world model with 720p and 360p tiers that runs locally on RTX 3090 through 5090 hardware at up to 60 FPS.
Falcon Perception is a 0.6B-parameter early-fusion Transformer that reaches 68.0 Macro-F1 on SA-Co versus 62.3 for SAM 3, and Falcon OCR is a 0.3B model scoring 80.3 on olmOCR and 88.6 on OmniDocBench.
Granite 4.0 3B Vision is available on Hugging Face under Apache 2.0 and leads on PubTablesV2 cropped at 92.1 and full-page at 79.3, OmniDocBench at 64.0, and TableVQA at 88.1.
TRL v1.0 is released with a stable core following semantic versioning and an experimental layer, and the library is downloaded 3 million times a month.
OpenMed trained 4 production mRNA language models across 25 species in 55 GPU-hours for $165, with CodonRoBERTa-large-v2 reaching perplexity 4.10 and Spearman CAI correlation 0.40.
Gradio released gradio.Server, letting developers pair any custom frontend with Gradio's backend, including queuing and ZeroGPU support.
ServiceNow AI released EVA, an end-to-end evaluation framework for voice agents that jointly scores accuracy (EVA-A) and experience (EVA-X), with an airline dataset of 50 scenarios and results for 20 systems.
Holotron-12B, post-trained from NVIDIA Nemotron-Nano-2 VL, raises WebVoyager from 35.1% to 80.5% and achieves over 2x higher throughput than Holo2-8B on a single H100.
LeRobot v0.5.0 adds full Unitree G1 humanoid support, Pi0-FAST autoregressive VLAs, Real-Time Chunking, streaming video encoding with zero wait between episodes, and EnvHub for loading simulation environments from the Hub.
Modular Diffusers introduces a new way to build diffusion pipelines by composing reusable blocks, with a quickstart example running FLUX.2 Klein 4B and community pipelines including Krea Realtime Video achieving 11fps on a single B200 GPU.
GGML, creators of llama.cpp, are joining Hugging Face, with Georgi Gerganov and team dedicating 100% of their time maintaining llama.cpp with full autonomy and leadership on technical directions and the community.
Gradio 6 shipped gr.HTML with custom templates, scoped CSS, and JavaScript interactivity, letting Claude or any other frontier LLM generate frontend, backend, and state management in a single Python file with no build step.
Unsloth and Hugging Face Jobs enable fast LLM fine-tuning of LiquidAI/LFM2.5-1.2B-Instruct through coding agents like Claude Code and Codex, with Unsloth providing approximately 2x faster training and about 60% less VRAM usage compared to standard methods.
Hugging Face released an agent skill that teaches Codex and Claude to write production CUDA kernels, producing an RMSNorm kernel for Qwen3-8B with an average 1.94x speedup on H100.
Turing contributed a production-grade calendar management environment to OpenEnv, where agents achieved close to 90% success with explicit calendar identifiers but dropped to roughly 40% with natural language descriptions.
Transformers.js v4 is now available on NPM with a C++ WebGPU runtime, a 10x build time drop from 2 seconds to 200 milliseconds, and a default export that is 53% smaller.
Holo2-235B-A22B Preview achieves 78.5% on Screenspot-Pro in agent mode within 3 steps and 79.0% on OSWorld G, and is available on Hugging Face.
SyGra 2.0.0 introduces Studio, an interactive environment that turns synthetic data generation into a visual canvas with guided model forms, data previews, and live execution streaming.
Hugging Face released Daggr, an open-source Python library for building AI workflows that connect Gradio apps, ML models, and custom functions, with automatic visual canvas.
IBM Research released AssetOpsBench, a benchmark for industrial AI agents with 2.3M sensor telemetry points and 4.2K work orders.
Microsoft introduced Differential Transformer V2, improving inference speed and training stability for production LLMs, with experiments still running.
Overworld released Waypoint-1, a real-time interactive video diffusion model trained on 10,000 hours of game footage, with weights on the Hub.
NVIDIA showed how to build a personal office robot agent using DGX Spark with Reachy Mini, combining NVIDIA Nemotron 3 Nano for reasoning, Nemotron Nano 2 VL for vision, and ElevenLabs for text-to-speech.
Falcon-H1-Arabic 3B, 7B, and 34B models outperform all SOTA models of similar sizes and sometimes bigger, with the 34B reaching approximately 75% on OALL and surpassing Llama-3.3-70B.
NVIDIA released Cosmos Reason 2, an open reasoning vision language model for physical AI that tops the Physical AI Bench and Physical Reasoning leaderboards as the #1 open model for visual understanding.
ServiceNow released AprielGuard, an 8B parameter safety and security safeguard model that detects 16 safety risk categories and adversarial attacks including prompt injection, jailbreaks, chain-of-thought corruption, context hijacking, memory poisoning, and multi-agent exploit sequences across standalone prompts, multi-turn conversations, and agentic workflows.
Transformers v5 redesigns tokenizers, separating architecture from trained vocab, making them inspectable, customizable, and trainable from scratch.
NVIDIA released the full evaluation recipe for Nemotron 3 Nano 30B A3B using NeMo Evaluator, with scores like 78.3 on MMLU-Pro and 89.1 on AIME 2025.
CUGA, a configurable generalist agent, is now on Hugging Face Spaces, achieving #1 on AppWorld and top-tier on WebArena.
Codex can now run end-to-end ML experiments using Hugging Face Skills, including fine-tuning, evaluation, and report generation.
llama.cpp server now supports router mode, allowing dynamic loading, unloading, and switching between multiple models without restarting.
Intel's DeepMath agent, built on Qwen3-4B Thinking and fine-tuned with GRPO, reduces output lengths by up to 66% while improving accuracy on MATH500, AIME, HMMT, and HLE benchmarks.
Transformers v5.0.0rc-0 launches with over 400 model architectures, more than 750,000 compatible checkpoints on the Hub, and 1.2 billion total installs, up from 20,000 installs per day at v4.
swift-huggingface is a new Swift package providing a complete client for the Hugging Face Hub with Python-compatible cache.
FLUX.2 is a new open image generation model from Black Forest Labs with a 32B parameter DiT and a single Mistral Small 3.1 text encoder. It supports multiple reference images and runs on 24GB GPUs with 4-bit quantization.
OVHcloud is now a supported Inference Provider on the Hugging Face Hub, offering serverless access to open-weight models like gpt-oss, Qwen3, DeepSeek R1, and Llama with pay-per-token pricing starting at €0.04 per million tokens.
Hugging Face TRL now officially integrates with RapidFire AI to accelerate fine-tuning and post-training experiments, with internal benchmarks showing approximately 16-24x higher experimentation throughput than sequential config comparison.
AnyLanguageModel is a Swift package that provides a drop-in replacement for Apple's Foundation Models framework, supporting local and remote LLM providers including MLX, llama.cpp, and cloud APIs.
Hugging Face's kernels library now supports building and sharing ROCm kernels, with a guide using the RadeonFlow GEMM kernel for MI300X.
IBM released Granite 4.0 Nano models from 350M to 1.5B parameters under Apache 2.0, trained on over 15T tokens with native support on vLLM, llama.cpp, and MLX.
Hugging Face improved streaming datasets with 100x fewer startup requests, 10x faster data resolution, and up to 2x faster streaming speed, outrunning local SSDs when training on 64xH100 with 256 workers.
huggingface_hub reached v1.0 with 113.5 million monthly downloads, powering access to over 2 million public models, 0.5 million public datasets, and 1 million public Spaces.
LeRobot v0.4.0 introduces Datasets v3.0 with chunked episodes and streaming, integrates PI0, PI0.5, and GR00T N1.5 policies, adds LIBERO and Meta-World simulation support, and launches a plugin system for third-party hardware.
Sentence Transformers is transitioning from the UKP Lab at TU Darmstadt to Hugging Face, with over 16,000 models on the Hub serving more than a million monthly unique users, and Tom Aarsen continuing as maintainer.
Every one of the 2.2M+ public model and dataset repositories on the Hugging Face Hub is being continuously scanned with VirusTotal, comparing file hashes against its threat-intelligence database without sharing raw file contents.
Hugging Face released Awesome Food Allergy Datasets, the first open collection of datasets on food allergies, to accelerate AI research.
Intel and Hugging Face benchmarked GPT OSS on Google C4 VMs with Intel Xeon 6, finding a 1.7x TCO improvement over C3 VMs.
OpenVINO.GenAI accelerates Qwen3-8B generation by about 1.4x on Intel Core Ultra using speculative decoding with a depth-pruned Qwen3-0.6B draft model.
RTEB is a new retrieval embedding benchmark in beta that combines open and private datasets across 20 languages and domains like law, healthcare, code, and finance, using NDCG@10 as the default metric.
Hugging Face converted dots.ocr, a 3B parameter model that surpasses Gemini 2.5 Pro on OmniDocBench, to Core ML and MLX, but the initial conversion is over 5GB and slow.
LeRobotDataset v3.0 packs multiple episodes in a single file using relational metadata, and natively supports streaming mode to process large datasets on the fly without downloading prohibitively large collections onto disk.
Scaleway is now a supported Inference Provider on the Hugging Face Hub, offering serverless access to models like gpt-oss, Qwen3, DeepSeek R1, and Gemma 3 with pay-per-token pricing starting at €0.20 per million tokens.
Public AI is now a supported Inference Provider on Hugging Face, offering free access to public and sovereign models like Apertus-70B, with usage free of charge at the time of writing.
The edition every row came from arrives in your inbox, free.
Join free