Baseten CLI 1.1.0
Baseten CLI 1.1.0 adds model rename and volume sync, with sync supporting Hugging Face, S3, GCS, Azure, R2, or CoreWeave sources.
The Deploy Log | Platforms
39 rows in The Deploy Log name Hugging Face: 29 from Hugging Face and 10 from 9 other companies, InternLM, Baseten, Perplexity, Kyutai, Xiaomi MiMo, NVIDIA, Google Research, Mistral and OpenAI. Newest first, each linking the publisher's own page.
Baseten CLI 1.1.0 adds model rename and volume sync, with sync supporting Hugging Face, S3, GCS, Azure, R2, or CoreWeave sources.
Hugging Face Hub v2.1.0 adds Jobs retries, rescheduling, port exposure, ZeroGPU quota tracking, and a revamped Model Catalog.
Perplexity released pplx-decider-v1-27b, a decision model fine-tuned from Qwen3.8-27B, on Hugging Face.
AutoVerifier model weights, tokenizer, and evaluation code released on Hugging Face.
Open tabular foundation model now on Hugging Face in three sizes from 28M to 215M parameters.
Kyutai released a DL3DV fine-tuned OVIE checkpoint with 0.1B params on Hugging Face.
MiMo-V2.6-Flash-MOPD and MiMo-V2.6-Pro-MOPD released on Hugging Face with 1M token context and 309B total parameters.
InternLM releases Intern-Decision-4B, 2B, and 0.8B multimodal structured decision models on Hugging Face.
Hugging Face released @huggingface/kernels with 207 WebGPU kernels, 2.57x faster than ORT WebGPU by geometric mean across 809 matching test cases on an Apple M4 GPU.
NVIDIA agreed to acquire Hugging Face for $12,930,300,000, keeping the platform open and multi-cloud.
Baseten is now a supported Inference Provider on the Hugging Face Hub, launching support for conversational and text-generation tasks with open-weight LLMs such as Kimi K3, DeepSeek V4 Flash, and GLM-5.2.
Hugging Face and Cerebras demonstrate a real-time speech-to-speech pipeline pairing Gemma 4 31B on Cerebras with Nvidia's Parakeet for speech recognition and Alibaba's Qwen3TTS for text-to-speech, already powering more than 9,000 Reachy Mini robots.
TabFM, a zero-shot foundation model for tabular data, is now available on Hugging Face, GitHub, and within Google Cloud BigQuery.
PaddleOCR 3.5 brings OCR and document parsing tasks closer to the Hugging Face ecosystem, with supported models able to run with Transformers as an inference backend.
NVIDIA released Nemotron 3 Nano Omni, a 30B-A3B omni-modal model for documents, audio, video, and agentic computer use, with OCRBenchV2-En 65.8, MMLongBench-Doc 57.5, Video-MME 72.2, and VoiceBench 89.4, and checkpoints in BF16, FP8, and NVFP4 on Hugging Face.
DeepInfra is now a supported Inference Provider on Hugging Face Hub, offering over 100 models including DeepSeek V4, Kimi-K2.6, GLM-5.1, with PRO users getting $2 monthly Inference credits.
Hugging Face published a guide on using Transformers.js in a Chrome extension, with a demo powered by Gemma 4 E2B for local AI features.
Granite 4.0 3B Vision is available on Hugging Face under Apache 2.0 and leads on PubTablesV2 cropped at 92.1 and full-page at 79.3, OmniDocBench at 64.0, and TableVQA at 88.1.
GGML, creators of llama.cpp, are joining Hugging Face, with Georgi Gerganov and team dedicating 100% of their time maintaining llama.cpp with full autonomy and leadership on technical directions and the community.
Unsloth and Hugging Face Jobs enable fast LLM fine-tuning of LiquidAI/LFM2.5-1.2B-Instruct through coding agents like Claude Code and Codex, with Unsloth providing approximately 2x faster training and about 60% less VRAM usage compared to standard methods.
Hugging Face released an agent skill that teaches Codex and Claude to write production CUDA kernels, producing an RMSNorm kernel for Qwen3-8B with an average 1.94x speedup on H100.
Holo2-235B-A22B Preview achieves 78.5% on Screenspot-Pro in agent mode within 3 steps and 79.0% on OSWorld G, and is available on Hugging Face.
Voxtral Mini Transcribe V2 is available now via API at $0.003 per minute, and Voxtral Realtime is available via API at $0.006 per minute and as open weights on Hugging Face under Apache 2.0.
Hugging Face released Daggr, an open-source Python library for building AI workflows that connect Gradio apps, ML models, and custom functions, with automatic visual canvas.
CUGA, a configurable generalist agent, is now on Hugging Face Spaces, achieving #1 on AppWorld and top-tier on WebArena.
Codex can now run end-to-end ML experiments using Hugging Face Skills, including fine-tuning, evaluation, and report generation.
swift-huggingface is a new Swift package providing a complete client for the Hugging Face Hub with Python-compatible cache.
OVHcloud is now a supported Inference Provider on the Hugging Face Hub, offering serverless access to open-weight models like gpt-oss, Qwen3, DeepSeek R1, and Llama with pay-per-token pricing starting at €0.04 per million tokens.
Hugging Face TRL now officially integrates with RapidFire AI to accelerate fine-tuning and post-training experiments, with internal benchmarks showing approximately 16-24x higher experimentation throughput than sequential config comparison.
Hugging Face's kernels library now supports building and sharing ROCm kernels, with a guide using the RadeonFlow GEMM kernel for MI300X.
Hugging Face improved streaming datasets with 100x fewer startup requests, 10x faster data resolution, and up to 2x faster streaming speed, outrunning local SSDs when training on 64xH100 with 256 workers.
OpenAI released gpt-oss-safeguard-120b and 20b, open-weight reasoning models for safety classification under Apache 2.0, downloadable from Hugging Face.
Sentence Transformers is transitioning from the UKP Lab at TU Darmstadt to Hugging Face, with over 16,000 models on the Hub serving more than a million monthly unique users, and Tom Aarsen continuing as maintainer.
Every one of the 2.2M+ public model and dataset repositories on the Hugging Face Hub is being continuously scanned with VirusTotal, comparing file hashes against its threat-intelligence database without sharing raw file contents.
Hugging Face released Awesome Food Allergy Datasets, the first open collection of datasets on food allergies, to accelerate AI research.
Intel and Hugging Face benchmarked GPT OSS on Google C4 VMs with Intel Xeon 6, finding a 1.7x TCO improvement over C3 VMs.
Hugging Face converted dots.ocr, a 3B parameter model that surpasses Gemini 2.5 Pro on OmniDocBench, to Core ML and MLX, but the initial conversion is over 5GB and slow.
Scaleway is now a supported Inference Provider on the Hugging Face Hub, offering serverless access to models like gpt-oss, Qwen3, DeepSeek R1, and Gemma 3 with pay-per-token pricing starting at €0.20 per million tokens.
Public AI is now a supported Inference Provider on Hugging Face, offering free access to public and sovereign models like Apertus-70B, with usage free of charge at the time of writing.
The edition every row came from arrives in your inbox, free.
Join free