deployedbyai

The Deploy Log | Models and APIs

vLLM transformers backend

The transformers vLLM backend now meets or beats native throughput on Qwen3 4B dense, 32B dense, and 235B-parameter FP8 MoE models, using torch.fx static analysis and ast source rewriting to apply inference-specific layer fusions at runtime.

Hugging Face | In Edition 42, Saturday 11 Jul 2026

SOURCE Hugging Face, Native-speed vLLM transformers modeling backend

Added to The Deploy Log Tuesday 15 Sep 2026, updated Thursday 17 Sep 2026. Read Edition 42, the edition that carried it.

The call on this one

What shipped is free. The call on it, what to do about it and the condition on that, opens with a signup: free, no card, every edition in full.

It becomes $50 a year, and signing up now keeps your first year free.