The Deploy Log | Models and APIs
vLLM transformers backend
The transformers vLLM backend now meets or beats native throughput on Qwen3 4B dense, 32B dense, and 235B-parameter FP8 MoE models, using torch.fx static analysis and ast source rewriting to apply inference-specific layer fusions at runtime.
The call on this one
What shipped is free. The call on it, what to do about it and the condition on that, opens with a signup: free, no card, every edition in full.
It becomes $50 a year, and signing up now keeps your first year free.