This page collects what the sources on blogs to watch have published recently and groups them by subject rather than by publisher. It is rebuilt every hour, so a post usually appears here within an hour of going up. Nothing on this page is required reading.
Grouping by subject is the point. A reader gives you 45 separate streams and leaves you to notice that four companies wrote about the same scheduling problem this week; this page puts those four posts under one heading. Topics come from the filter terms already listed on the blogs-to-watch page, so the vocabulary is the course's.
How a post is filed. Each post is scored by keyword against every topic, using its title and its summary, and it is filed under the topic it scores highest against. A term that names a subject on its own counts for more than one that merely co-occurs with it, and a match in the title counts for three times a match in the summary. Any second topic a post also matches is shown as a label beside it. The method is keyword matching rather than a model, which makes it predictable, cheap, and occasionally wrong.
You are reading one topic, inference and serving, over the last 120 days. Show every topic.
Serving engines, request paths, and end-to-end latency and throughput.
Access gateways and serving gateways explained, and how they handle identity, tenancy, limits, and metering for teams using and serving AI models.
DeepSeek V4 Pro 0813 is a 1.7T-parameter open frontier model under the MIT license and is available for inference today on Baseten model APIs.
Baseten is proud to power the inference behind You.com's search and answer stack.
At Hot Chips 2026, Cerebras details CS-4 and Nexus, plus the CS-5 and CS-6 roadmap for faster, more efficient frontier AI inference.
Discover how faster AI inference helps cybersecurity teams improve threat detection, AI security, validation, and response without sacrificing speed.
Cerebras and Upstage bring ultra-fast AI inference to Korea, delivering up to 2,000 tokens per second for real-time enterprise AI applications.
Serving automatic speech recognition (ASR) models at scale is costly when each request uses only a fraction of a GPU. Learn how NVIDIA CUDA Multi-Process Service (MPS) with NVIDIA Triton Inference Server on Amazon EC2 GPU instances cuts…
How our new Rust driver load-balances DynamoDB-style requests across a ScyllaDB cluster, and how we extended Latte to measure its performance
Today, the Qwen team open-sourced Qwen3.8-Flash-Next, a multimodal MoE model and an early preview of the Qwen4 architecture. It plays the same role for Qwen4 that Qwen3-Next played for Qwen3.5. The Ga...
Jalapeño is a custom inference chip from OpenAI that delivers faster, more power-efficient AI inference, with higher throughput and lower latency for modern models.
With ScyllaDB, a small engineering team could focus on building their product instead of battling their databases.
Error parsing meta tag attribute “keywords”: No content.
We implement a native sharded weight transfer engine in vLLM utilizing Ray Direct Transport (RDT), achieving weight transfer for the Kimi K2 model in BF16 on 48 8xH100 nodes in 7.53s
IsoExec unifies numerical execution across SkyRL's vLLM and Megatron runtimes, reducing the average rollout-versus-training logprob difference below 1e-6 on Qwen3.5-35B-A3B with 25% overhead.
Nowadays, State-of-the-Art (SOTA) models are getting much bigger and reloading the model service after a crash is very expensive. Therefore, we are introducing the Weight Cache Daemon, a persistent GP...
Research: A shot-scraper-style JSON API on Bun 1.4's new Bun.WebView Today saw the long awaited release of Bun 1.4 , the first stable version since the infamous Rust rewrite a few months ago . Interestingly, the Rust rewrite was…
Distributed Layerwise Offload shards and streams DiT weights across devices, serving a measured 124 GB Cosmos3 model on 64 GB HBM and estimating a path toward 200B+ models.
Sizing the DSpark draft-verification budget from per-request confidence instead of verifying every drafted token, so one configuration holds the throughput/latency frontier from batch size 1 to 256.
Day-0 vLLM support for Qwen3.8-2.4T-A95B: a 2.4-trillion-parameter hybrid MoE model served out of the box, with FP8/BF16 checkpoints plus NVFP4 and MXFP4 quantized weights, and co-developed kernels on NVIDIA and AMD hardware.
Cloudflare for Government achieves FedRAMP Class D (High) Certified status. We also announce our commitment to pursue DoD IL4 authorization. Cloudflare brings world-class security, performance, and developer products to the public sector.
How vLLM serves NVIDIA Nemotron 3.5 Lightning with OpenAI-compatible APIs, speculative decoding, and BF16/NVFP4 checkpoints across NVIDIA GPUs and edge systems.
Decode Context Parallelism (DCP) in vLLM shards KV cache across GPUs by sequence dimension, enabling 3× higher throughput on long-context agentic workloads compared to standard tensor parallelism.
How vLLM reaches 25K total TPS/GPU on Qwen3.5-397B-A17B-NVFP4 with GB200 NVL72 disaggregated serving, Blackwell GDN kernels, HMA cache transfer, async scheduling fixes, and srt-slurm recipes.
Benchmarking disaggregated VLM serving of Kimi-VL-A3B-Instruct with 4 Intel Arc Pro B60 vision encoder and 1 NVIDIA H200 language model worker, delivering 2.4x-2.8x higher throughput and 69%-80% lower TTFT than collocated serving.
An overview of Arm CPU enablement and inference performance optimizations in vLLM.
vLLM delivers day-0 Kimi K3 serving with hybrid KDA prefix caching, DSpark speculative decoding, production-scale disaggregation, and optimized kernels across NVIDIA and AMD GPUs.
How we took GLM-5.2-NVFP4 from 40 ms to 17 ms mean TPOT on 24 B300 GPUs with vLLM: P/D disaggregation, MTP speculative decoding, Model Runner V2, and the SLA-first trade-offs behind the final configuration.
What vLLM AFD Plugin adds to the vLLM ecosystem: Attention–FFN disaggregation for MoE serving, GPU and Ascend NPU backends, connector-based execution, and graph and ubatching support.
Run open models in production with full control over performance, cost, and quality. Deploy in minutes, roll out safely, and scale to your SLOs.
A preview of production-scale Kimi K3 support in vLLM, including KDA-aware prefix caching, fused kernels, optimized MXFP4 MoE, multimodal integration, and initial NVIDIA and AMD paths.
vLLM Semantic Router is expanding from intelligent routing into a system for building, evaluating, and running Mixture-of-Models.
By AI Platform’s Model Runtime team and Inference team Introduction Most organizations consume LLMs through hosted APIs. Netflix went further — we run the full stack ourselves, from model deployment through inference, inside our…
vLLM brings day-0 support to TML Inkling, a 1T-parameter multimodal model, with MTP, long-context serving, parallelism, and up to 380 tokens per second per user on NVIDIA GB200 GPUs.
vLLM prefill paired with TileRT decode through vLLM V1's connector interface: a specialized, latency-optimized decode engine that coexists with native vLLM decode behind one shared serving layer, with zero changes to vLLM.
How AMD Quark trains, quantizes, and serves EAGLE3 speculative-decoding drafts with vLLM on AMD Instinct GPUs, delivering up to 2.00x throughput gains for Kimi-K2.5 and 1.79x for MiniMax-M2.5.
Provisioned Throughput gives you reserved inference capacity for frontier open models like MiniMax M3 and GLM-5.2. Token-based pricing, a 99% uptime SLA, and up to 90% lower cost than proprietary APIs. No GPU-hour math, no…
How HPC-Ops integrates Hopper-optimized attention and FP8 MoE backends into vLLM for Tencent Hunyuan Hy3, improving mixed-length decode, MoE latency, TTFT, and TPOT on NVIDIA H20.
How vLLM-Omni serves and optimizes Qwen3-Omni with staged Thinker-Talker-Code2Wav execution, batching, CUDA Graphs, async chunk, async output, replicas, hot-path cleanup, and perf validation.
By transitioning from separate summary and index files to a prefix tree, we optimized cache efficiency, reduced disk I/O, and reduced memory overhead
llm-d v0.8 graduates Flow Control, Batch Gateway, and multi-modal serving to production, extends beyond Kubernetes with RL/Slurm support, and aligns with upstream vLLM — sharpening the project's identity as an inference control plane…
LLM inference at SotA speeds and Modal quality, now available to everyone.
llm-d extends vLLM's Hybrid Memory Allocator across KV offloading to CPU and storage and KV-aware routing, making the offload connector HMA-aware - for 1.8–1.9x faster KV loads and about 115% higher throughput at high request rates with…
Benchmarking llm-d's prefix-cache-aware routing across anonymized GPU pools on the NxtGen sovereign cloud, showing how one routing layer improves throughput and TTFT across single-vendor and heterogeneous fleets.
A modified salting technique that cuts P99 write latency 22x for large blobs
How Together served MiniMax-M3 efficiently with KV-block-major sparse attention, paged MSA decode, optimized index scoring, and a Rust-based multimodal gateway.
If we've said it once, we've said it once per millisecond: never block the GPU.
The index covers 45 sources from the watchlist. 37 publish a feed and are read from it; the other 8 publish none, so their index pages are scraped and each new post's own page supplies the title and date its card omits.
Last run finished 28 Aug 2026 at 00:08 UTC. The index holds 710 posts, keeps them for 120 days, and shows 51 posts on this page.
2 sources failed on the last run. Netflix TechBlog (HTTP 429); Replit (HTTP 403)
The last run read 833 posts across every source and added nothing new.
4 sources on the watchlist are not aggregated here, so check them by hand.
A date shown as "first seen" is not a publication date. Some sources publish no date at all, so the page records when the post entered the index instead of guessing.