This page collects what the sources on blogs to watch have published recently and groups them by subject rather than by publisher. It is rebuilt every hour, so a post usually appears here within an hour of going up. Nothing on this page is required reading.
Grouping by subject is the point. A reader gives you 45 separate streams and leaves you to notice that four companies wrote about the same scheduling problem this week; this page puts those four posts under one heading. Topics come from the filter terms already listed on the blogs-to-watch page, so the vocabulary is the course's.
How a post is filed. Each post is scored by keyword against every topic, using its title and its summary, and it is filed under the topic it scores highest against. A term that names a subject on its own counts for more than one that merely co-occurs with it, and a match in the title counts for three times a match in the summary. Any second topic a post also matches is shown as a label beside it. The method is keyword matching rather than a model, which makes it predictable, cheap, and occasionally wrong.
You are reading one topic, evaluation and reliability, over the last 120 days. Show every topic.
Measuring whether an agent works, and keeping it working.
The most comprehensive document extraction benchmark: 14 systems scored on accuracy, completeness, grounding, and cost across 370 enterprise documents.
Learn what impacts OCR accuracy, how it’s measured, and how to improve performance in real-world document processing workflows.
Discover how OCR for invoices streamlines finance operations, automates data extraction, reduces errors, and speeds up accounts payable with LlamaParse.
Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems.
Demystifying evals for AI agents
Moving live voice agents from demo to production requires rigorous, automated testing to handle the unpredictability of real multi-turn conversations. ADK now provides native live evaluation, allowing developers to test graph-based…
Agent Platform's evaluation service is now generally available, providing developers with a unified engine to measure agent quality consistently across local development experiments and live production traffic. You can evaluate agents…
Stop autoresearch loops from “cheating” by enforcing strict evaluation, isolating experiments, and designing metrics that prevent shortcuts and false gains.
Deep dive into self-improving evaluators in LangSmith, motivated by the rise of LLM-as-a-Judge evaluators plus research on few-shot learning and aligning human preferences.
LangSmith's homepage is now organized into Observability, Evaluation, and Prompt Engineering. Learn why we organized the homepage like this. Plus, see our latest Resource Tags updates.
A recap of Sai Srirampur's POSETTE 2026 talk on storage performance in Postgres, with a local NVMe vs. EBS benchmark and the production setup behind it.
To satisfy the demands of enterprise-grade agentic AI applications, underlying vector databases often struggle to scale effectively as modern use cases can scale to billions of vectors. As a fully managed PostgreSQL-compatible database…
Reinforcement learning (RL) for large language models (LLMs) alternates between two phases: generation (rollout), where the current policy produces responses, and training, where those responses are used to update the policy. In verl,…
XGBoost (Extreme Gradient Boosting) is an open-source library that implements gradient-boosted decision trees, an ensemble method that builds an additive sequence of trees where each new tree is fit to the gradient of the loss left by…
Today, we are releasing OfficeQA Pro V2, a new benchmark designed to evaluate whether...
We ran 900 DeepSWE rollouts on DeepSeek-V4 Flash and GPT-5.6 Luna. Luna leads pass@1 by 14 points; DeepSeek delivers 4.8x the solves per dollar.
A benchmark on how a single analytical query behaves in Redpanda SQL as the dataset and the cluster grow.
We ran 452 DeepSWE rollouts on Kimi K3 and Claude Fable 5. Fable leads pass@1 by 1.4 points; Kimi K3 wins pass@4 and delivers 2.8x the solves per dollar.
We moved low-level I/O execution off application cores to dedicated networking cores using Seastar’s new asymmetric_io_uring backend. Explore the architecture design, trade-offs, and benchmark results.
Mike Shaver joins Tailscale to scale engineering quality.
How vLLM maintains production quality with extensive CI across diverse accelerators, nightly performance benchmark and accuracy evaluation, and a two-week release process.
Reliability numbers are easy to publish. We break down what 99%, 99.9%, and 99.99% uptime actually require, the failure domains each tier has to survive, and the questions to ask any inference provider before you commit.
See how Together AI is improving production GPU clusters with passive health checks, node repair, stronger Slurm reliability, OIDC, and startup scripts.
We're proud to announce that the Center for Internet Security (CIS) has published the CIS CockroachDB v25.x Benchmark.
We built an evaluation suite to assess model trustworthiness. Our results indicate that models developed from open-source models can be trusted, provided…
Compare Kimi K3 with GLM 5.2 and other leading LLMs using current benchmark results and deployment options, with weight and license status called out.
We used DSPy to improve LLM judges and optimize our chat experience, creating an evaluation-driven feedback loop that produced better outputs.
By Celina Amados At Netflix, our catalog metadata is crucial to our member experience, and a single corrupted data state can impact millions of viewers immediately. To protect streaming reliability, we built an automated data canary…
We shipped a lot at Build 2026: hosted agents, Toolboxes, Foundry IQ, Memory, Managed Compute, fine‑tuning, Frontier Tuning, and a new evaluation and optimization stack. Read as a feature list, it is a lot to hold in your head. So here…
SPEC’s CPU benchmark suite has been a long established industry standard, and is almost impossible to miss when reading through various publications.
Requirements and ticket volumes keep growing while engineering capacity hasn't kept up. Through deployments with RV Tech and Mercedes, we share how AI is…
The index covers 45 sources from the watchlist. 37 publish a feed and are read from it; the other 8 publish none, so their index pages are scraped and each new post's own page supplies the title and date its card omits.
Last run finished 28 Aug 2026 at 01:08 UTC. The index holds 711 posts, keeps them for 120 days, and shows 33 posts on this page.
2 sources failed on the last run. Netflix TechBlog (HTTP 429); Replit (HTTP 403)
The last run read 833 posts across every source and added 1 post.
4 sources on the watchlist are not aggregated here, so check them by hand.
A date shown as "first seen" is not a publication date. Some sources publish no date at all, so the page records when the post entered the index instead of guessing.