This page collects what the sources on blogs to watch have published recently and groups them by subject rather than by publisher. It is rebuilt every hour, so a post usually appears here within an hour of going up. Nothing on this page is required reading.
Grouping by subject is the point. A reader gives you 45 separate streams and leaves you to notice that four companies wrote about the same scheduling problem this week; this page puts those four posts under one heading. Topics come from the filter terms already listed on the blogs-to-watch page, so the vocabulary is the course's.
How a post is filed. Each post is scored by keyword against every topic, using its title and its summary, and it is filed under the topic it scores highest against. A term that names a subject on its own counts for more than one that merely co-occurs with it, and a match in the title counts for three times a match in the summary. Any second topic a post also matches is shown as a label beside it. The method is keyword matching rather than a model, which makes it predictable, cheap, and occasionally wrong.
You are reading one topic, coding agents, over the last 120 days. Show every topic.
Agents that read, write, and review code, and the benchmarks that grade them.
Explore Claude's breakthrough performance on SWE-Bench, demonstrating advanced software engineering capabilities and code generation accuracy. Learn about our technical evaluation methods.
We put Poolside’s new Laguna S 2.1 model to the test, tasking it with a repository-scale transformation of the open-source game Hypersomnia.
The AI SDK harness layer now supports Cursor through the official @ai-sdk/harness-cursor adapter. The harness layer lets your application run different coding agents through the same HarnessAgent interface, so you can switch agents…
Managing library updates can be tedious at times. Learn how the GitHub Copilot app can handle this type of repetitive task. The post GitHub Copilot app for Beginners: Automate Dependabot pull request triage appeared first on The GitHub…
Discover how loveholidays uses OpenAI Codex to make software development accessible across the business, helping teams turn ideas into products faster.
Qwen 3.8 Flash from Alibaba is now available on AI Gateway. It takes text and images as input, serves a context window of 1 million tokens, and can return up to 65k tokens in a response. Alibaba recommends it for coding, tool use, and…
Use the Admin plugin for ChatGPT Work and Codex to analyze workspace usage, manage members and permissions, adjust limits, and act on admin requests.
The key skill required to make productive use of coding agents is being able to confidently instruct them on how to make changes and then confidently verify that those changes have been applied in the correct way. Sometimes this…
With a fixed deadline and design resources committed elsewhere, Stampli used Codex and ChatGPT Work to compress weeks of launch production into days.
If you’re juggling multiple Copilot sessions, use the My work pane to track what's in flight, what's done, and what's next. The post GitHub Copilot app for Beginners: Managing your work appeared first on The GitHub Blog .
Asana used OpenAI Codex to replace an outdated testing system in two weeks, completing work expected to take five years for about $12K.
Congrats to the team!
OpenHands joins the NVIDIA-led Open Secure AI Alliance to advance open, inspectable infrastructure for building and running secure AI agents.
Claude Code vs Cursor compared on interface, autonomy, models, context, execution, and cost, plus where OpenHands fits.
LTM has partnered with Cognition to deploy Devin, the AI software engineer, across its global client base and cybersecurity practice serving over 260…
Say you forked the OpenHands app twelve months ago and never merged from upstream. You would now be 2,600 merged PRs behind, including 866 bug fixes you do not have.
Devin, built by Cognition, is an AI software engineer: it plans, writes, tests, and ships code semi-autonomously. With Outposts, Devin can now run its work in Modal sandboxes.
Anhang and Yun have gone deep on everything that keeps software running once it ships, and we are excited to bring their work on automations into Devin.
Someone's subsidizing your coding agent. Aperture shows whether it's you.
Cognition’s entire platform is now FedRAMP Class D (High) In-Process and listed on the FedRAMP Marketplace, giving engineering teams working in federal…
The Devin Security Vulnerability Remediation Program helps organizations clear their vulnerability backlog and set up continuous remediation.
Devin Security Swarm finds vulnerabilities across the codebase, validates exploitability at runtime, and ships remediation PRs.
Introducing Devin Fusion: a hybrid-model harness that keeps frontier-level coding intelligence while cutting costs with sidekick agents and dynamic…
Engineering leaders want to know how much value AI is actually providing. We built a system that measures the number of human-equivalent hours of Devin's…
Introducing the AI Productivity Guarantee for enterprise customers. If Devin delivers less engineering value than you’re paying for, Cognition will fund…
The next generation of Windsurf, built around Devin Cloud, the Agent Command Center, and a full IDE for when you need to jump into the code.
Devin now builds, runs, and tests natively in Windows VMs, bringing the full power of autonomous AI engineering to the world's most mature developer…
Devin can monitor for bugs, alerts, and incidents. When something breaks, Devin responds immediately, investigates with your tools, connects related…
Devin can now spin up an Android Virtual Device (AVD), enabling autonomous development for Android applications.
The index covers 45 sources from the watchlist. 37 publish a feed and are read from it; the other 8 publish none, so their index pages are scraped and each new post's own page supplies the title and date its card omits.
Last run finished 28 Aug 2026 at 01:08 UTC. The index holds 711 posts, keeps them for 120 days, and shows 29 posts on this page.
2 sources failed on the last run. Netflix TechBlog (HTTP 429); Replit (HTTP 403)
The last run read 833 posts across every source and added 1 post.
4 sources on the watchlist are not aggregated here, so check them by hand.
A date shown as "first seen" is not a publication date. Some sources publish no date at all, so the page records when the post entered the index instead of guessing.