Friday, September 25, 2026
Microsoft consolidates Copilot into agentic cloud workspaces with dedicated code sandboxes, while software engineers formalize review barriers against uncomposed agentic code accumulation.
In this issue · 8
Cursor has shared a prompt for improving the token efficiency of AI agent harnesses. One round of prompt trimming, tool offloading, and cache layout changes cut one production team's overall token cost by about 7% with no loss in quality. Offloading non-core tools cut tool-description tokens by 60%, and making Model Context Protocol schemas discoverable via grep or jq cut total session tokens by 46.9% in sessions that used them.
Whiteboard released an open-source desktop application that bridges visual software architecture and code review for agents like Claude Code and Codex. It includes a Rust AST diff viewer and an agent canvas Software Development Kit.
Anthropic published an open-source suite of reference workflow agents and Model Context Protocol data connectors. The repository provides ready-to-run slash commands and dual deployment for Claude Cowork and headless APIs.
Driftproof benchmarked Claude Opus 5.5 within 15 hours of release across three core agent skills from addyosmani/agent-skills. Without explicit SKILL.md guidance, Opus 5.5 regressed on conventional commit prefixes and review severity labels compared to Opus 5. Calling the new model in Claude Code also strictly requires upgrading to version 2.1.280 or newer.
A new blog post condenses lessons from building production-ready Model Context Protocol (MCP) servers into a pre-release checklist, moving beyond quick terminal prototypes. It focuses on per-call authorization verification, session lifecycle cleanup, and model-friendly error semantics.
Google Research unveiled Co-Director and CANVAS, a multi-agent framework orchestrating Gemini and Veo models to mitigate character and environment drift in minutes-long AI videos. It uses persistent visual memory and multi-armed bandit optimization.
WHO AFRO data teams deployed custom Claude skills to extract, diff, and summarize case metrics across decentralized PowerPoint decks in the DRC. The pipeline cuts reporting time from a full day to under an hour.
Audit logging in autonomous systems frequently captures successful tool runs while discarding rejected operations, masking malicious replays. Furthermore, lossy context summarization can drop single negation tokens, silently converting restrictive policies into permission escalations.
Email digest
One email a day — the stories that matter for engineers, founders and tech leads. Human-edited, with links to primary sources.
By subscribing you agree to the privacy policy.