Wednesday, September 2, 2026
Today covers Google Antigravity multi-agent engineering workflows, Anthropic's customer-controlled Enterprise Frontier Safeguards, and AllenAI's BenchMIRT framework for auditing evaluation benchmarks.
In this issue · 8
Google has launched agentic video understanding across Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite. The feature scans video segments dynamically, reducing token consumption by up to 88% and API costs by up to 66%.
Ollama has introduced transparent per-token billing paired with monthly usage credits for its Pro, Max, and Team tiers. The service provides zero data retention and direct API integration with coding agents like Claude Code and Codex.
Google Workspace has launched Google Pics, a web and integrated image editor based on the Nano Banana model. It offers targeted object isolation, in-image text translation, and multi-user collaborative editing.
Hugging Face released @huggingface/kernels, an Apache-2.0 library providing 207 individual, versioned WebGPU compute kernels on the Hugging Face Hub. Benchmarks show a 2.57x geometric mean speedup over ONNX Runtime Web on Apple M4 hardware.
Inspection of OpenAI Codex desktop app cache revealed a 1.7GB primary runtime bundling complete installations of LibreOffice, Python, Node.js, Poppler, and git. Preconfigured agent skills direct the local AI agent on discovering and executing these binaries.
Google updated its Teamwork framework in Antigravity, pairing Gemini 3.7 Flash across autonomous agent swarms to tackle long-horizon technical problems. The system built a cycle-accurate RISC-V simulator from scratch and upstreamed optimizations to core open-source libraries.
Anthropic announced Enterprise Frontier Safeguards, allowing enterprises to maintain zero data retention on provider servers by storing monitoring logs in their own cloud storage. The framework will be supported across Claude Code, Claude Enterprise, Amazon Bedrock, Google Cloud, and Microsoft Azure Foundry.
The Allen Institute for AI released BenchMIRT, an open-source framework applying multidimensional Item Response Theory to evaluate what benchmark questions actually measure. Testing on 100 LLMs across 34,000 questions showed that retaining just 10% of high-discrimination questions preserves benchmark accuracy.
Email digest
One email a day — the stories that matter for engineers, founders and tech leads. Human-edited, with links to primary sources.
By subscribing you agree to the privacy policy.