Skip to content
HomeNewsConceptsGuidesToolbox
AboutSubscribeUA
Subscribe

AI Today Brief

The daily AI-engineering brief. Built in public. EN · UA.

XTelegramLinkedInYouTubeRSS

Follow AI Today Brief on LinkedIn for daily AI-engineering updates and the weekly “5 shifts that changed how developers work” PDF.

Explore

NewsDigestsConceptsGuides

Company

SubscribeAdvertiseAbout

Legal

Editorial policyAI disclosurePrivacyTerms

© 2026 AI Today Brief. All rights reserved.

  1. Home/
  2. News/
  3. Multi-Agent Systems, Frontier Safeguards, and Benchmark Psychometrics

Wednesday, September 2, 2026

Multi-Agent Systems, Frontier Safeguards, and Benchmark Psychometrics

Show moreShow less+

Today covers Google Antigravity multi-agent engineering workflows, Anthropic's customer-controlled Enterprise Frontier Safeguards, and AllenAI's BenchMIRT framework for auditing evaluation benchmarks.

AI-assisted · editor-reviewed·How we use AI

In this issue · 8

  1. 1
    Token & cost optimization

    Gemini Adds Agentic Video Processing to Cut Token Usage by 88 Percent

    Google has launched agentic video understanding across Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite. The feature scans video segments dynamically, reducing token consumption by up to 88% and API costs by up to 66%.

    Open full story
  2. 2
    Tools & releases

    Ollama Shifts Cloud Tiers to Usage Credits with Agent IDE Support

    Ollama has introduced transparent per-token billing paired with monthly usage credits for its Pro, Max, and Team tiers. The service provides zero data retention and direct API integration with coding agents like Claude Code and Codex.

    Open full story
  3. 3
    Creative AI

    Google Launches Pics Image Editor with Object Segmentation and Text Translation

    Google Workspace has launched Google Pics, a web and integrated image editor based on the Nano Banana model. It offers targeted object isolation, in-image text translation, and multi-user collaborative editing.

    Open full story
  4. 4
    Token & cost optimization

    Hugging Face Drops 200+ WebGPU Kernels for Accelerated Browser Inference

    Hugging Face released @huggingface/kernels, an Apache-2.0 library providing 207 individual, versioned WebGPU compute kernels on the Hugging Face Hub. Benchmarks show a 2.57x geometric mean speedup over ONNX Runtime Web on Apple M4 hardware.

    Open full story
  5. 5
    Agents & MCP

    OpenAI Codex Desktop Embeds LibreOffice, Node.js, and Python in Local Runtime

    Inspection of OpenAI Codex desktop app cache revealed a 1.7GB primary runtime bundling complete installations of LibreOffice, Python, Node.js, Poppler, and git. Preconfigured agent skills direct the local AI agent on discovering and executing these binaries.

    Open full story
  6. 6
    Agents & MCP

    Google Antigravity and Gemini 3.7 Flash Solve Multi-Agent Engineering Workflows

    Google updated its Teamwork framework in Antigravity, pairing Gemini 3.7 Flash across autonomous agent swarms to tackle long-horizon technical problems. The system built a cycle-accurate RISC-V simulator from scratch and upstreamed optimizations to core open-source libraries.

    Open full story
  7. 7
    Tools & releases

    Anthropic Launches Enterprise Frontier Safeguards with Customer-Owned Data Retention

    Anthropic announced Enterprise Frontier Safeguards, allowing enterprises to maintain zero data retention on provider servers by storing monitoring logs in their own cloud storage. The framework will be supported across Claude Code, Claude Enterprise, Amazon Bedrock, Google Cloud, and Microsoft Azure Foundry.

    Open full story
  8. 8
    Models & research

    AllenAI BenchMIRT Uses Item Response Theory to Audit LLM Benchmarks

    The Allen Institute for AI released BenchMIRT, an open-source framework applying multidimensional Item Response Theory to evaluate what benchmark questions actually measure. Testing on 100 LLMs across 34,000 questions showed that retaining just 10% of high-discrimination questions preserves benchmark accuracy.

    Open full story

Concepts in this brief

GeminiClaude CodeCodex
Browse all news

Email digest

Get the morning AI brief

One email a day — the stories that matter for engineers, founders and tech leads. Human-edited, with links to primary sources.

  • ✓120+ sources scanned daily
  • ✓Edited by a human
  • ✓1 email per day
  • ✓EN + UA

By subscribing you agree to the privacy policy.