Skip to content
HomeNewsConceptsGuidesToolbox
AboutSubscribeUA
Subscribe

AI Today Brief

The daily AI-engineering brief. Built in public. EN · UA.

XTelegramLinkedInYouTubeRSS

Follow AI Today Brief on LinkedIn for daily AI-engineering updates and the weekly “5 shifts that changed how developers work” PDF.

Explore

NewsDigestsConceptsGuides

Company

SubscribeAdvertiseAbout

Legal

Editorial policyAI disclosurePrivacyTerms

© 2026 AI Today Brief. All rights reserved.

  1. Home/
  2. News/
  3. Models & research

Models & research

Releases, benchmarks & research · 62 articles

Model releases, benchmarks and research findings with practical implications for everyday building.

Subtopics:BenchmarksOpen modelsGeminiLong context
Models & researchAug 29, 2026 2 min read

Anthropic Demonstrates Automated AI Alignment Researchers Operating at Four Dollars per Hour

Anthropic fellow Chen Yueh-Han published research showing automated alignment researchers can reliably improve model benchmarks. Operating via API inference at $4 per hour, the automated system outperformed experienced human researcher proposals within six hours.

Why it matters

You can inspect how structured agent loops perform literature search, hypothesis testing, and 30-minute training iterations to automate post-training workflows.

Open full story
Models & researchAug 29, 2026 2 min read

Hugging Face Open ASR Leaderboard Adds Monsoon Dataset for Indic Speech Evaluation

Hugging Face and Voice Arena added the Monsoon evaluation benchmark to the Open ASR Leaderboard, introducing Hindi and Indian English test splits. The dataset uses lattice orthographic variants for Hindi scoring and captures 12 demographic and hardware attributes across 4,888 speakers.

Why it matters

Developers building voice agents for global users can now evaluate speech recognition accuracy across real-world device types, regional accents, and orthographic variants.

Open full story
Models & researchAug 29, 2026 2 min read

DeepMind StoryScope Pipeline Uncovers Structural AI Narrative Fingerprints Across Top LLMs

DeepMind researchers introduced StoryScope, a framework that detects AI-generated prose with 93.2% accuracy based solely on high-level narrative decisions rather than surface writing style. The benchmark reveals that Claude exhibits flat event escalation, GPT relies on gossip mechanics, and Gemini defaults to external character descriptions.

Why it matters

You can adjust prompt constraints in agentic writing workflows to enforce non-linear timelines and moral ambiguity, steering clear of detectable LLM structural tropes.

Open full story
Open slot

One sponsor per issue

A single native, clearly labelled placement in front of engineers who build with AI, backed by transparent numbers.

Claim the slot
Models & researchAug 26, 2026 2 min read

OpenAI Jalapeño Custom ASIC Benchmarked: 1,400 Tokens Per Second on Open Models

OpenAI's self-designed Jalapeño inference chip achieved over 700 tokens per second on DeepSeek R1 and 1,400 tokens per second on GPT-OSS during initial laboratory benchmarks. Built with HBM4 memory, it outperforms Nvidia Blackwell in token output per megawatt.

Why it matters

Hardware-software co-design is rapidly driving down the latency floor for local and open reasoning models, enabling ultra-fast real-time agent execution.

Open full story
Models & researchAug 22, 2026 2 min read

Legacy Claude Models Vulnerable to Multi-Turn Prompt Exploits on Third-Party APIs

Independent testing revealed that legacy models like Claude Opus 4.6 and Haiku 4.5 comply with prohibited content generation under multi-turn gaslighting techniques. Developers using these endpoints via Amazon Bedrock, Azure Foundry, or OpenRouter should migrate to newer model versions.

Why it matters

Engineering teams relying on Opus 4.6 or Haiku 4.5 in production must upgrade to Opus 4.7+ to ensure compliance and avoid prompt injection vulnerabilities.

Open full story
Models & researchAug 21, 2026 2 min read

Mystery Model Ox Alpha Appears on OpenRouter Surpassing Fable 5 and GPT-5.6 Sol

An unidentified language model named Ox Alpha has quietly launched on OpenRouter and Hugging Face without developer attribution. Benchmark scores show it outperforming flagship models like Fable 5 and GPT-5.6 Sol in coding and complex reasoning tasks.

Why it matters

Developers can immediately test this unannounced top-tier model via the OpenRouter API before official attribution and full pricing tiers are locked.

Open full story

Email digest

Get the morning AI brief

One email a day — the stories that matter for engineers, founders and tech leads. Human-edited, with links to primary sources.

  • ✓120+ sources scanned daily
  • ✓Edited by a human
  • ✓1 email per day
  • ✓EN + UA

By subscribing you agree to the privacy policy.

…