Posts tagged “ai”
Anthropic’s Embedded Evaluator Pledge Needs More Than Access
Anthropic’s embedded-evaluator pledge could shift AI assurance from model cards to operational audits. Here is what builders should demand next.
OpenAI Agents Were Linked to RubyGems. The Confirmed Failure Is Permission Design
Reports link OpenAI agents to May’s RubyGems abuse. The confirmed lesson is to gate identities, registry writes, and downstream execution.
GPT-Rosalind Gets a Price, Not a Public-App License
OpenAI priced GPT-Rosalind for October 5, but its internal-use API and preview Workbench make access, evidence, and governance the real product.
Fugu Max Turns Model Orchestration Into a Spending Policy
Sakana’s Fugu Max and Ultra v2 turn learned routing into an API product. Builders should measure accepted-task cost, drift, and opacity.
Anthropic’s AI Misuse Report Makes Model Provenance a Supply-Chain Problem
Anthropic’s September report shows why model identity, transcript custody, agent traces, and account revocation now belong in one security review.
Cognition’s SWE-2 Trains Devin to Spend Effort Where It Pays
Cognition’s SWE-2 turns Kimi K3 into a cost-aware Devin worker. Here’s how to read its benchmarks, effort levels and Terminal-Bench 4 gap.
OpenAI’s Agents API Rents You the Codex Harness. Keep Your Own Ledger.
OpenAI’s Agents API turns the Codex harness into managed infrastructure. Here is what it owns, what builders retain, and how to pilot it safely.
DeepSeek V4.1 Flash Reprices Long Context and Schedules Pro-to-Flash Convergence
DeepSeek V4.1 Flash cuts input prices and will temporarily converge Pro routing. Here is what builders must re-test across caches, agents and APIs.
OpenAI’s Navier–Stokes Claim Is a Proof-Factory Stress Test
OpenAI says an AI swarm found a Navier–Stokes blowup proof. The bigger story is proof factories, verification debt, and research provenance.
DeepMind Precomputed 9 Billion DNA Variants. The Lookup Layer Is the Breakthrough
DeepMind precomputed 9 billion DNA variant effects. The opportunity is faster genomic triage; the trap is mistaking a rank for a diagnosis.
Mistral’s €3B Round Turns Open-Weight AI Into an Infrastructure Business
Mistral’s €3B Series D funds models, regional compute, and private deployment. Here’s what sovereign open-weight AI now demands from builders.
GPT-6 Astra Is Live at $10/$50. Treat It as an Escalation Model
OpenAI’s GPT-6 Astra is live at $10/$50 with staged access, a 272K pricing cliff, and agent gains that make routing and recovery essential.