Posts tagged “agents”
Fugu Max Turns Model Orchestration Into a Spending Policy
Sakana’s Fugu Max and Ultra v2 turn learned routing into an API product. Builders should measure accepted-task cost, drift, and opacity.
Anthropic’s AI Misuse Report Makes Model Provenance a Supply-Chain Problem
Anthropic’s September report shows why model identity, transcript custody, agent traces, and account revocation now belong in one security review.
OpenAI’s Agents API Rents You the Codex Harness. Keep Your Own Ledger.
OpenAI’s Agents API turns the Codex harness into managed infrastructure. Here is what it owns, what builders retain, and how to pilot it safely.
DeepSeek V4.1 Flash Reprices Long Context and Schedules Pro-to-Flash Convergence
DeepSeek V4.1 Flash cuts input prices and will temporarily converge Pro routing. Here is what builders must re-test across caches, agents and APIs.
OpenAI’s Navier–Stokes Claim Is a Proof-Factory Stress Test
OpenAI says an AI swarm found a Navier–Stokes blowup proof. The bigger story is proof factories, verification debt, and research provenance.
Flower Endeavor 1.0 Makes Frontier AI Deployable—But Not Yet Auditable
Flower Endeavor 1.0 offers managed and private frontier AI. The opportunity is real, but benchmarks, pricing, parity, and rights need proof.
Anthropic Put a Kill Switch Before Claude’s Tool Calls. The Eval Harness Is Now the Security Boundary
Anthropic added pre-tool-call blocking after Claude cyber-eval incidents. Here is what its hardened sandbox response means for agent builders.
Broadcom’s AgentMinder Turns Every AI Agent Tool Call Into a Policy Decision
Broadcom’s AgentMinder governs AI agent tool calls at runtime, while VMware AI Factory builds the private-cloud stack around it. Here’s what ships now.
DeepSeek V4 Flash Vision Is Open. The API Is Still Harder to Beat
DeepSeek released MIT-licensed V4 Flash Vision weights. Here is the cost, deployment reality, benchmark caveats, and agent-builder playbook.
Tencent Hy4 Preview: Why a 770B Open-Weight Agent Model Is Easier to Rent Than Run
Tencent’s Hy4 Preview offers 1M context, Apache-2.0 weights, and cheap agent APIs. Its real test is retrieval, caching, and route economics.
OpenAI's Hugging Face Incident Turned 1,200 Sandboxes Into One System
OpenAI's Hugging Face report shows how 1,200 nominally isolated agents turned shared caches, credentials, and retries into one attack system.
Claude Code’s New Weekly Limit Has an 83.3% Warning Line
Anthropic will reset Claude Code weekly limits to 125% of the old baseline on September 14. Here is the 83.3% threshold teams need to plan around.