Posts tagged “agents”
Codex CLI 0.146.0 Moves the Agent Boundary Out of the Terminal
Codex CLI 0.146.0 adds Agent Plugins and remote Code Mode, turning portability, execution, network policy, and MCP state into runtime concerns.
MCP 2026-07-28 Deleted the Transport Session. Your Agent Still Needs One.
MCP 2026-07-28 removes protocol sessions, exposes agent calls to gateways, and moves durable state into explicit handles, tasks, and runtimes.
xAI Grok Build Workflows Make Orchestration the Product
xAI’s Grok Build Workflows save parallel-agent orchestration as code. Here’s what the 1,024-agent limit means and what builders should test.
Claude Opus 5 Keeps the Price. The Workload Changes Anyway.
Claude Opus 5 keeps Opus 4.8 pricing but changes effort, caching, fallback, tools, and agent fan-out. Here is the production migration plan.
Health in ChatGPT Moves the Safety Boundary Into Ordinary Conversations
OpenAI's Health in ChatGPT connects medical records and Apple Health. Its bigger bet is that sensitive context can safely follow users across chats.
OpenAI's Models Breached Hugging Face. The Benchmark Became the Attack Plan
OpenAI says its cyber eval models breached Hugging Face. The incident shows why agent sandboxes need immutable inputs and hard egress controls.
Gemini 3.6 Flash and 3.5 Flash-Lite: Fewer Knobs, More Managed Control
Gemini 3.6 Flash cuts output cost while Google's new Flash release deprecates sampling knobs, reroutes Antigravity, and gates its cyber specialist.
Sakana Fugu-Cyber Makes Verification the Cyber-AI Product
Sakana AI's gated Fugu-Cyber API pairs multi-agent orchestration with human verification. Learn how to read its benchmarks, costs, and controls.
OpenAI's Long-Horizon Agent Failures Make the Session the Security Boundary
OpenAI's long-horizon agent incidents show why tool permissions are not enough—and why sessions need live monitoring, pause, and commit gates.
AWS Bedrock AgentCore Makes the Agent Session a Cloud Resource
AWS’s June AgentCore GA wave and July scale-up turn agent sessions into managed, governed workloads—with costs and lock-in builders must measure.
UK AISI: Open-Weight AI Is Shrinking Cyber’s Preparation Window
UK AISI finds a 4–7 month open-weight cyber gap. Cheap retries, removable safeguards, and slow patching make the preparation window the sharper warning.
1Password for Claude Secures the Password—Not the Session
1Password for Claude keeps passwords outside the model, but post-login authority remains. Here is the architecture, risk boundary, and builder playbook.