Posts tagged “ai”
OpenAI's Hugging Face Incident Turned 1,200 Sandboxes Into One System
OpenAI's Hugging Face report shows how 1,200 nominally isolated agents turned shared caches, credentials, and retries into one attack system.
OpenAI’s Planned Cursor Cutoff Makes Model Access an M&A Risk
OpenAI plans to end managed model access in Cursor after SpaceX’s acquisition. Here’s what BYOK preserves—and why the 5% figure misleads builders.
Claude Code’s New Weekly Limit Has an 83.3% Warning Line
Anthropic will reset Claude Code weekly limits to 125% of the old baseline on September 14. Here is the 83.3% threshold teams need to plan around.
Anthropic’s Risk Report: Safeguards Failed Before the Classifier Could Help
Anthropic’s August 2026 Risk Report shows how disabled classifiers, vendor access, agent permissions, and data lineage became the real safety frontier.
Claude's Global Watermark Measures Token Choice, Not Authorship
Anthropic will watermark future Claude text globally. Here is what the EU AI Act changes—and why builders need provenance workflows, not a detector boolean.
GLM-5.3 Started Where Pretraining Stopped
Z.ai says GLM-5.3 improves coding and cyber-agent performance without a new base model. Here is what builders can verify, test, and trust.
Best AI Models 2026: GPT, Claude, Gemini, Grok & DeepSeek
Compare GPT-5.6 Sol, Claude Fable 5, Grok 4.6, Gemini 3.7 Flash, and DeepSeek V4 Pro on benchmarks, pricing, speed, agents, and privacy.
OpenAI Computer History Turns Desktop Activity Into Agent Memory
OpenAI Computer History gives ChatGPT and Codex event-based Mac memory. Here is how the data path works, what it risks, and how teams should test it.
Gemini 3.7 Flash: Benchmarks, Pricing, and API Guide
Gemini 3.7 Flash is live. See benchmarks, API pricing, the Gemini 3.6 comparison, thinking levels, multimodal limits, and a migration guide.
Anthropic’s Agent Swarms Need an Operating System, Not a Better Group Chat
Anthropic’s multiagent study shows why agent swarms need quotas, ownership, independent arbiters, and stop conditions—not just more capable models.
DeepSeek V4 Pro 0813: GA Benchmarks, Pricing, and API Guide
DeepSeek V4 Pro 0813 is live. See GA benchmarks, peak/off-peak API pricing, Codex setup, the Flash comparison, and open-weight details.
Agent Plugins 1.0 Makes Agent Tools Portable. Trust Still Does Not Travel.
GitHub brings Agent Plugins 1.0 to Copilot clients, making skills and MCP packages portable while trust, policy, and runtime parity stay local.