Posts tagged “coding-agents”
Mistral Patched Six Vibe Shell Bypasses. The Approval Prompt Saw the Wrong Action
Mistral Vibe 2.25.4 fixes six shell-permission CVEs. The deeper lesson is how coding agents must bind approvals to actual runtime effects.
Cognition’s SWE-2 Trains Devin to Spend Effort Where It Pays
Cognition’s SWE-2 turns Kimi K3 into a cost-aware Devin worker. Here’s how to read its benchmarks, effort levels and Terminal-Bench 4 gap.
Gemini 3.8 Flash Is Live. The Coding-Agent Cost War Just Got More Complicated
Gemini 3.8 Flash is live at 3.7 pricing. Its coding gains come from more agent work, changing how builders should test cost and reliability.
Alibaba’s Qoder Agent Desktop Puts the Harness Above the Model
Alibaba’s Qoder Agent Desktop turns coding agents into a task control plane. See why the harness, permissions and review layer matter.
OpenAI’s Planned Cursor Cutoff Makes Model Access an M&A Risk
OpenAI plans to end managed model access in Cursor after SpaceX’s acquisition. Here’s what BYOK preserves—and why the 5% figure misleads builders.
GLM-5.3 Started Where Pretraining Stopped
Z.ai says GLM-5.3 improves coding and cyber-agent performance without a new base model. Here is what builders can verify, test, and trust.
DeepSeek V4 Pro 0813: GA Benchmarks, Pricing, and API Guide
DeepSeek V4 Pro 0813 is live. See GA benchmarks, peak/off-peak API pricing, Codex setup, the Flash comparison, and open-weight details.
Grok 4.6’s 500K Context Window Has a 200K Toll Booth
xAI’s Grok 4.6 targets coding agents with 500K context and xhigh reasoning, but its 200K price cliff makes runtime design the deciding factor.
Meta’s Muse Code Bets Cheap Tokens Will Train a Better Agent
Meta’s Muse Code pairs Muse Spark 1.2 with persistent agents and cheap tokens. The deeper play is a coding-agent data and distribution flywheel.
VS Code 1.132 Gives Agent Host Sessions an Attention Layer
VS Code 1.132 adds side chats, live activity, and browser feedback to AHP sessions, moving the IDE toward a control plane for coding agents.
Poolside Laguna S 2.1 Makes Open-Weight Coding an Operations Problem
Laguna S 2.1 pairs cheap hosted inference with open weights. The tradeoff is a deployment contract that changes by checkpoint, runtime, and mode.
Qwen3.8-Max-Preview: Alibaba Subsidizes Agent Adoption Before the Proof Arrives
Qwen3.8-Max-Preview pairs 2.4T parameters and 1M context with a 0.01x Qoder coefficient, while reproducible benchmarks, weights, and production terms lag.