Posts tagged “model-evaluation”
Gemini 3.8 Flash Is Live. The Coding-Agent Cost War Just Got More Complicated
Gemini 3.8 Flash is live at 3.7 pricing. Its coding gains come from more agent work, changing how builders should test cost and reliability.
OpenAI's Models Breached Hugging Face. The Benchmark Became the Attack Plan
OpenAI says its cyber eval models breached Hugging Face. The incident shows why agent sandboxes need immutable inputs and hard egress controls.