Posts tagged “models”
JoyAI-Video-Edit Hits 30 FPS. The Stream Still Runs on Five Clocks.
JD’s open-weight JoyAI-Video-Edit reports 30 FPS on one B200. Here’s what latency, stream state, and single-session serving mean for builders.
DeepSeek V4-Flash-0731 Makes the Harness Part of the Model
DeepSeek V4-Flash-0731 adds major agent gains, Codex-ready Responses API support, and open weights. Here is how builders should evaluate it.
MiniMax H3 Gives Media Agents a 2K Render Queue
MiniMax H3 brings 2K video, native audio, mixed references, and async API jobs. Here is the architecture, pricing, and open-weights catch.
LG’s K-EXAONE 2.0 Makes Open Weights a Cluster Procurement Decision
LG’s 750B K-EXAONE 2.0 is Apache 2.0, agent-ready, and built for eight H200s. Here is what its open-weight economics mean for builders.
Inkling-Small Puts Open Weights Back in the Agent Router
Thinking Machines’ Inkling-Small pairs 12B-active MoE compute with open weights. Here is where it fits—and where it should escalate.
Poolside Laguna S 2.1 Makes Open-Weight Coding an Operations Problem
Laguna S 2.1 pairs cheap hosted inference with open weights. The tradeoff is a deployment contract that changes by checkpoint, runtime, and mode.
Gemini 3.6 Flash and 3.5 Flash-Lite: Fewer Knobs, More Managed Control
Gemini 3.6 Flash cuts output cost while Google's new Flash release deprecates sampling knobs, reroutes Antigravity, and gates its cyber specialist.
Sakana Fugu-Cyber Makes Verification the Cyber-AI Product
Sakana AI's gated Fugu-Cyber API pairs multi-agent orchestration with human verification. Learn how to read its benchmarks, costs, and controls.
GPT-Realtime-2.1 Makes Voice Agents a Production API Bet
OpenAI's GPT-Realtime-2.1 release gives API voice agents a sharper cost, routing, and eval model across WebRTC, WebSocket, and SIP.
Google LiteRT.js Makes Browser AI a Runtime Choice
Google LiteRT.js brings fast .tflite inference to browsers with WebGPU, WebAssembly, and WebNN, reshaping private local AI apps.
GPT-5.6 GA Turns OpenAI's Agent Stack Into a Product
GPT-5.6 GA, ChatGPT Work, Codex desktop, and API agents show OpenAI turning frontier models into a governed work layer for teams.