MiniMax just open-weighted H3, and it's worth paying attention to. The model takes text, images, video, and audio as a single context window. It outputs four to fifteen second clips at up to 2K resolution, 24 frames per second, with native stereo audio already baked into the H.264 file. No post-processing step to bolt audio on afterward. You get text-to-video, first-and-last-frame control, reference-based generation for subjects, motion, and voice, and in-place editing. All in one model.
ComfyUI workflows dropped on day zero. vLLM-Omni got an OpenAI-compatible slash-v1-slash-videos endpoint the same day. On Artificial Analysis's leaderboard it sits at number one for video editing, number two for text-to-video, number three for image-to-video. It's the new open-weights leader, pushing LTX-2.3 out of that spot.
Pricing is where it gets interesting. Two dollars and thirteen cents per second at 2K. That works out to seven dollars and eighty cents per minute. Seedance 2.0 at 1080p costs twenty-two dollars and forty-five cents per minute. Kling 3.0 at 1080p runs twenty dollars and sixteen cents. The commercial license covers organizations under twenty million in annual revenue with attribution required.
One more detail on H3. The Hailuo team ran the chalkboard "write hi" prompt on it. That's the prompt that broke Veo 2 and every other model that tried it. H3 passed.
Alibaba dropped Qwen3.8-Max. Two-point-four trillion parameters. Alongside it came open weights for the 27B variant. The benchmark they're leading with is autonomous coding that ran for more than ten days straight, self-evolving from an empty folder to a full production GitHub trace. They're also showing five hundred plus turns of chip design optimization and a 365-day e-commerce strategy generated end to end. Native multimodal with a continuous vision feedback loop. Pricing is two dollars per million input tokens, six dollars per million output, and twenty-five cents per million for implicit caching.
DeepSeek cut the price of one-million-token context to twenty-eight cents. That's not a typo. The argument that context costs were making self-hosted agent loops impractical is gone. If your agents were underperforming, the cost of context is no longer a viable explanation.
A few things moved on OpenRouter worth flagging. Deepseek-v4-flash-0731 got listed there first. The Cloudflare team used OpenRouter to test it before deciding whether to commit it to Workers AI. Modal listed Kimi K3 on OpenRouter as their first model on the platform. Poolside's Laguna S 2.1 hit two hundred fifty billion tokens per day on OpenRouter after a ten-times rate limit increase and a ten percent price cut on its dedicated one-million-context endpoint. Laguna S 2.1 is also live on Vercel AI Gateway now, with retry logic, timeout handling, circuit-breaking, caching, and per-request data residency records built in.
Ant Group's Ling-3.0-flash showed up on OpenRouter for free. It's a 124 billion parameter model with 5.1 billion active parameters. Someone ran a 963-line single-file SaaS landing page in one shot with every prompt constraint met.
Alibaba open-sourced a tool called Open Code Review. It's sitting at over seventeen thousand stars on GitHub under Apache-2.0. The benchmark they published shows higher precision than Claude Code on the same tasks. It drops exact-line comments, has built-in detection for null pointer exceptions, thread safety issues, XSS, and SQL injection, handles batching for large changesets, and has an OCR mode for full-file audits. It works with Claude Code, Codex, Cursor, and OpenCode.
Teknium from Nous Research updated the Hermes Agent. The changes were traced across two hundred fifty thousand conversations. Fewer turns to complete tasks, schema fixes, and better token efficiency for smaller and local models. The update includes Nvidia Nemo Relay integration. MiaAI lab tested DeepSeek v4 Flash inside the Hermes agent harness and found the output files outperformed every other harness they had tried.
One last thing. Someone ran Fable on high in Cursor. The agent told them that markdown files lived in the production database. The files were sitting in a public GitHub repo. Worth keeping in mind when you're delegating file system decisions to an agent.
That's your AI digest for 03 Aug 2026.