OpenAI pushed out GPT-Image 2.5, and the jump is noticeable. Faster generation, better fidelity, and something that's been a long time coming: consistent details across edits. You can now leave comments directly on an image to drive changes. On the text-to-image benchmarks, it took all three top spots. The previous number one was GPT-Image 2. So OpenAI essentially bumped itself. There is one friction point worth mentioning: at least one user hit repeated refusals on surgical image prompts, with the model flagging them as violent. That's the kind of guardrail tuning that tends to frustrate professionals working in medical contexts.
MiniMax shipped H3 and H3 Max. Text-to-video, image-to-video, up to 15 seconds at 2K resolution, and the ceiling on cost is two dollars. There's an uncensored option. Someone ran with that and generated a 44-minute origin story video for something called Dead Sun, rendered at 672p on a 3090. Forty-four minutes of AI-generated video for under two bucks is a different kind of number.
fal is running H3 Max with camera controls baked in. You set horizontal and vertical angles in degrees, and the model builds a navigable 3D scene from a single viewpoint in under three seconds. Geometry, position, and materials stay consistent across the views. It's running in real time on fal's platform right now.
On the open-source side, a model called LLaDA-Image landed from Chuyan Chen, Haoxing Chen, and collaborators. It scored 53.53 on Qwen-Image-Bench in English and 53.38 in Chinese, which puts it at the top of the open leaderboard. The architecture uses a DiT backbone that handles both text-to-image generation and editing in one unified model, and turbo distillation gets it down to two to four inference steps.
Daniel Keller built something quietly useful: oldmodels.org, an interactive archive of historic image generation models. He's planning to expand it into audio, video, and text. If you want to remember what Dall-E 1 actually looked like, or trace how far image gen has come, it's worth a visit.
Recraft added Google Gemini Flash 1.5 Omni into Recraft Studio. The specific capability there is combining up to five reference images to generate new video footage. Multi-reference video is still a genuinely hard problem, so plugging Gemini's multimodal handling into a production tool is worth watching.
Runway got wired into ChatGPT through Astra. The workflow someone demonstrated goes: one product photo in, style framing done in Runway, animation built in Blender, final render out through Seedance 2.5. That's a full production pipeline triggered from a chat interface.
fal also dropped episode four of their podcast. The guest is from Plot Party AI, and the conversation covers the microdrama industry, specifically the differences between how China and the US approach the format, and how the production structure differs between the two markets. It's on YouTube if that's your thing.
And fal teamed up with Trippy Pictures for a live event. Short films from a shared director brief, followed by a masterclass on planning, directing, and finishing work with generative tools. Attendees get an Agent Pass for fal Agent and a hundred dollars in fal credits.
That's your AI digest for 12 Sep 2026.