Pull to refresh
Logo
fal launches H3 Max, a post-trained video model that tops independent benchmarks

fal launches H3 Max, a post-trained video model that tops independent benchmarks

New Capabilities

Built on MiniMax's open-weights H3, fal's model renders a five-second clip in about three seconds.

Yesterday: fal officially launches H3 Max

Overview

Updated 47 minutes ago

fal released H3 Max on September 1, its first post-trained video model. Built from MiniMax's open-weights H3, it generates a five-second clip in about three seconds of wall time.

That combination is rare. Labs usually trade quality for speed, but fal redesigned the model and its serving stack together, keeping gains in prompt adherence and visual quality while cutting latency. H3 Max ranks first on Design Arena and Artificial Analysis image-to-video leaderboards and runs roughly 35 times faster than MiniMax's official H3 endpoint.

Why it matters

If fal's approach holds, production video generation gets dramatically faster and cheaper, opening real-time and high-volume uses.

Questions about this story

Free account needed to ask — your question is kept and asked for you right after sign-up. Answers are public.

No questions yet — be the first to ask.

Key Indicators

~3 seconds
Time to generate a 5-second video
H3 Max's claimed wall-clock speed for a five-second clip at 768p.
35x
Throughput vs. official MiniMax H3 endpoint
fal says H3 Max is roughly 35 times faster than the original model's official endpoint.
1,341
Design Arena image-to-video Elo rating
Ranked #1 on Design Arena's recent image-to-video leaderboard, ahead of MiniMax H3 and Seedance 2.5.
1,201
Artificial Analysis image-to-video-with-audio Elo
Ranked #1 across 2,177 samples, ahead of Veo 3.1, Kling 3.0 and Gemini Omni Flash.
$0.04/sec
Launch promo price at 768p
50% off through September 7, after which generation costs $0.08 per second.

Voices

Curated perspectives — historical figures and your fellow readers.

Ever wondered what historical figures would say about today's headlines?

Sign up to generate historical perspectives on this story.

Play

Exploring all sides of a story is often best achieved with Play.

Most of these play right now — no account needed. Sign up to save scores, keep a streak, and unlock Debate and Predict. Log in Sign Up
Predict 3 ways this could play out. Back the one you believe — contrarian picks score more when a scenario has a resolution date. Log in to play

People Involved

Organizations Involved

Timeline

August 2026 September 2026

2 events Latest: Yesterday
  1. fal officially launches H3 Max

    Latest Launch

    fal releases H3 Max with #1 rankings on Design Arena and Artificial Analysis, a 3-second render time for 5-second clips, and a week of 50% off pricing.

  2. fal announces H3 Max

    Product announcement

    fal's blog introduces H3 Max, a post-trained version of MiniMax H3 developed by fal Research.

Historical Context

3 moments from history that rhyme with this story — and how they unfolded.

March 2023

Stanford Alpaca and the Llama fine-tune wave (2023)

Stanford researchers fine-tuned Meta's open-weights Llama model on 52,000 self-instruct examples for under $600. The tuned model chased GPT-3.5's quality at a fraction of the cost.

Then

A wave of fine-tuned Llama variants appeared within weeks, reshaping how open models were built.

Now

It proved that post-training open weights can approach frontier capability cheaply, a pattern now repeating in video.

Why this matters now

H3 Max is the video-era Alpaca: fal took MiniMax's open weights, added post-training data, and pushed quality past the base model without building from scratch.

Late 2023

SDXL Turbo and the real-time image race (2023-2024)

Stability AI's SDXL Turbo and latent consistency models cut image generation from dozens of sampling steps to one or a few, enabling near-real-time output by co-optimizing the model with the sampling method.

Then

Real-time image generation became standard on consumer hardware.

Now

It established that co-designing the model and the inference path beats bolting optimizations onto a finished model.

Why this matters now

fal applied the same principle to video: it optimized H3 and the serving stack together so speed gains didn't destroy quality gains.

2023

Llama fine-tune hosting by inference providers (2023)

GPU cloud and inference platforms like Together AI and Replicate began hosting fine-tuned variants of open-weights Llama, selling community-tuned models alongside the originals.

Then

Fine-tuning moved from academic experiments to commercial products hosted by third parties.

Now

Hosting providers became gatekeepers of which open-model variants reached production workloads.

Why this matters now

fal is extending that hosting-plus-fine-tuning playbook from text to video, post-training MiniMax's weights and serving them as its own product.

Sources

(6)