Field map · Data as of August 19, 2026

The Generative Media Ecosystem — An Architectural Map

A working map of how the generative media stack is actually structured: which layers are consolidating, which are commoditizing, and where durable control points appear to be forming. Treat it as a set of evolving hypotheses grounded in shipped systems rather than a settled verdict. Compiled from six research passes across roughly 150 sources, weighted toward January–August 2026 developments; figures marked est. are third-party estimates, not company-reported.

$5.6B
AI-video funding, 2026 YTD — +43% vs all of 2025
~$475M
Kling run-rate, Q2 2026 (filed Aug 19) — the adoption benchmark: filed, not claimed
>10x
Video inference cost decline since 2024
9 of 10
Top video models that are Chinese (AA arena, Aug 2026)

§ 01

What Changed, 2023 → 2026: From Single-Turn Generation to Continuous Creative Systems

In 2023 generative media ran on a three-step loop: prompt → model → output. You typed, a diffusion model dreamed, and you got one artifact — impressive, hard to steer, disposable. By August 2026 the loop that matters is intent → plan → generate → evaluate → edit → compose → collaborate → publish, and nearly every important product decision in the industry is about owning more of it. Node canvases (Flora, Krea, ComfyUI, Figma Weave) turned generation into pipelines; storyboard and character-persistence systems (LTX Studio, Higgsfield’s Soul ID, Popcorn) turned pipelines into projects; and in 2026, creative agents (Adobe Firefly Assistant, fal Agent, FAUNA, Creatify Agent) began turning projects into delegated work.[11][17]

The competitive battleground moved with it. In 2023–24 the fight was raw model capability. In 2024–25 it was generation products. In 2025–26 it became creative workflows — and the 2026 working hypothesis, held simultaneously by incumbents, startups, and infrastructure companies, is that the end state is multimodal creative agents: systems of creation, not tools of generation. Raw capability is now the least defensible layer in the stack: video-model leadership now turns over in weeks, and the top eleven models sit within ~150 Elo of each other.[28] The interaction model shifted with it: pure text-to-video has become an onboarding feature, while production work runs on image-to-video, multi-reference chaining, and keyframe conditioning — the same reference shift this site’s own prompt dataset surfaced a year ago, now built into the models themselves.

Two events define the year. On March 24, 2026, OpenAI announced the shutdown of Sora — the best-known consumer video product in the West, dead in six months on inference costs estimated anywhere from $1M to $15M a day against roughly $2.1M of lifetime in-app revenue (both third-party estimates).[1] That same quarter, Kuaishou disclosed that Kling grew revenue 300% year-over-year, ~70–75% of it from outside China (as of Q1) — the best-grounded number in the industry, because it sits in a listed company’s (unaudited interim) filings.[3] The Q2 filing (August 19) shows the curve still climbing on a consistent quarterly-annualized basis — roughly $360M (Q1) to ~$475M (Q2), on over RMB 850M of quarterly revenue up more than 200% year-over-year.[47] Frontier quality with no distribution surface to amortize its serving costs did not survive, while good-enough quality inside an owned funnel became the industry’s filed-revenue benchmark — a difference of system architecture, not of model capability. If the 36Kr-lineage reporting on ByteDance holds — Seedance API revenue past RMB 1B a month by June, unaudited — the same architecture is working at ~3.5x that scale inside China.[53]

FactStandalone AI-video feeds haven't held

Sora shut down (app Apr 26, API Sep 24, 2026); Meta's Vibes limps at ~2M DAU (Nov 2025) with weak retention. AI video as a tool — inside CapCut, Shorts, ad platforms — is where the revenue actually is.

FactVideo quality converged; image held a frontier

~150 Elo covers the top 11 video models and 9 of the top 10 are Chinese. In image, GPT Image 2 still holds every #1 — though MAI-2.6's August debut cut the Arena lead from ~83 to 45 Elo — the one modality where frontier capability still differentiates.

FactGeneration went free at the point of distribution

Veo free in YouTube Shorts, Seedance in CapCut, Adobe's 12-month unlimited-generations promo, Amazon giving ad creative away, Apple shipping photorealistic generation in iOS 27. The standalone generation button is rapidly losing pricing power.

FactMusic flipped from lawsuits to licenses in nine months

WMG and BMG settled with Suno while UMG and WMG captured Udio as a label-controlled walled garden; UMG's suit against Suno is still live. Suno settled from strength at $5.4B, and licensed catalogs are now the durable advantage.

PatternEveryone shipped a creative agent in 2026

Adobe, fal, Flora, Krea, Creatify, Amazon — every layer of the stack converged on the same product. The durable advantage is shifting from model access to creative state: characters, brand constraints, project memory.

PatternOmni absorption is compressing the specialist tail

Native audio, lip-sync, video editing, and camera control folded into frontier models in 18 months. Specialists survive only behind hard workflow or real-time constraints: 3D rigging, fidelity upscaling, <500ms avatars.

PatternWorld models became the capital magnet — and left entertainment

Over $3B into World Labs, Decart, and Odyssey in 2026, with monetization pivoting from playable worlds to AV/robotics simulation. GenMedia technology is exiting media at the frontier.

ThesisThe market bifurcated by geography and business model

The West monetizes GenMedia as enterprise software and platform features; China monetizes it as direct consumer and creator revenue. What holds up under churn is distribution, creative state, and licenses — not model checkpoints.

§ 02

The Market Map

Eight categories, organized by job-to-be-done rather than model modality, holding only strategically meaningful companies. The foundation-model layer is mapped separately in §06, and who-owns-which-layers in §03.

FIG. 01The Generative Media Ecosystem Map, August 2026

Eight categories by job-to-be-done, architecturally meaningful companies only — with a one-line technical thesis (Durable / Fragile / Unproven) inline on every entry.

startup incumbent frontier lab momentum 25figures dated per entry · amber date = older than 90 days · Durable / Fragile / Unproven = our read on how each position holds up
General Creative Workspaces
One surface from ideation to publish
Canva
Leonardo-lineage models + licensed Veo 3 — multi-model under the hood, no user-facing picker; Affinity now free
~$4B+ ARR · 265M MAUas of Jan 2026
Durable265M users make models interchangeable inputs
Adobe Firefly
Own commercially-safe models + partner-model hub; AI-first ARR up 3x YoY
over $500M AI-first ARRas of Jun 2026[16]
Durablegoverned workflow over everyone’s models is working
Google Flow
Flow+Whisk+ImageFX unified Feb 2026, 140+ countries, bundled with AI subscriptions
1.5B+ creations (claimed)as of Feb 2026[18]
Durablebundled frontier gen rides a 1B-MAU assistant
Magnific (Freepik)
Every model under one roof; half of revenue from video
$230M ARRas of Apr 2026[4]
Durableintegration speed and SEO distribution beat model ownership
Krea
Real-time canvas, 60+ models; K2 open weights cracked the AA text-to-image top 10 (Jun 2026); enterprise logos incl. LEGO, Samsung, Microsoft — but no new capital since Apr 2025
$500M (Apr 2025) val · 30M+ usersas of Jun 2026[57]
Durablereal-time canvas is a real paradigm; state must follow
Figma Weave
Weavy acquired for over $200M; shareable AI workflows inside Figma
acq. >$200M (Figma) val · $4M seed (as Weavy) raised · Entrée Capital, Designer Fund, Founder Collectiveas of Oct 2025
DurableCommunity makes workflows a network effect
OpenArt
Shareable Recipes
$70M ARR · 8M MAU
UnprovenRecipes are early state; free platform bundles loom
Recraft
Design-grade image models, shifting from model shop to workspace
$42M raised · 4M+ users (claimed) · Accel, Khosla Ventures, Madronaas of May 2025
Unprovenmodel-shop-to-workspace pivot must outrun the frontier
Reve
Sutter Hill incubation, ex-Adobe research founders; layout-first "images as code"; #2 on AA image; the Series B was never press-announced
$1.9B val · ~$390M raised · Sutter Hill Ventures, Top Harvest Capital, Basis Set Venturesas of Nov 2025[59]
Unprovena $1.9B quiet bet on frontier image quality; demand evidence still absent
Lovart
"AI design agent" archetype; ex-ByteDance founder; no disclosed funding
~$30M annualized (claimed) · 10M+ users (claimed)as of Aug 2026
Unprovenagent-first design has no demand proof yet
Advertising & Commerce
Performance creative at auction speed
Meta (GEM / Advantage+)
Targeting fully automated ad creation by end-2026; Muse model (Jul 2026) set to replace Midjourney/BFL licensing — Advantage+ advertiser rollout still pending as of mid-Aug
$75B+ Advantage+ run-rate (co-claimed) · 4M+ advertisersas of Jul 2026
Durablethe ad auction is a closed loop no startup can enter
Google Asset Studio
Veo in Google Ads; ~70M Gemini-generated assets in Q4 2025 across AI Max + PMax
Durablecreative folds into media spend
Amazon Creative Agent
Free conversational ad agent (Feb 2026), monetized via media spend
Durablefree creative subsidized by retail media
Creatify
Creatify Agent (May 2026) trained on 15M+ ads; locked brand facts as constraints
$23M+ raised · $9M ARR (May 2025) · 1M+ marketers (claimed) · WndrCo, Kindred Ventures, NFDGas of Sep 2025
Unproventhe 15M-ad corpus is real; platform absorption risk is too
Typeface
Enterprise brand governance, Arc Agents
$1B (Jun 2023) val · $165M raised · Salesforce Ventures, Lightspeed, GVas of Jun 2023
Unprovenbrand governance sells; suite bundling squeezes
Jasper
900+ enterprise customers; GEO Agent (Jun 2026)
$1.5B (Oct 2022) val · $131M raised · Insight Partners, Coatue, Bessemeras of May 2025
Unprovenenterprise base real, differentiation thinning
Arcads
AI-actor performance ads
$25M raised · ~$15M ARR (est.)
Unprovenlean and niche; AI-actor ads commoditize fast
Photoroom
Product imagery leader; squeezed by free platform tools
$500M (Mar 2024) val · $64M raised · ~$150M ARR (est.) · Balderton, Y Combinator, Aglaé Venturesas of Jan 2026
Fragilefree platform tools are eating product imagery
Pencil (Brandtech)
Agency-embedded enterprise creative gen
acq. by Brandtech (Jun 2023) val · 35K+ teams (claimed)as of Aug 2026
Fragileagency-embedded delivery limits how far the product scales
Smartly
Synapse orchestration layer (Jun 2026) — adtech incumbents absorbing the agent pattern
Providence Equity majority (2019) val · ~$101–221M rev (est.) · 700+ brandsas of Aug 2026
Unprovenadtech incumbents absorb the agent pattern
Film / TV / Pro Video
Studio-grade production and post
Runway
Gen-4.5 + GWM world-model pivot; revenue undisclosed (est. $100–300M; trackers est. ~$300M annualized by late 2025)
$5.3B val · $315M Series E raisedas of Feb 2026[5]
DurableGWM opens a second act beyond per-generation media revenue
Luma
Ray3.2 HDR/EXR pro deliverables
$4B val · $900M raised · HUMAIN
Unprovenpro deliverables niche vs omni gravity; HUMAIN compute buys time
Moonvalley
Marey trained exclusively on licensed data — the clean-model studio wedge; absorbed into Reka AI (May 2026, all-share) for the physical-AI push
acq. by Reka AI (May 2026, all-share) val · $154M raised · General Catalyst, Khosla Ventures, CAAas of Jun 2026
Durableprovenance is what studio procurement actually buys
LTX Studio (Lightricks)
Open-sourced LTX-2 (Jan 2026); script-to-screen pipeline + own open models
$1.8B (Sep 2021) val · $335M raised · ~$250M ARR (Lightricks apps, 2025) · Insight Partners, Goldman Sachs Growth, Violaas of Dec 2025
Durablethe Linux-of-video position compounds
Adobe Premiere
Firefly in the timeline; Generative Extend on licensed data
Durablethe timeline is where pro video already lives
DaVinci Resolve
Local-first AI free tier — commoditizes assistive editing
Unprovenfree local AI commoditizes the assist layer
Autodesk Flow Studio
ex-Wonder Dynamics video→CG; $200M into World Labs
acq. by Autodesk (May 2024) valas of Aug 2025
Unprovenstrategic buyer more than standalone winner
Promise
AI-native studio; a16z, Google AI Futures Fund, Crossbeam
a16z, North Road, Google AI Futures Fundas of Jul 2026
UnprovenAI-native studio thesis unproven at feature length
Flawless
Consent-based visual dubbing for studios
~$33M raisedas of Nov 2021
Unprovenconsent-based dubbing is defensible but narrow
Deepdub
Production dubbing with voice-clone royalties
$26M raised · Insight Partnersas of Feb 2022
Unprovenroyalty model aligns talent; scale unproven
Social / UGC / Consumer
Short-form creation at meme speed
CapCut / Dreamina
Seedance 2.5 (public Jul 31) into a 300M+-MAU editor; Seedance API reportedly past RMB 1B/month (36Kr-lineage, unaudited)
300M+ MAU (a16z est. ~736M)as of Jun 2026[53]
Durablethe largest creation funnel on earth feeds its own model
YouTube Shorts + Veo
Free frontier video gen inside the feed — the biggest distribution event in GenMedia
200B+ daily Shorts viewsas of Apr 2026
Durablefree frontier video in the feed ends standalone apps
Grok Imagine
#1 on video arenas in early 2026; Video 1.5 now ~$4.20/min; Image 2.0 slipped to Arena #3 in image (MAI-2.6 took #2) but holds #2 in image-edit; X distribution
20M+ images/day (claimed, Aug 2025)as of Mar 2026
Unprovenarena wins and X distribution; retention unproven
Higgsfield
Soul ID character persistence is the accumulated state; revenue figures company-claimed, not audited
$5.4B val · $400M Series B raised · $700M annualized (claimed)as of Aug 2026[40]
Unprovenrevenue still company-claimed; Soul ID persistence is the real test
PixVerse
Real-time R1 model
over $2B val · $439M raised · 150M users (claimed)as of Jul 2026[25]
Unprovenhuge consumer reach; serving-cost structure opaque
Mirage (Captions)
Own short-form video models
$75M revenue-linked raisedas of Mar 2026
Unprovenrevenue-linked financing implies real recurring revenue
Viggle
Physics-aware character motion; meme-format engine
$19M raised · a16z, Google AI Futures Fundas of May 2025
Unprovenmeme engine with feature-absorption risk
Character.AI
Chat → feed → AvatarFX video arc
$1B (2023); ~$2.7B Google deal (2024) val · $193M raised · ~$30–50M ARR · 20M MAUas of Aug 2025
Unprovenengagement is real; monetization is not yet
Meta Vibes
~2M DAU (Nov 2025), weak retention; the surviving-but-limping AI feed
~2M DAU, decliningas of Nov 2025
FragileAI feeds need creators, not prompts
Music & Audio
Licensed sound at API speed
ElevenLabs
Voice → dubbing → agents → licensed music
$11B val · $600M ARRas of Jul 2026[6]
Durablethe cleanest scale-up in GenMedia; licensed-first compounds
Suno
Settled WMG (Nov 2025), licensed BMG (Aug 2026, never a plaintiff); UMG suit live, Sony ruling not before 2027
$5.4B val · $400M raised · ~$300M ARR (est.)as of Jun 2026[7]
DurableWMG/BMG licenses turned legal risk into catalog access; UMG looms
Udio
Absorbed into a UMG/WMG walled garden — the labels captured it
$10M seed raised · a16z, UnitedMastersas of Oct 2025
Fragilelabel-captured; roadmap and terms now set by UMG/WMG
Cartesia
Sonic 3.6 (Aug 18) leads both AA speech arenas; sub-100ms vendor-claimed TTFA
$191M raised · 50K+ business customers · Kleiner Perkins, Index, NVIDIAas of May 2026
Durablelatency is a constraint the frontier ignores
KLAY
First AI music co licensed by all three majors
~$10M raised · Magma Partnersas of Nov 2025
Unproventhree-major licensing; product still early
Descript
Underlord agentic editor is now the product center
~$550M (Nov 2022) val · $100M raised · $55M ARR (late 2024) · OpenAI Startup Fund, a16z, Redpointas of Dec 2024
Unprovenright agent direction, crowded editor field
Google Lyria
Music gen wired into Gemini Live API
Unprovenplatform feature, not a product
Stability Audio
Survival-mode audio pivot; UMG partnership post-Udio
~$1B (Jun 2024 recap) val · ~$225M+ raised · ~$50M rev (2024) · Coatue, Lightspeed, Greycroftas of Mar 2025
Fragilesurvival pivot in a licensed-catalog game
Gaming, 3D & World Models
Assets, characters, playable worlds
World Labs
Marble GA; mesh-native worlds
$1B round raised · Autodesk ($200M)as of Feb 2026[22]
Durablemesh-native worlds drop into real pipelines
Decart
Sub-40ms real-time; Anthropic acquisition reported near signing at ~$7B, mostly stock, chosen over a higher NVIDIA offer (Aug 16 — still unsigned)
~$4B val · $300M raisedas of Aug 2026[48]
Durablethe optimization stack is worth more than the demos
Google Genie
Project Genie shipped Jan 2026 to $250/mo Ultra subscribers, US-only
Durablefirst consumer world model; bundling does the rest
NVIDIA Cosmos
Cosmos 3 open omnimodel + Coalition (BFL, Runway, LTX) — Llama-izing world models
10M HF downloads (claimed)as of Jul 2026
Durablesupplying every builder beats depending on any one lab
Odyssey
Pivoting to simulation infrastructure
$1.45B val · $310M raised
Unprovensimulation pivot chases robotics budgets
Meshy
Committed capital runs ~50x ahead of revenue; 12x YoY growth claimed
$1.5B val · ~$400M Series B raised · ~$30M ARRas of Jul 2026[33]
Fragilecapital underwrites world-model optionality, not current usage
Tripo
Speed + topology leader in 3D assets
~$200M raised
Unprovenspeed lead in a commoditizing category
Rodin (Deemos)
Lowe’s 30k-item catalog at <$1/model — 3D pricing already commoditized
CNY 100Ms round (Jun 2026) raised · Cathay Capital, Lanchi Ventures, ByteDanceas of Jun 2026
Unprovensub-$1 pricing proves the commodity endgame
Inworld
NPC infra leader, diversifying into voice agents
>$500M (Aug 2023) val · ~$120M raised · Lightspeed, M12 (Microsoft), Samsung Nextas of Jul 2026
UnprovenNPC infra leader hedging into voice
Scenario
Style-locked game assets; little holds once frontier models ship the same control
$11M raised · Play Venturesas of Jan 2023
Fragilestyle-lock is becoming a frontier-model feature
Hidden Door
Licensed-IP interactive fiction
$9M raised · Northzone, Makers Fund, Betaworksas of Aug 2025
Unprovenlicensed-IP fiction, the games analog of music licensing
Ubisoft / EA (internal)
Ubisoft’s studio-wide generative pipeline (Ghostwriter, NEO NPCs); EA co-developing with Stability
Unproveninternal pipelines cut costs, not add revenue
Enterprise Video & Comms
Governed video for work
Synthesia
NRR over 140%, 90% of Fortune 100; expects $200M ARR during 2026
$4B val · $200M Series E raised · ~$150M ARRas of Jan 2026[9]
DurableNRR over 140% across 90% of the Fortune 100
HeyGen
Cash-flow break-even; burned only $25M of the $74M raised
~$74M total raised · $200M ARRas of Jun 2026[8]
Durablebreak-even on ~$25M burned; the app layer’s cost-discipline benchmark
Gamma
Profitable with ~50 people; MSFT/Google now ship native rivals
$2.1B val · over $100M ARRas of Nov 2025
Unprovenprofitable and lean; native rivals arrived
Gemini Notebook
ex-NotebookLM, 30M users; free doc→video caps the explainer category
30M+ users, 600K+ orgs (claimed)as of Jul 2026
Durablefree doc-to-video caps the category
Tavus
Real-time conversational avatars, <500ms end-to-end
~$250M (secondary est.) val · $64M raised · CRV, Sequoia, Scale Venture Partnersas of Aug 2026
Unprovenreal-time niche holds on latency; demand breadth unknown
Sync
Lip-sync API layer powering other platforms
~$5.5M seed raised · GV, Y Combinatoras of Aug 2026
Unproventhe API layer rivals quietly resell
Argil
Creator clone videos; notable European entrant
€4.9M raised · EQT Ventures, Seedcampas of Nov 2024
Unprovenmust outrun avatars-as-a-feature
Workflow, Agents & Orchestration
The connective tissue between apps and models
fal
Launched fal Agent Aug 2026; ~$8B round in talks since Mar, unclosed
$4.5B val · ~$400M annualizedas of Mar 2026[10]
Durablemedia-native inference compounds; agents add workflow state
ComfyUI
Workflows-as-JSON is the portable orchestration format
$500M val · $30M raised · 4M+ usersas of Apr 2026[13]
Durableworkflows-as-JSON is becoming the substrate
Flora
FAUNA agent wires node pipelines from a brief; Nike/Netflix/Pentagram run FAUNA, Lionsgate is a FLORA client
$42M Series A raisedas of Jan 2026[14]
UnprovenFAUNA plus marquee logos; needs enterprise state
Adobe Firefly Assistant
Creative agent across CC apps; embedded into ChatGPT and Claude (Jun 2026)
Durablethe agent already lives where work happens
Replicate (Cloudflare)
Acquired Nov 2025 — thin marketplaces get absorbed
acq. by Cloudflare (terms undisclosed) val · ~$58M pre-acq raised · 50K+ hosted models · a16z, Sequoia, NVIDIAas of Nov 2025[12]
Fragilethin marketplaces get absorbed; the deal proved it
Runware
1MW containerized inference pods; the EU champion
$50M Series A raised
UnprovenEU champion on owned hardware; scale gap vs fal
WaveSpeed
Singapore-based; fastest non-China access to ByteDance/Alibaba models
Unprovenfastest non-China shelf for Chinese models; fragile edge
Vercel AI Gateway
33 image + 32 video models (Aug 2026); gateways treat media as first-class now
Vercel $9.3B (parent) val · 200K+ teams (claimed)as of Aug 2026
Unprovengateways commoditize routing from above
Baseten / Modal / Together
General inference clouds serving media as catalog extension
Baseten $13B · Modal $4.65B · Together $8.3B valas of Jul 2026
Unprovenmedia as catalog extension, not focus
Notes, method & sources

Panel color groups the category; the dot beside each name marks company kind (violet startup, gray incumbent, pink frontier lab). ▲ marks membership in the Momentum 25 (§09). Figures are dated per entry — an amber date means the underlying number was more than 90 days old at publication.

Inclusion is editorial: companies appear only where they hold a strategically distinct position, so absence is not a judgment of quality. The foundation-model layer is mapped separately in §06.

§ 03

The Battle for the Surface: Model Ownership vs Distribution

TakeawayA model fused to owned distribution compounds data, cost, and default status; a frontier model without a surface has no loop to close.

Who owns which layers decides who keeps the margin when model quality converges. The 2x2 plots the two ownership axes that matter most — model ownership and distribution ownership — and the bars underneath give the full layer-by-layer detail; the incumbent and startup evidence follows from them.

FIG. 02Model ownership vs distribution ownership

Up-and-right compounds — a model fused to a billion-user surface closes its own feedback loop. The lower-right, frontier labs without a surface, has no loop to close.

Compounders — model + surface
Surface owners renting models
Labs without a funnel — the melting quadrant
The unowned middle
Google
ByteDance
Meta
Kling
OpenAI
xAI
Microsoft
Adobe
Canva
Magnific
Higgsfield
ElevenLabs
Midjourney
HeyGen
Runway
fal
Luma
Black Forest Labs
rents models
owns frontier models →
owns distribution →
Notes, method & sources

Positions are editorial judgments on a 0–100 scale, not measurements. Dot color follows the map legend (violet startup, gray incumbent, pink frontier lab). The washed quadrant marks the compounding position; Sora is the case study for what happens in the lower-right with no distribution surface to close the loop.[1]

FIG. 03Who owns which layers of the stack

Each color is a layer; an unbroken run of color is vertical integration — the compounding position. Quiet dots mark layers a company rents from someone else.

Company
Distribution
Application
Workflow
Model
Infra
Owns
Google
5/5
ByteDance
5/5
Adobe
4/5
Meta
3/5
OpenAI
3/5
Kuaishou / Kling
3/5
xAI
3/5
Microsoft
3/5
Canva
3/5
Runway
3/5
Lightricks / LTX
3/5
Magnific (Freepik)
3/5
ElevenLabs
2/5
Suno
2/5
Higgsfield
2/5
fal
2/5
NVIDIA
2/5
Black Forest Labs
1/5
Notes, method & sources

Layer ownership is an editorial judgment of where a company operates with strategic weight, not a product inventory — Adobe’s de-emphasized own models still count as a model layer; Vercel-style gateways don’t make everyone an infrastructure owner. Rows sort by layers owned.

Full-stack (5/5) demands frontier capital intensity — only Google and ByteDance sustain it. Deliberate single-layer specialists (Black Forest Labs licensing models, fal owning inference) trade ceiling for focus. The dangerous position is the unowned middle: an application renting models with no workflow state above and no cost advantage below.

Incumbent advantages bind in four places. Distribution and bundling: Google made video generation a feature of a 1B-MAU assistant and a roughly $8–200/month subscription ladder — after Sora’s exit it won the Western consumer field largely uncontested.[19] Ad-system data: Meta’s GEM models optimize creative against auction outcomes, a closed loop no startup can enter, now feeding its in-house Muse model and a stated goal of fully automated ad creation by end-2026.[31] Enterprise workflow and indemnification: Adobe’s AI-first ARR passed $500M growing 3x year-over-year — and notably, Adobe now monetizes other companies’ models through its surfaces.[16] Compute economics: Google’s TPUs and ByteDance’s scale run video inference at costs that killed Sora.

Incumbents also failed visibly: OpenAI exited consumer video; Meta’s Movie Gen never shipped and Vibes has no retention story; Microsoft has no video model (its MAI foray is image-only so far — though MAI-Image-2.6 debuted at Arena #2 in August[43]); Amazon’s Nova is an ads utility nobody picks on merit; Apple is two years behind on quality. xAI is the ambiguous case: Grok Imagine took both AA video arenas in late January, Video 1.5 now prices at ~$4.20 per minute (~86% below Sora), and it ships inside X — distribution plus a cheap in-house frontier model — yet it went paid-only in March, the Chinese wave has since pushed it down the video boards, and xAI discloses no usage or revenue.[52] The pattern: incumbency wins where an existing engine (ads, enterprise seats, OS distribution) absorbs generation as a feature — not where incumbents chase new consumer behavior.

Startups can still build $10B+ companies in five lanes: audio (ElevenLabs at $11B and $600M ARR proves a full-stack modality winner where incumbents under-invested); enterprise vertical video (Synthesia compounds at 140%+ NRR and HeyGen at break-even, beneath incumbent attention); cost-disciplined consumer video (Kling, PixVerse, Hailuo — the discipline Sora lacked); media-native infrastructure (fal — every new model widens the serving problem it is paid to solve); and open-weights and world models (BFL, Lightricks, World Labs — being the neutral standard as media and simulation converge).

§ 04

The Battle for State: Creative Agents & the Workflow Layer

TakeawayEvery layer of the stack shipped a creative agent in 2026; what accumulates is creative state. Demand-side proof that creators want delegation is still missing.

A genuine new layer formed between applications and foundation models, and 2026 is the year every player in the stack tried to claim it. The node-canvas cohort (Flora, Krea Nodes, Freepik Spaces, ComfyUI) made multi-model pipelines a first-class artifact; the defining pattern of 2026 is that each of them then shipped an agent that builds the workflow for you — Flora’s FAUNA, Krea’s Node Agent, Adobe’s Firefly Assistant orchestrating Photoshop-to-Premiere, Amazon’s free Creative Agent, and in August, fal Agent reaching up from the inference layer.[11][14]

What accumulates in this layer is not model access — every player rents the same shelf — but creative state: Higgsfield’s Soul ID carries a trained character identity across models and sessions; Creatify locks verified brand facts as generation constraints; fal Agent keeps persistent project memory; Figma turned Weave workflows into shareable community assets.[15] State means switching costs, and switching costs are what the model layer structurally lacks. The mechanism compounds: accumulated state raises first-pass success, fewer retries lower effective generation cost, better output feeds back into richer state — and each pass deepens the switching cost.

The acquisitions already register what buyers think accumulates here. Weavy raised ~$4M and sold to Figma for over $200M; Visual Electric’s team went to Perplexity and the product died; Leonardo disappeared into Canva. The layer’s most instructive shutdown argues the same case from the other side: Sora had the best-known model in the world and no workflow, no B2B motion, no state — and it’s gone.[1]

What no one in this layer has yet shown is demand-side proof. Every revenue figure here is earned by the surface underneath the agent — fal’s inference, Adobe’s suite, HeyGen’s avatar product — not by the agent itself: fal Agent is a week old, Adobe reports Firefly Assistant “traction” without disclosing usage, and FAUNA’s marquee logos come with no revenue attribution. The most telling signal may be Google, which shipped Flow Agent to every tier at I/O in May — free accounts included — while keeping Veo generation itself behind the paywall.[58] The player best positioned to charge for an agent chose to price it at zero and monetize the layer beneath it. Three months in, that is the honest read on the whole pattern: the agent is a funnel to generation spend, not yet a product anyone has demonstrated people will pay for.

Verdict: the workflow/agent layer is becoming the primary control point of GenMedia, but it is contested from both directions at once — Adobe, Google Flow, Canva, and Figma extend down into it from owned surfaces while fal and ComfyUI build up into it from the serving layer. The open question for independent workflow companies is whether 2026’s growth converts into accumulated enterprise state — project memory, brand constraints, reusable pipelines — before both fronts arrive. On current evidence the strongest positions belong to Adobe, Figma, and Google among incumbents; Higgsfield, Magnific, and Flora among startups; and fal as the infrastructure entrant working upward.

§ 05

The Battle Underneath: Infrastructure & Orchestration

TakeawayMedia serving is structurally different work from LLM inference — durable while the model zoo stays heterogeneous, compressed the day an open serving standard wins.

Media generation is not LLM inference with bigger outputs — it is structurally different work. A video job is a long-running batch process, not a token stream, which forces queueing, webhooks, retries, and preemption-tolerant scheduling. The model zoo is architecturally heterogeneous (DiT, autoregressive, GAN upscalers, TTS, 3D) with no shared serving standard — there is no dominant open serving standard for diffusion — vLLM-Omni (open-sourced Nov 2025) is the first credible contender[50] — which is why fal builds tracing compilers and Decart builds sub-40ms kernels by hand. Chaining those architectures in one pipeline spikes VRAM unpredictably, and a mid-render failure burns minutes of GPU time unless the stack does stateful checkpoint recovery — a failure mode token streaming simply doesn’t have. Caching differs in kind too: hot-swappable LoRA weights and reusable keyframe latents, not prefix caches. Intermediate assets are gigabyte-scale per job, making storage and egress a real serving-cost line. And evaluation is still blind human preference: nothing machine-scores temporal consistency, character permanence, or edit fidelity at scale.

Model proliferation created a real orchestration layer — but 2026 showed it splits into two fates. Media-native orchestration with hard engineering depth is durable and compounding: fal roughly doubled from ~$200M to ~$400M annualized revenue in months, raised at $4.5B in December, and was reportedly in talks at ~$8B by March.[10] Its drivers don’t reverse — a majority-Chinese video model supply that Western apps won’t integrate one-by-one, bursty GPU economics, and weekly model churn that makes single-vendor commitments irrational. Runware ($50M Series A, containerized 1MW inference pods) and WaveSpeed (fastest non-China access to Chinese models) are growing in its wake.

FIG. 04The model supply chain

Foundation models → aggregators → output surfaces. Band width marks how load-bearing a connection is; the heavy left-side bands are the point — the video shelf is majority-Chinese, and Western apps reach it through aggregators.

Foundation modelsAggregation & orchestrationWhere the output goesGemini Omni / VeoGoogleSeedance 2.5ByteDanceKling 3.0KuaishouHailuo H3MiniMaxFLUX.2 / 3Black Forest LabsWan 2.2 / 3.0AlibabaLTX-2.5Lightricks · openGPT Image 2OpenAIGoogle Flowfirst-party · Veo onlyfalinference + agentComfyUIworkflows-as-JSONKreareal-time canvasMagnificall-models workspaceHiggsfieldconsumer suite + Soul IDReplicateCloudflareVercel AI Gateway33 image + 32 video modelsAds & performanceauction-speed creativeSocial & UGCmeme-speed short formFilm & pro videostudio pipelinesEnterprise commsgoverned video at workE-commerce & 3Dcatalogs, product media
Notes, method & sources

Band color follows the source model; weights are editorial judgments of how load-bearing each integration is, not measured volume. Leverage sits at the two ends — frontier models and owned distribution — while the middle holds its position only by adding workflow state (fal Agent, ComfyUI JSON, Soul ID) on top of routing. Google Flow is the first-party exception: orchestration that only ever routes Veo, holding the column by owning both ends instead.[10][11][13]

Thin aggregation without that depth gets absorbed: Replicate — the #2 independent media marketplace — sold to Cloudflare.[12] Gateways (Vercel’s now lists 33 image and 32 video models) commoditize the unified-API surface from the side, and hyperscalers own regulated-enterprise workloads by default. Meanwhile ComfyUI’s JSON workflows are quietly becoming the portable orchestration format — the closest thing GenMedia has to Terraform.[13]

Acquirers have started registering these gaps. Cloudflare bought Replicate and rights-marketplace Human Native; Anthropic is reportedly closing in on acquiring Decart at ~$7B — advanced drafts exchanged by mid-August, mostly stock, with chosen over a higher NVIDIA offer — a frontier lab valuing a media inference-optimization stack at acquisition scale.[23][48] On the compliance side, EU AI Act Article 50 transparency obligations began enforcement August 2, 2026, mandating C2PA metadata plus imperceptible watermarking — 6,000+ organizations have adopted C2PA, with Midjourney the prominent holdout — while China’s labeling regime has been live since September 2025.[29][38] Demand for provenance now outruns the technology: metadata still doesn’t survive re-encoding and platform uploads.

What remains genuinely unsolved at this layer — a standard serving engine, automated evaluation, provenance that survives distribution, long-form continuity at viable unit cost, rights clearing — is cataloged with the rest of the open problems in the Saturated Zones & Open Problems section (§10).

Where the leverage sits: at the two ends of the pipeline rather than the middle, and the gradient looks structural rather than cyclical. Model owners set the marginal cost floor, while surfaces holding accumulated state — characters, brand constraints, project memory — set what users actually pay; a pure router between them competes on latency and price with no state of its own to defend. The most honest signal comes from fal itself, which launched an agent in August 2026 despite winning the orchestration layer — a working admission that raw routing, however well engineered, does not hold its margin structure once the serving problem is solved more than once.[11]

§ 06

The Foundation Model Landscape

TakeawayNo single model wins every workload, and each modality is converging toward a different economic structure; treat every ranking as dated the week it posts.

There is no single foundation-model market. Each modality is converging toward a different economic structure: video toward codec-like ubiquity — critical, everywhere, rarely paid for directly, monetized by whoever owns the surface it runs in; image holding an LLM-like frontier premium for now; audio behaving like creative software fused to content licensing; 3D defended by hard workflow constraints; and world models trading as simulation optionality.

The landscape has consolidated into three durable archetypes: distribution-owned omni models (Google, ByteDance, Kuaishou, xAI), independent pro-grade labs (Runway, BFL, Luma, ElevenLabs, Reve), and open or China-first price leaders (Alibaba, Tencent, MiniMax, Lightricks). A dated-as-of stamp matters more than any ranking: video leaderboard half-life ran one to two quarters through 2025 and is now compressing toward weeks.[28] Treat the arenas (Artificial Analysis, and Arena — formerly LMArena, rebranded January 2026[39]) with a second caveat: the literature shows models post inflated consistency scores on quasi-static scenes — motion magnitude trades off against temporal coherence — so professional buyers increasingly select on control surfaces (first/last-frame conditioning, motion masks, reference counts) rather than rank.

Video

The most contested modality. Only ~150 Elo separates #1 from #11 on the Artificial Analysis arena (Aug 19, 2026); 9 of the top 10 are Chinese. Native audio is now table stakes, and pure text-to-video has become an onboarding feature — production work runs on image-to-video, multi-reference chaining, and keyframe conditioning (Seedance 2.5 takes 50 reference inputs; Ray3.2 takes 16 keyframes). Leaderboard half-life is now measured in weeks — Gen-4.5 led in Dec 2025; Grok Imagine took both AA video arenas in late Jan at a fraction of rivals’ prices; Wan 3.0 debuted at #1 on AA text-to-video this week, pushing Gemini to #2; Veo 3.1 sits #9 (Arena) to #12 (Artificial Analysis); Seedance edged back ahead of Hailuo H3 on both image-to-video boards mid-August — the gap sits inside the error bars, and H3 remains the top open-weights entry.

Model
Developer
Weights
Capability & control
Pricing / adoption
Gemini Omni Flash
Google
closed
Still #1 on Arena text-to-video, but slipped to #2 on AA (Wan 3.0 debuted 10 Elo above it this week) and #3 on image-to-video — the lead is contested weekly; conversational editing without re-prompting
$0.10/sec · Gemini app, Flow, YouTube Shorts, Vertex
Hailuo H3
MiniMax
hybrid
#2 on Arena image-to-video (Seedance edged ahead mid-Aug, within error bars) and top open-weights entry on both boards; #3 on AA text-to-video, 11 Elo behind Gemini; 15s, 2K, native stereo (license excludes US/EU/UK/KR local deploy)
~1/3 of rivals (est.) · HK-listed; aggregator shelves everywhere
Seedance 2.5
ByteDance
closed
30s single-pass, 4K, 50 multimodal reference inputs; #2 on Arena image-to-video within error bars of #1, #1 on the new Video Edit board; Seedance 2.0 still holds AA image-to-video #1
$0.05–0.40/sec by tier · CapCut (400M+ MAU), Jimeng/Dreamina, Volcano API — reportedly over RMB 1B/month (36Kr-lineage, unaudited)
Kling 3.0 Omni
Kuaishou
closed
Native 4K/60fps, 15s, in-model lip-sync; image gen too; 3.0 Turbo speed tier (Jun 2026)
~$0.11–0.17/sec · Q2 revenue RMB 850M+, up over 200% YoY (filed Aug 19) — the filed-revenue leader, ~70–75% overseas (Q1); spun out at $18B post (Jul 2026)
Gen-4.5
Runway
closed
1-minute multi-shot, native audio, character consistency; GWM world-model track
~$0.12/sec (Gen-4.5 class) · Studio deals (Lionsgate, AMC); enterprise workflows
Veo 3.1
Google
closed
8s, up to 4K, native audio; slid to #9 (Arena) – #12 (AA) as Omni took over
$0.05–0.75/sec by tier · Google Ads, Shorts, API
Wan 3.0
Alibaba
hybrid
Debuted #1 on AA text-to-video (Aug 2026); 30s single-pass in public beta since Aug 6; open-weights line stops at Wan 2.2 — the flagship is closed
~$0.05/sec · Top open-video lineage in ComfyUI workflows
LTX-2.5
Lightricks
open
Native 4K + synced audio on one RTX 4090; open weights (free under $10M ARR)
Free <$10M ARR, licensed above · The open substrate for on-prem studio work
Ray3.2
Luma
closed
16-bit HDR, EXR export, 16 keyframes/clip — pro/VFX deliverables
n/a public · Hollywood pipeline; HUMAIN compute
Grok Imagine Video 1.5
xAI
closed
Took #1 on both AA video arenas in late Jan 2026; fast 5–30s clips at 720p; the spring Chinese wave has since pushed it out of the AA text-to-video top 10
~$4.20/min (86% below Sora 2) · Bundled into X/Grok apps, paid-only since Mar 2026; xAI API + aggregators
Sora 2
OpenAI
closed
Still strong; app dead Apr 26, API sunsets Sep 24, 2026
$0.10–0.50/sec until sunset · Exiting — the cycle’s cautionary tale

Image

The closest thing to a stable frontier: GPT Image 2 holds #1 on every arena (1368–1463 Elo across boards), but MAI-Image-2.6’s Aug 10 debut at Arena #2 (displacing Grok Image 2.0 to #3) cut the text-to-image lead to 45 Elo — a real capability lead, but no longer a widening one.

Model
Developer
Weights
Capability & control
Pricing / adoption
GPT Image 2
OpenAI
closed
#1 on every arena (1381 Arena T2I, 1463 image-edit, 1368 AA — lead narrowing); 2K native, text rendering, thinking mode
$0.03–0.08/image · ChatGPT + API; 53% of Vercel gateway image volume
MAI-Image-2.6
Microsoft
closed
Arena #2 at 1336 Elo (Aug 10); +79 Elo over 2.5 in one release; cut the GPT Image lead to 45 Elo
Copilot-bundled · Copilot/Bing; OpenAI-independence signal
Reve 2.1
Reve
closed
#2 on AA; layout-first (posters, typography, packaging); claims frontier quality on roughly a tenth of rivals’ compute
~$0.20/image API (Jul 2026) · Best-funded independent image lab — $350M Series B at $1.9B, never announced; on fal, Replicate, Krea; no disclosed revenue
Grok Image 2.0
xAI
closed
#2 on Arena image-edit (1439 Elo, behind only GPT Image 2); #3 Arena text-to-image since MAI-2.6’s debut
X subscription tiers · In-feed generation on X — 20M+ images/day claimed (Aug 2025); no public API
Nano Banana 2 (Gemini 3.1 Flash Image)
Google
closed
#3 AA; Imagen 4 Fast is the $0.02/image cost floor
$0.02–0.15/image · Gemini, Flow, Workspace, Vertex
FLUX.2 / FLUX 3
Black Forest Labs
hybrid
Dev (32B) is the open standard; Klein Apache-2.0; FLUX 3 goes omni — FLUX 3 Video GA Aug 5, #6 on Arena image-to-video and #2 on Arena text-to-video (not yet listed on AA)
~$0.03/MP API · ~$300M contract value: Meta ($140M), Adobe, Canva
Midjourney V8.2
Midjourney
closed
Aesthetics-led; still no API; video stuck at V1
Subscription only · ~$200–500M revenue est. (wide variance); Disney suit in discovery
HunyuanImage 3.0
Tencent
open
80B MoE — largest open image model; Instruct adds reasoning
Self-host · Open-ecosystem anchor
Seedream 4.5 / 5.0
ByteDance
closed
5.0 now publicly ranked — #8 on Arena T2I (Aug 2026)
Ark platform · Dreamina, Doubao, partner shelves

Audio, Music & Voice

Capability alone no longer differentiates here — what a model is licensed to train on and emit does. The unlicensed-training era ended commercially in a nine-month window (Oct 2025 – mid-2026).

Model
Developer
Weights
Capability & control
Pricing / adoption
Eleven v3 / Eleven Music
ElevenLabs
closed
Broadest suite: TTS, dubbing, SFX, agents, licensed music (Merlin/Kobalt)
~$0.10/1k chars; music $0.15/min · $600M ARR (Jul 2026); Spotify audiobooks; 41% of Fortune 500 claimed
Suno v5/v6
Suno
closed
V5/V5.5 still live; licensed V6 committed under WMG/BMG deals, not yet shipped
Subscription; paid downloads · 2M paid subs, 100M users (Feb 2026); UMG and Sony suits live, no fair-use ruling before 2027
Sonic 3.6
Cartesia
closed
#1 on both AA Speech Arenas (shipped Aug 18); sub-100ms vendor TTFA, ~166–190ms in independent tests
API · Voice-agent latency leader; AWS JumpStart
Udio (licensed)
Udio + UMG/WMG
closed
Walled-garden remix platform; creations can’t leave
Subscription · Label-captured; exited open generation
Lyria 3
Google
closed
Music gen in Gemini Live API
Bundled · Platform feature, not product
MiniMax Speech 2.6
MiniMax
closed
Expressive HD + <250ms Turbo, 40+ languages
Aggressive · Cost leader in voice APIs

3D

Workflow-dominated, not leaderboard-dominated. No omni model has absorbed rigging or retopology — the long tail here is durable, but unit pricing already commoditized (<$1/model in enterprise e-commerce).

Model
Developer
Weights
Capability & control
Pricing / adoption
Meshy-6
Meshy
closed
Most production-ready meshes + texturing; engine export
Subscription + API · ~$400M Series B @ $1.5B on ~$30–40M ARR (reports conflict)
Tripo H3.1 / P1.0
VAST
closed
Clean low-poly in ~2s; rigs bipeds and creatures
Subscription + API · ~$200M raised; 6.5M creators claimed (Mar 2026)
Rodin Gen-2.5
Deemos
closed
Sculpt-level detail, production topology controls
<$1/model at volume · Lowe’s 30k-item 2D→3D conversion
Hunyuan3D 2.5
Tencent
open
Best open 3D; near-proprietary fidelity
Self-host · Open-ecosystem anchor

World Models

The largest concentration of new capital in the stack (over $3B committed in 2026) — and the work is exiting entertainment for simulation infrastructure (AV, robotics). Media, gaming, and robotics requirements converge here.

Model
Developer
Weights
Capability & control
Pricing / adoption
Genie 3 (Project Genie)
Google DeepMind
closed
Real-time navigable worlds, 720p/24fps, minutes of consistency
$249.99/mo Ultra tier, US-only · First consumer world-model product (Jan 2026)
Marble
World Labs
closed
Worlds with exportable triangle + collider meshes; World API
Freemium + API · $1B round (Feb 2026), $200M from Autodesk
Oasis 3 / Lucy 2
Decart
closed
Sub-35ms real-time generation and live video editing
API · Anthropic acquisition reported near signing at ~$7B (still unsigned)
GWM-1
Runway
closed
Worlds / Avatars / Robotics SKUs atop Gen-4.5
Enterprise · Media lab bridging into robotics revenue
Cosmos 3
NVIDIA
open
Open omnimodel with action output for physical AI
Open · Cosmos Coalition: BFL, Runway, LTX, Skild
Odyssey-2
Odyssey
closed
Causal frame-streaming "interactive movie"
n/a public · $310M @ $1.45B; simulation pivot

Specialized Survivors

Omni models absorbed lip-sync, SFX, and editing as features. Specialists survive only where there is a hard workflow or real-time constraint the frontier models don’t touch.

Model
Developer
Weights
Capability & control
Pricing / adoption
Sync
Sync Labs
closed
Lip-sync API powering other platforms
API · Quality + developer leader
Tavus
Tavus
closed
Real-time conversational avatars, <500ms end-to-end
API · The real-time constraint niche
Viggle 2.5
Viggle
closed
Physics-aware motion transfer onto any character
Freemium · Huge short-form creator adoption
Topaz Bloom / Magnific
Topaz / Freepik
closed
Fidelity vs creative upscaling — frontier models still cap at 2–4K
Subscription · The standard two-upscaler stack

§ 07

Economics: Cost Structure & Adoption Signals

TakeawayThe durable revenue pools sit at the bottom of the stack and in owned weights; app-layer economics hinge on whether inference deflation accrues to margins or is competed away.

50–60%
Revenue kept after model-serving costs — apps on third-party models (Bessemer, Feb 2026)
2–5x
Effective cost vs list price once retries are counted
$0.02–0.75
Per-second video API price envelope, frontier to challenger
~$400M
fal annualized revenue — the orchestration proof point
FIG. 05The $100M+ ARR club

Annual recurring revenue in $M, sorted by midpoint. Fill saturation encodes evidence quality — from filed numbers down to claimed-and-unverified.

Higgsfield
$700M annualized (claimed)
ElevenLabs
$600M ARR (Jul 2026)
Adobe (AI-first)
over $500M AI-first ARR
Kling
~$475M run-rate (Q2 2026)
fal
~$400M annualized
Midjourney
$200–500M (est., wide variance)
Suno
~$300M ARR (est.)
Magnific
$230M ARR
HeyGen
$200M ARR
Runway
est. $100–300M
Synthesia
~$150M ARR
Gamma
over $100M ARR
OpenArt
$70M+ ARR
Meshy
~$30M ARR
auditedcompany-statedthird-party estimateclaimed, unverified
Notes, method & sources

Figures are as of the date on each entry in the map above. Canva (~$4B total ARR) is excluded — its revenue is not GenMedia-attributable and would break the scale. ByteDance’s Seedance is also off the chart: reportedly over RMB 1B a month (~$1.7B annualized) via 36Kr-lineage press[53], with no company-disclosed figure to plot. Lighter extensions mark estimate ranges (Runway, Midjourney). Evidence tiers: audited/filed figures (Kling, via Kuaishou’s interim filings[47]), company-stated, third-party estimates, and claimed (Higgsfield’s $700M annualized is company-claimed and unverified[40]).

The $100M+ ARR club as of August 2026, weighted by evidence quality: Higgsfield (claimed $700M annualized on the heels of its Aug 17 round — company figures, unverified), Kling (~$475M run-rate per the Aug 19 Q2 filing), ElevenLabs ($600M, company-stated Jul 2026), Adobe AI-first (over $500M, earnings), Canva (~$4B+ total ARR with AI as retention), fal (~$400M est.), Suno (~$300M est.), Runway (est. $100–300M with wide variance — trackers est. ~$300M annualized by late 2025), Magnific ($230M company-stated), HeyGen ($200M company-stated, near break-even), Synthesia (~$150M), Gamma ($100M+, profitable), Midjourney (~$200–500M est., wide variance).[3][8][16][40] Sector funding: AI video alone took $5.6B in 2026 year-to-date, 43% above all of 2025 — concentrated in model builders and world models, while thin-wrapper seed activity visibly cooled.[30] How that committed capital compares with demonstrated adoption, company by company, is charted in the closing section (Fig. 08).

Where does the money settle when an app calls someone else’s model? The durable pools sit at the bottom and at the optimization layer: GPU landlords (CoreWeave’s $21B Meta expansion), media-native inference (fal’s kernel spread), and apps that serve their own weights — the four strongest cost structures among GenMedia app companies (ElevenLabs, HeyGen, Synthesia, Midjourney) all own their models. Apps renting third-party video models send 40–50% of revenue back out as model serving cost and escape through credit-pricing breakage, retry reduction, and riding a cost curve that falls ~10x per 18 months while their credit prices fall slower.[36] That last point is the open variable rather than a settled tailwind: app-layer cost structures improve with every quarter of inference deflation only if credit prices keep falling slower than serving costs — and nothing in the 2026 data yet shows whether the deflation accrues to the apps or gets competed away as cheaper generations.

FIG. 06Where app margins settle as inference deflates

Three schematic paths for the revenue an app keeps after model-serving costs, 2026 → 2028. The driver is fixed — inference cost falls ~10x per 18 months — and the open variable is how fast competition passes it through to credit prices.

credit prices fall slower than costsprices fall in step with costsdeflation competed away in price
25%50%75%today: apps on third-party models keep 50–60% after serving costsCredit prices fall slower than costs→ keep-rate widens toward ~75%Prices fall in step with costs→ margin structure unchangedDeflation competed away in price→ compresses toward commodityAug 2026Aug 2027Aug 2028schematic — driver held fixed at ~10x inference-cost decline per 18 months
Notes, method & sources

An illustrative mechanism, not a forecast: only the starting band is measured — apps on third-party models keep 50–60% of revenue after serving costs (Bessemer, Feb 2026)[36] — and the three paths are editorial readings of one unresolved variable. The ~75% endpoint is the ceiling argued in §12; the bottom path is the consensus wrapper-compression case. Which path the market takes is the margin question the 2026 data does not yet answer.

Attractive economics: audio/voice, enterprise avatar video, media inference infrastructure, model-licensing-to-platforms (BFL’s ~$300M of contracts), and compliance/provenance tooling with regulatory forcing functions. Structurally difficult: consumer free-tier video (Sora’s shutdown), thin wrappers, frontier video labs without distribution, and licensing intermediaries with thin take rates. One threshold worth watching: if per-second video pricing breaks below ~$0.005, programmatic ad video at auction scale becomes economical — and the largest commercial use case moves from creative teams to ad servers.

§ 08

What Holds Up: Durability, Open Weights & Geography

TakeawayPositions anchored in accumulated state, enforceable rights, or physical capital hold under model churn; positions anchored in model capability alone do not.

Layer by layer: which forms of differentiation hold up under model churn, which architectural levers make them stick, and where the current advantage is likely to commoditize within a couple of years. The pattern that emerges is consistent with the rest of the map: durability comes from distribution, accumulated state, enforceable rights, hard workflow or real-time constraints, a proprietary cost advantage, or capital-intensive physical infrastructure — and model capability alone holds only while it is fused to one of them.

FIG. 07Durability and commoditization risk, layer by layer

Where durable advantages can still form — and which layers are already commodity. Low risk clusters where courts, capital intensity, or regulation enforce the position.

Layer
Current differentiation
Commodit. risk
Durability lever
Where leverage settles
Frontier video models
Eroding fast — ~150 Elo across the top 11; leadership turnover compressed from quarterly to weeks
high
Fusion with owned distribution and subsidized inference; otherwise none durable
Migrates to distribution owners; standalone labs pivot to pro niches or robotics
Image models
GPT Image 2 holds every #1, but MAI-2.6 cut the Arena lead to 45 Elo
medium
Reasoning-in-generation, text rendering; open FLUX.2 caps the price umbrella
OpenAI at the frontier; BFL via platform licensing; the cost-price spread thins below the top
Audio / voice / music
High — licensed catalogs, latency, enterprise trust
low
Label deals are legally enforceable; owned models keep serving costs low
The strongest cost structures in GenMedia — ElevenLabs, Suno, Synthesia-class
Thin generation apps
Minimal — same model shelf as everyone else
high
None; 40–50% of revenue goes back out as inference serving cost
Little hold on either side of the pipeline — the consolidations and shutdowns (Icon, Visual Electric) trace the failure mode
Creative workspaces & aggregators
Workflow depth, speed of model integration, enterprise governance
medium
Templates, brand kits, enterprise contracts, credit-pricing arbitrage
Strong today (Canva, Adobe, Magnific) but contested by free platform bundles
Creative agents & state
Emerging — everyone shipped v1 in 2026
medium
Accumulated creative state: characters, brand constraints, project memory
The prize; unproven demand-side, claimed from above and below simultaneously
Orchestration & inference
Kernel/compiler engineering, queueing, media-native DX
medium
Hand-built optimization (no dominant open standard yet — vLLM-Omni emerging); usage pricing
Real and compounding (fal), but thin aggregation gets absorbed (Replicate)
GPU / compute
Scale, contracts, energy access
low
Capital intensity + multi-year backlogs (CoreWeave’s $21B Meta deal)
Rent-setting position — a high share of every dollar of spend, at high capital intensity
Provenance / rights infra
Regulatory tailwind — EU Art. 50 live Aug 2026, China labeling since Sep 2025
low
Compliance mandates with fines attached; C2PA network effects
Small revenue today; strategic — and metadata still doesn’t survive re-encoding
Notes, method & sources

Risk is the likelihood that the layer’s current differentiation commoditizes within roughly 24 months. Assessments are editorial, synthesized from the leaderboard-turnover, pricing, and licensing evidence in §06, §07, and the open-weights record below.[28][36]

Open weights: what open is now for

The open-closed gap now differs sharply by modality. In video it nearly closed: MiniMax’s H3 put open weights at #3 overall — with the geopolitical caveat that its license excludes local deployment in the US, EU, UK, and Korea, a new "open for China and the rest-of-world" flavor. Truly permissive open video (LTX-2’s 4K-plus-audio on a single consumer GPU, Hunyuan, Wan ≤2.2) trails the frontier by a clear tier.[27] In image the gap is small — FLUX.2 Dev is the open standard — but watch the direction of travel: Alibaba, BFL, and MiniMax are all gating their newest tiers. Open weights are increasingly a trailing-edge distribution strategy, not a frontier strategy. In audio, weights are irrelevant — licensed catalogs are the durable advantage. In 3D, open (Hunyuan3D) is genuinely competitive. In world models, NVIDIA’s Cosmos 3 is the open anchor, deliberately arming the ecosystem the way Llama armed LLMs.[21]

Wan is the cleanest case study in what open weights are now for. Alibaba’s own figures put the series past 6.9M downloads by August 2025, and the 2.1/2.2 checkpoints remain the default base for ComfyUI video work and the dominant fine-tune target on Civitai — an installed base every Western aggregator resells.[61] But the open line quietly stopped there: every flagship since — 2.5 through 3.0 — ships API-only on Alibaba Cloud at per-second pricing, while Apache-licensed side models keep goodwill flowing to the ecosystem. Read as a funnel, the conversion looks complete — build the substrate open, sell the frontier closed — and the community has registered it: the loudest response to Wan 3.0’s #1 debut came from the open-source side that built on 2.2, treating the release as confirmation the open era is over.[62] The concentration underneath is easy to miss: counting the newer HappyHorse line from a second internal team, Alibaba holds five of the top eight slots on the AA text-to-video board. A workflow standardized on "open video" today is standardized on a line its owner has already stopped feeding.

The pricing implication holds across every modality open weights reach: they cap the price umbrella, which pushes closed labs toward distribution fusion, licensing, or robotics — exactly the pivots Runway, BFL, and Luma made this year.

Geography: who owns which revenue

China owns consumer GenMedia revenue and export. Kling is the global filed-revenue leader in video and ByteDance’s Seedance API reportedly runs at roughly 3.5x that scale, unaudited and almost entirely domestic — the figures are in §01[53]; MiniMax IPO’d in Hong Kong with a +109% debut — beating every US lab to public markets, though its prospectus is candid about how early the monetization is: US$79M of FY2025 company-wide revenue, with Hailuo an undisclosed slice of a $53.1M consumer bucket[64]; PixVerse raised $439M at a $2B+ valuation on 150M claimed registered users; ByteDance ships Seedance to emerging markets first through CapCut.[3][24][25] Alibaba and Tencent supply the open-weights substrate (Wan, Hunyuan) that runs half the world’s ComfyUI workflows. The constraint is trust: the Disney, Universal, and Warner suit against MiniMax survived its motion to dismiss in May 2026 and is heading into the merits,[63] and Western enterprise procurement mostly can’t adopt Chinese models — which bifurcates the market and protects Adobe/Runway/licensed-lane pricing in regulated segments.

The US owns platforms, enterprise monetization, and image. Google is the only player integrated from silicon to YouTube; OpenAI leads image; Adobe leads governed enterprise workflow. Europe owns durable verticals rather than platforms: audio (ElevenLabs, UK/Poland), open image (Black Forest Labs, Germany), enterprise video (Synthesia, UK), aggregation (Magnific, Spain; Runware, UK) — while the EU AI Act makes provenance a compliance advantage for whoever already has the machinery.[29] Israel punches above its weight in open video (Lightricks) and real-time inference (Decart). Regulation converging on mandatory provenance (§05) is a fixed compliance cost small consumer apps struggle to carry — and a tailwind for watermarking infrastructure.[38]

§ 09

Momentum 25

The companies with the strongest January–August 2026 evidence — product breakthroughs, filed or credibly reported revenue, funding at higher marks, enterprise wins, or strategic distribution. Where a figure is reported rather than filed — ByteDance at #1 — the entry says so. The top twelve are shown; ranks 13–25 expand below. Anti-momentum, for balance: OpenAI Sora (dead), Stability (survival mode), Getty–Shutterstock (merger terminated), Pika (quiet), Meta Vibes (no retention), Amazon Nova creative (no traction).

  1. 01
    ByteDance (Seedance / CapCut)model + funnel
    The biggest reported business in GenMedia: Seedance API revenue passed RMB 1B/month by June (~$1.7B annualized — 36Kr-lineage reporting, not a filing), over half of Volcano Engine’s RMB 15B MaaS target, ~95% of China’s short-drama industry. Seedance 2.0 holds #1 on AA image-to-video, 2.5 (public Jul 31) debuted #1 on Arena’s Video Edit board, and it all ships into a 300M+-MAU editor. The asterisks: monetization is China-domestic behind RMB 10M-minimum contracts, and the Hollywood deepfake fight (MPA demand, Disney/Paramount cease-and-desists) is unresolved.
  2. 02
    Kling (Kuaishou)video model + app
    The hardest evidence in GenMedia: Q2 revenue over RMB 850M, up over 200% YoY (filed Aug 19, 2026 — a ~$475M annualized run-rate), ~70–75% overseas as of Q1, spun out with a ~$3B round at $18B post (Jul 2026), HK IPO targeted 2027. Smaller than Seedance’s reported numbers — but this one sits in a listed company’s filings, and no rival matches the overseas mix.
  3. 03
    ElevenLabsaudio platform
    $600M ARR (Jul 2026), $11B Series D, Spotify distribution, licensed music expansion — the cleanest scale-up in all of GenMedia, and the largest company-stated figure that is directly GenMedia-attributable.
  4. 04
    Google (Gemini / Veo / Flow)full stack
    Gemini passed 1B MAU and Flow unified into a 140-country workspace — the widest distribution in the field, won largely by Sora’s forfeit. But Wan 3.0 just took the AA text-to-video #1 from Omni Flash, and none of the 1B-MAU scale converts to attributable GenMedia revenue.
  5. 05
    Adobeworkspace + agent
    AI-first ARR >$500M (+3x YoY); Firefly became a multi-model hub and its creative agent now lives inside ChatGPT and Claude — the incumbent that adapted.
  6. 06
    falinference / orchestration
    ~$400M annualized (doubled in months), Sequoia-led $4.5B with an ~$8B round in talks since March, and an Aug 2026 move up into agents — the tooling-layer winner.
  7. 07
    Runwayvideo lab
    $315M Series E at $5.3B with NVIDIA and Adobe as investors; Gen-4.5 plus the GWM world-model line gives it a second act beyond media.
  8. 08
    Sunomusic
    $400M at $5.4B raised mid-lawsuit; settled WMG and licensed BMG from a position of strength — licensing is turning its biggest legal risk into its most durable advantage, with UMG still litigating.
  9. 09
    MiniMaxmodel lab (public)
    HK IPO popped +109% (Jan 2026); Hailuo H3 put open-weights video at the frontier — briefly #1 on Arena image-to-video in early August, #2 within error bars since — the first public pure-play in GenMedia.
  10. 10
    Black Forest Labsimage + video models
    $300M at $3.25B; FLUX is the open image standard, ~$300M of licensing contracts (including Meta’s $140M) prove the sell-to-platforms model, and FLUX 3 Video went GA in August.
  11. 11
    Synthesiaenterprise video
    $200M at $4B, ~$150M ARR, NRR >140%, 90% of the Fortune 100 — governed enterprise video keeps compounding beneath the hype.
  12. 12
    HeyGenenterprise video
    $200M ARR near break-even on ~$74M raised — the sharpest cost-discipline datapoint in the application layer.
Show ranks 13–25 — the research-complete list
  1. 13
    World Labsworld models
    Marble went GA with a $1B round including $200M from Autodesk; mesh-native outputs make its worlds drop into real 3D pipelines.
  2. 14
    Decartreal-time inference
    Sub-40ms real-time generation, $300M at ~$4B, and an Anthropic acquisition reported near signing at ~$7B (mostly stock, over a higher NVIDIA bid) — external capital registering what a hand-built optimization stack is worth.
  3. 15
    Magnific (Freepik)aggregator workspace
    $230M ARR with zero frontier models of its own — proof that distribution plus integration speed out-monetizes model ownership at the app layer.
  4. 16
    Higgsfieldconsumer suite
    $400M Series B at $5.4B (Aug 17, DST-led) — 4x its valuation in eight months — now claiming $700M annualized (still company figures); Soul ID character persistence is the aggregator state-layer experiment to watch.
  5. 17
    PixVerseconsumer video
    $439M at >$2B (Jul 2026), 150M registered users (claimed), first real-time consumer video model — China’s consumer export machine at work.
  6. 18
    Lightricks / LTXopen video
    LTX-2 open-sourced in January (4K + audio on one consumer GPU) through LTX-2.5 in August — the "Linux of video" position, plus a Cosmos Coalition seat.
  7. 19
    Lumavideo lab
    $900M at $4B led by HUMAIN with 2GW of committed compute; Ray3.2’s HDR/EXR deliverables target the pro/VFX lane the omni models ignore.
  8. 20
    Midjourneyimage
    Still ~$200–500M revenue (est.) with zero funding and ~40 people — but no API, no C2PA, and Disney’s suit in discovery make it momentum with an asterisk.
  9. 21
    NVIDIA Cosmosopen world models
    Cosmos 3 (Jun 2026) plus the Coalition arms the whole ecosystem with open world models — the Llama play for physical AI, run by the compute monopolist.
  10. 22
    ComfyUIworkflow substrate
    $30M at $500M, 4M+ users, Comfy Cloud out of beta — its JSON workflows are becoming the portable orchestration format of the industry.
  11. 23
    xAI (Grok Imagine)model + X distribution
    Took #1 on both AA video arenas in late January and holds image podium spots (#2 image-edit, #3 text-to-image since MAI-2.6 arrived); Video 1.5 now prices at ~$4.20/min — 86% below Sora 2 — with own models bundled into X, but paid-only since March and no disclosed usage or revenue.
  12. 24
    Meta Museimage model
    Launched July 2026 into Meta AI, Stories, and WhatsApp at billion-user scale, with the Advantage+ advertiser rollout next; the "infinite creative" pipeline optimized against auction outcomes is a closed loop no startup can enter.
  13. 25
    Tencent Hunyuanopen ecosystem
    The most modality-diverse open lineage (80B image, video, 3D, world) — the open-weights supply chain under indie GenMedia tooling worldwide.
Honorable mentions
Real 2026 cases that missed on evidence — each entry says which.
Alibaba (Wan)open video models
Wan 3.0 took the AA text-to-video #1 in August, and the Wan lineage is (with Hunyuan) the open-weights substrate under half the world’s ComfyUI workflows — the strongest case for a 26th slot, held back only by zero disclosed revenue or funding events for the model line.
Kreareal-time canvas
The strongest pure-product year of anything off the list — K2 open weights cracked the AA text-to-image top 10, an open-sourced real-time video model, the first creative-facing node agent, 30M+ users. What’s missing is the money signal: no new capital since Apr 2025, no ARR disclosure, and a 7x traffic gap to Higgsfield in its own lane.
Mirage (Captions)creator video
$75M of revenue-linked financing implies real recurring revenue, and it renamed itself around its own short-form models — edged out once xAI’s leaderboard run demanded a slot.
Cartesiaspeech models
Sonic 3.6 tops both AA speech arenas (Aug 18) — but ElevenLabs owns the audio narrative, the enterprise logos, and the licensing story, so capability alone doesn’t move the audio hierarchy yet.
Microsoft (MAI-Image)image models
MAI-Image-2.6 launched at #2 on Arena (Aug 10) with Copilot and Bing as distribution — real capability arriving late, still a derivative GenMedia strategy rather than an owned lane.
Latent players
Positions strong enough to reshape the list, with nothing shipped or filed in 2026 that moves it yet.
OpenAIpost-Sora
GPT Image 2 still leads the image arena and ChatGPT remains the largest creative-adjacent surface in the West. Sora died of distribution economics, not capability — video re-entry is a unit-economics decision away, and the list would rearrange the week it happens.
AppleOS-level generation
Free photorealistic generation announced for iOS 27, shipping this fall — OS-level distribution that would reset consumer image economics. Announced is not shipped, so it stays latent by definition.
Anthropicacquirer-in-waiting
The reported ~$7B Decart acquisition was near signing as of Aug 16. If it closes, a frontier lab enters real-time generation by purchase — the largest single bet on the modality to date.
Udiolicensed music
Captured by the majors — UMG and WMG turned their lawsuit into ownership. A label-licensed music catalog product is sitting in inventory until relaunch; when it ships, it lands on Suno’s position from above.

§ 10

Saturated Zones & Open Problems

Saturated: consumer text-to-video apps (free platform bundles cap the ceiling), text-to-image workspaces, 3D asset generation (the two best-funded leaders will starve the long tail at sub-$1/model pricing), AI presentations (Gamma won just as Microsoft and Google shipped native equivalents), and ad-creative SaaS (crushed between platform giveaways and synthetic-UGC fatigue — Icon’s pivot from "AI Admaker" to "Human Admaker", followed by a reported March 2026 shutdown, is the era’s best tell).

Commoditizing: video model quality itself (the ~150-Elo pileup), assistive editing AI (free in DaVinci Resolve), product photography, standalone lip-sync and SFX (absorbed by omni models). Emerging control points: covered in §11 — distribution surfaces, the state layer, media-native inference, licensed data, world models.

White spaceAutomated media evaluation

Quality measurement is still blind human Elo. Agents can't self-evaluate, pipelines can't regression-test. Unsolved because taste resists metrics — but temporal consistency, character permanence, and edit fidelity are measurable. A trusted eval layer becomes the QA gate for every creative agent.

White spaceThe vLLM of diffusion

No open serving engine dominates the heterogeneous media-model zoo yet — vLLM-Omni (Nov 2025) is the first credible contender — so fal and Decart still hand-build kernels, and that gap is literally priced at ~$7B (the Decart talks). A winning open engine + hosted control plane would restructure the inference layer — and compress its pricing overnight.

White spaceRights, likeness & provenance clearing

Likeness detection exists (Loti, Vermillio); a rail that clears identity, style, and catalog rights at generation time does not — despite music proving rights holders will deal. Provenance is the same rail's flip side: Article 50 mandates watermark plus metadata, but re-encoding strips both and detection-at-consumption is unbuilt, so compliance demand exceeds technical capability with fines attached. Whoever builds the clearing-and-attestation layer collects a small percentage of an enormous base.

White spaceLong-form narrative generation

Retry-adjusted economics make minutes-long, character-consistent video 10–100x too expensive; the cost curve is solving seconds, not stories. Models generate inside isolated temporal windows with no global scene memory, so the likely winner is continuity middleware — converting rendered output into reusable 3D/keyframe state enforced across heterogeneous model APIs — combined with draft-then-upscale workflows and retry reduction.

White spacePortable creative memory

Characters, brand systems, and project state are locked inside each workspace (Soul ID in Higgsfield, Weave in Figma). A cross-platform asset/context layer — the creative equivalent of a password manager — doesn't exist, and whoever owns it owns switching costs across the whole map.

§ 11

Architectural Control Points, 2030

Derived from the research, not assumed: six places where ownership plausibly produces disproportionate leverage in 2030.

01

Distribution-owned generation surfaces

When quality converges, the default surface wins. Free generation inside YouTube Shorts, CapCut, Gemini, and Instagram decides what billions of people use without ever choosing a model, and subsidizes inference that kills standalone consumer apps.

Leads today
Google (Gemini 1B MAU, YouTube, Flow), ByteDance (CapCut/TikTok), Meta (ad system), xAI (Grok Imagine inside X), Apple (OS-level, arriving)
Why it holds
User bases measured in billions plus owned inference (TPUs, ByteDance scale) — customer-acquisition and serving-cost advantages no startup can match.
Absorption path
This IS the incumbents’ position; the question is whether they extend it from casual creation into professional work.
Commoditization path
Regulatory separation (EU, US-China), or creation shifting to agent interfaces that sit above any one surface.
Stakes
Largest in the map — consumer creation folds into existing attention/ads economics measured in tens of billions.
02

The creative agent & state layer

Whoever holds the project memory — characters, brand constraints, references, reusable workflows — owns switching costs the model layer can’t touch. In 2026 every stack layer shipped a creative agent because everyone believes the durable position lives here.

Leads today
Adobe (Firefly Assistant in ChatGPT/Claude), fal Agent, Higgsfield (Soul ID), Flora (FAUNA), Figma (Weave), Krea
Why it holds
Accumulated creative state = switching costs; workflow communities (Figma Community, ComfyUI JSON, OpenArt Recipes) add network effects.
Absorption path
High risk — Adobe, Google, Canva, and Figma are claiming it from above while fal and ComfyUI claim it from below. Independents must convert 2026 growth into enterprise state before the squeeze.
Commoditization path
If agents become thin shells over frontier models’ own memory, state migrates down to the model layer.
Stakes
Professional creative software — a workflow surface on the order of $50B in annual spend — re-platformed onto agents.
03

Media-native inference & optimization

There is no dominant open serving standard for diffusion yet — vLLM-Omni (Nov 2025) is the first credible contender. Serving 1,000+ heterogeneous models fast, cheap, and queued is still a hand-built kernel business, and video’s cost curve (down >10x since 2024) is set here.

Leads today
fal (~$400M annualized), Decart (sub-40ms real-time; Anthropic deal reported near signing at ~$7B), Cloudflare/Replicate, Runware
Why it holds
Compiler and kernel engineering compounds; every new model release widens the serving problem this layer is paid to solve; usage pricing scales with the whole category.
Absorption path
Partial — hyperscalers and frontier labs are buying in (Cloudflare/Replicate, the Decart talks) rather than out-building.
Commoditization path
An open-source standard serving engine for DiT models would compress the layer overnight — vLLM-Omni (Nov 2025) is the first credible attempt, not yet dominant.
Stakes
A single-digit-billions revenue pool growing with all media compute; strategic value far above the revenue.
04

Frontier omni models with owned distribution

Pure model quality has a half-life now measured in weeks to quarters, but a frontier model fused to a billion-user surface (Gemini Omni + YouTube, Seedance + CapCut) compounds data, cost, and default status simultaneously.

Leads today
Google, ByteDance, Kuaishou, xAI (Grok Imagine riding X); OpenAI in image post-Sora (a “Spud” video successor is reported); Microsoft entering via MAI + Copilot; MiniMax as the public-market pure-play
Why it holds
Capital intensity of frontier training + proprietary usage data + subsidized inference. Distribution is the differentiator, not the checkpoint.
Absorption path
This is incumbent territory already; independent labs (Runway, Luma, BFL) survive via pro niches, licensing, or robotics pivots.
Commoditization path
Open-weights frontier releases (H3-style) plus falling training costs erode the standalone model premium continuously.
Stakes
Winner-take-most per surface; monetizes as subscriptions, ads, and API — tens of billions but concentrated.
05

Licensed data & rights clearing

Music proved the sequence: lawsuits became licenses, and the licensed catalog became the durable advantage. Studio procurement now selects for provenance (Moonvalley, Adobe indemnification), and EU AI Act Article 50 enforcement (Aug 2, 2026) makes provenance infrastructure a compliance requirement.

Leads today
The majors (UMG/WMG capturing Udio), Suno post-WMG-settlement, ElevenLabs (licensed-first), Adobe, Moonvalley, C2PA ecosystem; Loti/Vermillio in likeness
Why it holds
Exclusive catalogs and consent frameworks are legally enforceable advantages — the only kind courts actively strengthen.
Absorption path
Rights holders themselves are the incumbents here; tech companies become licensees.
Commoditization path
Blanket statutory licensing would flatten the advantage; conversely, a Sony v. Suno fair-use win could weaken it.
Stakes
Sits in the request path of all commercial generation — a small percentage of an enormous base; likeness rights alone sized at ~$10B.
06

World models & the simulation bridge

The over-$3B that flowed into world models in 2026 rests on the hypothesis that generative media and robotics/AV simulation are one technology. Whoever owns the world model owns both the next entertainment format and the training ground for physical AI.

Leads today
Google (Genie 3), World Labs (Marble + Autodesk), Decart, NVIDIA (Cosmos 3 open standard + Coalition), Runway (GWM)
Why it holds
Frontier research talent and compute; NVIDIA’s open Cosmos strategy is deliberately commoditizing the closed labs’ edge.
Absorption path
Active — Autodesk bought in, Anthropic is reportedly bidding, NVIDIA is arming everyone. Expect the category to be absorbed into larger platforms by 2028.
Commoditization path
If Cosmos-style open models reach parity, value shifts to simulation data and integration, not the model.
Stakes
Speculative but potentially the largest: entertainment plus the simulation layer of the entire physical-AI economy.

§ 12

The State of Generative Media — August 2026

The wrong question. The industry entered 2026 still organized around a simple question — whose model makes the best pixels? — and exits August 2026 having concluded it was the wrong one. The best pixels changed hands five times in twelve months. What didn’t change hands: YouTube’s two billion users, Adobe’s enterprise contracts, CapCut’s creation funnel, the labels’ catalogs.

What is commoditizing: video model quality (a ~150-Elo pileup with sub-quarterly leadership turnover), raw generation interfaces, 3D asset pricing, assistive editing, standalone specialist models. What remains scarce: distribution measured in billions of users; creative state that accumulates switching costs; licensed catalogs and consent frameworks; media-native inference engineering; and — still — taste, the one input no model has commoditized. Where value is migrating: up from models into agents and state, down from models into inference and compute, and sideways into rights. The model layer itself is the valley: indispensable, expensive, and structurally the hardest place in the stack to hold a margin structure unless fused to distribution.

Who is best positioned. Google, because it is the only company integrated from silicon to a billion-user creation surface, and it won consumer video by forfeit. ByteDance, for the same integration in the world’s largest creation funnel. Adobe, which converted from disruption target to the governed gateway for professional work, selling everyone’s models through its own surfaces. ElevenLabs, the cleanest full-stack modality winner. fal, which owns the layer every new model release enriches. And the licensed-catalog holders — the majors, Suno post-WMG-settlement — who hold the only advantages courts actively enforce. The most interesting open cases are the state-layer startups: Higgsfield, Magnific, and Flora are racing to accumulate enough creative state before the incumbents arrive from both sides.

Where the consensus reading diverges from the evidence. Three places, each grounded in the same turnover data. Leaderboard position is still widely treated as an accumulating asset, but with leadership changing hands in weeks it behaves like a depreciating one — a recurring engineering expense that buys temporary placement rather than durable position. App-layer cost compression is read as permanent when the underlying curve says otherwise: inference cost falls roughly 10x every 18 months while credit prices have so far fallen slower, so an app that keeps 50–60% of revenue after serving costs today has room to widen toward 75% — if competition keeps letting that gap accrue to the app rather than pricing it away, the one part of the mechanism 2026 has not settled. And China registers as a threat to Western model labs when the adoption data points elsewhere — Chinese models pressure Western consumer apps while simultaneously supplying the model shelf that makes Western aggregators and workflow layers more capable.

The most important unanswered question: does the creative agent actually change user behavior? The entire industry committed 2026 to the hypothesis that delegation will replace direct manipulation, but no retention data yet proves creators want to hand off the loop rather than hold it. If agents win, the state layer is the biggest prize in creative software history. If they don’t, 2026’s agent build-out will read like 2021’s metaverse pivots — and the canvas owners keep everything. A second open question sits underneath the first: will spatial world models replace 2D frame rendering before 2D video reaches affordable temporal continuity? If real-time simulation hits cost parity first, the industry skips the long-form video problem entirely and renders live camera paths through explorable worlds instead.

The stack in 2030, as the current evidence points: a small set of full-stack distribution giants (Google, ByteDance, possibly Meta) serving casual creation as a free feature; enterprise workflow concentrated around Adobe plus whoever wins the agent race — if the agent hypothesis survives contact with demand; a licensed-content regime sitting in the request path of commercial generation, unless a fair-use ruling breaks it; media-native inference held by a few platforms, or compressed outright if an open serving standard wins; a persistent open-weights substrate (Chinese labs plus NVIDIA’s Cosmos orbit) capping prices; and a handful of vertical modality winners — audio looks decided — with video’s independent labs absorbed, IPO’d as robotics companies, or gone. Each clause carries its condition; the watch list below is what would move them.

Five working conclusions

01Standalone video labs hold depreciating positions

By 2028, no independent video-only model company sustains a premium position without owned distribution or a robotics/simulation revenue line. Runway's GWM and Luma's pro-pipeline pivots are the leaders reading their own future.

02Creative state is what acquirers are buying next

The next wave of $1B+ GenMedia acquisitions will be workflow/state companies, not model labs (infrastructure like Decart excepted). Weavy at >$200M on ~$4M raised was the opening price, not the peak.

03China wins consumer; the West keeps enterprise

Compliance, IP litigation, and provenance mandates keep Western enterprise procurement in the licensed lane regardless of leaderboards — a durable price premium for Adobe, Moonvalley, and licensed-first labs that no Chinese model can compete away.

04App-layer cost structures have room to improve

Inference cost falls ~10x per 18 months; if credit prices keep falling slower — the unresolved variable — apps keeping 50–60% of revenue after serving costs reach 75%+ by 2028. The consensus 'wrapper compression' fear is backward-looking either way: the squeeze already happened.

05An open serving standard compresses orchestration pricing

A vLLM-of-diffusion becomes the standard within 24 months — vLLM-Omni already shipped (Nov 2025) and too much value is pooled behind hand-built kernels for open source to ignore. fal's move up into agents is the incumbent hedging its own commoditization.

Five things to watch

W1Sony and UMG v. Suno

The two unsettled major-label suits — no fair-use ruling expected before 2027 (dispositive motions due April). A fair-use win for Suno weakens the licensed-catalog advantage across all modalities; a loss cements licensing as a permanent per-generation fee.

W2Does Anthropic–Decart close?

A frontier LLM lab paying ~$7B for media inference optimization would confirm that real-time media serving is strategic infrastructure — and start a bidding war for the remaining independents.

W3The Kling IPO

The spin-out closed in July — roughly $3B at $18B post, with Tencent and Alibaba among the investors — and a Hong Kong listing is targeted for 2027. A public listing for a Chinese video unit would put audited disclosure behind the consumer side of the bifurcated ecosystem for the first time.

W4Adobe's agent inside ChatGPT and Claude

The first real test of whether creative agents can live inside general assistants. If usage migrates there, the chat surface — not the creative suite — becomes the distribution layer for creative work.

W5Meta's end-2026 full ad automation

If advertisers hand Meta a URL and a budget and get campaigns back, the third-party ad-creative category collapses into the platforms — and the largest commercial GenMedia use case disappears into an ad auction.

FIG. 08Committed capital vs demonstrated adoption

Latest valuation against revenue, log-log. Dashed guides mark capital-to-revenue ratios; most of the $100M+ club clusters between 10x and 30x, and the outliers say the most about where expectations run ahead of evidence.

$50M$100M$200M$500M$1B$1B$2B$5B$10B$20B10x30x100xHiggsfield · $5.4BKling · $18B postElevenLabs · $11Bfal · $4.5BSuno · $5.4BRunway · $5.3BSynthesia · $4BGamma · $2.1BMeshy · $1.5BARR, $M · log scaleLatest valuation · log scale
Notes, method & sources

A closing cross-check on everything above: where external capital has committed relative to what usage demonstrates, with evidence grade encoded in the marks. Ranges plot at their midpoint. Solid dots are company-stated or audited figures, half-tone dots third-party estimates, and the dashed hollow dot (Higgsfield) is claimed and unverified. Meshy sits alone near the 35x guide — capital there is underwriting a world-model research program rather than current usage.[33] The chart also can’t plot the map’s most cost-disciplined company: HeyGen — $200M ARR on about $74M raised — has no disclosed valuation.[8]

The one-sentence thesis

The next era of generative media belongs to whoever holds the creative state — the characters, brands, and project memory that turn interchangeable models into irreplaceable workflows — and the distribution to put it in front of a billion people.

Appendix A

Ten Hypotheses, Tested

The ten hypotheses this research set out to test, scored against the evidence in the sections above. Most are settled; the two still genuinely open are H9 — whether creative agents change demand-side behavior, the map’s biggest unresolved question — and H8, whether hand-built media serving holds its position now that vLLM-Omni exists.

Show the full scorecard
H1
Raw media generation is becoming a feature rather than a standalone product.Strongly Supported
Veo is free inside YouTube Shorts, Seedance ships inside CapCut, Adobe is running a 12-month unlimited-generations promo, Apple announced free photorealistic generation with iOS 27 (shipping this fall), and Amazon gives ad creative away to drive media spend. The counter-examples that still monetize generation directly — Midjourney, Kling — do it through community or an owned distribution funnel, not the generation button itself. Sora, the best-known standalone generation product in the West, is dead.
H2
Model quality is converging faster than model economics.Supported
In video, ~150 Elo covers #1 through #11 and leaderboard leadership now turns over in weeks, while the price spread between frontier and challenger models is still 5–10x and inference cost fell >10x since 2024 — quality converged, economics did not. The caveat is image, where GPT Image 2 still holds every #1 (though MAI-2.6 just cut the lead nearly in half), and audio, where the licensed catalog (not the model) sets the economics.
H3
GenMedia will remain a heterogeneous multimodel ecosystem rather than collapsing around one dominant foundation model.Strongly Supported
No model wins every workload. The market consolidated into three durable archetypes — distribution-owned omni models (Google, ByteDance, Kuaishou, xAI), independent pro labs (Runway, BFL, Luma, ElevenLabs), and open/China-first price leaders (Alibaba, Tencent, MiniMax, Lightricks). Every aggregator shelf is majority-Chinese in video by usage; 3D and real-time niches resist absorption entirely.
H4
The winning application layer will increasingly own an agentic creative workflow rather than a generation interface.Supported
Every layer of the stack shipped a creative agent in 2026 — Adobe Firefly Assistant, fal Agent, Flora FAUNA, Krea Node Agent, Creatify Agent, Amazon Creative Agent — and the durable advantage is shifting from model access to creative state: persistent characters, brand constraints, project memory, reusable workflows. Not yet "strongly": today’s revenue leaders (Kling, Midjourney, Magnific) still monetize generation interfaces.
H5
Model routing/orchestration becomes more valuable as specialized models proliferate.Supported
fal doubled from ~$200M to ~$400M annualized in months and is reportedly raising at ~$8B — orchestration with media-native depth (kernels, queues, fine-tunes) compounds. But thin aggregation gets absorbed: Replicate sold to Cloudflare, gateways commoditize the unified-API part, and fal itself moving up into agents signals that raw routing alone doesn’t hold its margin structure forever.
H6
Distribution becomes more important as raw model differentiation declines.Strongly Supported
The two defining data points of 2026: Sora — frontier model, no distribution economics — shut down; Kling — good-enough model inside Kuaishou’s funnel and a global API — hit a ~$475M annualized run-rate (Q2 2026). Google won the Western consumer field by default through a 1B-MAU assistant and YouTube. Adobe monetizes other people’s models through workflow distribution.
H7
Incumbent creative platforms are advantaged in workflows, while startups remain advantaged in new AI-native interaction paradigms.Supported
Adobe’s AI-first ARR >$500M (+3x YoY) and Canva’s ~$4B ARR show workflow incumbency converts. Startups own the new paradigms — node canvases (Flora, Krea), real-time generation (Decart, Krea), character persistence (Higgsfield Soul ID). But the boundary is porous: incumbents buy the paradigm (Weavy → Figma for >$200M) and startups build workflow (LTX Studio), so this reads as tendency, not law.
H8
Video generation creates sufficiently different infrastructure requirements to support a distinct GenMedia infrastructure ecosystem.Supported
Long-running jobs, GB-scale intermediate assets, a heterogeneous model zoo, and stateful multi-model pipelines forced fal, Decart, and Runware to hand-build compilers, kernels, and queueing — though a dominant open standard is no longer absent: vLLM-Omni (official vLLM project, open-source since Nov 2025) now serves Wan, FLUX, and Hailuo-class diffusion models in production, which is exactly the commoditization trigger this hypothesis names. External capital still prices the hand-built layer highly — fal at $4.5B+ on ~$400M revenue, Anthropic reportedly near a ~$7B deal for Decart’s optimization stack — while Cloudflare’s $57.4M Replicate purchase (per its 10-Q) shows what absorption of thin aggregation looks like.
H9
Creative agents become the primary interface between humans and generative media models.Unclear
The supply side is unanimous — every major player shipped an agent in 2026, which is where the industry believes lock-in lives. But demand-side proof is young: no retention or revenue data yet shows creators preferring delegation to direct manipulation, and the biggest creative revenues still flow through canvases, timelines, and prompt boxes. This is 2026’s consensus hypothesis, not its confirmed behavior.
H10
The largest GenMedia company may ultimately look less like a model company and more like an operating system for creativity.Supported
The largest GenMedia businesses today — Canva (~$4B+ ARR), Adobe (>$5B AI-influenced ARR as of Sep 2025), Google — are surfaces that aggregate models, not model companies. The strongest startups (Magnific, Higgsfield, HeyGen) win on workflow and state over commodity models. The caveat: Google is both the OS and the frontier model owner, and in video the model+distribution combination is what actually wins.

§ 13

Sources

Key primary and reported sources. Leaderboard positions and private-company figures are as of their cited dates and decay quickly; estimates are labeled throughout.

  1. [1]TechCrunch — OpenAI shuts down Sora · Mar 24, 2026
  2. [2]OpenAI Help Center — Sora discontinuation timeline · 2026
  3. [3]Kuaishou IR — Q1 2026 results (Kling +300% YoY) · May 27, 2026
  4. [4]Fortune — Freepik becomes Magnific at $230M ARR · Apr 28, 2026
  5. [5]TechCrunch — Runway raises $315M at $5.3B · Feb 10, 2026
  6. [6]TechCrunch — ElevenLabs raises $500M at $11B · Feb 4, 2026
  7. [7]Variety — Suno raises $400M at $5.4B · Jun 2026
  8. [8]HeyGen — $200M ARR announcement · Jun 2026
  9. [9]CNBC — Synthesia Series E at $4B · Jan 26, 2026
  10. [10]Bloomberg — Sequoia-led round values fal at $4.5B · Dec 9, 2025
  11. [11]PR Newswire — fal launches fal Agent · Aug 12, 2026
  12. [12]Cloudflare — agreement to acquire Replicate · Nov 17, 2025
  13. [13]TechCrunch — ComfyUI hits $500M valuation · Apr 24, 2026
  14. [14]The AI Insider — Flora raises $42M Series A · Jan 30, 2026
  15. [15]Figma — Config 2026 / Weave rollout · Jun 2026
  16. [16]Adobe Q2 FY26 earnings call (AI-first ARR >$500M) · Jun 11, 2026
  17. [17]Forbes — Adobe Firefly agent inside ChatGPT and Claude · Jun 19, 2026
  18. [18]Google — Flow, Whisk and ImageFX unified · Feb 25, 2026
  19. [19]Google — Gemini passes 1B monthly users · 2026
  20. [20]9to5Google — Project Genie launches for Ultra subscribers · Jan 29, 2026
  21. [21]NVIDIA — Cosmos 3 open frontier model for physical AI · Jun 1, 2026
  22. [22]The AI Insider — World Labs raises $1B · Feb 19, 2026
  23. [23]Bloomberg — Anthropic in talks to buy Decart for ~$6B · Aug 13, 2026
  24. [24]WinBuzzer — MiniMax raises $619M in Hong Kong IPO · Jan 8, 2026
  25. [25]TechFundingNews — PixVerse lands $439M · Jul 2026
  26. [26]TechCrunch — Black Forest Labs raises $300M at $3.25B · Dec 1, 2025
  27. [27]GlobeNewswire — Lightricks open-sources LTX-2 · Jan 6, 2026
  28. [28]Artificial Analysis — Video Generation Arena leaderboard · Aug 2026
  29. [29]Greenberg Traurig — EU AI Act Article 50 transparency obligations · Jun 2026
  30. [30]PitchBook — AI video investment reaches $5.6B in 2026 YTD · 2026
  31. [31]CNBC — Meta launches in-house Muse image model · Jul 7, 2026
  32. [32]Billboard — what the Suno/Udio licensing deals mean · 2025–26
  33. [33]TechFundingNews — Meshy raises ~$400M at $1.5B · Jul 21, 2026
  34. [34]SEC — Getty terminates Shutterstock merger (board action Jul 7) · Aug 13, 2026
  35. [35]TechCrunch — Decart’s Oasis 3 world model · Jun 10, 2026
  36. [36]Dodo Payments / Bessemer — AI-native gross margin benchmarks · Feb 2026
  37. [37]Atlas Cloud — cheapest AI video generation APIs 2026 · 2026
  38. [38]CGTN — China enforces AI-content labeling rules · Sep 1, 2025
  39. [39]Arena — LMArena is now Arena · Jan 28, 2026
  40. [40]TechCrunch — Higgsfield raises $400M Series B at $5.4B · Aug 17, 2026
  41. [41]BlueFive Capital (co-lead) — Kling AI ~$3B round at $18B post ($2B initial close per Bloomberg, Jul 2) · Jul 2026
  42. [42]Music Business Worldwide — Suno inks global licensing deal with BMG · Aug 12, 2026
  43. [43]Microsoft AI — MAI-Image-2.6 launches at #2 on Arena · Aug 10, 2026
  44. [44]Black Forest Labs — FLUX 3 Video release notes · Aug 5, 2026
  45. [45]The Information — fal in funding talks at ~$8B valuation · Mar 2026
  46. [46]Cartesia — Sonic 3.6 tops both AA speech arenas · Aug 18, 2026
  47. [47]Kuaishou IR — Q2 2026 results (Kling RMB 850M+, up over 200% YoY) · Aug 19, 2026
  48. [48]Calcalist — Anthropic closing in on Decart at ~$7B · Aug 16, 2026
  49. [49]TechNode Global — Alibaba releases Wan 3.0 in public beta · Aug 10, 2026
  50. [50]vLLM-Omni — open-source diffusion/omni model serving (vLLM project) · Nov 2025
  51. [51]ElevenLabs reaches $600M ARR · Jul 2026
  52. [52]Tech Times — Grok Imagine Video 1.5 tops AI video leaderboard at 86% below Sora · Jun 18, 2026
  53. [53]KrAsia (36Kr) — ByteDance raises Volcano Engine MaaS target on Seedance 2.0 growth (>RMB 1B/month) · Jun 5, 2026
  54. [54]BigGo (36Kr-lineage) — Seedance contract minimums, margins, ~95% short-drama penetration · Jul 7, 2026
  55. [55]Arena — Seedance 2.5 #2 image-to-video, #1 on the new Video Edit board · Aug 2026
  56. [56]Artificial Analysis — image-to-video leaderboard (Seedance 2.0 #1; Kling 3.0 Pro #12) · Aug 19, 2026
  57. [57]VentureBeat — Krea 2 Raw/Turbo open weights; 30M+ users, enterprise logos · Jun 2026
  58. [58]Google — I/O 2026 roundup: Flow Agent available to all Flow users globally; multi-step tasks, batch edits, Flow Tools · May 2026
  59. [59]PostRound (PitchBook data) — Reve $350M Series B at $1.9B, Top Harvest Capital (never press-announced) · Nov 2025
  60. [60]Reve — launching Reve 2.1 (#2 on AA image; independent-lab framing, compute-efficiency claim) · Jul 9, 2026
  61. [61]Alibaba Cloud blog — Wan series passes 6.9M downloads (HF + ModelScope) · Aug 2025
  62. [62]Hugging Face community analysis — Wan 3.0 and the end of the open Wan line · Jul 29, 2026
  63. [63]MLex — MiniMax fails to dismiss Disney/Universal/WBD copyright claims (C.D. Cal.) · May 26, 2026
  64. [64]MiniMax — FY2025 results (US$79M revenue; AI-native products US$53.1M incl. Hailuo + Talkie) · Mar 2, 2026