Council Post: AI Video Flips The Economics Of Inference: Compute Once, Earn Forever

Jerry Tang is CEO of Atlas Cloud, providing enterprises and creators access to the leading generative AI models across all modalities.

getty

​How many times can an AI application get paid before it has to compute again? The answer places every AI company on a spectrum, somewhere between computing an evergreen asset and running a service that keeps a costly model in the serving path.

I worked in investment banking building a commercial mortgage-backed securities business and now invest as a founding partner of VCV Digital alongside operating Atlas Cloud. I can tell you the world is looking for winning AI businesses on the wrong side of this spectrum. One end turns compute into ongoing revenue-producing assets, and the other must outrun the meter every time the product works.​

Finance teams already treat inference as the cost of goods sold (COGS) based on a simple test: If a customer triggered the API call, it’s COGS. A customer-service agent books an inference expense every time it works, and that cost recurs for as long as customers keep using the service.

AI content businesses don’t carry inference as COGS. They spend it once to build an asset and amortize it against everything that asset later earns. Netflix already accounts for its library this way, amortizing more than 90% of a title’s cost within four years of release. The relevant financial ratio is production compute cost per monetized use: For a finished asset, the production bill stays fixed while the audience grows.

In other words, compute once, earn forever.

Caching is the Band-Aid for the other end of this spectrum. Prefill caching and cached input rates make each run cheaper, but the next customer still triggers a model run. A finished video isn’t a cache with a short-term expiration. It renders once and plays for years, the same way Netflix doesn’t pay for the actors to come back into the studio each time viewers hit play.

Atlas Cloud sells inference, so a business that reruns the model for every customer is good for my revenue. As an investor, I prefer the opposite curve. Our own data shows the difference in usage: Generative media widens across new accounts, whereas language-model use deepens within existing ones.​​

Turning Inference Into Revenue-Generating Inventory​

Versatile Media, a studio running on our infrastructure, operates this model today. The company runs about 30 short dramas per month through its production system, with a reported cost of about $35 for a completed two-minute video. At $35 per video, if a finished video produces 1,000 monetized views, its production cost spreads to 3.5 cents per view. At 10,000, it falls to 0.35 cents. The denominator can keep growing without another production cycle for the same asset, exactly what finance likes to see.​

The market for this content is growing. According to TheWrap, China’s short-drama market increased from $500 million in 2021 to about $7 billion (paywall) in 2024, so AI producers don’t have to generate demand. They enter it with a different cost base. TheWrap also reports that a live-action vertical series typically costs $150,000 to $300,000, and Holywater told the publication that AI lets it make comparable productions for about one-tenth the cost. That’s a 90% cut to the largest line item in the business. Audiences aren’t rejecting the result either: Holywater’s AI app MyMuse converts to paid at 23.6%, against 23.9% for its live-action app MyDrama.​​

This is arriving faster than the people who study it expected. Doug Shapiro, the former Turner strategy chief and author of Infinite Content, put blockbuster below-the-line cost at $1 million to $2 million per finished minute, reaching $10 to $20 per minute only in his most speculative 2030 scenario. Versatile’s $35 two-minute episode is $17.50 per finished minute today. Vertical drama isn’t blockbuster film, so that compares cost, not quality. But the cost he reserved for 2030 is in market now, years ahead of schedule.

Cheap production will flood the market, and audiences will pick the winners. A content mill spends compute and earns nothing, so its cost per monetized use runs to infinity no matter how cheap each render gets. A creator with taste spends the same compute and earns for years. This ratio rewards work that gets shared and rewatched. As I wrote before: Falling costs lower barriers for creators with good ideas and no capital.​

The Meter Gets Harder To Outrun

This isn’t to say that serving-path businesses can’t earn strong margins if they price well above their inference COGS. They are, however, running an exposed business where they don’t have full control over their underlying costs. The most famous model provider, OpenAI, told investors that second-quarter 2026 revenue rose 18% to $6.7 billion while its operating loss widened to $12.3 billion. It can close that gap through efficiency, scale or product mix, but I don’t expect losses of that magnitude to coexist with today’s pricing forever, and I expect some of the pressure to reach their API customers.

This isn’t speculative doomsaying when it’s already happened. OpenAI launched GPT-5.5 at $5 per million input tokens and $30 per million output tokens, twice GPT-5.4’s $2.50 and $15 rates. Builders can route work to cheaper models, but that option narrows when a product depends on one model’s performance. When the model a product needs costs more, the application’s cost floor rises with it.

On the other side of the spectrum, studios still spend on new titles as attention fades. But an existing asset locks its production inference cost when it renders. Tomorrow’s compute cost can’t reach backward into yesterday’s catalog.

That’s what I underwrite: Does each new customer spread a fixed production bill or restart the meter?

My bet is that AI video content will produce some of the application layer’s strongest businesses. A serving-path product can make each run cheaper, but it can’t take compute out of its cost to serve. A finished asset can keep earning after its production meter stops.

Beat that cost floor.​​


Forbes Technology Council is an invitation-only community for world-class CIOs, CTOs and technology executives. Do I qualify?