BusinessIssue #7

The $1,000-Per-Minute Video Era Is Ending

AI video generation costs collapsed 65% in a year—a sign the industry's structure, not just its tech, is shifting.

The $1,000-Per-Minute Video Era Is Ending

Opening

How much do you think it costs to make a one-minute corporate promo video?

As of 2025, hiring a freelancer runs $1,000 to $5,000; going through an agency starts at $15,000. Add crew wages, equipment rental, location fees, and post-production, and you quickly blow past tens of thousands of dollars. In won terms, that’s a minimum of ₩1.3 million per minute — agency-tier work can top ₩20 million.

Meanwhile, today’s AI video generation tools cost $0.01 to $0.50 per second. Convert that to a one-minute clip and you get $0.6 to $30. Even the priciest premium model runs at about 0.2% of traditional production costs. And this price has dropped another 65% in just the past year.

Last week, ByteDance’s Seedance 2.0 became the newest chapter in this cost-curve collapse. But what made the model go viral wasn’t the price — it was an incident where a single photo was used to clone someone’s voice. What happens when technology gets cheaper and more powerful without safeguards? We found out, bluntly, within three days of launch.

Today I want to talk about the structural shift in AI video costs, the technical turning point Seedance 2.0 represents, and why all of this is dismantling the old formula that “you need money to make video.”

How Far Has Video Production Cost Actually Fallen?

Let’s start with the numbers.

In 2024, AI video generation services cost $50–$200 per minute. Back then, the story was simply “cheaper than traditional production.” But between 2025 and 2026, the whole board flipped.

Looking at today’s leading models’ per-second generation costs, the cheapest — Vidu 2.0 — runs about $0.0375. The industry average sits around $0.084 per second, meaning Vidu is roughly 55% below average. Some platforms offering bulk generation via API have pushed costs down to $0.01 per second. According to WaveSpeedAI’s early-2026 breakdown, pricing across platforms and features ranges from $0.01 to $0.50 per second.

Here’s a comparison that makes it tangible. If traditional corporate video production runs $1,000–$5,000 per minute, AI generation costs $0.6–$30 per minute. Even comparing the most expensive AI model against the cheapest traditional production, you’re looking at over 97% cost savings. One analysis found that producing 10 social media campaign videos with AI costs about $89, versus $100,000+ for the same work through an agency.

Industry adoption backs up this trend. Adoption of AI video generation tools grew 300% year-over-year, while per-second generation costs fell 65% over the same period. It’s a self-reinforcing loop: falling costs bring in more users, more users intensify competition, and competition pushes costs down further.

This isn’t simply “AI is cheap.” It means the barrier to entry for video production itself is disappearing. Previously, you needed capital to make video. Now you just need an idea.

What Seedance 2.0 Revealed — It’s Not Just About Being Cheap

A natural question follows falling costs: “Doesn’t quality drop along with the price?” ByteDance’s Seedance 2.0 is the model that answered this question head-on.

Most existing AI video tools operate through a single pathway — text goes in, video comes out, or an image goes in, video comes out. Seedance 2.0 flips this entirely. It adopted a unified multimodal1 architecture that processes text, images, video, and audio simultaneously as inputs. It can reference up to 12 files at once. “Use this image as the character, this video for camera-work reference, this audio for the background rhythm” — you can specify each element separately.

If existing AI video tools were vending machines (press a button, get a result), Seedance 2.0 is more like a kitchen (you pick ingredients, specify a recipe, and a chef makes it for you).

Technically, what stands out is its dual-branch diffusion transformer2 architecture. Visual and audio elements are generated simultaneously in separate branches, but precisely synchronized in time. Lip-sync3 supports 8 or more languages, and stereo audio (background music + ambient sound + dialogue) is generated simultaneously.

Swiss consulting firm CTOL tested the model directly and called it “the most advanced AI video generation model in existence.” That kind of assessment deserves scrutiny for independence, but the specs themselves are objectively impressive: 2K resolution output, 30% faster generation than the previous version, and multi-shot4 sequences generated within 60 seconds.

And crucially — no watermark. That’s a convenience in one sense, but it’s also the seed of a problem I’ll get to shortly.

The Big Four — There’s No “Best,” Only “Best Fit”

The AI video generation market is currently consolidating into a big-four structure. Each model has different strengths, so the more realistic question isn’t “which is best” but “which model fits this particular shot.”

OpenAI’s Sora 2 dominates physics simulation — it can realistically render shattering glass fragments. But it requires a $200/month (Pro) subscription and includes a watermark. Google’s Veo 3.1 stands out for cinematic quality. It supports broadcast-grade 24fps footage with up to 4 reference images, but also embeds a SynthID metadata watermark. Kuaishou’s Kling 3.0 offers the best value for money — about $0.50 for a 10-second clip, and it’s specialized for Asian content. And ByteDance’s Seedance 2.0 differentiates itself through multimodal control and editing flexibility, and is the only model that supports audio reference input.

There’s evidence this competition is generating real industrial impact. As of December 2025, Kuaishou’s Kling reported roughly 12 million monthly active users and $20 million in monthly revenue. AI video is no longer a lab experiment — it’s an actual money-making industry.

Right after Seedance 2.0’s launch, Chinese equities reacted too. On February 10, 2026, COL Group hit its 20% daily limit, while Perfect World and Shanghai Film each rose 10%. The CSI 300 index climbed 1.4%. Kaiyuan Securities analyst Fang Guangzhao called it “a potential singularity moment for the film and TV industry.”

One Photo, Cloned Voice — As Costs Fall, Risk Rises

Here’s where the story shifts. What happens when a technology that’s rapidly getting cheaper and more powerful spreads without safeguards? Seedance 2.0 showed us, in stark terms, within three days of its launch.

Chinese tech blogger Pan Tianhong uploaded just one photo of himself. No voice sample, no text prompt. Yet Seedance 2.0 produced a video that precisely replicated his voice tone, pace, and inflection. In his reaction, the word “terrifying” (恐怖) appeared six times.

Technically, this is quite significant. Existing voice-cloning tools required at least 30 seconds of voice sample. Seedance 2.0 inferred vocal characteristics from a facial image alone. The wall that used to separate visual likeness (face) from acoustic identity (voice) — a kind of “digital air gap” — has collapsed.

ByteDance responded immediately. It suspended the real-person reference feature and introduced live verification (face + voice recording) procedures for its Jimeng and Doubao apps. This is a textbook example of China’s “build first, regulate later” AI development pattern, but the response speed itself was fast.

But this isn’t a ByteDance-only problem. The EU’s investigation into X (formerly Twitter) over xAI’s Grok generating non-consensual sexual content belongs to the same context. As AI video generation costs fall, so does the cost of abuse. We’ve already arrived at a world where someone’s face and voice can be cloned for a few cents per second.

Oz’s Lens

I see this phenomenon not as “the democratization of video production” but as “the commoditization of video production.”

There’s a pattern I’ve watched play out countless times over 20 years of building GTM strategies. When a technology’s cost drops sharply, that technology itself stops being a differentiator. It happened with cloud computing. It happened with website building. AI video is walking the same path.

The fact that you can make a video for $0.01 per second means, conversely, that video alone can no longer secure competitive advantage. When anyone can make it, what matters becomes “what” you make — planning ability and storytelling. When the tool gets democratized, perspective becomes the scarce resource.

As someone who works with data, I’d add one more thing: when you see a figure like “65% cost decline,” check the measurement basis. Is it per-second API cost? A subscription-model conversion? What resolution and features are included? The numbers shift dramatically depending on the answer. In my assessment, pure generation cost has dropped dramatically, but real-world production cost — including prompt design, iteration, and post-editing — hasn’t fallen nearly as much.

And as the voice-cloning incident shows, the cheaper and more powerful technology gets, the higher the cost of safeguards ought to rise. The fact that the cost curve for generation and the cost curve for safety are moving in opposite directions is this industry’s structural dilemma. Seedance 2.0’s decision to skip a watermark was probably made for user convenience, but the social bill for that choice hasn’t arrived yet.


Closing

Here’s the summary.

First, as the per-second cost of AI video generation falls to $0.01–$0.50, video production is no longer a game of capital. It’s becoming a game of ideas and planning ability.

Second, Seedance 2.0’s unified multimodal architecture — handling text, image, video, and audio simultaneously — marks a turning point where AI video evolves from “generation” to “production.”

Third, as costs fall, the cost of abuse falls with them. The reality that a single photo can clone a voice is a warning that the pace of technological innovation has already outrun ethical safeguards.

Wu Di, head of ByteDance’s diffusion-engine AI algorithms, said in a recent interview that “there will be one or two more major leaps in AI video technology in 2026.” The cost curve will keep falling. What we’ll need then isn’t a cheaper tool — it’s a better question: “What should we make with this technology?” And, “What shouldn’t we make?”


📎 References & Further Reading

Author Kwangseob Ahn is a professor in the Department of Business Administration at Sejong University and lead consultant at OBF (Oswarld Boutique Consulting Firm). He teaches statistics and data analysis — including business data management and business analytics — at university, while leading GTM strategy and AI strategy consulting in the field, designing the intersection of technology and business. He has published an academic paper on memory architecture for AI conversation systems (HEMA) and runs Daily Arxiv, a project curating global AI papers every day. He holds a master’s from Korea University’s Graduate School of Technology Management and completed KMBA. He is the author of Those Who Outsource Their Thinking: Homo Brainless.

Footnotes

  1. Multimodal: The ability to understand and process multiple forms of data — text, image, audio, video — at the same time. Similar to how humans simultaneously use sight, hearing, and language.

  2. Diffusion Transformer: An AI model structure that starts from noise and progressively generates a clean video. Like carving a shape out of marble, it “carves” a video out of noise.

  3. Lip-sync: Technology that precisely matches an AI-generated character’s mouth movements to the audio. Synchronization happens at the level of individual phonemes.

  4. Multi-shot: A feature that generates multiple connected scenes (shots) from a single prompt. The key is maintaining consistent character appearance and personality even as scenes change.