Hunyuan Video
Back to models

Hunyuan Video API Generator

Resolution
1080p
Max duration
10 seconds
Latency
60 to 120 seconds
Pricing
~$0.04 / second

What the Hunyuan Video API Does

Every capability is accessible from a single Oracium endpoint. No Tencent Cloud account, no China-region credentials, no self-hosting required.

Hunyuan Video text to video

Convert any text prompt into a 1080p video clip with authentic cinematic motion. The dual-stream transformer architecture processes your text description with significantly higher fidelity than models using separate text encoder add-ons, enabling complex multi-element scene generation that holds together across the full clip duration.

Image-to-video generation

HunyuanVideo-I2V, released in March 2025, animates any reference image into a fluid video using a token-replacement technique that reconstructs and incorporates the reference image information into the generation process. Visual consistency from the first frame is maintained throughout, with a bug fix released on March 7, 2025, ensuring full identity preservation.

Multi-style generation

Hunyuan Video generates across a wide range of visual styles from a single unified model architecture: realistic documentary, cinematic film, anime, illustration, and stylized artistic output. No separate model switching per style. Specify the aesthetic in your prompt, and the architecture adapts the output accordingly.

Cinematic camera motion

The model supports authentic cinematic camera movements specified directly in prompts: pans, dolly pushes, tracking shots, and depth shifts. These are not post-processing effects; they are generated as part of the 3D-aware temporal modeling that Hunyuan builds into its diffusion transformer architecture.

HunyuanVideo-Avatar

Released in May 2025, HunyuanVideo-Avatar is an audio-driven human animation model built on the Hunyuan Video foundation. It generates realistic human video from an audio input and a reference image, in which the character speaks, moves, and gestures, driven entirely by the audio signal without additional keyframe input.

HunyuanCustom multimodal input

Released in May 2025, HunyuanCustom enables customized video generation driven by multiple input modalities simultaneously; text, image, audio, and video reference inputs can all contribute to a single generation. It solves the subject consistency problem across custom video workflows that single-input models cannot address.

From API key to Hunyuan Video in 60 seconds

One Oracium key unlocks Hunyuan Video, Runway Gen-4, Seedance Pro, Veo 3, Kling 2.0, FLUX Pro 1.1, and 100+ models. No Tencent Cloud. No China-region setup.

01

Create your Oracium account

Sign up at oracium.io. Instant access to Hunyuan Video, both text-to-video and image-to-video modes, from the same account and key that access every other model on the platform.

02

Get your API key

Generate from the Oracium dashboard. One key grants full access to the original Hunyuan model family, I2V, Avatar, and Custom, without separate Tencent credentials or WeChat verification requirements that restrict direct access outside China.

03

POST to the Hunyuan Video endpoint

Send your text prompt or image URL, along with the model version, resolution, and style parameters. Poll the returned task ID. Video URL is typically ready within 60 to 120 seconds, depending on clip length and resolution.

04

Ship to production

Oracium handles infrastructure, scaling, and quota management. Build without the self-hosting requirements of running a 13B-parameter model ( minimum 24GB VRAM for the original model, 14GB for 1) on your own hardware.

What developers build with the Hunyuan Video API

Open-source licensing, multi-style output, and competitive pricing make Hunyuan Video the right model for applications where portability, customization, and cost control matter as much as quality.

01

High-volume content generation pipelines

At competitive per-second pricing, Hunyuan Video enables content pipelines that generate hundreds of video clips per day without the cost floor of premium closed models. The multi-style unified architecture means one integration handles realistic, cinematic, and stylized content without switching models.

02

Anime and stylized content platforms

Hunyuan Video's multi-style architecture achieves higher-fidelity anime generation than models optimized solely for photorealistic output. Platforms serving creative communities that produce anime, illustration, and stylized video content benefit from a single model that handles all aesthetic categories without separate fine-tuning.

03

Digital avatar and virtual human products

HunyuanVideo-Avatar enables audio-driven human animation from a reference image and audio input. Applications building talking-head products, virtual presenters, and digital spokespeople can access the Avatar model through the same Oracium endpoint without separate Tencent enterprise agreements.

04

Custom model fine-tuning pipelines

As a fully open-source model under the Apache 2.0 license, Hunyuan Video can be fine-tuned on proprietary datasets. Developers who access it via Oracium for prototyping retain the option to fine-tune and self-host the model later without licensing restrictions, an escape hatch unavailable with closed commercial models.

05

E-commerce and product visualization

HunyuanVideo-I2V generates product video from still images with high first-frame consistency. For e-commerce applications that generate videos for large product catalogs, the combination of image-to-video capability and per-second pricing makes Hunyuan Video a cost-effective choice at scale.

06

Multimodal creative tools

HunyuanCustom accepts text, image, audio, and video reference inputs simultaneously in a single generation call. Creative tool builders who need to let users specify a character's appearance, a scene's atmosphere, and a musical tone in a single prompt have a generation endpoint that accepts all these inputs natively.

Hunyuan Video vs Kling 2.0 vs Runway Gen-4 vs Seedance Pro

Hunyuan Video's open-source license is its clearest differentiator against commercial-only models. All available via Oracium with unified billing.

Parameters

Hunyuan Video (Oracium)

13B (open)

Kling 2.0

Undisclosed

Runway Gen-4

Undisclosed

Seedance Pro

Undisclosed

Open source license

Hunyuan Video (Oracium)

Apache 2.0

Kling 2.0

Closed

Runway Gen-4

Closed

Seedance Pro

Closed

Max resolution

Hunyuan Video (Oracium)

1080p

Kling 2.0

1080p

Runway Gen-4

Native 4K

Seedance Pro

1080p

Text-to-video

Hunyuan Video (Oracium)

Yes

Kling 2.0

Yes

Runway Gen-4

Yes

Seedance Pro

Yes

Image-to-video

Hunyuan Video (Oracium)

HunyuanVideo-I2V

Kling 2.0

Yes

Runway Gen-4

Yes

Seedance Pro

Yes

Audio-driven avatar

Hunyuan Video (Oracium)

HunyuanVideo-Avatar

Kling 2.0

No

Runway Gen-4

Act-Two

Seedance Pro

No

Multi-style (anime, illustration)

Hunyuan Video (Oracium)

Native unified model

Kling 2.0

Limited

Runway Gen-4

Photorealistic focus

Seedance Pro

Limited

Fine-tuning allowed

Hunyuan Video (Oracium)

Yes, open weights

Kling 2.0

No

Runway Gen-4

No

Seedance Pro

No

Cost per second (Oracium)

Hunyuan Video (Oracium)

~$0.04

Kling 2.0

~$0.056

Runway Gen-4

$0.05–$0.12

Seedance Pro

~$0.04

One Oracium API key

Hunyuan Video (Oracium)

Yes

Kling 2.0

Yes

Runway Gen-4

Yes

Seedance Pro

Yes

Start generating in under 60seconds.

Get an API key, install the SDK, and ship your first generation today.

FAQ

Start generating in
under 60 seconds.

Get an API key, install the SDK, ship your first generation today.