Live AI video model catalogue

60+ generation-ready AI video models. One studio.

INXSocial brings multiple AI video providers and model families into one production workspace. Let AI Recommended choose a suitable model for your brief, or choose the model yourself and use the controls that model actually supports.

63 available now104 tracked in the live catalogueCatalogue refreshes automatically
63Generation ready
61Text-to-video
58Image-to-video
42Audio-capable
Why a multi-model studio matters

Use the latest suitable model without rebuilding your workflow.

AI video models specialise in different things: speed, cinematic motion, reference consistency, audio, longer duration, resolution or cost efficiency. INXSocial keeps those differences behind one consistent Video Studio.

The public catalogue below only showcases models that pass the current INXSocial generation-readiness checks. If a provider retires a model, pricing becomes unsafe or required configuration is unavailable, it can be removed from the ready catalogue automatically.

Available in INXSocial

Latest generation-ready AI video models

Sorted by the release metadata available in the live catalogue. Capabilities vary by model.

Latest17 Sept 2026

P-Video-2-Pro

prunaai

P-Video-2-Pro is Pruna AI's quality-tier video generation model built on MiniMax H3, creating clips from text or from a first frame with an optional last frame. It handles multi-beat camera moves, heavy physics like water, fire, and fabric, coherent full-body motion, and two-shot dialogue with lip sync, generating audio with every clip. It offers a speed or quality recipe, three levels of prompt expansion, durations from 5 to 15 seconds, and 480p or 768p output at 24 FPS.

Text to videoImage to videoAudio480p768p
Live pricing in studioTry model →
Latest10 Sept 2026

P-Video-2

prunaai

Quality-focused generation for everyday social video, with native audio, draft previews and frame guidance.

Text to videoImage to videoAudio to videoAudio720p
From ~15 creditsTry model →
Latest8 Sept 2026

MiniMax H3 Fast

minimax

Fast reference-driven video with strong visual continuity for product shots and rapid creative iteration.

Text to videoImage to videoVideo to videoAudio to video480p
From ~27 creditsTry model →
Available2 Sept 2026

MiniMax H3 Max Turbo

minimax

MiniMax H3 Max Turbo is a distilled, speed-focused variant of H3 Max designed to preserve its prompt adherence and visual character at higher throughput. It generates 5 to 15 second videos at 480p or 768p from text or an opening image, with optional end-frame guidance for controlled transitions and synchronized audio generated with the video. It is well suited to rapid creative iteration, interactive experiences, and high-volume production workflows.

Text to videoImage to video480p768p
Live pricing in studioTry model →
Available27 Aug 2026

Gemini Omni Flash 1.1

google

Gemini Omni Flash 1.1 is Google's updated multimodal video generation and editing model in the Gemini Omni family. It generates native synchronized audio from the prompt, and expands the original Omni Flash workflow with additional 1080p and 4K output modes, a 360p draft mode, scene extension in 3 to 10 second increments up to 30 seconds total, start-to-end frame interpolation for fluid transitions, and reference-to-video generation guided by both images and short video clips. It is built for teams that need stronger continuity, higher-resolution delivery, and more controllable multi-input video creation than the first Omni Flash release.

Text to videoImage to videoVideo to videoAudio1080p4K
From ~58 creditsTry model →
Available26 Aug 2026

MiniMax H3 Max

minimax

MiniMax H3 Max is a performance-tuned variant of MiniMax H3 built for faster video generation while preserving strong prompt following and polished audiovisual results. It supports text-to-video and image-to-video workflows, including optional first-and-last-frame guidance for more controlled motion between key images, and is well suited to rapid ideation, high-volume content production, and interactive creative workflows that need lower latency without giving up visual quality.

Text to videoImage to videoAudio480p768p
Live pricing in studioTry model →
Available24 Aug 2026

Wan 3.0

alibaba

High-fidelity generation for polished hero Reels, product storytelling and stronger motion quality.

Text to videoImage to videoVideo to videoAudio to videoAudio720p1080p
From ~58 creditsTry model →
Available24 Aug 2026

Wan3.0 Prime

alibaba

Wan3.0 Prime is the faster inference variant of Alibaba's Wan3.0 video model. It keeps the same API shape, the same multimodal creation and editing workflows, and the same output quality as Wan3.0, while trading up to a higher price tier for lower latency. It supports text-to-video, first- and last-frame image-to-video, reference-driven video generation, localized video editing, and temporal extension, with support for image, video, audio, document, and webpage inputs in the broader Wan3.0 workflow.

Text to videoImage to videoVideo to videoAudio to videoAudio720p1080p
From ~82 creditsTry model →
Available11 Aug 2026

LTX-2.5 Fast

lightricks

LTX-2.5 Fast is the speed-focused variant in the LTX 2.5 video family, built for high-throughput text-to-video and image-to-video generation. It supports output from 720p through 4K, longer durations at HD resolutions, optional native audio generation, and first-to-last-frame image guidance for more directed motion and shot planning. It is well suited to rapid creative iteration, storyboard development, social and advertising content, and other video workflows that need faster turnaround without dropping to low-end resolutions.

Text to videoImage to videoAudio
Live pricing in studioTry model →
Available11 Aug 2026

LTX-2.5 Pro

lightricks

Production-focused video with excellent turnaround, synchronized audio and strong image-to-video control.

Text to videoImage to videoVideo to videoAudio to videoAudio720p1080p
From ~497 creditsTry model →
Available7 Aug 2026

Seedance 2.5

bytedance

Premium multimodal generation for complex branded stories, longer clips and demanding creative direction.

Text to videoImage to videoAudio720p1080p
From ~135 creditsTry model →
Available4 Aug 2026

FLUX 3 Video

black-forest-labs

FLUX 3 Video is Black Forest Labs' multimodal foundation model for video generation with synchronized audio. It generates clips from 5 to 20 seconds across text-to-video, image-to-video, and video-to-video modes on one architecture, with keyframe control to pin an opening image or interpolate motion across pinned frames, chained continuations for arcs beyond 20 seconds, multi-shot sequences with hard cuts inside one generation, and native multilingual dialogue. A draft mode returns a fast low-resolution preview and a cache that a follow-up call enhances at full quality, tightening iteration loops. Style range spans candid camcorder footage, animation, motion design, and cinematic photoreal, character consistency holds across scenes within one generation, and in-video typography renders cleanly for titles and animated designs.

Text to videoImage to videoVideo to videoAudio720p1080p
Live pricing in studioTry model →
Available30 Jul 2026

MiniMax H3

minimax

MiniMax H3 is a multimodal video generation model that supports text-to-video, first-frame and keyframe-guided generation, multi-reference conditioning, and audio-video continuation in a single workflow. It accepts text together with images, videos, and audio references to keep subjects, voice, motion, and scene identity more consistent across shots, while generating synchronized sound natively rather than as a separate dubbing pass. It is well suited to cinematic multi-shot generation, reference-driven character performance, instruction-based video editing, and continuation workflows that extend an existing clip or audio segment into a seamless new video.

Text to videoImage to videoVideo to videoAudio to videoAudio768p1440p
Live pricing in studioTry model →
Available30 Jun 2026

Gemini Omni Flash

google

Gemini Omni Flash is Google's multimodal video generation and editing model in the Gemini Omni family. It turns text, photos, and video into 10-second clips with native audio generation, supports photo-to-video creation from up to five reference images, and adds video-to-video plus multi-turn editing workflows. Google positions it as the Gemini app successor to Veo 3.1, combining Gemini's world understanding with conversational control for video creation and editing.

Text to videoImage to videoVideo to videoAudio
From ~290 creditsTry model →
Available23 Jun 2026

Seedance 2.0 Mini

bytedance

Seedance 2.0 Mini is a lighter variant in ByteDance's Seedance 2.0 video model family. It is positioned for teams that want the cinematic prompt understanding and multimodal generation style of Seedance 2.0 in a smaller, more iteration-friendly package. It is suited to fast concepting, quicker turnaround, and broader throughput when the full flagship model is not required.

Text to videoImage to videoAudio480p720p
From ~235 creditsTry model →
Available22 Jun 2026

HappyHorse 1.1

alibaba

HappyHorse 1.1 is Alibaba's upgraded multimodal video model for text-to-video, image-to-video, and reference-to-video generation. It improves motion continuity, prompt following, character consistency, facial texture quality, cinematic shot logic, and audio-visual synchronization over HappyHorse 1.0, making it better suited to multi-shot storytelling, multi-character scenes, close-up performance, and reference-driven production workflows.

Text to videoImage to videoAudio720p1080p
Live pricing in studioTry model →
Available17 Jun 2026

Kling VIDEO 3.0 Turbo

klingai

Kling VIDEO 3.0 Turbo is a speed-optimized multimodal video generation model in the Kling 3.0 family. It is built for high-volume production workflows that need faster turnaround without giving up stable motion, prompt adherence, multi-shot consistency, or audio-visual alignment. It supports text-to-video and image-to-video generation, with particular emphasis on improved lip-sync quality and more efficient large-scale content creation.

Text to videoImage to videoAudio720p1080p
Live pricing in studioTry model →
Available9 Jun 2026

Ray3.2

luma

Ray3.2 is Luma's flagship video model for turning creative direction into controllable production workflows. It supports text-to-video, image-to-video, and video-to-video generation, with stronger continuity, motion transfer, camera motion transfer, character transformation, relighting, environment change, and product-swap workflows. It is built for cinematic-quality output, multi-keyframe control inside a single clip, and Modify Video V2 workflows that preserve performance, lighting, and scene structure while transforming existing footage.

Text to videoImage to videoVideo to video720p1080p
From ~173 creditsTry model →
Available30 May 2026

Grok Imagine Video 1.5

xai

Grok Imagine Video 1.5 is xAI's newer image-to-video model. It is positioned above the earlier Grok Imagine Video release with higher per-second pricing, supports durations up to 15 seconds, and generates 480p or 720p video from a single still-image starting frame for cinematic clips, animated visuals, and prompt-guided short-form video creation.

Image to videoAudio720p1080p
From ~420 creditsTry model →
Available27 Apr 2026

HappyHorse-1.0

alibaba

HappyHorse-1.0 is a video generation model for text-to-video and image-to-video workflows. It supports output at 720p or 1080p, clip durations from 3 to 15 seconds, seeded generation, watermark control, and first-frame image conditioning for image-to-video generation.

Text to videoImage to videoAudio720p1080p
From ~84 creditsTry model →
Available24 Apr 2026

SkyReels V4

skywork

SkyReels V4 is a unified multimodal video foundation model for joint video-audio generation, inpainting, and editing. It accepts text, images, video clips, masks, and audio references, and supports cinematic outputs up to 1080p, 32 FPS, and 15 seconds with synchronized audio, making it suitable for prompt-driven generation as well as guided editing workflows.

Text to videoImage to videoVideo to videoAudio to videoAudio720p1080p
From ~81 creditsTry model →
Available23 Apr 2026

Kling VIDEO 3.0 4K

klingai

Kling VIDEO 3.0 4K is the 4K variant of Kling VIDEO 3.0 for text-to-video and image-to-video generation. It extends the 3.0 series from 720p Standard and 1080p Pro into 4K output while keeping the same multimodal strengths: native audio generation, multi-shot sequencing, element consistency, prompt-driven scene control, and stable temporal coherence across longer clips.

Text to videoImage to videoAudio
Live pricing in studioTry model →
Available23 Apr 2026

Kling VIDEO 3.0 Omni 4K

klingai

Kling VIDEO 3.0 Omni 4K is the 4K variant of Kling VIDEO 3.0 Omni for text-to-video and image-to-video workflows. It raises the 3.0 Omni line from 720p Standard and 1080p Pro to 4K output while preserving the series strengths: native audio generation, reference-guided video creation, prompt-based editing, multi-shot structure, and stable subject consistency for more demanding cinematic and advertising workflows.

Text to videoImage to videoAudio4k
From ~242 creditsTry model →
Available3 Apr 2026

Seedance 2.0

bytedance

Seedance 2.0 is a unified multimodal audio-video generation model from ByteDance that accepts text, image, audio, and video inputs in combination, supporting up to 9 images, 3 video clips, and 3 audio clips as reference. It generates multi-shot videos up to 15 seconds with dual-channel synchronized audio including dialogue, ambient sound, and effects. It features physics-aware motion, improved controllability for video extension and editing, and strong instruction following for complex scene composition.

Text to videoImage to videoVideo to videoAudio to videoAudio1080p4k
Live pricing in studioTry model →
Available3 Apr 2026

Seedance 2.0 Fast

bytedance

Seedance 2.0 Fast is a speed-optimized variant of ByteDance's unified multimodal audio-video generation model. It accepts text, image, audio, and video inputs in combination, like Seedance 2.0, but targets shorter wall-clock times and higher throughput for iterative workflows. It produces multi-shot videos with dual-channel synchronized audio including dialogue, ambient sound, and effects, with physics-aware motion and editing controls, while prioritizing responsiveness over the last increment of visual refinement so teams can preview and ship ideas faster.

Text to videoImage to videoVideo to videoAudio to videoAudio480p720p
From ~76 creditsTry model →
Available3 Apr 2026

Wan2.7

alibaba

Wan2.7 is Alibaba's next-generation multimodal video model supporting text-to-video, image-to-video, reference-to-video, and video editing. It features multi-shot storytelling, subject-consistent multi-character generation, first-and-last-frame interpolation, video continuation, style transfer, instruction-based editing, and audio-conditioned generation with auto-dubbing. Output at 720p or 1080p, 30 FPS in multiple aspect ratios.

Text to videoImage to videoVideo to videoAudio720p1080p
Live pricing in studioTry model →
Available31 Mar 2026

Veo 3.1 Lite

google

Veo 3.1 Lite is the most cost-effective model in the Veo 3.1 family, designed for high-volume applications requiring rapid iteration. It supports text-to-video and image-to-video generation at 720p or 1080p in landscape and portrait formats, with customizable duration of 4, 6, or 8 seconds. It maintains the same generation speed as Veo 3.1 Fast at less than 50% of the cost, and includes native synchronized audio generation.

Text to videoImage to videoAudio720p1080p
From ~23 creditsTry model →
Available30 Mar 2026

PixVerse V6

pixverse

PixVerse V6 is a video generation model focused on multi-shot storytelling with native synchronized audio. It provides over 20 cinematic camera controls including focal length, aperture, depth of field, lens distortion, and vignetting. It features improved character consistency across shots using multi-image references, supports 1080p output at up to 15 seconds, and includes multilingual text rendering in frames.

Text to videoImage to videoAudio720p1080p
From ~26 creditsTry model →
Available26 Feb 2026

P-Video

prunaai

Pruna P-Video is a real-time AI video generation model designed for fast creative iteration and production workflows. It supports text-to-video, image-to-video, and audio-to-video through a unified endpoint, delivering up to 1080p at 48 FPS with integrated dialogue generation and audio import. The model emphasizes speed, cost efficiency, sequencing consistency across clips, and stable subject identity, making it well suited for brand content, multi-format distribution, and rapid draft-to-refine pipelines.

Text to videoImage to videoAudio to videoAudio720p1080p
From ~21 creditsTry model →
Available9 Feb 2026

Vidu Q3 Turbo

vidu

Vidu Q3 Turbo is a speed-optimized multimodal video generation model that produces short video clips with synchronized audio directly from text or images. It prioritizes fast inference and responsive iteration while preserving stable motion, coherent composition, and reliable audio alignment, making it suitable for rapid prototyping and production workflows where latency is critical.

Text to videoImage to videoAudio
From ~21 creditsTry model →
Available5 Feb 2026

Kling VIDEO 3.0

klingai

A balanced social-video model with stable motion, strong prompt following and optional synchronized audio.

Text to videoImage to videoAudio720p
From ~73 creditsTry model →
Available5 Feb 2026

Kling VIDEO 3.0 Omni Pro

klingai

Kling VIDEO 3.0 Omni Pro is a unified multimodal video model that generates HD clips from text or images with native audio output. It prioritizes detail, motion realism, and stable subject identity, and it supports reference-driven generation plus prompt-based video editing with strong temporal consistency.

Text to videoImage to videoAudio1080p
From ~81 creditsTry model →
Available5 Feb 2026

Kling VIDEO 3.0 Omni Standard

klingai

Kling VIDEO 3.0 Omni Standard is a cost-efficient version of the 3.0 Omni generation that produces HD video from text or images with native audio. It balances quality with speed and price, and it supports reference-based generation plus prompt-based video edits that preserve temporal stability across the clip.

Text to videoImage to videoAudio720p
From ~65 creditsTry model →
Available5 Feb 2026

Kling VIDEO 3.0 Pro

klingai

Kling VIDEO 3.0 Pro is a unified multimodal video model that generates high-quality video with synchronized audio from text or images. It supports reference-guided generation, prompt-based editing, fine control over motion and pacing, and stable temporal coherence for cinematic and narrative clips. Native audio output includes dialogue, ambient sound, and effects aligned to the visuals.

Text to videoImage to videoAudio
From ~117 creditsTry model →
Available1 Feb 2026

HeyGen Video Agent

heygen

HeyGen Video Agent is an AI video production model that generates complete, multi-scene videos from a single text prompt. It automates the full production pipeline — scriptwriting, avatar selection, shot planning, B-roll integration, motion graphics, captions, and editing — producing broadcast-ready videos with consistent branding. The agent supports customizable avatars, voice cloning, and iterative editing without full regeneration, enabling scalable video content creation for marketing, training, and social media.

Text to videoAudio
Live pricing in studioTry model →
Available30 Jan 2026

Vidu Q3

vidu

Vidu Q3 is a multimodal video generation model that creates video with synchronized audio directly from text or images, supports intelligent multi-shot sequencing, and produces complete outputs with stable visuals and embedded subtitles without post-processing.

Text to videoImage to videoAudio to videoAudio720p1080p
From ~38 creditsTry model →
Available29 Jan 2026

Grok Imagine Video

xai

Grok Imagine Video is a multimodal generative video model that produces short video clips with native audio from text descriptions or static images. It supports text-to-video and image-to-video generation with synchronized sound effects and dialogue, enabling developers to animate scenes with motion, camera dynamics, and audio in a single API workflow.

Text to videoImage to videoVideo to videoAudio480p720p
From ~41 creditsTry model →
Available26 Jan 2026

PixVerse V5.6

pixverse

PixVerse V5.6 is an upgraded video generation model that improves visual stability, motion clarity, and audio-visual alignment over previous versions. It supports text-to-video and image-to-video generation with optional native audio, delivering more accurate multi-character lip-sync, cleaner motion in complex scenes, and more natural speech and environmental sound for single-shot cinematic outputs.

Text to videoImage to videoAudio720p1080p
From ~31 creditsTry model →
Available16 Dec 2025

Wan2.6

alibaba

Wan2.6 is a multimodal video model for text to video and image to video generation with support for multi-shot sequencing and native sound. It emphasizes temporal stability, consistent visual structure across shots, and reliable alignment between visuals and audio in short form video generation.

Text to videoImage to videoAudio720p1080p
From ~288 creditsTry model →
Available11 Dec 2025

Kling VIDEO 2.6 Pro

klingai

Kling VIDEO 2.6 Pro is a full audio-visual AI video model that combines cinematic-quality video generation with native audio (dialogue, sound effects, ambience). It supports flexible workflows from text or image input, delivering synchronized video and sound in one pass with strong consistency and creative control. Via the API, Motion Control enables creators to guide character movement using a reference video for more realistic and physically grounded motion.

Text to videoImage to videoAudio
From ~81 creditsTry model →
Available11 Dec 2025

Kling VIDEO 2.6 Standard

klingai

Kling VIDEO 2.6 Standard is a high-quality AI video generation model focused on producing visually coherent short clips with stable motion, expressive camera movement, and strong prompt adherence. It generates video from text prompts or an optional input image, making it suitable for cinematic previews, social content, and creative prototyping where audio is not required.

Text to videoImage to video
From ~38 creditsTry model →
Available1 Dec 2025

Kling VIDEO O1 Pro

klingai

Kling VIDEO O1 Pro is a unified multimodal video foundation model for controllable generation and instruction based editing. It supports text prompts, visual references, and video input so developers can build high control pipelines for pacing, transitions, object changes, and style revisions.

Text to videoImage to videoVideo to video
From ~65 creditsTry model →
Available1 Dec 2025

Kling VIDEO O1 Standard

klingai

Kling VIDEO O1 Standard is a unified multimodal video model for controllable generation and instruction-based editing. It supports text prompts, image references, and video input to enable precise control over motion, transitions, object changes, and visual adjustments within short-form video workflows.

Text to videoImage to videoVideo to video
From ~49 creditsTry model →
Available1 Dec 2025

PixVerse V5.5

pixverse

PixVerse V5.5 is a director focused video model for story driven clips. It supports multi image fusion for character continuity, multi shot sequences, and native audio. It delivers smooth motion, refined cinematic control, and precise text guided video generation for complex scenes.

Text to videoAudio720p1080p
Live pricing in studioTry model →
Available1 Dec 2025

Runway Gen-4.5

runway

Cinematic realistic motion and strong composition for premium visual storytelling.

Text to videoImage to video720p
From ~70 creditsTry model →
Available28 Oct 2025

MiniMax Hailuo 2.3

minimax

MiniMax Hailuo 2.3 is a cinematic video model for short form production. It accepts text prompts or image inputs and outputs 6 or 10 second clips at 768p or 1080p. It focuses on consistent motion, strong physics, and stable scenes for ads, social content, and creative shots.

Text to videoImage to video
Live pricing in studioTry model →
Available25 Oct 2025

Seedance 1.0 Pro Fast

bytedance

Seedance 1.0 Pro Fast accelerates the core Seedance pipeline for expressive dance and performance clips. It turns text prompts or reference images into smooth, cinematic motion with strong temporal consistency. Ideal for rapid iteration in creative tools and production workflows.

Text to videoImage to video720p1080p
Live pricing in studioTry model →
Available23 Oct 2025

KlingAI 2.5 Turbo Standard

klingai

KlingAI 2.5 Turbo Standard is a streamlined image to video model tuned for speed and cost efficiency. It generates smooth cinematic clips with strong motion control and clear frames at up to 720p. Ideal for rapid iteration in creative pipelines and production tests.

Image to video
From ~25 creditsTry model →
Available15 Oct 2025

Veo 3.1

google

Veo 3.1 is a cinematic video generation model for developers. It turns text prompts or reference images into high fidelity scenes with richer native audio, better prompt adherence, and granular shot control. Use it for story driven clips with smoother motion and consistent style.

Text to videoImage to videoAudio to videoAudio1080p4K
Live pricing in studioTry model →
Available15 Oct 2025

Veo 3.1 Fast

google

Veo 3.1 Fast is a high speed variant of Veo 3.1 for rapid creative iteration. It supports text prompts, image prompts, and reference images. It targets low latency workflows while keeping cinematic quality for short form and multi shot video generation with native audio.

Text to videoImage to videoAudio1080p4K
From ~69 creditsTry model →
Available28 Sept 2025

Wan2.5-Preview

alibaba

Wan2.5-Preview is Alibaba’s multimodal video model in research preview. It supports text to video and image to video with native audio generation for clips around 10 seconds. It offers strong prompt adherence, smooth motion, and multilingual audio for narrative scenes.

Text to videoImage to videoAudio to videoAudio720p1080p
From ~53 creditsTry model →
Available23 Sept 2025

KlingAI 2.5 Turbo Pro

klingai

KlingAI 2.5 Turbo Pro is a high performance video generation model for cinematic work. It converts prompts or stills into smooth 1080p clips with strong motion, precise camera control and tight prompt adherence. Ideal for creative tools, ads, trailers and sports scenes.

Text to videoImage to video
From ~41 creditsTry model →
Available16 Sept 2025

PixVerse V5 Fast

pixverse

PixVerse V5 Fast is an optimized variant of PixVerse v5 designed for faster video generation and lower latency. It supports text to video and image to video workflows while prioritizing speed and responsiveness, making it suitable for rapid iteration and preview-focused pipelines where audio, templates, and advanced controls are not required.

Text to videoImage to video720p1080p
Live pricing in studioTry model →
Available29 Aug 2025

PixVerse V5

pixverse

PixVerse V5 generates high fidelity video from text prompts or single images. It delivers smooth motion and sharp cinematic frames with strong prompt alignment. Ideal for creators who need fast iteration, keyframe control, and consistent style across shots.

Text to videoImage to video720p1080p
Live pricing in studioTry model →
Available27 Aug 2025

HeyGen Avatar IV

heygen

HeyGen Avatar IV is a photorealistic AI avatar generation model that creates talking videos from a single image and a script or audio input. The model synchronizes voice with facial motion, expressions, and gestures to produce lifelike avatar performances. It supports multilingual speech, realistic lip synchronization, and expressive body language, enabling scalable production of presenter-style videos without cameras, actors, or studio setups.

Text to videoImage to videoAudio to videoAudio720p1080p
From ~233 creditsTry model →
Available19 Jun 2025

MiniMax Hailuo 02

minimax

MiniMax Hailuo 02 is a 1080p AI video model for cinematic, high motion scenes. It converts text prompts or still images into short, polished clips with strong instruction following and realistic physics. Ideal for commercial spots, trailers, music promos, and social shorts.

Text to videoImage to video
Live pricing in studioTry model →
Available1 Jun 2025

Seedance 1.0 Pro

bytedance

Seedance 1.0 Pro is a ByteDance video model for 5 to 10 second clips at up to 1080p. It supports text prompts and image first frames. It delivers smooth motion with strong temporal consistency. Ideal for multi shot storytelling, ads, and design previews in real time pipelines.

Text to videoImage to video480p1080p
Live pricing in studioTry model →
Available14 May 2025

PixVerse V4.5

pixverse

PixVerse V4.5 generates stylized cinematic video from text prompts or reference images. It adds refined camera motion control, multi image fusion, and faster modes for iteration. Ideal for creators who need dynamic shots, complex motion, and consistent stylized outputs.

Text to videoImage to video720p1080p
From ~35 creditsTry model →
Available21 Apr 2025

Vidu Q1

vidu

Vidu Q1 is a generative video model that preserves visual fidelity from multiple reference images. It supports character, scene and prop control with smooth transitions and 1080p clips. Ideal for ads, story sequences and animation workflows that need tight visual continuity.

Text to videoImage to video
From ~26 creditsTry model →
Available1 Mar 2025

PixVerse V3.5

pixverse

PixVerse V3.5 provides basic text to video generation with support for visual effects and limited subject motion. It targets short clips for experiments or prototypes. Camera movement is not available, which simplifies control and integration in pipelines.

Text to video720p1080p
Live pricing in studioTry model →
Available24 Feb 2025

PixVerse V4

pixverse

PixVerse V4 is a generative video model for text prompts or source images. It improves motion quality and complex camera movement. It adds motion modes, sound effect sync, and style transfer. Ideal for short cinematic clips and rapid creative iteration in production pipelines.

Text to videoImage to videoVideo to video720p1080p
From ~35 creditsTry model →
Available28 Jan 2025

MiniMax 01 Director

minimax

MiniMax 01 Director generates short cinematic video clips from text prompts with director level control. It supports detailed camera movement instructions, stable framing, and reduced motion randomness. Ideal for film previz, ads, and story beats inside production tools.

Text to video
From ~33 creditsTry model →
Available1 Sept 2024

MiniMax 01

minimax

MiniMax 01 is a compact text to video model for short clips. It turns simple prompts into 720p videos with smooth motion and cinematic framing. It targets fast iteration and stable output so developers can prototype interactive video features and creative tools with low latency.

Text to video
From ~33 creditsTry model →
One workflow

Choose a model—or let INXSocial choose for you.

01

Describe the video

Start with the creative brief and add source or reference imagery when the selected mode supports it.

02

AI Recommended or Choose Model

Use AI Recommended for quality-to-cost routing or browse the generation-ready catalogue yourself.

03

Generate, save and publish

Use model-specific settings, monitor background generation, keep the result in Media Library and move it into Posts or scheduling.

AI Video Studio + social publishing

Stop switching tools every time a new video model launches.

INXSocial keeps model discovery, generation, media management, scheduling and publishing in one connected product.