Models & Pricing
Find the right model for your use case. Transparent pricing — pay only for what you use.
Model List
Featured models
Take a look at these crowd-favorite models.
Z-Image-Turbo
Fast text-to-image with sub-second inference and accurate in-image text rendering.
Motif-Video-T2V-2B
Generate videos from text prompts with Motif-Video-2B.
Motif-Video-I2V-2B
Generate videos from an input image + text prompt with Motif-Video-2B.
SAM-3D-Object
Reconstruct 3D object models from images using SAM 3D Object.
Cosmos-Transfer2.5 General
Multi-control video-to-video generation for Sim2Real transformation. Transform simulation videos to photorealistic using depth, edge, segmentation, and visibility controls.
Supertonic2 Multilingual TTS
Multilingual text-to-speech synthesis supporting Korean, English, Spanish, Portuguese, and French with multiple voice styles.
VOID-Inpaint-V1
Remove objects from video using VOID's quadmask flow or auto SAM3 + Gemini masking.
Model list
Qwen-Image-Edit-LoRA
LoRA-enhanced image editing with precise control and flexible style adaptation.
Qwen-Image
Text‑centric text‑to‑image with sharp glyph rendering and stable layout.
Qwen-Image-LoRA
LoRA-tuned text-to-image with enhanced style control and fine-grained customization.
Qwen-Image-Edit
Delivers precise image edits guided by text prompts, supporting nuanced visual modifications. Output image size can be slightly resized from the input image.
Whisper Large V3 Turbo (Korean ASR)
Automatic speech recognition model fine-tuned for Korean. Transcribes audio files to text with high accuracy.
SAM-3D-Body
Reconstruct 3D human body models from images using SAM 3D Body.
SAM3-Auto-Image
Automatic Segmentation for images using SAM3.
SAM3-PVS-Image
Promptable Visual Segmentation for images using SAM3.
SAM3-PVS-Video
Promptable Visual Segmentation for videos using SAM3.
Wan2.2-Animate-Replace
Replace characters in videos while preserving the original background and environment.
Wan2.2-Animate-Move
Transfer a reference character into a motion video, replacing both character and scene.
Wan2.2-I2V-A14B
High-quality 14B MoE image-to-video with dual-expert denoising architecture.
Wan2.2-T2V-A14B
High-quality 14B MoE text-to-video with dual-expert denoising architecture.
Wan2.2-FLF2V-A14B
Interpolate smooth video between first and last keyframe images with text guidance.
Wan2.2-I2V-5B
Lightweight 5B image-to-video optimized for fast inference on consumer GPUs.
Prompt-Enhancer
Expand short text prompts into detailed, optimized descriptions for image generation.
SAM3-PCS-Image
Promptable Concept Segmentation for images using SAM3.
SAM3-PCS-Video
Promptable Concept Segmentation for videos using SAM3.
ERNIE-Image-Turbo
Fast text-to-image generation by Baidu. DMD and RL optimized 8B DiT model achieving high-quality results in only 8 inference steps. Excels at complex instruction following, text rendering, and structured image generation.
Nucleus-Image
Sparse MoE text-to-image with 17B total parameters (2B active per pass).