Match your specific use case
Text generation
From most capable to most cost-effective — pick what fits your use case. More
Qwen models
qwen3.8-max
Complex reasoning and coding
qwen3.7-plus
Balanced performance, speed, and cost
qwen3.7-flash
Fast and cost-effective
Third-party models
Same API format as Qwen models.
deepseek-v4-pro-0813
Text generation
deepseek-v4-flash-0731
Text generation
kimi-k2.7-code
Text generation
glm-5.2
Text generation
Image & video
Understanding
Extract text descriptions or structured data from images and videos. More
qwen3.8-max
Visual understanding
qwen3.7-plus
Visual understanding
qwen3.5-omni-plus
Visual understanding
kimi-k2.7-code
Visual understanding
Generation
Generate images and videos from text or images, with support for editing, reference, and high-resolution output. More
qwen-image-3.0-pro
Image generation
wan2.7-image-pro
Image generation
happyhorse-1.1-t2v
Video generation
happyhorse-1.1-i2v
Video generation
happyhorse-1.1-r2v
Video generation
happyhorse-1.0-video-edit
Video generation
Audio & speech
Text-to-speech
For audiobook reading, voice broadcasting, virtual avatars, and more. More
Speech recognition
Dedicated ASR and LLM-based approaches — choose based on accuracy and flexibility. More
qwen-audio-3.0-asr-flash-streaming
Speech recognition
qwen-audio-3.0-asr-flash-filetrans
Speech recognition
qwen3.5-omni-plus-realtime
Speech recognition
qwen3.5-omni-plus
Speech recognition
Speech-to-speech
End-to-end voice conversation without separate ASR and TTS calls. More
Omni
Integrates understanding and generation capabilities across text, image, audio, and video modalities. More
Embeddings & reranking
Convert text or multimodal content into vectors, combined with reranking to improve retrieval accuracy. More
View all models
Model Catalog
Explore our complete collection of text, image, video, audio, and embedding models with detailed specifications and pricing.