Skip to main content
Getting Started

Choose models

Match your specific use case

Text generation

From most capable to most cost-effective — pick what fits your use case. More

Qwen models

Third-party models

Same API format as Qwen models.

Image & video

Understanding

Extract text descriptions or structured data from images and videos. More

Generation

Generate images and videos from text or images, with support for editing, reference, and high-resolution output. More

Audio & speech

Text-to-speech

For audiobook reading, voice broadcasting, virtual avatars, and more. More

Speech recognition

Dedicated ASR and LLM-based approaches — choose based on accuracy and flexibility. More

Speech-to-speech

End-to-end voice conversation without separate ASR and TTS calls. More

Omni

Integrates understanding and generation capabilities across text, image, audio, and video modalities. More

Embeddings & reranking

Convert text or multimodal content into vectors, combined with reranking to improve retrieval accuracy. More

View all models

Model Catalog

Explore our complete collection of text, image, video, audio, and embedding models with detailed specifications and pricing.