Skip to main content
Getting Started

Choose models

Match your specific use case

Text generation

From most capable to most cost-effective — pick what fits your use case. More

Qwen models

Third-party models

Same API format as Qwen models.

Image & video

Understanding

Extract text descriptions or structured data from images and videos. More

Generation

Generate images and videos from text or images, with support for editing, reference, and high-resolution output. More

World models

Build interactive digital worlds with Adventure, Directing, and Acting modes. More

Audio & speech

Text-to-speech

For audiobook reading, voice broadcasting, virtual avatars, and more. More

Speech recognition

Dedicated ASR and LLM-based approaches — choose based on accuracy and flexibility. More

Speech-to-speech

End-to-end voice conversation without separate ASR and TTS calls. More

Omni

Integrates understanding and generation capabilities across text, image, audio, and video modalities. More

Embeddings & reranking

Convert text or multimodal content into vectors, combined with reranking to improve retrieval accuracy. More

Decision models

Structured decision models for high-frequency business decisions. A single forward pass performs classification, yes-or-no decisions, and scoring, returning probability distributions and confidence.

View all models

Model Marketplace

Explore our complete collection of text, image, video, audio, and embedding models with detailed specifications and pricing.