AI Models Directory
GPT-4o
OpenAIOpenAI's flagship multimodal model that can reason across text, images, and audio. Offers strong performance across a wide range of tasks with fast response times.
GPT-4.1
OpenAIThe latest evolution of the GPT-4 family with improved instruction following, coding, and long-context performance. Supports up to 1M token context window.
o3
OpenAIOpenAI's advanced reasoning model using chain-of-thought to solve complex problems in math, science, and coding.
DALL-E 3
OpenAIOpenAI's latest image generation model integrated with ChatGPT. Produces high-quality images from text descriptions with improved prompt following.
Whisper
OpenAIOpen-source speech recognition model trained on 680,000 hours of multilingual audio. Capable of transcription and translation across 99 languages.
Claude Mythos Preview
AnthropicAnthropic's most powerful frontier model to date, internally codenamed 'Capybara'. Uses Mixture of Experts architecture with an estimated ~10 trillion parameters. Restricted to vetted organizations through Project Glasswing due to extraordinary cybersecurity capabilities deemed too dangerous for public release.
Claude Opus 4
AnthropicAnthropic's most capable model, excelling at complex analysis, nuanced writing, coding, and multi-step reasoning. Known for strong safety alignment.
Claude Sonnet 4
AnthropicAnthropic's balanced model offering strong performance at lower cost than Opus. Ideal for most production workloads.
Claude 3.5 Haiku
AnthropicAnthropic's fastest and most affordable model. Designed for high-throughput, low-latency tasks like classification and extraction.
Gemini 2.5 Pro
GoogleGoogle's most capable model with a massive 1M token context window. Excels at long-document analysis and multimodal tasks.
Gemini 2.5 Flash
GoogleGoogle's fast, efficient model optimized for speed and cost while maintaining strong quality.
Gemma 3
GoogleGoogle's open-weight model family derived from Gemini research. Available in sizes from 1B to 27B for on-device and self-hosted deployments.
Llama 4 Maverick
MetaMeta's latest open-weight Mixture-of-Experts model with 400B total parameters (17B active). Excellent multilingual performance.
Llama 3.3 70B
MetaMeta's high-performance dense model offering near-405B quality at a fraction of the cost. Popular for self-hosting and fine-tuning.
Mistral Large
Mistral AIMistral's flagship model with 128K context window. Strong at multilingual tasks, code generation, and reasoning.
Mistral Small
Mistral AIMistral's cost-efficient open-weight model optimized for simple tasks and high-throughput applications. Apache 2.0 license.
Codestral
Mistral AIMistral's specialized code generation model optimized for code completion and understanding across 80+ programming languages.
DeepSeek V3
DeepSeekPowerful open-source MoE model with 671B total parameters. Known for exceptional cost-efficiency and competitive performance.
DeepSeek R1
DeepSeekDeepSeek's reasoning model rivaling OpenAI o1. Uses chain-of-thought for complex math, science, and coding. Fully open source.
Grok 3
xAIxAI's flagship model with strong reasoning and real-time information access via X (Twitter). Known for less restrictive responses.
Command R+
CohereCohere's enterprise-focused model optimized for RAG, tool use, and business applications. Strong at grounded generation with citations.
Stable Diffusion 3.5
Stability AIThe most popular open-source image generation model with extensive customization through LoRAs, ControlNet, and community models.
Midjourney v6
MidjourneyLeading AI image generation service known for stunning artistic quality. Widely used by artists, designers, and creative professionals.
FLUX.1
Black Forest LabsState-of-the-art open-source image generation model from the creators of Stable Diffusion. Excellent prompt following and quality.
Sora
OpenAIOpenAI's video generation model capable of creating realistic scenes from text. Can generate videos up to 60 seconds.
text-embedding-3-large
OpenAIOpenAI's most capable embedding model for converting text into vectors. Used for semantic search, clustering, and RAG.
Qwen 2.5 72B
Alibaba CloudAlibaba's flagship open-source model with strong multilingual performance, excelling in Chinese and English.
Phi-4
MicrosoftMicrosoft's compact but powerful model designed to maximize quality-per-parameter. Excels at reasoning and coding despite its small size.
GPT-4o Mini
OpenAIOpenAI's small, affordable, and fast model. Outperforms GPT-3.5 Turbo at a fraction of GPT-4o's cost. Ideal for lightweight tasks.
GPT-4.1 Mini
OpenAISmaller, faster, cheaper variant of GPT-4.1 balancing cost and performance. Strong instruction following at a fraction of the cost.
GPT-4.1 Nano
OpenAIThe smallest and fastest GPT-4.1 variant. Optimized for speed and cost, suitable for simple tasks at massive scale.
o4-mini
OpenAIOpenAI's cost-efficient reasoning model. Brings chain-of-thought reasoning at a fraction of o3's cost, ideal for applications that need reasoning without the premium price.
o3-mini
OpenAISmaller, cheaper version of OpenAI's o3 reasoning model. Good balance of reasoning ability and cost for everyday tasks requiring logical thinking.
Claude Haiku 4
AnthropicAnthropic's fastest and most affordable model. Compact yet capable, ideal for lightweight tasks, real-time responses, and cost-sensitive deployments.
Gemini 2.0 Flash
GoogleGoogle's fast, efficient model optimized for high-throughput applications. Strong multimodal capabilities with very low latency and generous free tier.
Gemma 2
GoogleGoogle's open-source lightweight model family. Efficient and capable, designed for on-device and resource-constrained deployments.
Llama 4 Scout
MetaMeta's efficient open-source model using mixture-of-experts architecture. 17B active parameters from 109B total, offering strong performance at low compute cost.
Grok 3 Mini
xAIxAI's compact reasoning model with chain-of-thought capabilities. Fast and affordable alternative to full Grok 3 with built-in reasoning.
ElevenLabs
ElevenLabsLeading AI voice synthesis platform offering realistic text-to-speech, voice cloning, and dubbing in 32 languages.
Perplexity Sonar
PerplexityAI search model that combines LLM capabilities with real-time web search. Returns answers with citations and up-to-date information.
Recraft V3
RecraftState-of-the-art image generation model excelling at design-quality output. Top performer on ELO rankings for image generation quality.
Mistral Medium
Mistral AIMistral's balanced model offering strong performance between Small and Large. Good all-rounder for diverse enterprise tasks with European data residency.
Jina Embeddings v3
Jina AIHigh-performance multilingual embedding model supporting 89 languages. Optimized for search, retrieval, and RAG applications.
Kimi
Moonshot AIChinese AI model with one of the longest context windows available — up to 2M tokens. Strong multilingual capabilities with a focus on long-document understanding.
Kling 2.5 Turbo Pro
KuaishouLatest video generation model from Kuaishou with 1080p HDR output. Produces high-quality videos from text or images with superior motion, physics, and facial expression understanding.
Google Imagen 3
GoogleGoogle's highest quality image generation model. Produces photorealistic images with excellent text rendering and prompt adherence.
Amazon Nova Pro
AmazonAmazon's highly capable multimodal model offering strong performance across text, image, and video understanding. Optimized for accuracy, speed, and cost on AWS.
Yi-Lightning
01.AI01.AI's fast and capable model rivaling GPT-4o class performance. Strong on benchmarks with competitive pricing and fast inference.
Qwen 3
AlibabaAlibaba's latest open-source model family with dense and MoE variants. Strong multilingual performance rivaling frontier closed-source models.
GPT-5.2
OpenAIOpenAI frontier model for complex professional work with a 400K token context window and 128K max output. Strong reasoning, coding, and multimodal capabilities.
Claude Opus 4.5
AnthropicAnthropic model that handles ambiguity and complex multi-system bugs without hand-holding. Excels at long-horizon autonomous tasks with sustained reasoning and fewer dead-ends.
Gemini 3 Pro
GoogleGoogle's reasoning-focused model with internal chain-of-thought. Breaks problems into steps, checks its own logic, and refines output. Supports text, vision, and native image generation.
Gemini 3 Nano Banana
Google DeepMindGoogle DeepMind image generation model built on Gemini 3. Offers personalized image generation using Google Photos context and high-fidelity text rendering for professional asset production.
GPT-Image-1
OpenAIOpenAI's natively multimodal image generation model. Accepts both text and image inputs through a unified transformer, enabling seamless text-to-image and image editing.
FLUX.2 Pro
Black Forest LabsNext-generation 32B parameter image generation model with up to 4 megapixel output. Features the Kontext Engine for natural language image editing and dramatically improved text rendering.
Grok 4
xAIxAI's model with video and image generation capabilities via the Imagine API. Supports text-to-video, image-to-video, and synchronized audio generation at 720p.
Gemini Flash (Audio)
GoogleGoogle Gemini Flash optimized for audio processing tasks. Transcribes, analyzes, and reasons about audio content with fast response times and competitive pricing.
Gemini Live API
GoogleGoogle's real-time multimodal streaming API for voice-in, voice-out conversational AI. Supports low-latency audio interaction with function calling mid-conversation.
OpenAI TTS
OpenAIOpenAI's text-to-speech API with multiple quality tiers and 13+ voices. Converts text to natural-sounding speech with real-time streaming support.
Hume
Hume AIEmotionally intelligent voice AI platform with the Empathic Voice Interface (EVI). Offers expressive TTS with emotion detection, natural language tone control, and under 200ms latency.
Kokoro-TTS
hexgradLightweight open-source TTS model with 82M parameters delivering quality comparable to commercial models. Supports 9 languages, 54+ voices, and runs on both GPU and CPU.
Riverflow 2.0 Pro
SourcefulTop-ranked image generation model on Artificial Analysis leaderboard. Produces up to 4K resolution images with precise text rendering, custom fonts, and self-correction pipeline.
Seedream 4
ByteDanceByteDance's fast and cost-effective image generation model. Generates images at ~1.8 seconds with strong text rendering accuracy and multi-reference support.
Veo
Google DeepMindGoogle DeepMind's video generation model family. Generates high-quality videos from text prompts with audio, available in multiple tiers from Lite to Standard quality.
GPT-5.4
OpenAIOpenAI's most capable frontier model optimized for agentic, coding, and professional workflows. Features native computer-use capabilities, 1M token context, and the most token-efficient reasoning of the GPT-5 family.
GPT-5.3-Codex
OpenAIOpenAI's most capable agentic coding model combining frontier coding performance with deep reasoning. Handles long-running tasks involving research, tool use, and complex execution. Can be steered interactively while working.
text-embedding-3-small
OpenAIOpenAI's efficient and affordable embedding model for converting text into 1536-dimensional vectors. Ideal for semantic search, clustering, and RAG with a 5x price reduction over ada-002.
Gemini 3.1 Flash Image Preview
GoogleGoogle's high-efficiency image generation and editing model delivering Pro-level visual quality at Flash speed. Supports text-to-image, image-to-image editing, multi-turn conversational editing, and up to 4K resolution output.
Claude Fable 5
AnthropicAnthropic's most capable widely available model, built for the most demanding reasoning and long-horizon agentic work. State-of-the-art across nearly all AI benchmarks, with exceptional performance in software engineering, scientific research, and vision tasks.
Claude Opus 4.8
AnthropicAnthropic's recommended model for complex agentic coding and enterprise work. Offers a 1M context window and adaptive thinking that defaults to high effort, making it the current flagship for demanding multi-step tasks.
Claude Sonnet 5
AnthropicAnthropic's best combination of speed and intelligence. 1M token context with adaptive thinking and fast response times, at $3 per million input tokens (introductory $2 through Aug 31, 2026).
GPT-5.5
OpenAIOpenAI's most capable model, designed for long-horizon agentic tasks. Smarter and more token-efficient than GPT-5.4, excelling at coding, research, data analysis, and computer use.
Gemini 3.5 Flash
GoogleGoogle's high-efficiency multimodal model bringing near-Pro level coding and reasoning at Flash-tier cost. Optimized for agentic execution loops, supports text, image, video, audio, and PDF inputs with configurable thinking levels.