Skip to content
  • Models
  • Rankings
  • Ori
ElevenLabs launch offer: every ElevenLabs model is 50% off through October 19, 2026. See ElevenLabs models
OpenRouterOpenRouter
© 2026 OpenRouter, Inc

Product

  • Chat
  • Rankings
  • Benchmarks
  • Apps
  • Discover
  • Models
  • Ori
  • Collections
  • Providers
  • Tools
  • Pricing
  • Business
  • Enterprise
  • Labs

Company

  • About
  • Blog
  • Careers
    Hiring
  • Privacy
  • Terms of Service
  • Trust Center
  • Support
  • Works With OR
  • Data
  • Brand

Developer

  • Documentation
  • API Reference
  • Developer Platform
  • Status
  • AI Site Map

Connect

  • Discord
  • GitHub
  • LinkedIn
  • X
  • YouTube

Models

CompareDiscover Models
Favicon for anthropic
Favicon for openai

Models

CompareDiscover Models
Favicon for anthropic
Favicon for openai
  • Favicon for microsoft
    Microsoft: Microsoft-Decision-1Microsoft-Decision-1
    213M tokens

    Microsoft-Decision-1 is a small model built for fast decision-making. Instead of generating text, it reads the provided content and returns a calibrated probability for each fixed answer option, so the response carries its own confidence and can determine when an application acts, defers, or asks for review. It is post-trained from Qwen3.5-9B for single-pass scoring and suited to classification, routing, prioritization, verification, workflow control, agent guardrails, and AI judging. It is not intended for open-ended generation, conversation, translation, or summarization. Weights are updated continually while the API shape stays the same. Learn more in Microsoft's announcement: https://commandline.microsoft.com/microsoft-decision-1-model-foundry/

    by microsoftOct 9, 202633K context$0.042/M input tokens$0/M output tokens
  • Favicon for nace-ai
    Nace.AI: Drex v1.5Drex v1.5
    250M tokens

    Drex 1.5 is a small decision model from Nace AI, served over the same /v1/systemone contract as TypeSafe's Jev. Send a state and typed questions (yes/no, multiple choice, or score) and it returns a probability for every option in one forward pass, with no generated text. With under 10B parameters and 128K tokens of context, it suits routing, classification, and policy checks over long documents that need a fast, scored answer instead of prose. Open weights and runtime: https://github.com/nace-ai/drex-decision-models

    by nace-aiOct 9, 2026131K context$0.04/M input tokens$0/M output tokens
  • Favicon for cloudflare
    Cloudflare: Clef OmniClef Omni
    54.3M tokens

    Clef Omni is the mixture-of-experts member of Cloudflare's open-source Clef decision model family, a fine-tune of Qwen3-Omni-30B-A3B (30B total, 3B active parameters) served on Workers AI. It turns a state plus a schema of typed questions into decisions, returning a calibrated probability for every allowed option of every question in a single forward pass instead of generating tokens. Use it for classification, routing, scoring, and guardrails over text, JSON, and images through the Decisions API.

    by cloudflareOct 9, 202666K context$0.15/M input tokens$0/M output tokens
  • Favicon for stepfun
    StepFun: Step 5 PreviewStep 5 Preview
    3.2T tokens
    Programming (#31)

    Step 5 Preview is StepFun's flagship model for agentic work, built on a sparse Mixture-of-Experts architecture (27B active / 600B total parameters). It performs strongly in software engineering and professional knowledge work, with particular strength in finance. It is designed for extended tasks that span large codebases and documents, using tools and refining results over multiple steps.

    by stepfunOct 8, 20261M context$1/M input tokens$2.70/M output tokens
  • Favicon for upstage
    Upstage: Solar Decide FlashSolar Decide Flash
    50% off
    1.91B tokens

    Solar Decide Flash is Upstage's low-latency structured decision model, a faster variant of Solar Decide built on Solar Mini 4 and served through the System One (/v1/systemone) API. Instead of generating text, it reads a state and answers typed questions, returning a choice, a score, or a yes/no answer, each with a probability taken directly from the model. Each decision is a single forward pass, and output tokens are free. It is tuned for consistently fast response times in routing, classification, and policy checks. It keeps the 512K context window, so a full document can serve as the state, and carries Solar Mini 4's Korean-language strength.

    by upstageOct 8, 2026524K context$0.05/M input tokens$0/M output tokens
  • Favicon for anthropic
    Anthropic: Claude Haiku 5.5 (batch)Claude Haiku 5.5 (batch)Batch variant
    745M tokens

    Claude Haiku 5.5 is Anthropic's small, fast model for high-volume, cost-sensitive work such as summarization, subagents, and browser use. It succeeds Claude Haiku 4.5 with stronger coding, computer use, and knowledge work, and accepts text and image input with a 1M-token context window. It is the first Haiku model with adjustable effort. Thinking is adaptive and on by default, with effort as the main lever for trading off depth, latency, and cost; it can be turned off at low, medium, and high effort.

    by anthropicOct 7, 20261M context$0.05/M input tokens$0.25/M output tokens
  • Favicon for anthropic
    Anthropic: Claude Haiku 5.5Claude Haiku 5.5
    655B tokens
    Marketing (#41)
    SEO (#23)
    Programming (#44)
    Translation (#45)

    Claude Haiku 5.5 is Anthropic's small, fast model for high-volume, cost-sensitive work such as summarization, subagents, and browser use. It succeeds Claude Haiku 4.5 with stronger coding, computer use, and knowledge work, and accepts text and image input with a 1M-token context window. It is the first Haiku model with adjustable effort. Thinking is adaptive and on by default, with effort as the main lever for trading off depth, latency, and cost; it can be turned off at low, medium, and high effort.

    by anthropicOct 7, 20261M context$0.10/M input tokens$0.50/M output tokens
  • Favicon for perplexity
    Perplexity: Decider V1.1 27BDecider V1.1 27B
    34.3B tokens

    Decider V1.1 27B is a new checkpoint of Perplexity's decision model, succeeding Decider V1 27B with the same API contract. Instead of generating text, it reads content passed as state (text, JSON, or images) and returns typed, probabilistic answers to one or more named questions in a single request: the probability of yes for a yes/no question (noul), a probability for every option plus the most likely one (choice), or a probability for every level of an ordered rubric plus the expected score (score). It is built for classification, routing, moderation, and rubric grading where application code thresholds the returned numbers rather than parsing a chat reply. A request can carry up to 128 questions about the same content.

    by perplexityOct 7, 2026262K context$0.02/M input tokens$0/M output tokens
  • Favicon for elevenlabs
    ElevenLabs: Eleven v4Eleven v4
    50% off
    2.55M tokens

    Eleven v4 is a text-to-speech model from ElevenLabs. It is ElevenLabs' most expressive model, with inline audio tags for emotional and delivery control, support for 90+ languages, and a 10,000-character request limit.

    by elevenlabsOct 7, 2026$40/M characters
  • Favicon for elevenlabs
    ElevenLabs: Eleven v4 TurboEleven v4 Turbo
    50% off
    1.55M tokens

    Eleven v4 Turbo is a low-latency text-to-speech model from ElevenLabs. It keeps Eleven v4's expressive delivery and audio tags while being tuned for faster generation, with support for 90+ languages and a 10,000-character request limit.

    by elevenlabsOct 7, 2026$20/M characters
  • Favicon for elevenlabs
    ElevenLabs: Eleven v3Eleven v3
    50% off
    311K tokens

    Eleven v3 is a text-to-speech model from ElevenLabs. It produces emotionally rich, highly expressive speech with inline audio tags, supports 70+ languages, and has a 5,000-character request limit.

    by elevenlabsOct 7, 2026$40/M characters
  • Favicon for elevenlabs
    ElevenLabs: Eleven v3 ConversationalEleven v3 Conversational
    50% off
    71K tokens

    Eleven v3 Conversational is a text-to-speech model from ElevenLabs, a variant of Eleven v3 optimized for natural dialogue in conversational agents. It supports 70+ languages and has a 5,000-character request limit.

    by elevenlabsOct 7, 2026$20/M characters
  • Favicon for elevenlabs
    ElevenLabs: Eleven Flash v2.5Eleven Flash v2.5
    50% off
    319K tokens

    Eleven Flash v2.5 is an ultra-low-latency text-to-speech model from ElevenLabs. It is suited for conversational and real-time use cases, supports 32 languages, and has a 40,000-character request limit.

    by elevenlabsOct 7, 2026$20/M characters
  • Favicon for elevenlabs
    ElevenLabs: Eleven Turbo v2Eleven Turbo v2
    50% off
    15K tokens

    Eleven Turbo v2 is an English-only, low-latency text-to-speech model from ElevenLabs. It is suited for developer use cases where speed matters and only English is needed, and has a 30,000-character request limit.

    by elevenlabsOct 7, 2026$20/M characters
  • Favicon for elevenlabs
    ElevenLabs: Eleven Multilingual v2Eleven Multilingual v2
    50% off
    366K tokens

    Eleven Multilingual v2 is a text-to-speech model from ElevenLabs. It is suited for lifelike, consistent long-form narration such as voice-overs and audiobooks, supports 29 languages, and has a 10,000-character request limit.

    by elevenlabsOct 7, 2026$40/M characters
  • Favicon for elevenlabs
    ElevenLabs: Eleven Turbo v2.5Eleven Turbo v2.5
    50% off
    1.26M tokens

    Eleven Turbo v2.5 is a low-latency text-to-speech model from ElevenLabs. It balances quality and speed for developer use cases that need non-English languages, supports 32 languages, and has a 40,000-character request limit.

    by elevenlabsOct 7, 2026$20/M characters
  • Favicon for elevenlabs
    ElevenLabs: Eleven Flash v2Eleven Flash v2
    50% off
    13K tokens

    Eleven Flash v2 is an English-only, ultra-low-latency text-to-speech model from ElevenLabs. It is suited for conversational and real-time English use cases, and has a 30,000-character request limit.

    by elevenlabsOct 7, 2026$20/M characters
  • Favicon for elevenlabs
    ElevenLabs: Scribe v2Scribe v2
    50% off
    174M characters

    ElevenLabs Scribe v2 is a speech-to-text model that transcribes audio in 90+ languages with word-level timestamps, optional speaker diarization, and audio-event tagging.

    by elevenlabsOct 7, 2026$0.000031/second
  • Favicon for elevenlabs
    ElevenLabs: Scribe v2 MedicalScribe v2 Medical
    50% off
    670K characters

    ElevenLabs Scribe v2 Medical is a speech-to-text model tuned for clinical and medical terminology, with word-level timestamps, optional speaker diarization, and audio-event tagging.

    by elevenlabsOct 7, 2026$0.000031/second
  • Favicon for openai
    OpenAI: GPT-6 Luna DecisionsGPT-6 Luna Decisions
    34.5B tokens
    Translation (#50)
    Trivia (#10)

    GPT-6 Luna Decisions is GPT-6 Luna served through OpenAI's Decisions API. Instead of generating text, it reads the content passed as state (text, JSON, or images) and returns typed, probabilistic answers to named questions in a single request: the probability of yes for a yes/no question (noul), a probability for every option plus the most likely one (choice), or a probability for every level of an ordered rubric plus the expected score (score). It is built for classification, routing, moderation, and rubric grading where application code thresholds the returned numbers rather than parsing a chat reply. A request can carry up to 200 questions about the same content.

    by openaiOct 6, 20261.05M context$0.10/M input tokens$0/M output tokens