Skip to content
  • Models
  • Rankings
  • Ori
OpenRouterOpenRouter
© 2026 OpenRouter, Inc

Product

  • Chat
  • Rankings
  • Benchmarks
  • Apps
  • Discover
  • Models
  • Collections
  • Providers
  • Tools
  • Pricing
  • Business
  • Enterprise
  • Labs

Company

  • About
  • Blog
  • Careers
    Hiring
  • Privacy
  • Terms of Service
  • Trust Center
  • Support
  • Works With OR
  • Data
  • Brand

Developer

  • Documentation
  • API Reference
  • Developer Platform
  • Status
  • AI Site Map

Connect

  • Discord
  • GitHub
  • LinkedIn
  • X
  • YouTube
Collections/Roleplay

Best Roleplay (RP) & Creative Writing AI Models by Usage

Model rankings updated October 2026 based on real usage data.

Models in this roleplay collection are ranked by total prompt and completion tokens processed through the OpenRouter API over the trailing 7 days. Showing the top 10 models. Rankings measure usage on OpenRouter and reflect adoption, not model quality or benchmark performance.

The best AI models for roleplay, character chat, and creative writing are ranked below by real usage on OpenRouter over the past week. The current top models are DeepSeek V4.1 Flash, GLM 5.3 Flash, and MiMo-V2.6-Flash. Users pick them for consistent personas, natural dialogue, and coherence across long sessions, whether through Janitor AI, SillyTavern, another frontend, or their own character chatbot. Every model is available through a single OpenRouter API key, so you can try several without changing your setup.

Browse All ModelsCompare Models

LLM Leaderboard for Roleplay Models

1.
Favicon for deepseek
Deepseek V4 Flash
by deepseek
1.03T
12.3%
2.
Favicon for stealth
Space Bunny Alpha
by stealth
995B
11.9%
3.
Favicon for deepseek
Deepseek V4.1 Flash
by deepseek
649B
7.7%
4.
Favicon for tencent
Hy4 Preview
by tencent
518B
6.2%
5.
Favicon for deepseek
Deepseek V4 Flash
by deepseek
434B
5.2%
6.
Favicon for deepseek
Deepseek V3.2
by deepseek
373B
4.5%
7.
Favicon for xiaomi
Mimo V2.6 Flash
by xiaomi
329B
3.9%
8.
Favicon for z-ai
GLM 5.3 Flash
by z-ai
312B
3.7%
9.
Favicon for google
Gemini 2.5 Flash Lite
by google
265B
3.2%
10.
Favicon for unknown
Others
3.48T
41.5%

Top Roleplay Models on OpenRouter

Favicon for deepseek

DeepSeek: DeepSeek V4.1 Flash

31.3T tokens
Academia (#1)
Finance (#1)
Health (#1)
Legal (#4)
Marketing (#1)

DeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the first built on the company's Causal Encoder-Decoder (CED) architecture. It activates 8B parameters on input and 16B on output from a 552B-parameter backbone, an asymmetric split that keeps per-token compute low relative to the model's total size. Image understanding is native to the architecture, with visual and text embeddings trained jointly from the start of pre-training rather than added afterward as in the earlier experimental V4 Flash Vision Exp.

It is suited for coding, terminal, and computer-use agents, along with long-horizon tasks that must run to completion across many steps and long-context analysis. Compressed KV caching cuts cache memory to roughly a quarter of the previous Flash generation, significantly reducing costs on agentic workloads. DeepSeek positions it as the cost-efficient tier of the V4.1 family and reports that it exceeds V4 Pro on performance, speed, and task completion time.

by deepseek1.05M context$0.003/M input tokens$2.40/M output tokens
Favicon for z-ai

Z.ai: GLM 5.3 Flash

11.1T tokens
Academia (#3)
Finance (#3)
Health (#5)
Legal (#5)
Marketing (#2)

GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while reducing compute overhead.

by z-ai1.05M context$0.0352/M input tokens$0.50/M output tokens
Favicon for xiaomi

Xiaomi: MiMo-V2.6-Flash

11T tokens
Academia (#17)
Finance (#13)
Health (#30)
Marketing (#35)
SEO (#41)

MiMo-V2.6-Flash is an open-source foundation model developed by Xiaomi. Built on a Mixture-of-Experts architecture with 309B total parameters and 15B activated per token, it employs a hybrid attention mechanism for greater computational efficiency. The model features a 1M-token context window and native multimodal capabilities. Optimized for agentic workflows, it delivers strong performance across coding, visual, general, and research scenarios, excelling at complex, long-horizon tasks with robust generalization across a diverse range of agent harnesses.

by xiaomi1.05M context$0.10/M input tokens$0.28/M output tokens
Favicon for tencent

Tencent: Hy4 preview

7.52T tokens
Academia (#13)
Finance (#28)
Legal (#24)
Marketing (#19)
SEO (#30)

Tencent: Hy4 preview is a mixture-of-experts model from Tencent, with 49B active parameters out of 770B total. It is designed for coding agents, complex tool-use workflows, and productivity tasks that require planning, context continuity, and sustained multi-step execution.

by tencent1.05M context$0.7506/M input tokens$2.251/M output tokens
Favicon for openai

OpenAI: GPT-6 Luna

6.89T tokens
Academia (#2)
Finance (#6)
Health (#3)
Legal (#3)
Marketing (#3)

GPT-6 Luna is the fast, cost-efficient model in OpenAI's GPT-6 series, positioned below GPT-6 Sol. It is suited for high-volume and latency-sensitive workloads such as chat, classification, and lightweight agentic tasks, and at higher reasoning effort it can take on complex software engineering and computer-use tasks that previously called for a Sol-tier model. It shares the GPT-6 family's gains in factual reliability and its clearer, more concise communication style.

by openai1.05M context$0.10/M input tokens$0.50/M output tokens
Favicon for deepseek

DeepSeek: DeepSeek V4 Flash 0731

6.74T tokens
Academia (#6)
Finance (#2)
Health (#9)
Legal (#8)
Marketing (#5)

DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows. This is the GA release of DeepSeek V4 Flash.

by deepseek1.05M context$0.0152/M input tokens$1.28/M output tokens
Favicon for openai

OpenAI: GPT-5.6 Luna

4.62T tokens
Academia (#7)
Finance (#9)
Health (#6)
Legal (#6)
Marketing (#7)

GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat, classification, and lightweight agentic workflows, providing capable reasoning for its price tier.

by openai1.05M context$0.20/M input tokens$1.20/M output tokens
Favicon for deepseek

DeepSeek: DeepSeek V4 Flash 0423

3.63T tokens
Academia (#4)
Finance (#5)
Health (#10)
Legal (#7)
Marketing (#4)

DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-token context window. It is designed for fast inference and high-throughput workloads, while maintaining strong reasoning and coding performance.

The model includes hybrid attention for efficient long-context processing. Reasoning efforts high and xhigh are supported; xhigh maps to max reasoning. It is well suited for applications such as coding assistants, chat systems, and agent workflows where responsiveness and cost efficiency are important.

by deepseek1.05M context$0.03/M input tokens$1.28/M output tokens
Favicon for typesafe

TypeSafe: Jev 1.13

3.54T tokens
Academia (#5)
Finance (#4)
Health (#2)
Legal (#1)
Marketing (#6)

Jev is a structured decision model from TypeSafe, and the first of its System One models. System One models make fast, structured decisions for software, returning a typed choice rather than free-form text. It is suited for routing, classification, and other decision points inside an application where a fast, predictable answer matters more than generated prose.

Learn more in TypeSafe's docs: https://docs.typesafe.ai/concepts/system-one

by typesafe32K context$0.042/M input tokens$0/M output tokens
Favicon for z-ai

Z.ai: GLM 5.3

3.34T tokens
Academia (#32)
Finance (#11)
Health (#27)
Legal (#20)
Marketing (#23)

GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and output with a 1M-token context window, and improves on GLM-5.2 in coding and in the balance between performance and token efficiency.

Reasoning is always on and cannot be disabled. Reasoning efforts low, high, and max are supported; max is the default.

by z-ai1.05M context$0.03/M input tokens$12/M output tokens

Explore more collections

  • Free Models
  • Discounted Models
  • Coding
  • Vision Models
  • Tool Calling
  • OpenClaw
  • Image Models
  • Video Models
  • Audio Models
  • Text-to-Speech
  • Speech-to-Text
  • Embedding Models
  • Rerank Models
  • Distillable Models
  • All collections