Rapid.Coder

Curated AI models, tested and ranked for coding.

Rapid.Coder connects to OpenRouter and gives you access to the best models for writing code. Every model on this page is benchmark-tested, cost-compared, and ready to use.

How Rapid.Coder picks models

We curate a shortlist so you can focus on building, not model shopping. Every model is scored against three indices from Artificial Analysis, and priced live from OpenRouter.

1

Benchmark

Human preference decides who gets in. A model qualifies on how it places in blind head-to-head voting — WebDev Arena for general web work, Design Arena for the visual categories. Automated benchmarks tell you whether code passed a test; arenas tell you whether a person preferred the result. For dashboards, the second question is the one that counts.

2

Price

Cost per million tokens (input and output) is pulled directly from OpenRouter. We divide blended cost by the coding index to get cost of intelligence — dollars per point of coding capability — so the value trade-off is a number, not a vibe.

3

Curate

Only models that pass our quality threshold appear in Rapid.Coder. We maintain a free tier, a balanced default, and frontier options so every use case has a sensible starting point.

Model catalog

Prices and benchmarks last updated August 6, 2026. Model IDs are OpenRouter slugs. Cost is USD per 1M tokens.

Free Models
Nemotron 3 Ultra (Free)
Free • 1M context
Free

NVIDIA's free flagship. No cost, large context window. Good for everyday tasks and prototyping.

Code
49.3#6
Design
36.6#5
Intelligence 37.8 · Agentic 27.4
Input $0.00 / Output $0.00
Cost of intelligence Free — prototype here
View on OpenRouter
Preferred Models
DeepSeek V4 Flash
Best cost of intelligence
Best value

Frontier-adjacent benchmark scores at throwaway prices, and near-tied with the default in general webdev voting. Human voters rank it well behind on visual work, so keep design-heavy builds on the default.

Code
69.1#4
Design
45.1#4
Intelligence 49.9 · Agentic 45.7
Input $0.09 / Output $0.18
Cost of intelligence $0.0039 / coding pt
View on OpenRouter
Claude Opus 5
Frontier. Top quality
Frontier

Highest intelligence and coding scores in the catalog. Reach for it when a build has already failed twice, then switch back down.

Code
78.0#1
Design
62.0#2
Intelligence 60.7 · Agentic 55.3
Input $5.00 / Output $25.00
Cost of intelligence $0.3846 / coding pt
View on OpenRouter

Code or design? They are not the same ranking.

Writing an application that works and designing one that looks right are different skills, and the models are not equally good at both. Most business work should follow the left column. Use the right one when the output is client-facing.

Best for code

Does the application work
1
Claude Opus 5
78.0
2
Kimi K3
76.2
3
Qwen3.8 Max
71.8
4
DeepSeek V4 Flash
69.1
5
GLM 5.2Default
68.8
6
Nemotron 3 UltraFree
49.3

Artificial Analysis coding index, 0–100. Automated test suites — whether code runs is objectively checkable.

Best for design

Does it look client-ready
1
Kimi K3
64.6
2
Claude Opus 5
62.0
3
GLM 5.2Default
58.4
4
DeepSeek V4 Flash
45.1
5
Nemotron 3 UltraFree
36.6

Design Arena win rate across website, UI component and dataviz — the share of blind head-to-head matchups real people gave it. Qwen3.8 Max is too new to be rated.

Two places the order flips

Kimi K3 out-designs Claude Opus 5 while ranking below it on code — at 60% of the price. If the deliverable is a client-facing screen, it is the better buy.

DeepSeek V4 Flash and GLM 5.2 are tied on code (69.1 vs 68.8) but 13 points apart on design. That gap is the entire reason GLM 5.2 is the default and DeepSeek is the bulk-work option, despite DeepSeek being 12× cheaper.

How people rank them, and what they cost

Both arena columns are blind human votes. The last column is blended cost divided by the coding index — dollars per point of coding ability, useful for separating models that humans rate similarly.

Model Code Design WebDev Arena Blended $/M $ / code pt Best for
DeepSeek V4 Flash 69.1 #4 45.1 #4 #8 prelim $0.27 $0.0039 Bulk and repetitive jobs. Codes as well as the default, looks markedly worse.
GLM 5.2 Default 68.8 #5 58.4 #3 #7 $3.18 $0.0462 Everyday dashboards. The cheapest model still rated highly for visual work.
Qwen3.8 Max 71.8 #3 not rated #4 prelim $8.00 $0.1114 Long unattended builds. Highest agentic score in the catalog.
Kimi K3 76.2 #2 64.6 #1 #2 $18.00 $0.2362 Client-facing screens. Best design score at any price.
Claude Opus 5 78.0 #1 62.0 #2 #1 $30.00 $0.3846 Hard problems that have already failed twice.

Suffixes you see on arena leaderboards — -max, -high, -xhigh — are reasoning-effort settings rather than separate models, so they resolve to the same OpenRouter slug and the same price shown here. Claude Opus 5 holds both #1 and #3 on WebDev Arena for this reason. Entries marked prelim are still accumulating votes and their ranking may move.

What the scores measure

Two headline scores, measured two different ways — because they are two different questions.

Code score

Does the application actually work — generation, debugging, refactoring, completion accuracy. This is the Artificial Analysis coding index, 0 to 100, from automated test suites. Whether code runs is objectively checkable, so a machine is the right judge. Weight this one for most business work.

Design score

Does the result look like something you would put in front of a client. This is the Design Arena win rate across website, UI component and dataviz — the share of blind head-to-head matchups real people picked it in. Whether a screen looks good is not testable, so only human votes measure it.

The three Artificial Analysis indices behind the Code score, all 0 to 100. Higher is better.

Intelligence Index

General reasoning, comprehension, and multi-step problem solving. A broad measure of how well the model understands complex prompts and produces coherent, accurate responses.

Coding Index

Code generation, debugging, refactoring, and completion accuracy across languages. This is the primary metric for Rapid.Coder users. Weight this against cost for your budget.

Agentic Index

Multi-turn tool use, planning, and autonomous task completion. Higher scores mean the model is better at long chains of tool calls and self-correcting when a step fails.

And one number we calculate ourselves: cost of intelligence

Blended cost (input + output, USD per 1M tokens) divided by the coding index. It answers the question that matters when two models look similar: what does a point of coding ability actually cost here? A frontier model can be 100× the price of a budget one for 13% more capability.

It is deliberately a second question, not the first. Its denominator is an automated benchmark, and benchmarks reward passing tests rather than producing something a person wants to look at. So human votes decide which models are good enough, and this number decides between the survivors. Dividing cost by an arena rating instead would not work: Elo has an arbitrary zero and this whole catalog sits inside a 130-point band, so the ratio would collapse into "cheapest wins" and tell you nothing.

Sources and further reading

Rapid.Coder is built on top of open infrastructure. These are the projects and tools we rely on.

Model catalog JSON

Rapid.Coder fetches its model list from https://www.rapiddashboard.ai/rapid.coder/model.json. The file uses OpenRouter model slugs and is updated as pricing and benchmarks change. If you are building on top of the catalog, set an Access-Control-Allow-Origin: * header so browser-based clients can fetch it cross-origin.

See RapidDashboard in action.

Rapid.Coder is one piece of the platform. Schedule a demo to see how prompt-to-dashboard, prompt-to-report, and AI-powered coding work together, grounded in your own data.

Schedule a demo