Curated AI models, tested and ranked for coding.
Rapid.Coder connects to OpenRouter and gives you access to the best models for writing code. Every model on this page is benchmark-tested, cost-compared, and ready to use.
How Rapid.Coder picks models
We curate a shortlist so you can focus on building, not model shopping. Every model is scored against three indices from Artificial Analysis, and priced live from OpenRouter.
Benchmark
Human preference decides who gets in. A model qualifies on how it places in blind head-to-head voting — WebDev Arena for general web work, Design Arena for the visual categories. Automated benchmarks tell you whether code passed a test; arenas tell you whether a person preferred the result. For dashboards, the second question is the one that counts.
Price
Cost per million tokens (input and output) is pulled directly from OpenRouter. We divide blended cost by the coding index to get cost of intelligence — dollars per point of coding capability — so the value trade-off is a number, not a vibe.
Curate
Only models that pass our quality threshold appear in Rapid.Coder. We maintain a free tier, a balanced default, and frontier options so every use case has a sensible starting point.
Model catalog
Prices and benchmarks last updated August 6, 2026. Model IDs are OpenRouter slugs. Cost is USD per 1M tokens.
NVIDIA's free flagship. No cost, large context window. Good for everyday tasks and prototyping.
Frontier-adjacent benchmark scores at throwaway prices, and near-tied with the default in general webdev voting. Human voters rank it well behind on visual work, so keep design-heavy builds on the default.
Strong coding at a fraction of frontier cost, and unusually good at building web UI. The value default for Rapid.Coder.
The long-run workhorse. Highest agentic score in the catalog — above Claude Opus 5 — at a quarter of the price. Best pick for overnight app builds.
The rare model both arenas agree on: #2 on WebDev Arena, and #1 for websites and UI components on Design Arena — at 60% of frontier price.
Highest intelligence and coding scores in the catalog. Reach for it when a build has already failed twice, then switch back down.
Code or design? They are not the same ranking.
Writing an application that works and designing one that looks right are different skills, and the models are not equally good at both. Most business work should follow the left column. Use the right one when the output is client-facing.
Best for code
Does the application workArtificial Analysis coding index, 0–100. Automated test suites — whether code runs is objectively checkable.
Best for design
Does it look client-readyDesign Arena win rate across website, UI component and dataviz — the share of blind head-to-head matchups real people gave it. Qwen3.8 Max is too new to be rated.
Two places the order flips
Kimi K3 out-designs Claude Opus 5 while ranking below it on code — at 60% of the price. If the deliverable is a client-facing screen, it is the better buy.
DeepSeek V4 Flash and GLM 5.2 are tied on code (69.1 vs 68.8) but 13 points apart on design. That gap is the entire reason GLM 5.2 is the default and DeepSeek is the bulk-work option, despite DeepSeek being 12× cheaper.
How people rank them, and what they cost
Both arena columns are blind human votes. The last column is blended cost divided by the coding index — dollars per point of coding ability, useful for separating models that humans rate similarly.
| Model | Code | Design | WebDev Arena | Blended $/M | $ / code pt | Best for |
|---|---|---|---|---|---|---|
| DeepSeek V4 Flash | 69.1 #4 | 45.1 #4 | #8 prelim | $0.27 | $0.0039 | Bulk and repetitive jobs. Codes as well as the default, looks markedly worse. |
| GLM 5.2 Default | 68.8 #5 | 58.4 #3 | #7 | $3.18 | $0.0462 | Everyday dashboards. The cheapest model still rated highly for visual work. |
| Qwen3.8 Max | 71.8 #3 | not rated | #4 prelim | $8.00 | $0.1114 | Long unattended builds. Highest agentic score in the catalog. |
| Kimi K3 | 76.2 #2 | 64.6 #1 | #2 | $18.00 | $0.2362 | Client-facing screens. Best design score at any price. |
| Claude Opus 5 | 78.0 #1 | 62.0 #2 | #1 | $30.00 | $0.3846 | Hard problems that have already failed twice. |
Suffixes you see on arena leaderboards — -max, -high, -xhigh — are reasoning-effort settings rather than separate models, so they resolve to the same OpenRouter slug and the same price shown here. Claude Opus 5 holds both #1 and #3 on WebDev Arena for this reason. Entries marked prelim are still accumulating votes and their ranking may move.
What the scores measure
Two headline scores, measured two different ways — because they are two different questions.
Code score
Does the application actually work — generation, debugging, refactoring, completion accuracy. This is the Artificial Analysis coding index, 0 to 100, from automated test suites. Whether code runs is objectively checkable, so a machine is the right judge. Weight this one for most business work.
Design score
Does the result look like something you would put in front of a client. This is the Design Arena win rate across website, UI component and dataviz — the share of blind head-to-head matchups real people picked it in. Whether a screen looks good is not testable, so only human votes measure it.
The three Artificial Analysis indices behind the Code score, all 0 to 100. Higher is better.
Intelligence Index
General reasoning, comprehension, and multi-step problem solving. A broad measure of how well the model understands complex prompts and produces coherent, accurate responses.
Coding Index
Code generation, debugging, refactoring, and completion accuracy across languages. This is the primary metric for Rapid.Coder users. Weight this against cost for your budget.
Agentic Index
Multi-turn tool use, planning, and autonomous task completion. Higher scores mean the model is better at long chains of tool calls and self-correcting when a step fails.
And one number we calculate ourselves: cost of intelligence
Blended cost (input + output, USD per 1M tokens) divided by the coding index. It answers the question that matters when two models look similar: what does a point of coding ability actually cost here? A frontier model can be 100× the price of a budget one for 13% more capability.
It is deliberately a second question, not the first. Its denominator is an automated benchmark, and benchmarks reward passing tests rather than producing something a person wants to look at. So human votes decide which models are good enough, and this number decides between the survivors. Dividing cost by an arena rating instead would not work: Elo has an arbitrary zero and this whole catalog sits inside a 130-point band, so the ratio would collapse into "cheapest wins" and tell you nothing.
Sources and further reading
Rapid.Coder is built on top of open infrastructure. These are the projects and tools we rely on.
Model catalog JSON
Rapid.Coder fetches its model list from https://www.rapiddashboard.ai/rapid.coder/model.json. The file uses OpenRouter model slugs and is updated as pricing and benchmarks change. If you are building on top of the catalog, set an Access-Control-Allow-Origin: * header so browser-based clients can fetch it cross-origin.
See RapidDashboard in action.
Rapid.Coder is one piece of the platform. Schedule a demo to see how prompt-to-dashboard, prompt-to-report, and AI-powered coding work together, grounded in your own data.