mintok.ai · self-service
Benchmark-as-a-Service
Your prompts across 12 leading models — measured cost, latency, tokens/sec, and judged quality. Cheapest-that-clears-your-bar, highlighted.
Prices have moved since this report: 9 price changes across 7 models since 2026-07-28.
This page is a snapshot. On the platform, these numbers reprice daily against the live catalogue.
Start a 14-day trial →Cheapest that clears your quality floor
Cerebras · gemma-4-31b
$0.00079/run avg · quality 72/100 · 802.5 tok/s
| # | Model | Tier | $/run avg | Quality | TTFT-ish (ms) | tok/s | OK |
|---|---|---|---|---|---|---|---|
| 1 | Cerebras · gemma-4-31bcerebras | T3 | $0.00079 | 72 | 638 | 802.5 | 1/1 |
| 2 | Anthropic · claude-haiku-4-5-20251001anthropic | T5 | $0.00199 | 72 | 4061 | 97 | 1/1 |
| 3 | Anthropic · claude-sonnet-4-6anthropic | T6 | $0.00774 | 78 | 10996 | 46.6 | 1/1 |
| 4 | Anthropic · claude-opus-4-8anthropic | T7 | $0.0129 | 72 | 8093 | 63.3 | 1/1 |
| — | Groq · openai/gpt-oss-120bgroq | T3 | $0.00000 | 55 | 1314 | 389.6 | 1/1 |
| — | Cerebras · gpt-oss-120bcerebras | T3 | $0.00041 | 55 | 510 | 1003.9 | 1/1 |
| — | Groq · llama-3.3-70b-versatilegroq | T4 | $0.00043 | 65 | 1528 | 335.1 | 1/1 |
Greyed rows fall below your quality floor. Costs are measured from real usage, latency is full-response wall clock, quality is LLM-judged 0–100 and averaged across prompts.
Per-prompt detail
Prompt 1: judge a good LLM from a bad LLM
| Model | $ | Quality | Response (start) |
|---|---|---|---|
| Anthropic · claude-haiku-4-5-20251001 | $0.00199 | 72 | # Judging Good LLMs vs Bad LLMs ## **Output Quality** | Good LLM | Bad LLM | |----------|---------| | Accurate, well-researched answers | Hallucinations & false confidence | | Clear, coherent reasoning | Contradictory or nonsensical responses | | Nuanced, acknowledges limitations | Oversimplified or misleading | | Contextually appropriate tone | Tone-deaf or inappropriate | ## **Reliability** |
| Cerebras · gpt-oss-120b | $0.00041 | 55 | ## How to Tell a Good LLM from a Bad One Below is a practical, checklist‑style guide you can use when you need to **compare two (or more) large‑language models (LLMs)** and decide which one is “good” and which one is “bad.” The criteria are grouped into four high‑level buckets that together cover the most important technical, safety, and usability dimensions: | Category | What to Look For | Wh |
| Cerebras · gemma-4-31b | $0.00079 | 72 | Judging whether an LLM is "good" or "bad" depends entirely on your **use case**. A model that is world-class at writing poetry might be terrible at writing Python code. To judge an LLM objectively, you should evaluate it across these five primary dimensions: --- ### 1. Factuality and Hallucination Rate A "bad" LLM confidently lies. A "good" LLM provides accurate information or admits when it do |
| Anthropic · claude-sonnet-4-6 | $0.00774 | 78 | # Judging a Good LLM from a Bad LLM ## Key Quality Indicators ### ✅ Signs of a **Good** LLM #### Accuracy & Knowledge - Gives **factually correct** answers - **Acknowledges uncertainty** ("I'm not sure, but...") - Knows its **knowledge cutoff** and states it - Cites **reasoning**, not just conclusions #### Honesty - **Admits mistakes** when corrected - Says **"I don't know"** rather than fabri |
| Groq · openai/gpt-oss-120b | $0.00000 | 55 | Below is a practical “cheat‑sheet” you can use to tell whether a language model (LLM) you’re evaluating is **good** (i.e., reliable, safe, and useful) or **bad** (i.e., fragile, risky, or low‑quality). The criteria are grouped into five high‑level buckets that together give you a well‑rounded picture of the model’s overall health. --- ## 1️⃣ Core Performance (What the model *knows* and *does*) |
| Anthropic · claude-opus-4-8 | $0.0129 | 72 | # Judging a Good LLM from a Bad One Here's a practical framework for evaluating LLM quality: ## Core Quality Indicators **1. Accuracy & Factual Reliability** - Good: Admits uncertainty, avoids confident fabrication - Bad: Invents facts, citations, or statistics ("hallucinations") **2. Instruction Following** - Good: Addresses what you actually asked, respects constraints - Bad: Ignores parts o |
| Groq · llama-3.3-70b-versatile | $0.00043 | 65 | Here are some key characteristics that can help you distinguish a good LLM (Large Language Model) from a bad one: **Good LLM:** 1. **Contextual understanding**: A good LLM can understand the context of the input text, including nuances, idioms, and figurative language. 2. **Coherence**: The LLM's responses are coherent, well-structured, and logically connected to the input text. 3. **Accuracy**: |
Keep this current
Model prices and quality change monthly. Model Watch re-benchmarks every new model and price change against these exact prompts and emails you when something beats your current choice. Methodology: how we measure.
Delivered by Mintok — AI infrastructure economics. This link is private to whoever holds it.