mintok.ai · self-service

Benchmark-as-a-Service

Your prompts across 12 leading models — measured cost, latency, tokens/sec, and judged quality. Cheapest-that-clears-your-bar, highlighted.

Delivered2026-07-28

Prices have moved since this report: 9 price changes across 7 models since 2026-07-28.

This page is a snapshot. On the platform, these numbers reprice daily against the live catalogue.

Start a 14-day trial →
1/1 prompts7 modelsQuality floor: 70/100Judge: Anthropic · claude-opus-4-8

Cheapest that clears your quality floor

Cerebras · gemma-4-31b

$0.00079/run avg · quality 72/100 · 802.5 tok/s

#ModelTier$/run avgQualityTTFT-ish (ms)tok/sOK
1Cerebras · gemma-4-31bcerebrasT3$0.0007972638802.51/1
2Anthropic · claude-haiku-4-5-20251001anthropicT5$0.00199724061971/1
3Anthropic · claude-sonnet-4-6anthropicT6$0.00774781099646.61/1
4Anthropic · claude-opus-4-8anthropicT7$0.012972809363.31/1
Groq · openai/gpt-oss-120bgroqT3$0.00000551314389.61/1
Cerebras · gpt-oss-120bcerebrasT3$0.00041555101003.91/1
Groq · llama-3.3-70b-versatilegroqT4$0.00043651528335.11/1

Greyed rows fall below your quality floor. Costs are measured from real usage, latency is full-response wall clock, quality is LLM-judged 0–100 and averaged across prompts.

Per-prompt detail

Prompt 1: judge a good LLM from a bad LLM

Model$QualityResponse (start)
Anthropic · claude-haiku-4-5-20251001$0.0019972# Judging Good LLMs vs Bad LLMs ## **Output Quality** | Good LLM | Bad LLM | |----------|---------| | Accurate, well-researched answers | Hallucinations & false confidence | | Clear, coherent reasoning | Contradictory or nonsensical responses | | Nuanced, acknowledges limitations | Oversimplified or misleading | | Contextually appropriate tone | Tone-deaf or inappropriate | ## **Reliability**
Cerebras · gpt-oss-120b$0.0004155## How to Tell a Good LLM from a Bad One Below is a practical, checklist‑style guide you can use when you need to **compare two (or more) large‑language models (LLMs)** and decide which one is “good” and which one is “bad.” The criteria are grouped into four high‑level buckets that together cover the most important technical, safety, and usability dimensions: | Category | What to Look For | Wh
Cerebras · gemma-4-31b$0.0007972Judging whether an LLM is "good" or "bad" depends entirely on your **use case**. A model that is world-class at writing poetry might be terrible at writing Python code. To judge an LLM objectively, you should evaluate it across these five primary dimensions: --- ### 1. Factuality and Hallucination Rate A "bad" LLM confidently lies. A "good" LLM provides accurate information or admits when it do
Anthropic · claude-sonnet-4-6$0.0077478# Judging a Good LLM from a Bad LLM ## Key Quality Indicators ### ✅ Signs of a **Good** LLM #### Accuracy & Knowledge - Gives **factually correct** answers - **Acknowledges uncertainty** ("I'm not sure, but...") - Knows its **knowledge cutoff** and states it - Cites **reasoning**, not just conclusions #### Honesty - **Admits mistakes** when corrected - Says **"I don't know"** rather than fabri
Groq · openai/gpt-oss-120b$0.0000055Below is a practical “cheat‑sheet” you can use to tell whether a language model (LLM) you’re evaluating is **good** (i.e., reliable, safe, and useful) or **bad** (i.e., fragile, risky, or low‑quality). The criteria are grouped into five high‑level buckets that together give you a well‑rounded picture of the model’s overall health. --- ## 1️⃣ Core Performance (What the model *knows* and *does*)
Anthropic · claude-opus-4-8$0.012972# Judging a Good LLM from a Bad One Here's a practical framework for evaluating LLM quality: ## Core Quality Indicators **1. Accuracy & Factual Reliability** - Good: Admits uncertainty, avoids confident fabrication - Bad: Invents facts, citations, or statistics ("hallucinations") **2. Instruction Following** - Good: Addresses what you actually asked, respects constraints - Bad: Ignores parts o
Groq · llama-3.3-70b-versatile$0.0004365Here are some key characteristics that can help you distinguish a good LLM (Large Language Model) from a bad one: **Good LLM:** 1. **Contextual understanding**: A good LLM can understand the context of the input text, including nuances, idioms, and figurative language. 2. **Coherence**: The LLM's responses are coherent, well-structured, and logically connected to the input text. 3. **Accuracy**:

Keep this current

Model prices and quality change monthly. Model Watch re-benchmarks every new model and price change against these exact prompts and emails you when something beats your current choice. Methodology: how we measure.

Delivered by Mintok — AI infrastructure economics. This link is private to whoever holds it.