mintok.ai · self-service

Instant Sizing Report

Describe your AI workload in plain English. Get hardware options, TCO, $/M-token, and break-even volume.

Delivered2026-07-28

Prices have moved since this report: 9 price changes across 7 models since 2026-07-28.

This page is a snapshot. On the platform, these numbers reprice daily against the live catalogue.

Start a 14-day trial →

Chatbot

Customer support chatbot . 100 users . inpute/output =100/200

100 events/day100 in / 200 out per eventpeak 0 req/s → 1 tok/s target0M tokens/daylatency-sensitive
Your customer support chatbot workload processes 100 events daily with 100 input tokens and 200 output tokens per event, targeting 1 token per second throughput. Across three evaluated self-hosted options on MI500 hardware—Qwen2.5 3B, Gemma 2 2B, and Llama 3.2 1B—bandwidth emerges as the binding constraint, limiting inference performance despite single-chip configurations. The annual token costs range from 0.0244 to 0.0733 dollars per million tokens depending on model selection. Against the hosted alternative (Kimi K3 at 10 dollars monthly), self-hosting requires approximately 24,907 events daily to break even; your current 100 daily events make hosted deployment substantially more economical at this scale. The practical takeaway is clear: remain with managed inference until event volume increases 249 times, at which point self-hosting becomes cost-competitive and delivers operational control over your latency-sensitive chatbot.

Self-host options

ModelChipChipsRacksBindingCapExTCO/yr$/M-tokResponse
Qwen2.5 3BT1MI50011Bandwidth$0.1M$0.03M$0.07138 ms
Gemma 2 2BT1MI50011Bandwidth$0.1M$0.03M$0.0598 ms
Llama 3.2 1BT1MI50011Bandwidth$0.1M$0.03M$0.0259 ms

Best $/M-token highlighted. Binding = which physical constraint sets the chip count.

Hosted-API comparison (cheapest per tier band)

BandModel$/event$/month at your volume
Mid (T5–6)Kimi K3Moonshot$0.00330$10

Build vs buy

Best self-host option: $2,500/mo · hosted on Kimi K3: $10/mo. Self-hosting breaks even at ≈24,907 events/day (you stated 100).

Assumptions
  • · pue: 1.3
  • · mbuPct: 50
  • · mfuPct: 45
  • · acqMode: CapEx
  • · kwhCost: 0.07
  • · utilPct: 70
  • · batchSize: 8
  • · amortYears: 4
  • · peakFactor: 3

Go deeper

Validate model quality on your real prompts with a benchmark run, or design the agent layer with the Agent Cost Designer. Methodology: how we size.

Delivered by Mintok — AI infrastructure economics. This link is private to whoever holds it.