Home / Rank #9
Cheaper & open models / self-hosting
Frontier-model prices keep falling, and small models (Haiku-class, GPT-mini-class, Nova, open Llama/Qwen/Mistral weights) now handle classification, extraction, and routine drafting at a tiny fraction of frontier price. For high-volume, well-scoped tasks, a fine-tuned small model regularly beats a prompted frontier model on cost and matches it on quality.
Self-hosting open weights (with quantization) makes sense past sustained volume thresholds — but be honest about GPU, ops, and eval costs; the API price war means the crossover point is higher than most teams assume.
How to do it
- Benchmark your top-volume tasks on one tier down (and two tiers down) from your current model.
- Fine-tune a small model on tasks with clear ground truth and high volume.
- For self-hosting, price the full picture: GPUs, autoscaling headroom, ops time, and eval maintenance.
- Re-benchmark quarterly — model prices and quality shift fast enough to change the answer.
Frequently asked questions
When does self-hosting pay off?
Rules of thumb vary, but sustained six-figure annual API spend on stable workloads is where serious evaluation starts. Below that, falling API prices usually beat owning GPUs.
Tools for this method
OpenRouter
One API and one bill across hundreds of models from every major lab — the fastest way to arbitrage the model price war and A/B che…
Next method: #10 LLM gateways & spend-tracking tools