Free Tools / Cost & Savings

LLM Cost Arbitrage Calculator

Compare OpenAI, Anthropic, and Gemini bills against self-hosting an open model, with growth and break-even charted over 12 months.

Current published API rates 12-month cumulative chart Editable GPU assumptions

Workload assumptions

Prices as of 20 July 2026

Model tier
Latency need
Self-host assumptions

Your cost comparison will appear here

Enter your workload assumptions, then calculate. We will ask you to verify your work email once before showing the result.

What changes the answer

Utilisation

A GPU billed all month is only economical when requests keep it busy. Bursty traffic wastes reserved capacity.

Output ratio

Generated tokens cost substantially more than input tokens on most hosted models, so response length matters.

Operations

Self-hosting adds deployment, scaling, security, monitoring, model upgrades, and an on-call responsibility.

Pricing sources: official OpenAI, Anthropic, and Google developer pricing pages. Prices as of 20 July 2026; verify before procurement.

Self-Hosting Cost Blueprint

The full worksheet behind the calculator, with cloud sizing guidance.

Sent to the work email you verified for tool access. No newsletter, no spam — ever.

Frequently asked questions

Which model prices does the calculator use?+

It uses standard list prices published by OpenAI, Anthropic, and Google as of 20 July 2026. It excludes batch discounts, prompt caching, negotiated pricing, regional premiums, and tool-call charges.

How is self-hosting cost estimated?+

You control GPU hourly price, monthly GPU hours, and one-off setup cost. Realtime workloads receive a capacity multiplier and batch workloads a lower utilisation multiplier. Add your own staff, storage, network, and observability costs before making a decision.

Does the cheapest model provide equivalent quality?+

Not necessarily. Tier grouping is an economic comparison, not a claim of benchmark equivalence. Evaluate your real prompts for quality, latency, context limits, safety, and structured-output reliability.

When does self-hosting usually make sense?+

It becomes plausible at sustained high utilisation, when data control is strategically important, or when a suitable open model meets your quality target. Low and bursty workloads usually favour hosted APIs.

Are my usage figures stored?+

No. The calculator runs entirely in your browser and sends no token volumes or cost assumptions to Framz.

Related free tool

Terraform Blueprint Generator

Turn the infrastructure decision into a production-structured AWS starting point.

Generate Terraform

Need a defensible AI infrastructure decision?

We benchmark real workloads, model quality, latency, and total operating cost before recommending hosted, hybrid, or self-managed AI.

Discuss your workload