Free tool

Private LLM TCO Calculator

Compare a hosted API against self-hosting over a year, with the lines most business cases leave out. Every price is an input you can edit, because unit costs move and a calculator with them baked in is wrong within months.

Your workload

Model realistic figures rather than hoped-for ones. Retrieval-augmented prompts are frequently much larger than teams assume, and input volume usually dominates cost.

including retrieved context

Hosted API pricing

These are inputs, not constants. Token pricing moves constantly, so put in the figures from your own agreement rather than trusting a default.

typically 15 to 30

Self-hosted costs

Include everything, not just hardware. Staffing is the line most often left out, and it is usually the largest single item.

including host share

owned kit costs the same idle

Annual cost

Hosted API
$24,750
Token consumption$20,625
Retries and evaluation traffic (20%)$4,125
Self-hosted
$257,000
Hardware, amortised$40,000
Facilities: power, cooling, rack$12,000
Staffing$180,000
Software and tooling$25,000
Breakeven51,919/day
API per request$0.0200
Self-hosted per request$0.2100

At 5,000 requests a day, the hosted API is cheaper. Self-hosting starts to win above roughly 51,919 requests a day on these inputs. Before treating that as the answer, ask how confident you are of reaching that volume and by when, because you will spend the first year well below it.

Check these assumptions

  • At 30% utilisation you are paying for capacity that sits idle most of the time. Owned hardware costs the same whether it is saturated or not, which is why a daytime-only internal workload reaches breakeven far later than a continuous one.

Want this model in writing to take to finance?

Cost should be the second question, not the first. For most large organizations the deployment model is settled by a compliance or contractual constraint before economics get a vote, and this comparison only matters once no constraint binds. It also compares a single year at steady state: model the same three-year window with the same adoption assumptions on both sides before deciding, and check the picture at half your assumed volume, because that is where you will be while adoption builds.

About this calculator

At what volume does self-hosting an LLM become cheaper?

There is a real crossover, but where it sits depends on your own inputs rather than on a universal number. Hosted APIs are priced per token so cost starts near zero and rises with use, while self-hosting is mostly fixed cost that barely moves. The crossover shifts substantially with utilisation, staffing cost and how large your prompts actually are.

What costs do private LLM business cases usually leave out?

Staffing first, and it is often the largest single line. Someone has to patch, monitor, tune, manage capacity and respond outside business hours. After that: facilities power and cooling, redundancy appropriate to your availability target, the serving and observability stack, and the evaluation traffic you will run continuously in production.

Why does utilisation matter so much?

Because owned hardware costs the same whether it is saturated or idle. A workload that is busy six hours a day on weekdays has very different unit economics from one running continuously, and ignoring that is the most common way these comparisons are made to flatter self-hosting.

What is the most common mistake in these comparisons?

Comparing steady-state self-hosted economics against first-year API consumption. That is not a like for like comparison. Model both over the same window with the same adoption assumptions, then check the picture at half your assumed volume, because that is where you will be while adoption builds.

Should cost decide whether we self-host?

Usually not on its own, and usually not first. For most large organizations the deployment model is settled by a compliance or contractual constraint before economics get a vote. Run the constraint test first: if something requires the data to stay inside a boundary you control, the cost comparison is not what decides it.

Why are the prices in this calculator editable rather than fixed?

Because token pricing and hardware costs move constantly. A calculator with current prices baked in becomes wrong within months and, worse, becomes wrong silently. Putting every unit cost in front of you as an input means the output reflects your actual agreement rather than someone's snapshot.