Pricing

Pay for the pour you reserved.

Rates below are current as of 13 August 2026. The dashboard invoice is canonical. GPU hours bill in 60-second increments after the first minute.

Serverless

Spout

Pay per million

  • Shared capacity in four regions
  • Streaming, batch, structured decode
  • Email support, 08:00 to 20:00 UTC
  • Cold starts are part of this class
Class In / 1M Out / 1M
8B $0.12 $0.36
32B $0.40 $1.20
70B $0.80 $2.40
Embed $0.02 $0.02

SKU prefix spout-

Dedicated

Caster

From $18,000 / month

  • 2x H100 80GB in us-east-1 to start
  • Written p95 and QPS on the order
  • VPC or single-tenant cell
  • TAM and a forward-deployed engineer
  • 24-hour paging on the replica
Add-on Rate
Committed H100 hour $3.40
Extra region pair $4,800 / mo
Customer-managed keys $900 / mo

SKU prefix caster-

Fine-tune hours

Jobs bill on the GPU that ran them.

LoRA

Included on Ladle up to 40 hours per month. Extra LoRA hours follow the table above. A typical 8B adapter on 50k rows finishes in 3.2 H100 hours in the July 2026 lab pass.

Full weights

Full fine-tunes run on Caster or on purchased H100 hours. 70B full jobs start at 96 committed H100 hours. Data never leaves the project without the written job ID.

FAQ

What the invoice actually means.

Cold starts on Spout?

Real. A shared cell can take 4 to 18 seconds to place an 8B and 20 to 45 seconds for a 70B after idle. Ladle and Caster exist for the latency class that cannot wait.

Is Tundrel a GPU broker?

Tundrel sells the serving stack: kernels, packing, and the API. GPU hours appear on the invoice because the pour sits on those trays.

How do regions price?

Spout token rates are the same in all four live cells. Ladle and Caster replicas are priced per region. A second region on Caster is $4,800 per month for a matching pair.

Who signs a Caster SLA?

Rahul countersigns the latency page. Target example on 2x H100: 180 ms TTFT p95 on 8B and 900 ms on 70B at the reserved QPS written on the order.

Talk to an engineer