SEC.01 — Pricing

Pricing that follows your work­load.

Per-token rates depend on the model tier, your region, and whether you're on shared or dedicated capacity. We quote against the shape you actually run, lock it for the term, and report the applied rate on every response.

SEC.02 — PLANS

Three ways to buy.

Every plan reaches the same catalog through the same schema. What changes is the rate, the capacity you can reserve, and how much paperwork your security team gets.

PLAN 01 — Pay-as-you-go

Developer

Free to start / no base fee

Per-token billing across the full catalog. No commitment, no minimums — start with a trial key, scale when you're ready.

Included
All catalog models
Shared inference pools
Streaming, tools, vision, JSON mode
Email support, business hours
10K req/min default ceiling

PLAN 02 — Most teams

Scale

Volume commitment / monthly term

Volume commitment unlocks priority routing, lower per-token rates, and reserved capacity headroom for peak traffic.

Included
Locked per-token rates for the term
Priority routing & failover
Per-tenant rate budgets
SLO with credit-back guarantee
Shared Slack channel

PLAN 03 — Reserved capacity

Enterprise

Custom / monthly term

Reserved replicas, BYO-provider, data-residency, zero-retention routes, and the paperwork your security team requires.

Included
Dedicated GPU replicas
Custom regions & data residency
Zero-retention inference routes
BYO-provider key support
SOC 2 type II artifacts on request

SEC.03 — PER-MODEL RATES

How the meter works.

The rate you pay is set by your contract and the route your call takes — region, capacity type, and load all move it. Whatever it lands at, it comes back in the response headers of every request, so nothing has to be reconstructed later.

FIG.01 — METERING PATH

METERING PATH · PER-TOKEN DWG IDC-MT-04 REQUEST TOKENS IN [ METER ] IN 000 214 618 OUT 000 041 302 RESPONSE X-IDC-Rate-In X-IDC-Rate-Out INVOICE LINE METERED PER TOKEN · NO ROUNDING-UP

Per-model rate bands → MODEL CATALOG · DWG models.html

Open model catalog

Committed customers run on a single flat rate for the whole term

The exact rate for your term is set in your order form

Dedicated capacity is priced per replica-month and quoted separately

SEC.04 — ADD-ONS

Pricing for the things around the model.

Most workloads only need the per-token meter. These line items show up for teams who need more control over where their inference runs.

Add-on What it covers How it's priced
Dedicated replica Reserved GPU capacity for a single model, in a region of your choice. Quoted per deploymentper replica · month
Zero-retention route Prompt and completion bytes never leave volatile memory. Audit log only. Uplift on token ratesset in your order form
Data-residency lock Requests pinned to a specific region with hard failover off. Small uplift on token ratesvaries by region
Premium support 24/7 paging, named contact, quarterly architecture review. Flat monthly feescoped to your team size
BYO-provider routing Use your own provider key under our schema, routing, and observability layer. Metered per requestvolume-tiered

Add-ons are quoted with your token rates · what applies to your account is set in your order form

SEC.05 — PRICING FAQ

Common questions from finance reviews.

Q.01Why don't you publish rates?

Because the rate genuinely varies. A request routed to a US-West shared pool at off-peak isn't priced the same as a Tokyo dedicated replica under peak load, and commitment moves it again. Publishing one number would mean either undercharging us out of business or overcharging customers who don't need the premium path. So we quote against your workload — and every response carries the rate that was actually applied, in its X-IDC-Rate-* headers, matching your invoice line for line.

Q.02How do committed rates work?

You commit to a monthly token volume (or a spend floor) for a term of 3, 6, or 12 months. In exchange, your per-token rate is locked at a single number for the duration, no matter how the underlying route moves. Unused commitment doesn't roll over by default; we can quote a roll-over clause if it matters to your forecasting.

Q.03How is billing measured?

Per token, as reported by the upstream provider, with no rounding-up. Cached prompt tokens are billed at a discounted cached rate that varies by model. Tool-use and structured-output overhead is included in the token count; we don't charge a separate "agent surcharge."

Q.04Do you offer free credits or a trial?

Yes — every new account starts with a trial credit large enough to validate a real workload, not just hello-world. Email sales@idclinktech.com with a short description of what you'd test and we'll size the credit accordingly.

Q.05What payment methods do you accept?

Credit card, ACH/wire, and invoice (net-30) for committed customers. We can support most procurement systems and master service agreements; talk to sales for the paperwork.

SEC.06 — QUOTATION

Want a quote
for your shape?

Send us your monthly token estimate, model mix, and region. We'll come back with a locked-rate offer for that exact shape — usually within one business day.