[ BILLING ]
How billing works
Short version: you pay your model provider for your model. We charge a flat subscription for the testing, the judging, and the signed artifact. There are no per-token charges from SichGate, and no invoice you didn't see coming.
[ 01 ]
WHAT A
RUN IS
A SichGate run fires a battery of safety probes at your model and classifies every response it gets back.
The size of a run is simple math:
probes × model states = generations
A typical quantization-drift run tests three states of the same model — base, fine-tuned, and quantized — so a 500-probe battery produces 1,500 generations and 1,500 classifications.
You'll see the exact generation count before you start a run, so you always know what you're committing to.
[ 02 ]
THE THREE
COSTS
Three costs, and who covers each.
| Cost | What it is | Who pays |
|---|---|---|
| Target model inference | Generating responses from your model | You |
| Judging and classification | Scoring every response for safety behavior | SichGate |
| Platform | Orchestration, storage, signing, AI-BOM, audit log | SichGate |
You pay for your model
Your model runs on your infrastructure or your provider account. Two ways to connect:
- —Bring your own endpoint. Point SichGate at a model you already host or an API you already pay for. Test traffic bills to your account at your existing rates.
- —Run the SichGate runner locally. For local weights and quantized artifacts, the runner executes on your own hardware. Nothing leaves your environment.
This isn't only a billing decision. It means your weights, your keys, and your test traffic stay under your control — which is usually the first thing a security team asks about.
We pay for the verdict
Judging is what you're actually buying, so it's ours. Classification of every response, the drift diff against your certified baseline, the AI-BOM, the ed25519-signed audit log, and the certification tier are all included in your subscription. No metering, no per-token line item.
Optional: we run the model too
If you'd rather not stand up compute for local weights, we can host and run them for you. That's a managed add-on, priced separately by GPU-hour, and it's always opt-in. It is never billed silently against a standard plan.
[ 03 ]
WHAT A RUN
COSTS YOU
Real numbers for the same 500-probe, 3-state run — 1,500 generations, roughly 600K tokens total.
[ fill in before publishing ]
These figures are illustrative and based on typical published rates. Recalculate against current provider pricing and your own measured token averages before this goes live.
| How you run it | Approximate cost to you |
|---|---|
| Small hosted model API (~$0.10/M in, $0.40/M out) | ~$0.20 per run |
| Mid-tier hosted model | ~$1–2 per run |
| Frontier model API (~$3/M in, $15/M out) | ~$6 per run |
| Self-hosted 7B on a single L4/A10G | under $0.50 per run |
For most teams testing small language models, a full certified run costs less than a coffee. The reason we don't meter it is that metering it would cost more to administer than the tokens are worth.
[ 04 ]
LIMITS AND
OVERAGES
Every plan includes a set number of certified runs per month. Allowances and overage rates live in the app, alongside the plan you're on, so the number you read is always the number you're billed.
COMPARE PLANS IN THE APP →Three things we commit to:
- 01Hard cap on by default. When you hit your included runs, we stop and ask. You opt in to overages — they never just happen.
- 02Published overage pricing. One number, on your plan page. Not “contact sales.”
- 03No surprise invoices. If you never enable overages, your SichGate bill is the same every month.
[ 05 ]
PARTNERS AND
RESELLERS
If you're embedding SichGate in your own offering, the model is the same but the billing relationship shifts:
- —You connect your customer’s endpoint, or your own. Model inference bills to whoever owns that endpoint.
- —We bill you per certified run at partner rates.
- —You set your own end-customer price. We don’t constrain your margin.
Partner terms are separate from this page — talk to us about partnering.
[ 06 ]
FAQ
Does SichGate charge per token?
No. We charge a flat subscription plus per-run overages if you enable them. Token costs are between you and your model provider.
Can my bill increase without me doing anything?
No. Runs are triggered by you or by a schedule you configure, and hard caps are on by default.
Do you ever see my model weights?
Only if you explicitly choose the managed hosting add-on. On the default path, weights and test traffic stay in your environment.
What about scheduled monitoring runs?
Same model. Each scheduled re-test counts as a run against your allowance, and each one generates inference cost on your endpoint. You set the cadence and the canary subset size, so you control both.
Who pays for a failed or cancelled run?
If a run fails on our side, it doesn't count against your allowance. If you cancel mid-run, you'll have already incurred whatever inference cost your provider metered up to that point — we can't refund that, but we won't count the run.
[ NEXT ]
Questions this page didn't answer?