The OpenAI-compatible /v1 API is pay-as-you-go, billed to
your Privateer account credit — no separate plan. Chat is priced per
token; image, video, music, sound effects, voices and 3D meshes per unit
of output. Packing a
sprite sheet from frames you already have calls no model, so it is free.
Every rate on this page is what your account is charged. Chat is priced in USD per 1 million tokens, input and output billed separately.
Rates for the whole catalogue — hundreds of chat models plus image, video, music, sound effects, voices and 3D — live on the models page, each with its id and the unit it is metered in, along with context windows, filters and any provider discounts running right now. It quotes the same developer-API rate this page bills at, and it is rebuilt from the live provider catalogue on every deploy.
These models run inside a hardware enclave rather than on a normal
inference host, and they reach /v1 through their own
providers rather than through the shared catalogue — so they get a rate
card of their own here. They are in the
models catalogue too, tagged Confidential. See
how confidential mode works for what the
guarantee does and doesn’t cover.
| Model | Input / 1M tokens | Output / 1M tokens |
|---|---|---|
| Qwen3-VL-30B-A3B-Instruct | $0.21 | $0.77 |
| GLM 5.3 Flash | $0.21 | $0.70 |
| Qwen 3.6 35B A3B FP8 | $0.24 | $1.54 |
| Qwen 3.8 27B | $0.62 | $4.62 |
| Kimi K2.6 | $1.13 | $5.39 |
| Kimi K3 | $4.62 | $23.10 |
| Model | Input / 1M tokens | Output / 1M tokens |
|---|---|---|
| Nemotron 3.5 Lightning | $0.10 | $0.28 |
| Qwen2.5 7B Instruct | $0.14 | $0.28 |
| Gemma-4 26B-A4B Uncensored (Heretic) | $0.21 | $0.98 |
| Qwen3.8 27B | $0.28 | $3.50 |
| Muse Glimmer 30B | $0.42 | $1.54 |
| Qwen3.8 27B Uncensored (Aggressive) | $0.42 | $2.10 |
| Qwen3.6 27B | $0.45 | $3.78 |
| GLM 5.3 | $1.96 | $6.16 |
| Model | Input / 1M tokens | Output / 1M tokens |
|---|---|---|
| GPT-OSS 120B | $0.21 | $0.84 |
| Gemma 4 31B | $0.56 | $1.40 |
| GLM 5.3 Flash | $0.56 | $1.75 |
| DeepSeek V4.1 Flash | $0.91 | $2.03 |
| Llama 3.3 70B | $2.45 | $3.85 |
| GLM 5.3 | $2.52 | $8.05 |
| Kimi K3 | $5.60 | $28.00 |
Confidential rates as of 2026-10-05 and may change with the underlying providers — the app always shows current per-model pricing before you spend.
GET /v1/models3d returns
the price for every model and every option, and a submit reports the
estimate for your exact request before you commit to it.
POST /v1/audio/music) is billed flat per
track, or per minute of the length you ask for, depending on the model —
every rate is on the models page alongside
everything else. It is also the one thing we cannot route to a
zero-retention endpoint — no provider we reach offers one, so unlike
sound effects there is no other model to switch to and no gate in front
of it. A music prompt should never be treated as private.
POST /v1/sprites/generate bills one video generation per
facing that isn't mirrored, at the video rate above, and says so at
submit: the response carries billed_facings alongside the
animations you'll get back. Because left is right flipped, an
eight-way set costs five generations rather than eight. Each clip is
reserved at submit and settled on completion; a failed one releases its
hold. Past directions: "one" there is a second, much smaller
charge: the picture you send is turned to face each billed direction by an
image model first, one still per billed facing after the first — two
for a four-way set, four for an eight-way — each billed at the image rate.
GET /v1/sprites/models reports both numbers per direction set
(billed and billed_turn_stills).
POST /v1/sprites — packing frames you already hold into
a sheet and a Godot resource — calls no model and is
not billed at all, with no daily cap and no balance check, so
re-packing at a different frame count, cell size or trim costs nothing.