Developer API pricing

The OpenAI-compatible /v1 API is pay-as-you-go, billed to your Privateer account credit — no separate plan. Chat is priced per token; image, video, music, sound effects, voices and 3D meshes per unit of output. Packing a sprite sheet from frames you already have calls no model, so it is free.

Base URL https://api.privateer.pro/v1 Auth Bearer sk-priv-… Billing pay-as-you-go

Every rate on this page is what your account is charged. Chat is priced in USD per 1 million tokens, input and output billed separately.

Looking to view specific model pricing?

Rates for the whole catalogue — hundreds of chat models plus image, video, music, sound effects, voices and 3D — live on the models page, each with its id and the unit it is metered in, along with context windows, filters and any provider discounts running right now. It quotes the same developer-API rate this page bills at, and it is rebuilt from the live provider catalogue on every deploy.

Browse all models →

Confidential compute

These models run inside a hardware enclave rather than on a normal inference host, and they reach /v1 through their own providers rather than through the shared catalogue — so they get a rate card of their own here. They are in the models catalogue too, tagged Confidential. See how confidential mode works for what the guarantee does and doesn’t cover.

NEAR AI

ModelInput / 1M tokensOutput / 1M tokens
Qwen3-VL-30B-A3B-Instruct $0.21 $0.77
GLM 5.3 Flash $0.21 $0.70
Qwen 3.6 35B A3B FP8 $0.24 $1.54
Qwen 3.8 27B $0.62 $4.62
Kimi K2.6 $1.13 $5.39
Kimi K3 $4.62 $23.10

Phala

ModelInput / 1M tokensOutput / 1M tokens
Nemotron 3.5 Lightning $0.10 $0.28
Qwen2.5 7B Instruct $0.14 $0.28
Gemma-4 26B-A4B Uncensored (Heretic) $0.21 $0.98
Qwen3.8 27B $0.28 $3.50
Muse Glimmer 30B $0.42 $1.54
Qwen3.8 27B Uncensored (Aggressive) $0.42 $2.10
Qwen3.6 27B $0.45 $3.78
GLM 5.3 $1.96 $6.16

Tinfoil

ModelInput / 1M tokensOutput / 1M tokens
GPT-OSS 120B $0.21 $0.84
Gemma 4 31B $0.56 $1.40
GLM 5.3 Flash $0.56 $1.75
DeepSeek V4.1 Flash $0.91 $2.03
Llama 3.3 70B $2.45 $3.85
GLM 5.3 $2.52 $8.05
Kimi K3 $5.60 $28.00

Confidential rates as of 2026-10-05 and may change with the underlying providers — the app always shows current per-model pricing before you spend.

Image, video, audio & 3D are billed the same way — from the same account credit, metered on each request — but per unit of output rather than per token, and the unit is whichever one the model's provider meters:
  • Images — per generated image, per megapixel, or per output token.
  • Video — per second of output, usually varying by resolution; a few models meter per video token.
  • Music — flat per track, per minute, or per second, depending on the generator.
  • Sound effects — flat per generation, or per second of the length you ask for.
  • Voices — per 1,000 characters of text submitted. Transcription is per minute of audio.
  • 3D meshes — per mesh, moving with the options you pick.
Every per-model rate is published on the models page, quoted at the developer-API rate and rebuilt from the live provider catalogue each deploy. For 3D the spread across the catalogue is more than tenfold and an option that is a surcharge on one model is free on another, so GET /v1/models3d returns the price for every model and every option, and a submit reports the estimate for your exact request before you commit to it.
Music (POST /v1/audio/music) is billed flat per track, or per minute of the length you ask for, depending on the model — every rate is on the models page alongside everything else. It is also the one thing we cannot route to a zero-retention endpoint — no provider we reach offers one, so unlike sound effects there is no other model to switch to and no gate in front of it. A music prompt should never be treated as private.
Sprite sheets have no rate of their own — they cost the models they run. POST /v1/sprites/generate bills one video generation per facing that isn't mirrored, at the video rate above, and says so at submit: the response carries billed_facings alongside the animations you'll get back. Because left is right flipped, an eight-way set costs five generations rather than eight. Each clip is reserved at submit and settled on completion; a failed one releases its hold. Past directions: "one" there is a second, much smaller charge: the picture you send is turned to face each billed direction by an image model first, one still per billed facing after the first — two for a four-way set, four for an eight-way — each billed at the image rate. GET /v1/sprites/models reports both numbers per direction set (billed and billed_turn_stills). POST /v1/sprites — packing frames you already hold into a sheet and a Godot resource — calls no model and is not billed at all, with no daily cap and no balance check, so re-packing at a different frame count, cell size or trim costs nothing.