Questions

How do I choose between runtimes?

For day-to-day workloads we suggest starting with the standard and compact runtimes. The standard runtime holds up better across a broad mix of jobs, while the compact one answers faster and costs less on the lighter requests teams send most often.

Our deliberative runtimes suit layered engineering and research work that benefits from longer internal reasoning. Reach for the compact deliberative build when response time matters more to you than the last few points of accuracy.

The quickest way to settle it is to send the same prompt through each option in the Sandbox and weigh cost against quality on your own traffic.

Are there volume commitments or support tiers?

Yes. Once monthly spend clears roughly $4,800 we can move your workspace onto a committed-throughput agreement with a named support contact, a written response window, and quarterly capacity planning. Write to the accounts desk to start that conversation.

Does testing in the Sandbox count toward billing?

It does. Prompts sent from the Sandbox bill at the same per-token rate as production traffic, against the same workspace balance. Every workspace also carries a small monthly evaluation credit that is consumed before any charge appears.

Where can I review my monthly token consumption?

The Usage view in your workspace breaks consumption down by day, by runtime, and by API key, and you can export any range as CSV. Figures refresh about every ten minutes, so a burst of traffic shows up well before the invoice does.

What controls exist for capping team spend?

Set a monthly ceiling for the whole workspace, then per-key limits underneath it. Notifications fire at 60% and 85% of the ceiling, and requests are declined once it is reached rather than silently rolling into overage.

Do workspace seats include developer platform access?

They are billed separately. A seat covers the assistant workspace your colleagues use day to day; platform usage is metered per token against a balance you top up on its own. One sign-in works across both.

How are image requests metered?

Images are converted into tokens before they are priced. A 1024 x 1024 frame resolves to about 1,105 input tokens at standard fidelity and roughly a quarter of that in draft mode, so the rate card you already use still applies.

Put your first prompt
into production on Pindrift.