Serve a chat-completions API from any open model
Run large language models behind a drop-in REST endpoint
Billing period usage
2% of included credits used
Active resources
Browse a few examples to see what you could ship this afternoon.
Run large language models behind a drop-in REST endpoint
Adapt a diffusion checkpoint with Prism and a shareable web demo
Raise throughput on batched inference jobs