Streaming
Responses arrive token by token over server-sent events, with token usage reported at the end of the stream.
Platform
If your code talks to an OpenAI-style API, it already talks to MarsCompute. Keys, spending limits and metering are part of the platform, not something you add later.
Request path
Five steps, in this order. Select one to read it.
The key is verified and matched to its one model and one member.
The largest possible cost of the request is reserved against your monthly spending limit before anything is sent.
The request goes to a healthy endpoint for that model, within the model's limit on concurrent requests.
The input and output tokens the model reports are recorded with their price.
The reservation is settled to the actual cost, which appears in your usage about a minute later.
curl https://api.marscompute.ai/v1/chat/completions \
-H "Authorization: Bearer $MARSCOMPUTE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "qwen-sea-lion-v4.5-27b",
"stream": true,
"messages": [
{"role": "user", "content": "Say hello in Malay."}
]
}'
import os
import requests
response = requests.post(
"https://api.marscompute.ai/v1/chat/completions",
headers={
"Authorization": f"Bearer {os.environ['MARSCOMPUTE_API_KEY']}",
},
json={
"model": "qwen-sea-lion-v4.5-27b",
"messages": [
{"role": "user", "content": "Say hello in Malay."}
],
},
timeout=120,
)
response.raise_for_status()
print(response.json()["choices"][0]["message"]["content"])
const response = await fetch(
"https://api.marscompute.ai/v1/chat/completions",
{
method: "POST",
headers: {
Authorization: "Bearer " + process.env.MARSCOMPUTE_API_KEY,
"Content-Type": "application/json",
},
body: JSON.stringify({
model: "qwen-sea-lion-v4.5-27b",
messages: [{ role: "user", content: "Say hello in Malay." }],
}),
},
);
const completion = await response.json();
console.log(completion.choices[0].message.content);
Built for production use
Responses arrive token by token over server-sent events, with token usage reported at the end of the stream.
Every key belongs to one member and one model. Revoke a key without touching the others.
Each request is metered by input and output tokens and listed with its cost. Export usage as CSV whenever you need it.
Set a monthly limit for your workspace. Cost is reserved before a request is sent, so the limit holds.
Invite members to your workspace. Each member signs in with their own email and manages their own keys.
Send an Idempotency-Key header and a repeated
request is not run or charged a second time.
Pricing
Prices are per one million tokens, in US dollars, and exclude applicable taxes.
| Model | Model ID | Input / 1M | Output / 1M |
|---|---|---|---|
| Qwen-SEA-LION-v4.5-27B-IT | qwen-sea-lion-v4.5-27b |
$0.60 | $3.60 |
| GPT-6 Astra | gpt-6-astra |
$10.00 | $50.00 |
| Kimi K3 | k3 |
Coming soon | |
| GLM-5.3 | glm-5.3 |
Coming soon | |
Request access and we will set up your workspace.