Ask every model at once, from your own code
Chat with All is a thin, honest wrapper over one platform primitive: a multi-turn AI session whose model you choose per turn. The browser app runs several of those sessions side by side, one per lane, and prints what each reply cost. This page documents the same REST API the app itself calls — mint a token, read the live model catalogue, price a turn before you pay for it, open a session, stream a reply, and keep a little state in the per-user store. Everything is plain JSON over HTTPS. Pick a language once in any code group and the whole page follows.
Basics
Base URL: https://api.skillsafe.ai/v1/app-api. App slug:
chat-with-all. Every authenticated request sends
Authorization: Bearer <token>, and JSON bodies go with
Content-Type: application/json.
Responses are always wrapped in an envelope. Success:
{"ok": true, "data": { … }, "meta": { … }}
Failure:
{"ok": false,
"error": {"code": "…", "message": "…", "status": 400, "details": { … }},
"meta": { … }}
The HTTP status and error.status agree; error.code is the stable,
machine-readable string you should branch on (see the
error table). meta carries request bookkeeping and, on
list endpoints, meta.pagination.
/apps/{slug}/ path segment
This is the single most common mistake against this API, so it gets the loudest box on the
page. The slug is supplied once, in the body of
POST /v1/app-api/guest, and from then on it is bound to the
token. Every other call is slug-free. Concretely, verified against the live API:
POST https://api.skillsafe.ai/v1/app-api/apps/chat-with-all/estimate
→ 404 {"ok":false,"error":{"code":"not_found", … }} wrong
POST https://api.skillsafe.ai/v1/app-api/estimate
→ 200 {"ok":true,"data":{"hold_credits":1367, … }} right
If you are getting 404 not_found from an endpoint you are sure exists, this is
almost always why.
One endpoint on this page lives outside the app API and outside auth entirely:
GET https://api.skillsafe.ai/v1/models is public — no
Authorization header, no /app-api prefix. See
step 3.
Browsers enforce CORS on this API, so run these examples from a terminal, script or server rather than from another site's frontend. The app itself is exempt because it is served from the app's own origin.
Step 1 — Get a token (and a tiny client)
Send the slug in the body — {"slug":"chat-with-all"}. There is no slug
header and no slug path segment.
201 Created
{"ok": true,
"data": {
"token": "…",
"guest_id": "gst_…",
"expires_at": "…"
},
"meta": { … }}
For a token tied to your own account and credits, don't script this at all:
open the token page, sign in with SkillSafe, and
copy the shell export it puts on your clipboard. That is the no-developer-tools way to get a
browser-minted token into your shell, and it is what the examples below assume is in
$SKILLSAFE_TOKEN.
/me and
/estimate, and /v1/models needs no token at
all. But session turns are metered and need credits, and this app is
published with guest_enabled: 0 — guests get no starting wallet. So a
guest can browse the model catalogue and price a question all day, and will be refused the
moment it posts a message. To actually run turns, sign in and use a user token.
Every call below is one HTTP request, so start with a short helper that adds the bearer
header, sends JSON and unwraps the data envelope. The later steps reuse it.
export API="https://api.skillsafe.ai/v1/app-api"
export SKILLSAFE_TOKEN="YOUR_TOKEN" # from /tokens.html, or the POST /guest below
# every authenticated call looks like:
# curl -s "$API/..." -H "Authorization: Bearer $SKILLSAFE_TOKEN" \
# -H "Content-Type: application/json" [-d '{json}']
# jq pulls fields out of the {"ok":true,"data":...} envelope
# a headless guest token (browse + estimate only on this app):
curl -s -X POST "$API/guest" \
-H "Content-Type: application/json" \
-d '{"slug":"chat-with-all"}' | jq -r '.data.token'
import os, requests
API = "https://api.skillsafe.ai/v1/app-api"
TOKEN = os.environ.get("SKILLSAFE_TOKEN", "YOUR_TOKEN")
def api(method, path, body=None, **headers):
res = requests.request(
method, API + path, json=body,
headers={"Authorization": f"Bearer {TOKEN}", **headers})
payload = res.json()
if not res.ok:
err = payload.get("error", {})
raise RuntimeError(f"{err.get('code', res.status_code)}: {err.get('message', res.reason)}")
return payload["data"]
# a headless guest token (browse + estimate only on this app):
guest = requests.post(API + "/guest", json={"slug": "chat-with-all"}).json()["data"]
print(guest["token"], guest["guest_id"], guest["expires_at"])
// Node 18+ (built-in fetch)
const API = "https://api.skillsafe.ai/v1/app-api";
const TOKEN = "YOUR_TOKEN"; // paste from /tokens.html - never commit it
async function api(method, path, body, extraHeaders = {}) {
const res = await fetch(API + path, {
method,
headers: {
Authorization: `Bearer ${TOKEN}`,
"Content-Type": "application/json",
...extraHeaders,
},
body: body === undefined ? undefined : JSON.stringify(body),
});
const json = await res.json();
if (!res.ok) {
const e = new Error(json.error?.message ?? res.statusText);
e.code = json.error?.code;
e.status = res.status;
throw e;
}
return json.data;
}
// a headless guest token (browse + estimate only on this app):
const guestRes = await fetch(API + "/guest", {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({ slug: "chat-with-all" }),
});
const { token, guest_id, expires_at } = (await guestRes.json()).data;
package main
import (
"bytes"
"encoding/json"
"fmt"
"net/http"
"os"
)
const API = "https://api.skillsafe.ai/v1/app-api"
var token = os.Getenv("SKILLSAFE_TOKEN")
func call(method, path string, body, out any, headers map[string]string) error {
var buf bytes.Buffer
if body != nil {
json.NewEncoder(&buf).Encode(body)
}
req, _ := http.NewRequest(method, API+path, &buf)
if token != "" {
req.Header.Set("Authorization", "Bearer "+token)
}
req.Header.Set("Content-Type", "application/json")
for k, v := range headers {
req.Header.Set(k, v)
}
res, err := http.DefaultClient.Do(req)
if err != nil {
return err
}
defer res.Body.Close()
var env struct {
Data json.RawMessage `json:"data"`
Error *struct {
Code string `json:"code"`
Message string `json:"message"`
} `json:"error"`
}
json.NewDecoder(res.Body).Decode(&env)
if res.StatusCode >= 400 {
return fmt.Errorf("%s: %s", env.Error.Code, env.Error.Message)
}
if out == nil {
return nil
}
return json.Unmarshal(env.Data, out)
}
// a headless guest token (browse + estimate only on this app):
func guest() (string, error) {
var g struct {
Token string `json:"token"`
GuestID string `json:"guest_id"`
ExpiresAt string `json:"expires_at"`
}
err := call("POST", "/guest", map[string]string{"slug": "chat-with-all"}, &g, nil)
return g.Token, err
}
// Java 17+, no dependencies. Pair it with your JSON library (Jackson, Gson...)
// to read fields out of the returned envelope.
import java.net.URI;
import java.net.http.HttpClient;
import java.net.http.HttpRequest;
import java.net.http.HttpResponse;
import java.util.Map;
public class ChatWithAll {
static final String API = "https://api.skillsafe.ai/v1/app-api";
static final String TOKEN = System.getenv("SKILLSAFE_TOKEN");
static final HttpClient HTTP = HttpClient.newHttpClient();
static String api(String method, String path, String jsonBody, Map<String, String> headers)
throws Exception {
var b = HttpRequest.newBuilder(URI.create(API + path))
.header("Content-Type", "application/json")
.method(method, jsonBody == null
? HttpRequest.BodyPublishers.noBody()
: HttpRequest.BodyPublishers.ofString(jsonBody));
if (TOKEN != null) b = b.header("Authorization", "Bearer " + TOKEN);
if (headers != null) for (var e : headers.entrySet()) b = b.header(e.getKey(), e.getValue());
var res = HTTP.send(b.build(), HttpResponse.BodyHandlers.ofString());
if (res.statusCode() >= 400) throw new RuntimeException(res.body());
return res.body(); // envelope: {"ok":true,"data":...}
}
// a headless guest token (browse + estimate only on this app):
static String guestEnvelope() throws Exception {
return api("POST", "/guest", "{\"slug\":\"chat-with-all\"}", null);
// token is at data.token, plus data.guest_id and data.expires_at
}
}
require "net/http"
require "json"
API = "https://api.skillsafe.ai/v1/app-api"
TOKEN = ENV.fetch("SKILLSAFE_TOKEN", "YOUR_TOKEN")
def api(method, path, body = nil, headers = {})
uri = URI(API + path)
req = Net::HTTP.const_get(method.capitalize).new(uri)
req["Authorization"] = "Bearer #{TOKEN}"
req["Content-Type"] = "application/json"
headers.each { |k, v| req[k] = v }
req.body = body.to_json if body
res = Net::HTTP.start(uri.host, uri.port, use_ssl: true) { |h| h.request(req) }
payload = JSON.parse(res.body)
unless res.is_a?(Net::HTTPSuccess)
raise "#{payload.dig('error', 'code')}: #{payload.dig('error', 'message')}"
end
payload["data"]
end
# a headless guest token (browse + estimate only on this app):
guest = api("POST", "/guest", { slug: "chat-with-all" })
puts guest["token"], guest["guest_id"], guest["expires_at"]
<?php
const API = "https://api.skillsafe.ai/v1/app-api";
$TOKEN = getenv("SKILLSAFE_TOKEN") ?: "YOUR_TOKEN";
function api(string $method, string $path, ?array $body = null, array $headers = []): mixed {
global $TOKEN;
$hdr = array_merge([
"Authorization: Bearer $TOKEN",
"Content-Type: application/json",
], $headers);
$ch = curl_init(API . $path);
curl_setopt_array($ch, [
CURLOPT_CUSTOMREQUEST => $method,
CURLOPT_RETURNTRANSFER => true,
CURLOPT_HTTPHEADER => $hdr,
CURLOPT_POSTFIELDS => $body === null ? null : json_encode($body),
]);
$payload = json_decode(curl_exec($ch), true);
$status = curl_getinfo($ch, CURLINFO_RESPONSE_CODE);
curl_close($ch);
if ($status >= 400) {
throw new Exception(($payload["error"]["code"] ?? "http_$status") . ": "
. ($payload["error"]["message"] ?? "request failed"));
}
return $payload["data"];
}
// a headless guest token (browse + estimate only on this app):
$guest = api("POST", "/guest", ["slug" => "chat-with-all"]);
echo $guest["token"], " ", $guest["guest_id"], " ", $guest["expires_at"], "\n";
// .NET 8+
using System.Net.Http.Json;
using System.Text.Json;
static class ChatWithAll
{
const string Api = "https://api.skillsafe.ai/v1/app-api";
static readonly HttpClient Http = new();
static ChatWithAll() =>
Http.DefaultRequestHeaders.Authorization =
new("Bearer", Environment.GetEnvironmentVariable("SKILLSAFE_TOKEN"));
public static async Task<JsonElement> ApiAsync(
HttpMethod method, string path, object? body = null,
(string Name, string Value)[]? headers = null)
{
var req = new HttpRequestMessage(method, Api + path);
if (body != null) req.Content = JsonContent.Create(body);
if (headers != null)
foreach (var h in headers) req.Headers.Add(h.Name, h.Value);
var res = await Http.SendAsync(req);
var json = await res.Content.ReadFromJsonAsync<JsonElement>();
if (!res.IsSuccessStatusCode)
{
var err = json.GetProperty("error");
throw new Exception($"{err.GetProperty("code")}: {err.GetProperty("message")}");
}
return json.GetProperty("data");
}
}
// a headless guest token (browse + estimate only on this app):
var guest = await ChatWithAll.ApiAsync(HttpMethod.Post, "/guest",
new { slug = "chat-with-all" });
Console.WriteLine(guest.GetProperty("token").GetString());
The browser app keeps this device's token in localStorage under
skillsafe_app_token:chat-with-all, on the app's own origin. The
token page reads and manages it for you. Treat a user token like a
password — it can spend your credits.
Step 2 — Who am I, and what's my balance
Returns the subject behind the token and its credit balance. Call it before a turn: a turn needs enough credits to cover the hold, not just the eventual charge.
{"ok": true,
"data": {
"subject_type": "user", // or "guest"
"subject_id": "…",
"credits": 48213
},
"meta": { … }}
curl -s "$API/me" -H "Authorization: Bearer $SKILLSAFE_TOKEN" | jq '.data'
me = api("GET", "/me")
print(me["subject_type"], me["subject_id"], me["credits"])
const me = await api("GET", "/me");
console.log(me.subject_type, me.subject_id, me.credits);
var me struct {
SubjectType string `json:"subject_type"`
SubjectID string `json:"subject_id"`
Credits int64 `json:"credits"`
}
err := call("GET", "/me", nil, &me, nil)
String envelope = api("GET", "/me", null, null);
// data.subject_type, data.subject_id, data.credits
me = api("GET", "/me")
puts "#{me["subject_type"]} #{me["subject_id"]}: #{me["credits"]} credits"
$me = api("GET", "/me");
echo "{$me['subject_type']}: {$me['credits']} credits\n";
var me = await ChatWithAll.ApiAsync(HttpMethod.Get, "/me");
Console.WriteLine($"{me.GetProperty("subject_type")}: {me.GetProperty("credits")} credits");
Credits are the platform's unit of account: $1.00 = 10,000 credits. A guest
token answers this call happily — it will just report
"subject_type": "guest" and, on this app, no usable balance.
Step 3 — List the live models
Two things make this endpoint unlike every other one on this page. It is on
/v1/models, not under /v1/app-api; and it is
public — verified to return 200 with no
Authorization header at all. It is the catalogue the app's lane pickers are
built from, and it is the only correct source of truth for which model ids exist right now.
{"ok": true,
"data": {
"default_model": "…",
"models": [
{
"id": "gpt-5.6-terra",
"label": "…",
"provider": "openai",
"available": true,
"unavailable_reason": null,
"rates": {"billed_usd_per_mtok": {"in": …, "out": …}},
"caps": {"max_input_tokens": …, "max_output_tokens": …, "wall_clock_ms": …}
}
],
"aliases": { … },
"pricing": { … }
}}
| Model field | Type | Notes |
|---|---|---|
id | string | The concrete model id. This is what goes in $model. |
label | string | Human-readable name, for your own UI. |
provider | string | openai, anthropic or workers-ai. The app colours a lane by this. |
available | boolean | Branch on this. false means the platform will refuse a run with this id. |
unavailable_reason | string or null | Why it is off, when available is false. Show it rather than inventing an explanation. |
rates.billed_usd_per_mtok.in | number | Billed USD per million input tokens. |
rates.billed_usd_per_mtok.out | number | Billed USD per million output tokens. Output is where the money goes. |
caps.max_input_tokens | number | Hard ceiling on the prompt. |
caps.max_output_tokens | number | Hard ceiling on the reply; the hold is sized from this. |
caps.wall_clock_ms | number | How long a single turn may run before the platform gives up. |
Aliases
data.aliases maps friendly names onto concrete ids, so a client can pin
claude-sonnet and follow the platform's choice of point release. The live map is
exactly:
| Alias | Resolves to |
|---|---|
gpt-sol | gpt-5.6-sol |
gpt-terra | gpt-5.6-terra |
gpt-luna | gpt-5.6-luna |
gemma-fast | @cf/google/gemma-4-26b-a4b-it |
claude-opus | claude-opus-5 |
claude-sonnet | claude-sonnet-5 |
claude-haiku | claude-haiku-4-5 |
claude-fable | claude-fable-5 |
"provider": "anthropic" currently reports
"available": false here — the claude-* ids and their aliases
are listed, but they will not run. Sending one to /estimate or to a session turn
returns 400 validation_error. Use a gpt-* model (for example
gpt-terra) instead. The lesson generalises: filter the catalogue on
available at runtime rather than hardcoding a list of ids, because
which models are live changes without a client release.
# public - no Authorization header, and NOT under /v1/app-api
curl -s "https://api.skillsafe.ai/v1/models" \
| jq '.data.models[] | select(.available) | {id, provider, out: .rates.billed_usd_per_mtok.out}'
# the alias map
curl -s "https://api.skillsafe.ai/v1/models" | jq '.data.aliases'
import requests
# public - no Authorization header, and NOT under /v1/app-api
cat = requests.get("https://api.skillsafe.ai/v1/models").json()["data"]
usable = [m for m in cat["models"] if m["available"]]
for m in usable:
print(m["id"], m["provider"], m["rates"]["billed_usd_per_mtok"]["out"])
for m in cat["models"]:
if not m["available"]:
print("skip", m["id"], "-", m["unavailable_reason"])
print("aliases:", cat["aliases"])
print("default:", cat["default_model"])
// public - no Authorization header, and NOT under /v1/app-api
const cat = (await (await fetch("https://api.skillsafe.ai/v1/models")).json()).data;
const usable = cat.models.filter((m) => m.available);
for (const m of usable) {
console.log(m.id, m.provider, m.rates.billed_usd_per_mtok.out);
}
// never hardcode ids - anthropic/* is available:false on this deployment
const pick = usable.find((m) => m.id === "gpt-5.6-terra") ?? usable[0];
type Model struct {
ID string `json:"id"`
Label string `json:"label"`
Provider string `json:"provider"`
Available bool `json:"available"`
UnavailableReason string `json:"unavailable_reason"`
Rates struct {
BilledUSDPerMTok struct {
In float64 `json:"in"`
Out float64 `json:"out"`
} `json:"billed_usd_per_mtok"`
} `json:"rates"`
Caps struct {
MaxInputTokens int `json:"max_input_tokens"`
MaxOutputTokens int `json:"max_output_tokens"`
WallClockMs int `json:"wall_clock_ms"`
} `json:"caps"`
}
// public - no Authorization header, and NOT under /v1/app-api
res, _ := http.Get("https://api.skillsafe.ai/v1/models")
defer res.Body.Close()
var env struct {
Data struct {
DefaultModel string `json:"default_model"`
Models []Model `json:"models"`
Aliases map[string]string `json:"aliases"`
} `json:"data"`
}
json.NewDecoder(res.Body).Decode(&env)
for _, m := range env.Data.Models {
if m.Available {
fmt.Println(m.ID, m.Provider, m.Rates.BilledUSDPerMTok.Out)
}
}
// public - no Authorization header, and NOT under /v1/app-api
var req = HttpRequest.newBuilder(URI.create("https://api.skillsafe.ai/v1/models")).build();
var res = HttpClient.newHttpClient().send(req, HttpResponse.BodyHandlers.ofString());
String envelope = res.body();
// data.models[] -> id, label, provider, available, unavailable_reason,
// rates.billed_usd_per_mtok.{in,out},
// caps.{max_input_tokens,max_output_tokens,wall_clock_ms}
// data.aliases -> {"gpt-terra":"gpt-5.6-terra", ...}
// Filter on `available` before offering a model in your UI.
require "net/http"
require "json"
# public - no Authorization header, and NOT under /v1/app-api
cat = JSON.parse(Net::HTTP.get(URI("https://api.skillsafe.ai/v1/models")))["data"]
cat["models"].select { |m| m["available"] }.each do |m|
puts "#{m["id"]} #{m["provider"]} #{m.dig("rates", "billed_usd_per_mtok", "out")}"
end
puts cat["aliases"]["gpt-terra"] # => "gpt-5.6-terra"
<?php
// public - no Authorization header, and NOT under /v1/app-api
$cat = json_decode(file_get_contents("https://api.skillsafe.ai/v1/models"), true)["data"];
foreach ($cat["models"] as $m) {
if (!$m["available"]) { continue; }
echo $m["id"], " ", $m["provider"], " ",
$m["rates"]["billed_usd_per_mtok"]["out"], "\n";
}
echo $cat["aliases"]["gpt-terra"], "\n"; // gpt-5.6-terra
// public - no Authorization header, and NOT under /v1/app-api
using var plain = new HttpClient();
var cat = (await plain.GetFromJsonAsync<JsonElement>("https://api.skillsafe.ai/v1/models"))
.GetProperty("data");
foreach (var m in cat.GetProperty("models").EnumerateArray())
{
if (!m.GetProperty("available").GetBoolean()) continue;
Console.WriteLine($"{m.GetProperty("id")} {m.GetProperty("provider")}");
}
var terra = cat.GetProperty("aliases").GetProperty("gpt-terra").GetString();
Step 4 — Price a turn before you pay for it
The body is the input object itself — there is no input
wrapper key. For this app the input is a message and a model:
POST /v1/app-api/estimate
{"content": "Explain CRDTs in two paragraphs.", "$model": "gpt-5.6-terra"}
Estimating is free: no job is created, no credits move, and a guest token is allowed. The response:
{"ok": true,
"data": {
"hold_credits": 1367,
"min_credits": 58,
"model": "gpt-5.6-terra",
"model_alias": null,
"markup_bps": 0,
"sponsor_enabled": false,
"byok": false
}}
| Field | Meaning |
|---|---|
hold_credits | The worst case: prompt tokens plus the model's full max_output_tokens. This is what gets held when you submit a turn, so this is the balance you need — not the eventual charge. |
min_credits | The floor: what the turn costs if the model answers in almost nothing. |
model | The concrete model id the run will use, after alias resolution. |
model_alias | The alias you asked for, or null. It is null for this app because the lanes pin concrete ids from /v1/models rather than aliases. |
markup_bps | The publisher's cut, in basis points. 0 for this app — turns bill at platform cost, the publisher takes nothing. |
sponsor_enabled | Whether the publisher is paying for runs instead of the caller. |
byok | Whether the run uses your own upstream provider key. |
Real measured numbers
For the exact input "Explain CRDTs in two paragraphs.", measured against the
live API. Note how far apart the models are — a hundredfold between the cheapest and
the dearest hold. This is precisely why the app shows a meter before you press Ask.
| Model | hold_credits | min_credits |
|---|---|---|
@cf/google/gemma-4-26b-a4b-it | 25 | 13 |
gpt-5.6-terra | 1367 | 58 |
gpt-5.6-sol | 2722 | 104 |
The hold is not the price. A two-paragraph answer from gpt-5.6-terra settles far
closer to min_credits than to hold_credits; the difference is
refunded when the job finishes. See step 6.
curl -s -X POST "$API/estimate" \
-H "Authorization: Bearer $SKILLSAFE_TOKEN" \
-H "Content-Type: application/json" \
-d '{"content":"Explain CRDTs in two paragraphs.","$model":"gpt-5.6-terra"}' \
| jq '.data'
# => {"hold_credits":1367,"min_credits":58,"model":"gpt-5.6-terra",
# "model_alias":null,"markup_bps":0,"sponsor_enabled":false,"byok":false}
# the body IS the input object - no {"input": ...} wrapper
est = api("POST", "/estimate", {
"content": "Explain CRDTs in two paragraphs.",
"$model": "gpt-5.6-terra",
})
print(est["hold_credits"], est["min_credits"], est["model"], est["markup_bps"])
# 1367 58 gpt-5.6-terra 0
me = api("GET", "/me")
if me["credits"] < est["hold_credits"]:
print("not enough credits to place the hold - the turn would 402")
// the body IS the input object - no { input: ... } wrapper
const est = await api("POST", "/estimate", {
content: "Explain CRDTs in two paragraphs.",
$model: "gpt-5.6-terra",
});
console.log(est.hold_credits, est.min_credits, est.model, est.markup_bps);
// 1367 58 gpt-5.6-terra 0
// price the whole lineup at once, like the app's meter does
const lineup = ["@cf/google/gemma-4-26b-a4b-it", "gpt-5.6-terra", "gpt-5.6-sol"];
const quotes = await Promise.all(lineup.map((m) =>
api("POST", "/estimate", { content: "Explain CRDTs in two paragraphs.", $model: m })));
console.log(quotes.reduce((sum, q) => sum + q.hold_credits, 0)); // total hold
// the body IS the input object - no {"input": ...} wrapper
var est struct {
HoldCredits int64 `json:"hold_credits"`
MinCredits int64 `json:"min_credits"`
Model string `json:"model"`
ModelAlias *string `json:"model_alias"`
MarkupBps int `json:"markup_bps"`
}
err := call("POST", "/estimate", map[string]any{
"content": "Explain CRDTs in two paragraphs.",
"$model": "gpt-5.6-terra",
}, &est, nil)
fmt.Println(est.HoldCredits, est.MinCredits, est.Model) // 1367 58 gpt-5.6-terra
// the body IS the input object - no {"input": ...} wrapper
String envelope = api("POST", "/estimate", """
{"content":"Explain CRDTs in two paragraphs.","$model":"gpt-5.6-terra"}""", null);
// data.hold_credits = 1367, data.min_credits = 58,
// data.model = "gpt-5.6-terra", data.model_alias = null, data.markup_bps = 0
# the body IS the input object - no {"input" => ...} wrapper
est = api("POST", "/estimate", {
"content" => "Explain CRDTs in two paragraphs.",
"$model" => "gpt-5.6-terra",
})
puts "hold #{est["hold_credits"]} / min #{est["min_credits"]} on #{est["model"]}"
# hold 1367 / min 58 on gpt-5.6-terra
<?php
// the body IS the input object - no ["input" => ...] wrapper
$est = api("POST", "/estimate", [
"content" => "Explain CRDTs in two paragraphs.",
"\$model" => "gpt-5.6-terra",
]);
printf("hold %d / min %d on %s (markup %d bps)\n",
$est["hold_credits"], $est["min_credits"], $est["model"], $est["markup_bps"]);
// hold 1367 / min 58 on gpt-5.6-terra (markup 0 bps)
// the body IS the input object - no { input = ... } wrapper.
// "$model" is not a legal C# identifier, so build the body as a dictionary.
var est = await ChatWithAll.ApiAsync(HttpMethod.Post, "/estimate",
new Dictionary<string, object> {
["content"] = "Explain CRDTs in two paragraphs.",
["$model"] = "gpt-5.6-terra",
});
Console.WriteLine($"hold {est.GetProperty("hold_credits")} / min {est.GetProperty("min_credits")}");
// hold 1367 / min 58
Step 5 — Open a session
A session is the server-kept conversation history. You post messages into it and the platform keeps the transcript, so a follow-up question costs a follow-up rather than a re-send of the whole thread. Creating a session is free — nothing is charged until you post a message into it (step 6).
POST /v1/app-api/sessions
{"title": "CRDT comparison"}
{"ok": true,
"data": {
"session": {
"session_id": "…",
"title": "CRDT comparison",
"created_at": "…"
}
}}
Note the extra nesting: the created session is at data.session, not at
data. The list endpoint mirrors it at data.sessions, with
meta.pagination alongside.
The browser app opens one session per lane, because a lane's history belongs to one model — switching a lane's model starts a fresh session rather than continuing a transcript half-written by a different model.
SESSION=$(curl -s -X POST "$API/sessions" \
-H "Authorization: Bearer $SKILLSAFE_TOKEN" \
-H "Content-Type: application/json" \
-d '{"title":"CRDT comparison"}' | jq -r '.data.session.session_id')
echo "$SESSION"
session = api("POST", "/sessions", {"title": "CRDT comparison"})["session"]
sid = session["session_id"]
const { session } = await api("POST", "/sessions", { title: "CRDT comparison" });
const sid = session.session_id;
var created struct {
Session struct {
SessionID string `json:"session_id"`
Title string `json:"title"`
} `json:"session"`
}
err := call("POST", "/sessions", map[string]string{"title": "CRDT comparison"}, &created, nil)
sid := created.Session.SessionID
String envelope = api("POST", "/sessions", """
{"title":"CRDT comparison"}""", null);
// the id is at data.session.session_id
session = api("POST", "/sessions", { title: "CRDT comparison" })["session"]
sid = session["session_id"]
<?php
$session = api("POST", "/sessions", ["title" => "CRDT comparison"])["session"];
$sid = $session["session_id"];
var created = await ChatWithAll.ApiAsync(HttpMethod.Post, "/sessions",
new { title = "CRDT comparison" });
var sid = created.GetProperty("session").GetProperty("session_id").GetString();
List, read and delete
GET /sessions returns data.sessions; GET /sessions/{id}
returns the session with its messages; DELETE /sessions/{id} drops it and its
transcript. All three are free.
curl -s "$API/sessions" -H "Authorization: Bearer $SKILLSAFE_TOKEN" | jq '.data.sessions'
curl -s "$API/sessions/$SESSION" -H "Authorization: Bearer $SKILLSAFE_TOKEN" | jq '.data'
curl -s -X DELETE "$API/sessions/$SESSION" -H "Authorization: Bearer $SKILLSAFE_TOKEN" | jq '.ok'
sessions = api("GET", "/sessions")["sessions"]
detail = api("GET", f"/sessions/{sid}")
api("DELETE", f"/sessions/{sid}")
const { sessions } = await api("GET", "/sessions");
const detail = await api("GET", `/sessions/${encodeURIComponent(sid)}`);
await api("DELETE", `/sessions/${encodeURIComponent(sid)}`);
var list struct {
Sessions []struct {
SessionID string `json:"session_id"`
Title string `json:"title"`
} `json:"sessions"`
}
call("GET", "/sessions", nil, &list, nil)
call("GET", "/sessions/"+url.PathEscape(sid), nil, nil, nil)
call("DELETE", "/sessions/"+url.PathEscape(sid), nil, nil, nil)
String list = api("GET", "/sessions", null, null); // data.sessions[]
String detail = api("GET", "/sessions/" + sid, null, null);
api("DELETE", "/sessions/" + sid, null, null);
sessions = api("GET", "/sessions")["sessions"]
detail = api("GET", "/sessions/#{sid}")
api("DELETE", "/sessions/#{sid}")
<?php
$sessions = api("GET", "/sessions")["sessions"];
$detail = api("GET", "/sessions/" . rawurlencode($sid));
api("DELETE", "/sessions/" . rawurlencode($sid));
var list = await ChatWithAll.ApiAsync(HttpMethod.Get, "/sessions");
foreach (var s in list.GetProperty("sessions").EnumerateArray())
Console.WriteLine(s.GetProperty("session_id"));
await ChatWithAll.ApiAsync(HttpMethod.Get, $"/sessions/{Uri.EscapeDataString(sid!)}");
await ChatWithAll.ApiAsync(HttpMethod.Delete, $"/sessions/{Uri.EscapeDataString(sid!)}");
Step 6 — Run a turn (this is the metered call)
Everything up to here was free. This is the one call that spends credits.
POST /v1/app-api/sessions/{session_id}/messages
Idempotency-Key: <your key>
{"content": "Explain CRDTs in two paragraphs.",
"$model": "gpt-5.6-terra",
"stream": true}
| Body field | Type | Notes |
|---|---|---|
content | string, required | The user's message. The session's earlier turns are supplied by the platform — do not re-send the transcript. |
$model | string | Per-turn model override: a concrete id (or alias) from /v1/models. An id with available:false is rejected with 400 validation_error. |
stream | boolean | true switches the response to text/event-stream (step 7). Omit it for the job + poll shape below. |
Without stream: submit, then poll
The POST returns immediately with a job handle:
{"ok": true, "data": {"job_id": "…", "session_id": "…"}}
Poll GET /v1/app-api/jobs/{job_id} (about once a second is plenty) until
status is succeeded or failed — those are the
only terminal states. The terminal job carries:
| Job field | Meaning |
|---|---|
status | succeeded or failed when terminal; anything else means keep polling. |
charged_credits | What the turn actually cost, after settlement. This is the number to show the user. |
output.output | The reply text. Note the double output — the job's output object has an output string inside it. |
truncated | true if the reply hit the output cap because the balance could not fund the full hold (see below). |
error | Present on failed; same {code, message} shape as the envelope error. |
Hold, settle, refund
At submit the platform places a hold for the worst case — the same
number /estimate reported as hold_credits,
i.e. your prompt plus the model's entire max_output_tokens. When the job
finishes it settles at the real usage, which lands in
charged_credits, and the rest of the hold is refunded. So a
balance of 1,400 credits is enough to start a gpt-5.6-terra turn that ends up
costing 90, but a balance of 1,000 is not — you get 402 before the model
ever runs.
If the balance can fund some of the hold but not all of it, the platform scales the
run's output cap down to what you can afford rather than refusing outright. The reply then
stops early and the job comes back with "truncated": true. That is not an error
and you were not overcharged — it means "your balance bought this much answer". The
browser app renders it as a "Response cut short by your balance" line with a top-up link;
yours should say something equally plain rather than silently presenting a half answer as
complete.
Idempotency-Key header
A turn is the only call here that costs money, which makes a retry the only dangerous retry.
With an Idempotency-Key, a repeat of the same request replays the
original turn — same job, same reply, charged once. Without it, a network
blip, a proxy timeout or a user double-click buys the same answer twice. Send it on
every turn, not only on retries: the key has to be present on the first
attempt for the replay to have anything to match. Derive it deterministically from what the
turn is — session id, message text, and the attempt number of the logical turn
— so that a retry recomputes the identical key while a genuinely new question does not.
# submit (no "stream" -> job + poll), with an idempotency key
KEY=$(printf '%s|%s|%s' "$SESSION" "Explain CRDTs in two paragraphs." 1 | shasum -a 256 | cut -c1-32)
JOB=$(curl -s -X POST "$API/sessions/$SESSION/messages" \
-H "Authorization: Bearer $SKILLSAFE_TOKEN" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: $KEY" \
-d '{"content":"Explain CRDTs in two paragraphs.","$model":"gpt-5.6-terra"}' \
| jq -r '.data.job_id')
# poll until terminal
while :; do
J=$(curl -s "$API/jobs/$JOB" -H "Authorization: Bearer $SKILLSAFE_TOKEN" | jq '.data')
S=$(echo "$J" | jq -r '.status')
[ "$S" = "succeeded" ] || [ "$S" = "failed" ] && break
sleep 1
done
echo "$J" | jq -r '.output.output'
echo "$J" | jq '{charged_credits, truncated}'
import hashlib, time
def turn_key(session_id, content, attempt=1):
raw = f"{session_id}|{content}|{attempt}".encode()
return hashlib.sha256(raw).hexdigest()[:32]
content = "Explain CRDTs in two paragraphs."
sub = api("POST", f"/sessions/{sid}/messages",
{"content": content, "$model": "gpt-5.6-terra"},
**{"Idempotency-Key": turn_key(sid, content)})
job = api("GET", f"/jobs/{sub['job_id']}")
while job["status"] not in ("succeeded", "failed"):
time.sleep(1)
job = api("GET", f"/jobs/{sub['job_id']}")
if job["status"] == "succeeded":
print(job["output"]["output"])
print("charged", job["charged_credits"], "credits")
if job.get("truncated"):
print("cut short by your balance - top up to continue")
else:
print("failed:", job["error"])
import { createHash } from "node:crypto";
const turnKey = (sessionId, content, attempt = 1) =>
createHash("sha256").update(`${sessionId}|${content}|${attempt}`).digest("hex").slice(0, 32);
const content = "Explain CRDTs in two paragraphs.";
const sub = await api("POST", `/sessions/${encodeURIComponent(sid)}/messages`,
{ content, $model: "gpt-5.6-terra" },
{ "Idempotency-Key": turnKey(sid, content) });
let job = await api("GET", `/jobs/${encodeURIComponent(sub.job_id)}`);
while (job.status !== "succeeded" && job.status !== "failed") {
await new Promise((r) => setTimeout(r, 1000));
job = await api("GET", `/jobs/${encodeURIComponent(sub.job_id)}`);
}
if (job.status === "succeeded") {
console.log(job.output.output);
console.log("charged", job.charged_credits, "credits");
if (job.truncated) console.log("cut short by your balance - top up to continue");
} else {
console.error("failed:", job.error);
}
func turnKey(sessionID, content string, attempt int) string {
sum := sha256.Sum256([]byte(fmt.Sprintf("%s|%s|%d", sessionID, content, attempt)))
return hex.EncodeToString(sum[:])[:32]
}
content := "Explain CRDTs in two paragraphs."
var sub struct {
JobID string `json:"job_id"`
SessionID string `json:"session_id"`
}
err := call("POST", "/sessions/"+url.PathEscape(sid)+"/messages", map[string]any{
"content": content,
"$model": "gpt-5.6-terra",
}, &sub, map[string]string{"Idempotency-Key": turnKey(sid, content, 1)})
var job struct {
Status string `json:"status"`
ChargedCredits int64 `json:"charged_credits"`
Truncated bool `json:"truncated"`
Output struct {
Output string `json:"output"`
} `json:"output"`
}
for {
if err := call("GET", "/jobs/"+url.PathEscape(sub.JobID), nil, &job, nil); err != nil {
break
}
if job.Status == "succeeded" || job.Status == "failed" {
break
}
time.Sleep(time.Second)
}
fmt.Println(job.Output.Output, job.ChargedCredits, job.Truncated)
static String turnKey(String sessionId, String content, int attempt) throws Exception {
var md = java.security.MessageDigest.getInstance("SHA-256");
var d = md.digest((sessionId + "|" + content + "|" + attempt).getBytes("UTF-8"));
var sb = new StringBuilder();
for (byte b : d) sb.append(String.format("%02x", b));
return sb.substring(0, 32);
}
String content = "Explain CRDTs in two paragraphs.";
String sub = api("POST", "/sessions/" + sid + "/messages",
"{\"content\":\"" + content + "\",\"$model\":\"gpt-5.6-terra\"}",
Map.of("Idempotency-Key", turnKey(sid, content, 1)));
// data.job_id -> poll GET /jobs/{job_id} once a second until
// data.status is "succeeded" or "failed"; then read
// data.output.output, data.charged_credits, data.truncated
require "digest"
def turn_key(session_id, content, attempt = 1)
Digest::SHA256.hexdigest("#{session_id}|#{content}|#{attempt}")[0, 32]
end
content = "Explain CRDTs in two paragraphs."
sub = api("POST", "/sessions/#{sid}/messages",
{ "content" => content, "$model" => "gpt-5.6-terra" },
{ "Idempotency-Key" => turn_key(sid, content) })
job = api("GET", "/jobs/#{sub["job_id"]}")
until %w[succeeded failed].include?(job["status"])
sleep 1
job = api("GET", "/jobs/#{sub["job_id"]}")
end
puts job.dig("output", "output")
puts "charged #{job["charged_credits"]} credits (truncated: #{job["truncated"]})"
<?php
function turn_key(string $sessionId, string $content, int $attempt = 1): string {
return substr(hash("sha256", "$sessionId|$content|$attempt"), 0, 32);
}
$content = "Explain CRDTs in two paragraphs.";
$sub = api("POST", "/sessions/" . rawurlencode($sid) . "/messages",
["content" => $content, "\$model" => "gpt-5.6-terra"],
["Idempotency-Key: " . turn_key($sid, $content)]);
do {
sleep(1);
$job = api("GET", "/jobs/" . rawurlencode($sub["job_id"]));
} while (!in_array($job["status"], ["succeeded", "failed"], true));
echo $job["output"]["output"], "\n";
printf("charged %d credits%s\n", $job["charged_credits"],
!empty($job["truncated"]) ? " (cut short by your balance)" : "");
using System.Security.Cryptography;
using System.Text;
static string TurnKey(string sessionId, string content, int attempt = 1)
{
var bytes = SHA256.HashData(Encoding.UTF8.GetBytes($"{sessionId}|{content}|{attempt}"));
return Convert.ToHexString(bytes)[..32].ToLowerInvariant();
}
var content = "Explain CRDTs in two paragraphs.";
var sub = await ChatWithAll.ApiAsync(HttpMethod.Post,
$"/sessions/{Uri.EscapeDataString(sid!)}/messages",
new Dictionary<string, object> { ["content"] = content, ["$model"] = "gpt-5.6-terra" },
new[] { ("Idempotency-Key", TurnKey(sid!, content)) });
var jobId = sub.GetProperty("job_id").GetString();
JsonElement job;
string status;
do
{
await Task.Delay(1000);
job = await ChatWithAll.ApiAsync(HttpMethod.Get, $"/jobs/{Uri.EscapeDataString(jobId!)}");
status = job.GetProperty("status").GetString()!;
} while (status != "succeeded" && status != "failed");
Console.WriteLine(job.GetProperty("output").GetProperty("output").GetString());
Console.WriteLine($"charged {job.GetProperty("charged_credits")} credits");
Deriving the key
A worked example, in words, of what the code above does. A lane is about to send message
number 1 of session ses_9f2c, with the text
"Explain CRDTs in two paragraphs.". Concatenate them with a separator that
cannot appear in an id — ses_9f2c|Explain CRDTs in two paragraphs.|1
— hash that with SHA-256, and take the first 32 hex characters. Retry that same turn
after a timeout and you recompute the same 32 characters, so the platform replays the
original job and charges nothing extra. Ask a different question, or deliberately
re-ask the same one as a new turn (bump attempt to 2), and the key changes, so
it bills as the new turn it is.
Idempotent replays also change the streaming response: a replayed turn may come back as plain
JSON rather than text/event-stream, because there is nothing left to stream. The
app's SDK handles this by checking the response Content-Type before deciding to
parse SSE — do the same.
Step 7 — The streaming format
{"stream": true}
With "stream": true in the body, the response is
text/event-stream instead of JSON. The billing is identical — same hold,
same settlement, same charged_credits — you just see the text as it is
produced. Check the response Content-Type first: a replayed
idempotent turn, or an early rejection, comes back as ordinary JSON with the usual envelope.
Frame syntax
Standard SSE. Frames are separated by a blank line (\n\n); inside a frame, the
line beginning event: names the event and the lines beginning data:
carry a JSON payload (concatenated and trimmed if there is more than one). A frame with no
data: line, or whose data is not valid JSON, is skipped rather than fatal —
so keepalive comments cannot break your reader.
event: job
data: {"job_id":"…","session_id":"…"}
event: delta
data: {"text":"A CRDT is a data structure "}
event: delta
data: {"text":"that can be replicated…"}
event: done
data: {"job_id":"…","status":"succeeded","charged_credits":91,
"output":{"output":"A CRDT is a data structure that can be replicated…"},
"truncated":false}
The events, exactly as the app's SDK reads them
These are the four event names handled by _readSse in the app's own
sdk.js. Anything else — including an unnamed frame, which SSE defaults to
message — is parsed and then ignored, so the platform can add events
without breaking existing clients.
| Event | Payload the SDK reads | What it does with it |
|---|---|---|
delta | text (string) | Appends data.text || "" to the answer — this is the incremental output. A missing text is treated as an empty string, not an error. |
job | the whole payload (carries job_id, session_id) | Handed to the caller's onJob callback. It means the hold is placed and the model is running; the app uses it to switch its lane status from "calling" to "waiting". |
done | the whole payload (fields below) | Becomes the stream's return value. Terminal. |
pending | the whole payload | Treated identically to done — it also becomes the return value. It is how a stream that could not finish inline hands you back a job handle to poll. |
error | message, code, job_id | Recorded, and thrown once the stream closes: an Error whose message is data.message (falling back to "job failed"), carrying code and job_id. |
The final done payload
| Field | Meaning |
|---|---|
job_id | The job this stream belonged to — poll it if you need to re-read the result later. |
status | succeeded or failed. |
charged_credits | The settled cost, after the unused hold is refunded. |
output | The job output object; the reply text is output.output. Use it as the source of truth if you did not accumulate the deltas — the app falls back to it whenever no delta ever arrived. |
truncated | true when the reply hit the output cap because the balance could not fund the full hold. |
curl -N -s -X POST "$API/sessions/$SESSION/messages" \
-H "Authorization: Bearer $SKILLSAFE_TOKEN" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: $KEY" \
-d '{"content":"Explain CRDTs in two paragraphs.","$model":"gpt-5.6-terra","stream":true}'
# -N disables curl's buffering so frames arrive as they are produced:
# event: job / data: {"job_id":"...","session_id":"..."}
# event: delta / data: {"text":"..."}
# event: done / data: {"job_id":"...","status":"succeeded","charged_credits":91,
# "output":{"output":"..."},"truncated":false}
import json, requests
def stream_turn(sid, content, model, key):
res = requests.post(
f"{API}/sessions/{sid}/messages",
json={"content": content, "$model": model, "stream": True},
headers={"Authorization": f"Bearer {TOKEN}",
"Idempotency-Key": key},
stream=True)
if "text/event-stream" not in res.headers.get("content-type", ""):
return res.json()["data"] # replay or early rejection: plain JSON
result, failure, frame = None, None, []
for line in res.iter_lines(decode_unicode=True):
if line:
frame.append(line)
continue
event, data_str = "message", ""
for ln in frame:
if ln.startswith("event:"):
event = ln[6:].strip()
elif ln.startswith("data:"):
data_str += ln[5:].strip()
frame = []
if not data_str:
continue
try:
data = json.loads(data_str)
except ValueError:
continue
if event == "delta":
print(data.get("text", ""), end="", flush=True)
elif event == "job":
pass # hold placed, model running
elif event in ("done", "pending"):
result = data
elif event == "error":
failure = data
if failure:
raise RuntimeError(f"{failure.get('code')}: {failure.get('message', 'job failed')}")
return result
done = stream_turn(sid, "Explain CRDTs in two paragraphs.", "gpt-5.6-terra", key)
print("\ncharged", done["charged_credits"], "truncated", done["truncated"])
// This is what sdk.js does internally (_readSse), unrolled.
async function streamTurn(sid, content, model, key, onDelta) {
const res = await fetch(`${API}/sessions/${encodeURIComponent(sid)}/messages`, {
method: "POST",
headers: {
Authorization: `Bearer ${TOKEN}`,
"Content-Type": "application/json",
"Idempotency-Key": key,
},
body: JSON.stringify({ content, $model: model, stream: true }),
});
const ctype = res.headers.get("content-type") || "";
if (ctype.indexOf("text/event-stream") === -1) {
const json = await res.json();
if (!res.ok) throw new Error(json.error?.message);
return json.data; // replay or early rejection
}
const reader = res.body.getReader();
const decoder = new TextDecoder();
let buffer = "", result = null, failure = null;
for (;;) {
const chunk = await reader.read();
if (chunk.done) break;
buffer += decoder.decode(chunk.value, { stream: true });
let idx;
while ((idx = buffer.indexOf("\n\n")) >= 0) {
const raw = buffer.slice(0, idx);
buffer = buffer.slice(idx + 2);
let event = "message", dataStr = "";
for (const line of raw.split("\n")) {
if (line.startsWith("event:")) event = line.slice(6).trim();
else if (line.startsWith("data:")) dataStr += line.slice(5).trim();
}
if (!dataStr) continue;
let data;
try { data = JSON.parse(dataStr); } catch { continue; }
if (event === "delta") onDelta(data.text || "");
else if (event === "done" || event === "pending") result = data;
else if (event === "error") failure = data;
}
}
if (failure) throw Object.assign(new Error(failure.message || "job failed"),
{ code: failure.code, job_id: failure.job_id });
return result; // { job_id, status, charged_credits, output, truncated }
}
func streamTurn(sid, content, model, key string, onDelta func(string)) (map[string]any, error) {
body, _ := json.Marshal(map[string]any{
"content": content, "$model": model, "stream": true,
})
req, _ := http.NewRequest("POST",
API+"/sessions/"+url.PathEscape(sid)+"/messages", bytes.NewReader(body))
req.Header.Set("Authorization", "Bearer "+token)
req.Header.Set("Content-Type", "application/json")
req.Header.Set("Idempotency-Key", key)
res, err := http.DefaultClient.Do(req)
if err != nil {
return nil, err
}
defer res.Body.Close()
if !strings.Contains(res.Header.Get("Content-Type"), "text/event-stream") {
var env struct{ Data map[string]any `json:"data"` }
json.NewDecoder(res.Body).Decode(&env)
return env.Data, nil // replay or early rejection
}
var result, failure map[string]any
sc := bufio.NewScanner(res.Body)
sc.Buffer(make([]byte, 1<<20), 1<<20)
event, dataStr := "message", ""
flush := func() {
defer func() { event, dataStr = "message", "" }()
if dataStr == "" {
return
}
var data map[string]any
if json.Unmarshal([]byte(dataStr), &data) != nil {
return
}
switch event {
case "delta":
s, _ := data["text"].(string)
onDelta(s)
case "done", "pending":
result = data
case "error":
failure = data
}
}
for sc.Scan() {
line := sc.Text()
switch {
case line == "":
flush()
case strings.HasPrefix(line, "event:"):
event = strings.TrimSpace(line[6:])
case strings.HasPrefix(line, "data:"):
dataStr += strings.TrimSpace(line[5:])
}
}
flush()
if failure != nil {
return nil, fmt.Errorf("%v: %v", failure["code"], failure["message"])
}
return result, nil // job_id, status, charged_credits, output, truncated
}
// Java 17+: stream the body line by line and reassemble SSE frames.
var body = "{\"content\":\"" + content + "\",\"$model\":\"gpt-5.6-terra\",\"stream\":true}";
var req = HttpRequest.newBuilder(URI.create(API + "/sessions/" + sid + "/messages"))
.header("Authorization", "Bearer " + TOKEN)
.header("Content-Type", "application/json")
.header("Idempotency-Key", turnKey(sid, content, 1))
.POST(HttpRequest.BodyPublishers.ofString(body))
.build();
var res = HTTP.send(req, HttpResponse.BodyHandlers.ofLines());
if (!res.headers().firstValue("content-type").orElse("").contains("text/event-stream")) {
// replay or early rejection: an ordinary {"ok":..,"data":..} envelope
}
var event = new StringBuilder("message");
var data = new StringBuilder();
res.body().forEach(line -> {
if (line.isEmpty()) {
// frame end: parse `data` as JSON and dispatch on `event`
// "delta" -> append data.text
// "job" -> hold placed, model running
// "done" | "pending"-> terminal result: job_id, status,
// charged_credits, output.output, truncated
// "error" -> data.message / data.code / data.job_id
event.setLength(0); event.append("message");
data.setLength(0);
} else if (line.startsWith("event:")) {
event.setLength(0); event.append(line.substring(6).trim());
} else if (line.startsWith("data:")) {
data.append(line.substring(5).trim());
}
});
require "net/http"
require "json"
def stream_turn(sid, content, model, key)
uri = URI("#{API}/sessions/#{sid}/messages")
req = Net::HTTP::Post.new(uri)
req["Authorization"] = "Bearer #{TOKEN}"
req["Content-Type"] = "application/json"
req["Idempotency-Key"] = key
req.body = { "content" => content, "$model" => model, "stream" => true }.to_json
result = nil
failure = nil
Net::HTTP.start(uri.host, uri.port, use_ssl: true) do |http|
http.request(req) do |res|
unless res["content-type"].to_s.include?("text/event-stream")
return JSON.parse(res.body)["data"] # replay or early rejection
end
buffer = ""
res.read_body do |chunk|
buffer += chunk
while (idx = buffer.index("\n\n"))
frame = buffer.slice!(0, idx + 2)
event = "message"
data_str = +""
frame.split("\n").each do |line|
if line.start_with?("event:") then event = line[6..].strip
elsif line.start_with?("data:") then data_str << line[5..].strip
end
end
next if data_str.empty?
data = (JSON.parse(data_str) rescue next)
case event
when "delta" then print data["text"].to_s
when "job" then nil # hold placed, model running
when "done", "pending" then result = data
when "error" then failure = data
end
end
end
end
end
raise "#{failure["code"]}: #{failure["message"] || "job failed"}" if failure
result
end
<?php
function stream_turn(string $sid, string $content, string $model, string $key): ?array {
global $TOKEN;
$result = null;
$failure = null;
$buffer = "";
$ch = curl_init(API . "/sessions/" . rawurlencode($sid) . "/messages");
curl_setopt_array($ch, [
CURLOPT_POST => true,
CURLOPT_HTTPHEADER => [
"Authorization: Bearer $TOKEN",
"Content-Type: application/json",
"Idempotency-Key: $key",
],
CURLOPT_POSTFIELDS => json_encode([
"content" => $content, "\$model" => $model, "stream" => true,
]),
CURLOPT_WRITEFUNCTION => function ($ch, $chunk) use (&$buffer, &$result, &$failure) {
$buffer .= $chunk;
while (($idx = strpos($buffer, "\n\n")) !== false) {
$frame = substr($buffer, 0, $idx);
$buffer = substr($buffer, $idx + 2);
$event = "message";
$dataStr = "";
foreach (explode("\n", $frame) as $line) {
if (str_starts_with($line, "event:")) {
$event = trim(substr($line, 6));
} elseif (str_starts_with($line, "data:")) {
$dataStr .= trim(substr($line, 5));
}
}
if ($dataStr === "") { continue; }
$data = json_decode($dataStr, true);
if ($data === null) { continue; }
if ($event === "delta") { echo $data["text"] ?? ""; }
elseif ($event === "done" || $event === "pending") { $result = $data; }
elseif ($event === "error") { $failure = $data; }
}
return strlen($chunk);
},
]);
curl_exec($ch);
curl_close($ch);
if ($failure) {
throw new Exception(($failure["code"] ?? "error") . ": "
. ($failure["message"] ?? "job failed"));
}
return $result; // job_id, status, charged_credits, output, truncated
}
// .NET 8+
static async Task<JsonElement?> StreamTurnAsync(
HttpClient http, string sid, string content, string model, string key, Action<string> onDelta)
{
var req = new HttpRequestMessage(HttpMethod.Post,
$"https://api.skillsafe.ai/v1/app-api/sessions/{Uri.EscapeDataString(sid)}/messages")
{
Content = JsonContent.Create(new Dictionary<string, object> {
["content"] = content, ["$model"] = model, ["stream"] = true,
}),
};
req.Headers.Add("Idempotency-Key", key);
var res = await http.SendAsync(req, HttpCompletionOption.ResponseHeadersRead);
if (res.Content.Headers.ContentType?.MediaType != "text/event-stream")
{
var plain = await res.Content.ReadFromJsonAsync<JsonElement>();
return plain.GetProperty("data"); // replay or early rejection
}
using var reader = new StreamReader(await res.Content.ReadAsStreamAsync());
JsonElement? result = null, failure = null;
string evt = "message", dataStr = "";
while (await reader.ReadLineAsync() is { } line)
{
if (line.Length == 0)
{
if (dataStr.Length > 0)
{
try
{
var data = JsonDocument.Parse(dataStr).RootElement.Clone();
if (evt == "delta")
onDelta(data.TryGetProperty("text", out var t) ? t.GetString() ?? "" : "");
else if (evt is "done" or "pending") result = data;
else if (evt == "error") failure = data;
}
catch (JsonException) { /* skip unparseable frame */ }
}
evt = "message"; dataStr = "";
}
else if (line.StartsWith("event:")) evt = line[6..].Trim();
else if (line.StartsWith("data:")) dataStr += line[5..].Trim();
}
if (failure is { } f)
throw new Exception(f.GetProperty("message").GetString() ?? "job failed");
return result; // job_id, status, charged_credits, output, truncated
}
Step 8 — The per-user key-value store
A small JSON store scoped to (this app, this subject). Chat with All keeps its
comparison history in it under the single key history, which is why your
lineups follow you from one device to the next once you are signed in.
| Call | Body | Returns |
|---|---|---|
PUT /data/{key} | {"value": <any JSON>} | Stores (and replaces) the value at that key. |
GET /data/{key} | — | {"data": {"value": …}}. A key that was never written is 404 — treat it as "empty", not as an error. |
DELETE /data/{key} | — | Removes the key. |
GET /data | — | {"data": {"keys": […]}} — the key names, not the values. |
Metered, but barely. These calls bill in nanodollars — roughly
400 nanos per write and 100 nanos per read. That is
negligible next to a model turn (a nanodollar is a billionth of a dollar), but it is not
literally free, so don't put a GET /data/{key} inside a render loop or poll it
every keystroke.
# write the comparison history
curl -s -X PUT "$API/data/history" \
-H "Authorization: Bearer $SKILLSAFE_TOKEN" \
-H "Content-Type: application/json" \
-d '{"value":[{"id":"cmp_x","ts":"2026-08-07T00:00:00Z","prompt":"Explain CRDTs in two paragraphs.","lanes":[{"model":"gpt-5.6-terra","label":"GPT-5.6 Terra","status":"succeeded","text":"...","chars":3,"elapsed_ms":4200,"charged_credits":190,"truncated":false,"error":null}]}]}'
# read it back (404 simply means "nothing stored yet")
curl -s "$API/data/history" -H "Authorization: Bearer $SKILLSAFE_TOKEN" | jq '.data.value'
# list keys, then drop one
curl -s "$API/data" -H "Authorization: Bearer $SKILLSAFE_TOKEN" | jq '.data.keys'
curl -s -X DELETE "$API/data/history" -H "Authorization: Bearer $SKILLSAFE_TOKEN" | jq '.ok'
lane = {"model": "gpt-5.6-terra", "label": "GPT-5.6 Terra", "status": "succeeded",
"text": "...", "chars": 3, "elapsed_ms": 4200,
"charged_credits": 190, "truncated": False, "error": None}
history = [{"id": "cmp_x", "ts": "2026-08-07T00:00:00Z",
"prompt": "Explain CRDTs in two paragraphs.", "lanes": [lane]}]
api("PUT", "/data/history", {"value": history})
try:
history = api("GET", "/data/history")["value"]
except RuntimeError:
history = [] # a missing key is a 404 - that just means "empty"
keys = api("GET", "/data")["keys"]
api("DELETE", "/data/history")
const lane = { model: "gpt-5.6-terra", label: "GPT-5.6 Terra", status: "succeeded",
text: "...", chars: 3, elapsed_ms: 4200,
charged_credits: 190, truncated: false, error: null };
const history = [{ id: "cmp_x", ts: "2026-08-07T00:00:00Z",
prompt: "Explain CRDTs in two paragraphs.", lanes: [lane] }];
await api("PUT", "/data/history", { value: history });
// a missing key is 404 - that just means "empty"
const stored = await api("GET", "/data/history")
.then((d) => d.value)
.catch((e) => (e.status === 404 ? [] : Promise.reject(e)));
const { keys } = await api("GET", "/data");
await api("DELETE", "/data/history");
type Lane struct {
Model string `json:"model"`
Label string `json:"label"`
Status string `json:"status"`
Text string `json:"text"`
Chars int `json:"chars"`
ElapsedMs int `json:"elapsed_ms"`
ChargedCredits *int `json:"charged_credits"` // nil = never settled
Truncated bool `json:"truncated"`
Error *string `json:"error"`
}
type Entry struct {
ID string `json:"id"`
Ts string `json:"ts"`
Prompt string `json:"prompt"`
Lanes []Lane `json:"lanes"`
}
cost := 190
history := []Entry{{ID: "cmp_x", Ts: "2026-08-07T00:00:00Z",
Prompt: "Explain CRDTs in two paragraphs.",
Lanes: []Lane{{Model: "gpt-5.6-terra", Label: "GPT-5.6 Terra", Status: "succeeded",
Text: "...", Chars: 3, ElapsedMs: 4200, ChargedCredits: &cost}}}}
call("PUT", "/data/history", map[string]any{"value": history}, nil, nil)
var got struct {
Value []Entry `json:"value"`
}
// a missing key returns 404 - treat that as an empty history
_ = call("GET", "/data/history", nil, &got, nil)
var keys struct {
Keys []string `json:"keys"`
}
call("GET", "/data", nil, &keys, nil)
call("DELETE", "/data/history", nil, nil, nil)
api("PUT", "/data/history", """
{"value":[{"id":"cmp_x","ts":"2026-08-07T00:00:00Z","prompt":"Explain CRDTs in two paragraphs.","lanes":[{"model":"gpt-5.6-terra","label":"GPT-5.6 Terra","status":"succeeded","text":"...","chars":3,"elapsed_ms":4200,"charged_credits":190,"truncated":false,"error":null}]}]}""", null);
String stored = api("GET", "/data/history", null, null); // data.value; 404 = never written
String keys = api("GET", "/data", null, null); // data.keys[]
api("DELETE", "/data/history", null, null);
lane = { "model" => "gpt-5.6-terra", "label" => "GPT-5.6 Terra", "status" => "succeeded",
"text" => "...", "chars" => 3, "elapsed_ms" => 4200,
"charged_credits" => 190, "truncated" => false, "error" => nil }
history = [{ "id" => "cmp_x", "ts" => "2026-08-07T00:00:00Z",
"prompt" => "Explain CRDTs in two paragraphs.", "lanes" => [lane] }]
api("PUT", "/data/history", { "value" => history })
stored = begin
api("GET", "/data/history")["value"]
rescue RuntimeError
[] # a missing key is a 404 - that just means "empty"
end
keys = api("GET", "/data")["keys"]
api("DELETE", "/data/history")
<?php
$lane = ["model" => "gpt-5.6-terra", "label" => "GPT-5.6 Terra", "status" => "succeeded",
"text" => "...", "chars" => 3, "elapsed_ms" => 4200,
"charged_credits" => 190, "truncated" => false, "error" => null];
$history = [["id" => "cmp_x", "ts" => "2026-08-07T00:00:00Z",
"prompt" => "Explain CRDTs in two paragraphs.", "lanes" => [$lane]]];
api("PUT", "/data/history", ["value" => $history]);
try {
$stored = api("GET", "/data/history")["value"];
} catch (Exception $e) {
$stored = []; // a missing key is a 404 - that just means "empty"
}
$keys = api("GET", "/data")["keys"];
api("DELETE", "/data/history");
var lane = new {
model = "gpt-5.6-terra", label = "GPT-5.6 Terra", status = "succeeded",
text = "...", chars = 3, elapsed_ms = 4200,
charged_credits = (int?)190, truncated = false, error = (string?)null,
};
var history = new[] {
new { id = "cmp_x", ts = "2026-08-07T00:00:00Z",
prompt = "Explain CRDTs in two paragraphs.", lanes = new[] { lane } },
};
await ChatWithAll.ApiAsync(HttpMethod.Put, "/data/history", new { value = history });
JsonElement? stored = null;
try { stored = (await ChatWithAll.ApiAsync(HttpMethod.Get, "/data/history")).GetProperty("value"); }
catch (Exception) { /* 404: nothing stored yet */ }
var keys = await ChatWithAll.ApiAsync(HttpMethod.Get, "/data");
await ChatWithAll.ApiAsync(HttpMethod.Delete, "/data/history");
Errors
Failures carry error.code, error.message,
error.status and sometimes error.details. Branch on
error.code — the message is written for humans and may be reworded.
| Status | code | What it means here, and what to do |
|---|---|---|
400 |
validation_error |
The body is malformed, a required field is missing, or — the common case on this app — $model names a model that is available: false. Every anthropic-provider id is currently in that state, so this is what you get for sending claude-opus-5. Re-read /v1/models and pick something with available: true. |
401 |
unauthorized |
No Authorization header, or a token that has expired (guest tokens carry expires_at). Mint a fresh one — step 1. Note that /v1/models needs no token at all, so a 401 from there means you are calling the wrong URL. |
402 |
insufficient_credits |
The balance cannot cover the hold for this turn — not the eventual charge. Compare /me's credits against /estimate's hold_credits before you submit, or offer a cheaper model. Guests on this app have no wallet, so every turn they attempt lands here. |
404 |
not_found |
Usually the wrong route — see the admonition in Basics: there is no /apps/{slug}/ segment. Otherwise: an unknown session_id or job_id, or a /data/{key} that was never written (which is normal and means "empty"). |
429 |
rate_limited |
Too many requests. The app-data endpoints share a budget of 120 requests per minute per IP. Back off exponentially and retry; batch your writes rather than storing on every keystroke. |
503 |
service_unavailable |
The platform is out of upstream capacity for that provider, or the model is briefly down. Nothing is wrong with your account and nothing was charged. Retry later, or switch that lane to another model — which is exactly what the app suggests. |
On a metered turn, a failure before the model runs releases the hold in full — you are
never charged for a 4xx. A turn that fails mid-generation settles at
what was actually produced, and the job's charged_credits tells you what that
was.
What this app actually sends
So the docs and the app cannot drift, here is the complete set of request bodies the browser app submits. There is nothing else — no hidden system prompt field, no extra options object.
The meter, once per lane, as you type:
POST https://api.skillsafe.ai/v1/app-api/estimate
Authorization: Bearer <token>
Content-Type: application/json
{"content": "<the text in the composer>",
"$model": "<that lane's model id>"}
The turn, once per lane, when you press Ask:
POST https://api.skillsafe.ai/v1/app-api/sessions/<that lane's session_id>/messages
Authorization: Bearer <token>
Content-Type: application/json
Idempotency-Key: cwa:<fnv1a32(session_id | model_id | content), base36>:<attempt>
{"content": "<the user's message>",
"$model": "<that lane's model id>",
"stream": true}
The key is derived from the session, the model and the exact message, so a double-click or a
double Enter replays the first turn rather than buying a second one. It is
not a cryptographic digest — it is a 32-bit FNV-1a hash rendered in
base 36, which is enough to key an idempotent turn and cheap enough to compute on every
keystroke-triggered re-render. The trailing <attempt> counter is
incremented only when the user presses "Retry this lane", which is deliberately a
new, billed turn; every automatic path reuses the same key.
The saved comparisons, written after every run:
PUT https://api.skillsafe.ai/v1/app-api/data/history
{"value": [
{"id": "cmp_…", "ts": "<ISO 8601>", "prompt": "<the question, after clipping>",
"lanes": [
{"model": "<model id>", "label": "<display name>",
"status": "succeeded" | "partial" | "failed",
"text": "<the reply, trimmed to 4000 chars>", "chars": 0,
"elapsed_ms": 0, "charged_credits": 0 | null,
"truncated": false, "error": null}
]}
]}
Newest entry first, at most 20, and the whole array is trimmed to stay under the platform's
64 KB per-document cap. charged_credits is null when a turn never
settled — it is deliberately not 0, because "we don't know" and "it was
free" are different facts. If you write this key yourself, keep the shape: the app renders
lanes[] directly.
Around those two, the app calls POST /guest or the SSO flow once for a token,
GET /me for the credit pill, GET /v1/models to build the lane
pickers, POST /sessions once per lane (and again when a lane's model changes),
and GET/PUT /data/history for the saved comparisons. All of those
are free or near-free; only the message POST spends real credits.