API docs

Change one line. Keep the rest.

An OpenAI-compatible gateway with an Anthropic-compatible twin. Point the SDK you already use at it and call any of the 65 models by name.

openapi.json Get a key

Quick start

from openai import OpenAI

client = OpenAI(
    base_url="https://ai.artur.work/api/v1",
    api_key="<gateway key>",  # issued on your profile
)

# any model from any provider, same call
r = client.chat.completions.create(
    model="claude-sonnet-5",
    messages=[{"role": "user", "content": "Explain reverse charge in one line"}],
)
print(r.choices[0].message.content)
print(r.usage.prompt_tokens, r.usage.input_tokens)  # both vendors' names are filled

img = client.images.generate(model="flux1.1-pro-ultra", prompt="a red bicycle",
                             extra_query={"wait": 1})  # async provider, sync answer
print(img.data[0].url)  # on the gateway CDN, permanent

Authentication

A gateway key, issued on your profile, in either header. The Anthropic SDK's header works unchanged.

Authorization: Bearer <gateway key>
x-api-key: <gateway key>

A key can be pinned to IP ranges. A request with no credential, or a bad one, gets 401 as JSON, never a redirect.

Base URLs

Every route is served under two prefixes, because the vendor SDKs assume them. /api/v1 and /v1 reach the same controllers.

OpenAI(base_url="https://ai.artur.work/api/v1")
Anthropic(base_url="https://ai.artur.work")  # adds /v1 itself

Two dialects, every model

Ask in either shape, get the answer in that shape. Both /chat/completions and /messages accept any model.

/chat/completions + claude-*

A system message is lifted into Anthropic's system, stop becomes stop_sequences, max_tokens defaults to 4096, and the reply is rebuilt in the OpenAI format with finish_reason.

/messages + gpt-5, grok-4.5, …

system becomes a leading system message, stop_sequences becomes stop, and the reply is rebuilt as a Messages response.

/messages + claude-*

Passed through whole: tool_use, thinking blocks, multi-block answers and stop_reason all survive.

usage carries both vendors' field names whatever answered, and the provider's nested detail (cached tokens, image tokens) is kept - that is what billing reads.

Endpoints

All under /api/v1 or /v1.

RouteBodyAnswerNote
POST/chat/completionsJSON, OpenAI dialectOpenAI formatany model
POST/messagesJSON, Anthropic dialectAnthropic formatany model
POST/images/generationsJSON{created, data:[{url}]} or {taskId}?wait=1 makes async look sync
POST/images/editsmultipartas above
POST/images/variationsmultipartas above
POST/images/upscalemultipart{taskId}gateway-only
GET/images/status/{model}/{taskId}-{created, data} when done; {} while runningMesh: {state, eta}
POST/audio/speechJSONbinary audio
POST/audio/transcriptionsmultipart{text}
POST/decisionsJSON{answers, model, usage}Jev · probabilities
POST/decisions/batchJSON{results:[{id, answers}]}up to 50 items, one charge
POST/tasks/{taskId}/cancel-{taskId, state}queued Mesh job only, else 409
GET/models-OpenAI {object, data} + rich cataloguepublic
GET/models/{model}/quote-price and latency of that requestpublic, also POST
GET/models/capabilities?need=…-{need, count, models}public · models that meet every need
POST/uploadsmultipart{id, sha256, expiresAt}private, owner-only, 24 h
GET/runs/{id}-{state, chargedEur, provenance}state, charge, provenance
POST/batchesJSON{id, state, counts}quote · run · results · cancel
GET/account-{balanceEur, lowBalance, key}balance, key cap and spend

Not offered: streaming, /embeddings, /responses, /files, /moderations, /audio/translations, and Anthropic's count_tokens and Batches.

Images: async providers and ?wait=1

Midjourney, Runway, Luma, Flux and NanoBanana answer {taskId} and are polled on /images/status/{model}/{taskId}. Add ?wait=1 and the gateway polls for you and answers {created, data}, so client.images.generate just works.

POST /api/v1/images/generations?wait=1
{ "model": "midjourney", "prompt": "a red bicycle", "n": 1 }

200 { "created": 1789572360, "data": [{ "url": "…/public/images/7f/99/7f99….png" }] }

A query parameter, not a body field: bodies are forwarded, and OpenAI rejects unknown fields. If the wait runs out you get the ordinary {taskId}. response_format=b64_json is honoured for OpenAI models.

Errors

One error format, readable by the OpenAI, xAI and Anthropic clients alike.

{ "type": "error", "error": {
    "type": "invalid_request_error",
    "message": "…",
    "code": "insufficient_quota"  # only on 402
} }
400invalid request, including "stream": true
401no or bad key · JSON, not a redirect
402the balance cannot cover the reservation · nothing is sent upstream
404unknown API route · JSON
5xxa provider's own error passes through; anything else is wrapped, with the original under error.provider

Billing and balance

Every call reserves an amount against your balance before the provider is contacted and settles with what the provider reports. Per-piece tariffs are exact; token-metered runs are capped by the remaining balance. A failed run releases the reservation and is not charged.

GET /models lists each model's price items and latency; the price list has the same in EUR with provider links.

Bring your own provider keys

Add a provider key on your profile; it is sealed with AES-256-GCM under one of your gateway keys. Only requests authenticated with that gateway key can open it, and it is never displayed again. Requests on your own key are not charged, only counted against a daily allowance.

Rotating the gateway key re-seals; deleting it deletes the sealed credentials.

Decisions

Typed decisions instead of text: yes/no (noul), one of several (choice) or an ordered score, each with probabilities you can put a threshold on. Served by Jev; ask for jev-1.13 or auto.

POST /api/v1/decisions
{ "model": "auto",
  "state": "Checkout throws a 500 when a card is declined.",
  "questions": {
    "is_bug":  { "type": "noul",   "instructions": "Is this a bug report?" },
    "team":    { "type": "choice", "criteria": { "payments": "money", "frontend": "UI" } },
    "urgency": { "type": "score",  "criteria": ["low", "medium", "high"] } } }

200 { "answers": {
        "is_bug":  { "type": "noul", "noul": 0.96 },
        "team":    { "type": "choice", "choice": "payments", "confidence": 0.75, "probabilities": {…} },
        "urgency": { "type": "score", "score": 1.99, "probabilities": {…} } },
      "model": "jev-1.13", "usage": { "input_tokens": 476 } }

Up to 20 questions per request, keys [a-z0-9_], 2-20 options per choice, 2-10 levels per score, and a state (text, object or array) of at most about 90 KB. Text only; English is the most accurate. Billed per input token; output is free.

POST /decisions/batch takes the same questions and up to 50 items of {id, state} and answers one result per item, in the order sent. An item that fails carries its own error; the others are answered, and the whole batch is one charge.

Callbacks, retries and caps

Add callback_url (https, a public address) to any request that starts a task - images, chat, audio, decisions, Mesh jobs - and when the task ends we POST the JSON the status endpoint would give, plus taskId, state, model and charged (EUR micros). Generate the webhook secret on your profile first: every callback is signed with it.

POST ai.artur.work/aimage/callback
X-AiArtur-Signature: t=1789572360,v1=5f0c…
{ "created": 1789572360, "data": [{ "url": "…" }],
  "taskId": "…", "state": "SUCCESS", "model": "midjourney", "charged": 44000 }

The signature is t=<unix time>,v1=<hex HMAC-SHA256 of t + "." + body>. Compute it over the raw body with your secret, compare in constant time, and refuse anything older than five minutes.

// PHP: verify a callback before trusting it
$body = file_get_contents('php://input');
parse_str(str_replace(',', '&', $_SERVER['HTTP_X_AIARTUR_SIGNATURE'] ?? ''), $sig);
$expected = hash_hmac('sha256', ($sig['t'] ?? '') . '.' . $body, $webhookSecret);
if (!hash_equals($expected, $sig['v1'] ?? '') || abs(time() - (int) ($sig['t'] ?? 0)) > 300) {
    http_response_code(400); exit;
}
$task = json_decode($body, true);  // answer 2xx quickly; work after

Answer with any 2xx to accept it. Anything else is retried after 1, 5, 15, 60 and 240 minutes, then given up. Asynchronous image tasks are finished for you, so a plugin never has to poll.

Send Idempotency-Key (up to 128 printable characters) with any POST that starts a task, and a retry cannot start - or charge for - the same work twice.

curl https://ai.artur.work/api/v1/images/generations \
  -H "Authorization: Bearer $AI_KEY" \
  -H "Idempotency-Key: order-4711-hero" \
  -d '{"model": "midjourney", "prompt": "a red bicycle", "callback_url": "https://shop.example/aimage/callback"}'

The same key with the same request gets the stored answer again (header Idempotent-Replayed: true); with a different request, 422 idempotency_key_reused; while the first is still running, 409 idempotency_in_progress. A 5xx answer is not stored, so retrying it is a real retry. Keys are kept for 24 hours.

POST /tasks/{taskId}/cancel takes back a Mesh job no worker has started: it becomes CANCELLED and nothing is charged. Anything already running answers 409 not_cancellable: a vendor bills for a started task whether we stop listening or not, and we do not pretend otherwise.

Give each site its own gateway key and set a spending cap per key, per day or per month, on your profile. A request that would take the key over its cap is refused before anything is sent: 402 key_spend_cap_reached. What a key has spent counts finished charges and tasks still running.

Plugins: connect, batches, assets

A CMS plugin should not ask anyone to copy a key. Send the manager's browser to /connect (OAuth2 code with PKCE S256): the account owner allows it once, with a monthly cap, and your backend trades the code at POST /connect/token for a key of its own. GET /account then gives the plugin the balance, the key's cap and spend, and whether the balance is low.

POST /batches/quote prices thousands of items exactly as single requests would, and its quoteToken holds those prices for an hour. POST /batches runs them through the same checks as single requests, stops at a spending ceiling and sends one signed callback at the end; /batches/{id}/results pages them out in the order sent.

Say what you want once, in options: aspectRatio, n, background, quality (draft, standard, high), outputFormat, seed, negativePrompt and references. The gateway maps them onto whichever model serves the request. Under strict (the default) an option the model cannot honour is a 400 naming the option; with "strict": false you get the nearest result and warnings. GET /models/capabilities lists the models that can.

POST /api/v1/images/generations
{ "model": "gpt-image-1", "prompt": "a glass bottle on white",
  "options": { "aspectRatio": "2:3", "background": "transparent", "outputFormat": "png", "n": 2 },
  "outputs": { "private": true } }

200 { "created": 1790000000,
      "data": [{ "url": "…/v1/media/9f2c…?exp=…&sig=…", "expires_at": 1790604800 }, …],
      "provenance": { "runId": "4711", "model": "gpt-image-1", "providerModel": "gpt-image-1",
                      "inputs": [], "options": {…}, "c2pa": "present" },
      "moderation": null }

inputs.private keeps the images you send off every public URL: a vendor gets a short-lived signed link or the bytes, and the copy is deleted after the run. outputs.private returns results as signed links that expire. Keys made by Connect have both on. POST /uploads keeps an image for 24 hours under an id only you can use - in options.references, upscales and batch edits.

Every result item carries expires_at (null when your retention keeps it), and the answer carries provenance: model, the vendor's own model id, input hashes, options and whether a C2PA manifest is present. Download and import before expires_at, never hotlink; GET /runs/{id} returns the same provenance later.

Providers that let us ask them not to train on your data per request are asked on every request, unless you untick that on your profile; the privacy page lists what each provider's terms say.

Models

Model names are the canonical ids from GET /models; the gateway routes each to its provider. The public catalogue also lists modalities, the actions each model supports, prices and latency with its sample size.

gpt-5o3claude-opus-5claude-sonnet-5claude-haiku-4.5gemini-3.1-progrok-4.5llama-4-maverickgpt-image-1.5flux1.1-pro-ultraflux1-kontext-maxmidjourneyideogram-2arecraft-3gen4-imagephoton-1seedream-3whisper-1gpt-4o-transcribetts-1-hd
GET /api/v1/models
{ "object": "list", "data": [{ "id": "claude-sonnet-5", "owned_by": "Anthropic" }, …],
  "status": "done", "count": 65, "legend": {…},
  "models": [{ "model": "claude-sonnet-5", "input": ["i", "t"], "output": ["t"], "price": {…}, "latency": {…} }, …] }

Machine-readable: openapi.json

Every endpoint and the public model catalogue, with the schemas above, including the two shapes Mesh status has to admit to.

ai.artur.work/openapi.json