An OpenAI-compatible gateway with an Anthropic-compatible twin. Point the SDK you already use at it and call any of the 65 models by name.
from openai import OpenAI client = OpenAI( base_url="https://ai.artur.work/api/v1", api_key="<gateway key>", # issued on your profile ) # any model from any provider, same call r = client.chat.completions.create( model="claude-sonnet-5", messages=[{"role": "user", "content": "Explain reverse charge in one line"}], ) print(r.choices[0].message.content) print(r.usage.prompt_tokens, r.usage.input_tokens) # both vendors' names are filled img = client.images.generate(model="flux1.1-pro-ultra", prompt="a red bicycle", extra_query={"wait": 1}) # async provider, sync answer print(img.data[0].url) # on the gateway CDN, permanent
import anthropic client = anthropic.Anthropic( base_url="https://ai.artur.work", # the SDK adds /v1 itself api_key="<gateway key>", ) # Messages for every model, not only Claude r = client.messages.create( model="gpt-5", max_tokens=1024, messages=[{"role": "user", "content": "Explain reverse charge in one line"}], ) print(r.content[0].text)
import OpenAI from "openai"; const client = new OpenAI({ baseURL: "https://ai.artur.work/api/v1", apiKey: "<gateway key>", }); const r = await client.chat.completions.create({ model: "grok-4.5", messages: [{ role: "user", content: "Explain reverse charge in one line" }], }); console.log(r.choices[0].message.content);
curl https://ai.artur.work/api/v1/chat/completions \ -H "Authorization: Bearer <gateway key>" \ -H "Content-Type: application/json" \ -d '{"model": "claude-sonnet-5", "messages": [{"role": "user", "content": "Explain reverse charge in one line"}]}'
A gateway key, issued on your profile, in either header. The Anthropic SDK's header works unchanged.
Authorization: Bearer <gateway key> x-api-key: <gateway key>
A key can be pinned to IP ranges. A request with no credential, or a bad one, gets 401 as JSON, never a redirect.
Every route is served under two prefixes, because the vendor SDKs assume them. /api/v1 and /v1 reach the same controllers.
OpenAI(base_url="https://ai.artur.work/api/v1") Anthropic(base_url="https://ai.artur.work") # adds /v1 itself
Ask in either shape, get the answer in that shape. Both /chat/completions and /messages accept any model.
/chat/completions + claude-*A system message is lifted into Anthropic's system, stop becomes stop_sequences, max_tokens defaults to 4096, and the reply is rebuilt in the OpenAI format with finish_reason.
/messages + gpt-5, grok-4.5, …system becomes a leading system message, stop_sequences becomes stop, and the reply is rebuilt as a Messages response.
/messages + claude-*Passed through whole: tool_use, thinking blocks, multi-block answers and stop_reason all survive.
usage carries both vendors' field names whatever answered, and the provider's nested detail (cached tokens, image tokens) is kept - that is what billing reads.
All under /api/v1 or /v1.
| Route | Body | Answer | Note |
|---|---|---|---|
| POST/chat/completions | JSON, OpenAI dialect | OpenAI format | any model |
| POST/messages | JSON, Anthropic dialect | Anthropic format | any model |
| POST/images/generations | JSON | {created, data:[{url}]} or {taskId} | ?wait=1 makes async look sync |
| POST/images/edits | multipart | as above | |
| POST/images/variations | multipart | as above | |
| POST/images/upscale | multipart | {taskId} | gateway-only |
| GET/images/status/{model}/{taskId} | - | {created, data} when done; {} while running | Mesh: {state, eta} |
| POST/audio/speech | JSON | binary audio | |
| POST/audio/transcriptions | multipart | {text} | |
| POST/decisions | JSON | {answers, model, usage} | Jev · probabilities |
| POST/decisions/batch | JSON | {results:[{id, answers}]} | up to 50 items, one charge |
| POST/tasks/{taskId}/cancel | - | {taskId, state} | queued Mesh job only, else 409 |
| GET/models | - | OpenAI {object, data} + rich catalogue | public |
| GET/models/{model}/quote | - | price and latency of that request | public, also POST |
| GET/models/capabilities?need=… | - | {need, count, models} | public · models that meet every need |
| POST/uploads | multipart | {id, sha256, expiresAt} | private, owner-only, 24 h |
| GET/runs/{id} | - | {state, chargedEur, provenance} | state, charge, provenance |
| POST/batches | JSON | {id, state, counts} | quote · run · results · cancel |
| GET/account | - | {balanceEur, lowBalance, key} | balance, key cap and spend |
Not offered: streaming, /embeddings, /responses, /files, /moderations, /audio/translations, and Anthropic's count_tokens and Batches.
?wait=1Midjourney, Runway, Luma, Flux and NanoBanana answer {taskId} and are polled on /images/status/{model}/{taskId}. Add ?wait=1 and the gateway polls for you and answers {created, data}, so client.images.generate just works.
POST /api/v1/images/generations?wait=1
{ "model": "midjourney", "prompt": "a red bicycle", "n": 1 }
200 { "created": 1789572360, "data": [{ "url": "…/public/images/7f/99/7f99….png" }] }A query parameter, not a body field: bodies are forwarded, and OpenAI rejects unknown fields. If the wait runs out you get the ordinary {taskId}. response_format=b64_json is honoured for OpenAI models.
One error format, readable by the OpenAI, xAI and Anthropic clients alike.
{ "type": "error", "error": {
"type": "invalid_request_error",
"message": "…",
"code": "insufficient_quota" # only on 402
} }| 400 | invalid request, including "stream": true |
| 401 | no or bad key · JSON, not a redirect |
| 402 | the balance cannot cover the reservation · nothing is sent upstream |
| 404 | unknown API route · JSON |
| 5xx | a provider's own error passes through; anything else is wrapped, with the original under error.provider |
Every call reserves an amount against your balance before the provider is contacted and settles with what the provider reports. Per-piece tariffs are exact; token-metered runs are capped by the remaining balance. A failed run releases the reservation and is not charged.
GET /models lists each model's price items and latency; the price list has the same in EUR with provider links.
Add a provider key on your profile; it is sealed with AES-256-GCM under one of your gateway keys. Only requests authenticated with that gateway key can open it, and it is never displayed again. Requests on your own key are not charged, only counted against a daily allowance.
Rotating the gateway key re-seals; deleting it deletes the sealed credentials.
Typed decisions instead of text: yes/no (noul), one of several (choice) or an ordered score, each with probabilities you can put a threshold on. Served by Jev; ask for jev-1.13 or auto.
POST /api/v1/decisions
{ "model": "auto",
"state": "Checkout throws a 500 when a card is declined.",
"questions": {
"is_bug": { "type": "noul", "instructions": "Is this a bug report?" },
"team": { "type": "choice", "criteria": { "payments": "money", "frontend": "UI" } },
"urgency": { "type": "score", "criteria": ["low", "medium", "high"] } } }
200 { "answers": {
"is_bug": { "type": "noul", "noul": 0.96 },
"team": { "type": "choice", "choice": "payments", "confidence": 0.75, "probabilities": {…} },
"urgency": { "type": "score", "score": 1.99, "probabilities": {…} } },
"model": "jev-1.13", "usage": { "input_tokens": 476 } }Up to 20 questions per request, keys [a-z0-9_], 2-20 options per choice, 2-10 levels per score, and a state (text, object or array) of at most about 90 KB. Text only; English is the most accurate. Billed per input token; output is free.
POST /decisions/batch takes the same questions and up to 50 items of {id, state} and answers one result per item, in the order sent. An item that fails carries its own error; the others are answered, and the whole batch is one charge.
Add callback_url (https, a public address) to any request that starts a task - images, chat, audio, decisions, Mesh jobs - and when the task ends we POST the JSON the status endpoint would give, plus taskId, state, model and charged (EUR micros). Generate the webhook secret on your profile first: every callback is signed with it.
POST ai.artur.work/aimage/callback
X-AiArtur-Signature: t=1789572360,v1=5f0c…
{ "created": 1789572360, "data": [{ "url": "…" }],
"taskId": "…", "state": "SUCCESS", "model": "midjourney", "charged": 44000 }The signature is t=<unix time>,v1=<hex HMAC-SHA256 of t + "." + body>. Compute it over the raw body with your secret, compare in constant time, and refuse anything older than five minutes.
// PHP: verify a callback before trusting it $body = file_get_contents('php://input'); parse_str(str_replace(',', '&', $_SERVER['HTTP_X_AIARTUR_SIGNATURE'] ?? ''), $sig); $expected = hash_hmac('sha256', ($sig['t'] ?? '') . '.' . $body, $webhookSecret); if (!hash_equals($expected, $sig['v1'] ?? '') || abs(time() - (int) ($sig['t'] ?? 0)) > 300) { http_response_code(400); exit; } $task = json_decode($body, true); // answer 2xx quickly; work after
Answer with any 2xx to accept it. Anything else is retried after 1, 5, 15, 60 and 240 minutes, then given up. Asynchronous image tasks are finished for you, so a plugin never has to poll.
Send Idempotency-Key (up to 128 printable characters) with any POST that starts a task, and a retry cannot start - or charge for - the same work twice.
curl https://ai.artur.work/api/v1/images/generations \ -H "Authorization: Bearer $AI_KEY" \ -H "Idempotency-Key: order-4711-hero" \ -d '{"model": "midjourney", "prompt": "a red bicycle", "callback_url": "https://shop.example/aimage/callback"}'
The same key with the same request gets the stored answer again (header Idempotent-Replayed: true); with a different request, 422 idempotency_key_reused; while the first is still running, 409 idempotency_in_progress. A 5xx answer is not stored, so retrying it is a real retry. Keys are kept for 24 hours.
POST /tasks/{taskId}/cancel takes back a Mesh job no worker has started: it becomes CANCELLED and nothing is charged. Anything already running answers 409 not_cancellable: a vendor bills for a started task whether we stop listening or not, and we do not pretend otherwise.
Give each site its own gateway key and set a spending cap per key, per day or per month, on your profile. A request that would take the key over its cap is refused before anything is sent: 402 key_spend_cap_reached. What a key has spent counts finished charges and tasks still running.
A CMS plugin should not ask anyone to copy a key. Send the manager's browser to /connect (OAuth2 code with PKCE S256): the account owner allows it once, with a monthly cap, and your backend trades the code at POST /connect/token for a key of its own. GET /account then gives the plugin the balance, the key's cap and spend, and whether the balance is low.
POST /batches/quote prices thousands of items exactly as single requests would, and its quoteToken holds those prices for an hour. POST /batches runs them through the same checks as single requests, stops at a spending ceiling and sends one signed callback at the end; /batches/{id}/results pages them out in the order sent.
Say what you want once, in options: aspectRatio, n, background, quality (draft, standard, high), outputFormat, seed, negativePrompt and references. The gateway maps them onto whichever model serves the request. Under strict (the default) an option the model cannot honour is a 400 naming the option; with "strict": false you get the nearest result and warnings. GET /models/capabilities lists the models that can.
POST /api/v1/images/generations
{ "model": "gpt-image-1", "prompt": "a glass bottle on white",
"options": { "aspectRatio": "2:3", "background": "transparent", "outputFormat": "png", "n": 2 },
"outputs": { "private": true } }
200 { "created": 1790000000,
"data": [{ "url": "…/v1/media/9f2c…?exp=…&sig=…", "expires_at": 1790604800 }, …],
"provenance": { "runId": "4711", "model": "gpt-image-1", "providerModel": "gpt-image-1",
"inputs": [], "options": {…}, "c2pa": "present" },
"moderation": null }inputs.private keeps the images you send off every public URL: a vendor gets a short-lived signed link or the bytes, and the copy is deleted after the run. outputs.private returns results as signed links that expire. Keys made by Connect have both on. POST /uploads keeps an image for 24 hours under an id only you can use - in options.references, upscales and batch edits.
Every result item carries expires_at (null when your retention keeps it), and the answer carries provenance: model, the vendor's own model id, input hashes, options and whether a C2PA manifest is present. Download and import before expires_at, never hotlink; GET /runs/{id} returns the same provenance later.
Providers that let us ask them not to train on your data per request are asked on every request, unless you untick that on your profile; the privacy page lists what each provider's terms say.
Model names are the canonical ids from GET /models; the gateway routes each to its provider. The public catalogue also lists modalities, the actions each model supports, prices and latency with its sample size.
GET /api/v1/models
{ "object": "list", "data": [{ "id": "claude-sonnet-5", "owned_by": "Anthropic" }, …],
"status": "done", "count": 65, "legend": {…},
"models": [{ "model": "claude-sonnet-5", "input": ["i", "t"], "output": ["t"], "price": {…}, "latency": {…} }, …] }Every endpoint and the public model catalogue, with the schemas above, including the two shapes Mesh status has to admit to.