The Lumeta API
Every Lumeta model over plain HTTPS: the same account, credits, Shelf and projects you use on the web. Ask what a run will cost before you spend, start it, and get the result by waiting, polling or a webhook. Everything you make lands in your Lumeta library with a small API tag.
Quick start
Make a key under Account → Developer (keys are part of the Pro and Ultimate plans), then:
export LUMETA_API_KEY=lm_live_...
# 1. What will it cost? (free, spends nothing)
curl https://api.lumeta.ai/v1/estimate \
-H "Authorization: Bearer $LUMETA_API_KEY" -H "Content-Type: application/json" \
-d '{"model": "nano-banana-2", "input": {"prompt": "a red panda astronaut, studio light"}}'
# 2. Make it
curl https://api.lumeta.ai/v1/generations \
-H "Authorization: Bearer $LUMETA_API_KEY" -H "Content-Type: application/json" \
-H "Idempotency-Key: red-panda-001" \
-d '{"model": "nano-banana-2", "input": {"prompt": "a red panda astronaut, studio light"}}'
# 3. Wait for the result (the id came back from step 2)
curl "https://api.lumeta.ai/v1/generations/gen_123?wait=45" -H "Authorization: Bearer $LUMETA_API_KEY"Step 2 answers 202 with the generation's id while it runs (or 200 when the tool answers at once). Step 3 holds the request until it finishes or 45 seconds pass, whichever comes first; call it again if it's still running. A finished generation looks like this (shortened):
{
"id": "gen_123",
"status": "succeeded",
"model": {
"id": 140,
"name": "Nano Banana 2"
},
"outputs": [
{
"index": 0,
"type": "image",
"url": "https://…/result.jpg",
"width": 1024,
"height": 1024
}
],
"credits_charged": 20,
"origin": "api",
"completed_in_seconds": 6.2
}The same in Python and JavaScript
import os, requests
API = "https://api.lumeta.ai/v1"
H = {"Authorization": "Bearer " + os.environ["LUMETA_API_KEY"]}
body = {"model": "nano-banana-2", "input": {"prompt": "a red panda astronaut, studio light"}}
print(requests.post(f"{API}/estimate", json=body, headers=H).json()["credits"], "credits")
gen = requests.post(f"{API}/generations", json=body, headers={**H, "Idempotency-Key": "red-panda-001"}).json()
while gen["status"] not in ("succeeded", "failed"):
gen = requests.get(f"{API}/generations/{gen['id']}", params={"wait": 45}, headers=H).json()
print(gen["outputs"][0]["url"] if gen["status"] == "succeeded" else gen["error"])const API = "https://api.lumeta.ai/v1";
const H = { Authorization: `Bearer ${process.env.LUMETA_API_KEY}`, "Content-Type": "application/json" };
const body = JSON.stringify({ model: "nano-banana-2", input: { prompt: "a red panda astronaut, studio light" } });
let gen = await (await fetch(`${API}/generations`, { method: "POST", headers: { ...H, "Idempotency-Key": "red-panda-001" }, body })).json();
while (!["succeeded", "failed"].includes(gen.status)) {
gen = await (await fetch(`${API}/generations/${gen.id}?wait=45`, { headers: H })).json();
}
console.log(gen.status === "succeeded" ? gen.outputs[0].url : gen.error);Keys and sign-in
Send Authorization: Bearer lm_live_... on every request. Keys are made under Account → Developer, and each one carries its own settings:
- Permissions (scopes):
read(look at models, balance, generations, Shelf and projects),generate(start generations, spends credits) andlibrary(upload files, change your Shelf and projects, manage webhooks). - Credit limits: an optional daily and monthly cap. A request that would go over it is refused with
402 token_cap_exceeded, and nothing is spent. - Automatic top-up: off by default. When it's off, a key that runs out of credits stops and tells you, instead of charging your card.
- Expiry: 30, 90, 365 days, or never.
A key is shown once, when you make it. We keep only a fingerprint of it, so treat it like a password: keep it on your server, never in a web page or an app you ship. Revoke it on the same page and it stops working at once.
Building an app other people sign into
Use OAuth 2.1 instead of asking people for keys: the authorization code flow with PKCE, for public clients. Clients can register themselves (dynamic client registration) or use a client ID metadata document. Everything a library needs is at /.well-known/oauth-authorization-server. People see a Lumeta consent page where they choose permissions, a daily limit and top-up, just like for a key. Access tokens (lma_...) last 1 hour; refresh tokens last 90 days and change with every use.
Models and inputs
GET /v1/models lists every model you can run with its starting price (base_credits). Filter with type=image|video|audio or q=, or describe the job with task= and get a recommendation. It also tells you which model your account opens with for each kind (defaults).
Each model names its inputs its own way, so read GET /v1/models/{slug} first: it includes the model's input_schema (JSON Schema 2020-12) with every field, its choices and defaults. Starting prices are a guide; POST /v1/estimate with your exact input is the price you'll pay.
Media inputs
Fields that take an image, video or audio take a reference, never a raw filename:
| Write | Means |
|---|---|
file_123 | A file you uploaded or imported (POST /v1/files, POST /v1/files/import) |
shelf_45 or shelf:Blue Hat | An item on your Shelf, by id or by the name you gave it |
gen_123#0 | Output 0 of one of your generations (counting from 0) |
voice_7 | One of your own voices (GET /v1/voices) |
https://… | A public link, imported for you |
In a prompt, @Blue Hat (or @[Blue Hat] for a name with spaces) attaches the Shelf item with that name where the model takes it. A word after @ that is no Shelf name is sent as text and noted in warnings, so a handle never breaks a script.
Generations
POST /v1/generations with {"model", "input", "project"} starts one and spends credits. A generation is queued, processing, succeeded or failed. A failed one is refunded by itself (credits_charged reads 0).
- Waiting.
GET /v1/generations/{id}?wait=45holds the request until it's done or the time is up. Ask again for long videos; they keep running whether you wait or not. - Safe retries. Send an
Idempotency-Keyheader on a create. Repeating the same key with the same body within 24 hours returns the first run instead of starting (and paying for) another. - Canceling.
POST /v1/generations/{id}/cancelstops a run and refunds it where the provider can really stop it (today, models that run on Replicate). Elsewhere it answers409 not_cancelableand the run finishes normally. A canceled run readsfailedwitherror.codecanceled. - Your history.
GET /v1/generationslists your runs, newest first, with filters forstatus,model,projectandorigin(web, api, mcp, cli). Lists takelimit(default 20, up to 100) and page withcursor.
Files, Shelf and projects
POST /v1/filesuploads a file (multipart, partfile) with the same checks and limits as the web uploader;POST /v1/files/importtakes a publichttps://link. Both answer with afile_…id to use as an input.GET /v1/shelveslists your Shelves (proj_0is All Projects);GET /v1/shelves/{project}/items?q=finds items by name.POSTpins an output or file under a name,PATCHrenames,DELETEremoves the pin (the file stays).GET /v1/projectsandPOST /v1/projects. A new project becomes your current one, as on the web.GET /v1/account: your plan, credits, and this key's permissions, limits and spend.
Webhooks
Rather than wait, let Lumeta tell you. Register a receiver once:
curl https://api.lumeta.ai/v1/webhook-endpoints \
-H "Authorization: Bearer $LUMETA_API_KEY" -H "Content-Type: application/json" \
-d '{"url": "https://example.com/lumeta-webhook", "events": ["*"]}'The answer includes the endpoint's signing secret (whsec_…), shown this once. From then on, every generation on your account that finishes, whichever way it was started (web, API, assistant or CLI; data.origin says which), is POSTed to your URL a few seconds later as generation.succeeded or generation.failed. data is the generation exactly as GET /v1/generations/{id} shows it. You can have up to 5 endpoints; POST /v1/webhook-endpoints/{id}/test sends a test event right away.
Check the signature
Every request carries Lumeta-Signature: t=<unix seconds>,v1=<hex>, where v1 is the HMAC-SHA256 of "<t>.<raw body>" keyed with your secret. Compute it over the raw bytes (before any JSON parsing) and refuse a t more than 5 minutes old:
import crypto from "node:crypto";
export function verifyLumeta(rawBody, header, secret) {
const p = Object.fromEntries(header.split(",").map((kv) => kv.split("=")));
if (!p.t || Math.abs(Date.now() / 1000 - Number(p.t)) > 300) return false;
const want = crypto.createHmac("sha256", secret).update(`${p.t}.${rawBody}`).digest("hex");
return typeof p.v1 === "string" && p.v1.length === want.length
&& crypto.timingSafeEqual(Buffer.from(p.v1), Buffer.from(want));
}import hashlib, hmac, time
def verify_lumeta(raw_body: bytes, header: str, secret: str) -> bool:
p = dict(kv.split("=", 1) for kv in header.split(","))
if abs(time.time() - int(p.get("t", 0))) > 300:
return False
want = hmac.new(secret.encode(), p["t"].encode() + b"." + raw_body, hashlib.sha256).hexdigest()
return hmac.compare_digest(want, p.get("v1", ""))function verify_lumeta(string $rawBody, string $header, string $secret): bool
{
parse_str(str_replace(',', '&', $header), $p);
if (abs(time() - (int)($p['t'] ?? 0)) > 300) return false;
return hash_equals(hash_hmac('sha256', $p['t'] . '.' . $rawBody, $secret), (string)($p['v1'] ?? ''));
}Answering, retries and order
- Answer with any
2xxwithin 5 seconds. Do the slow work after you answer. - Anything else (a timeout, a redirect, a
4xxor5xx) is tried again after 1 minute, 5 minutes, 30 minutes, 2 hours and 12 hours, with the same event id (Lumeta-Event-Id, alsobody.id). Use it to skip duplicates. - Events can arrive out of order, and a retry carries the generation as it is at that moment, so act on
data.status. - If an event fails all 6 attempts and your endpoint has not answered a
2xxonce in the meantime, the endpoint is switched off (enabled: false).PATCH {"enabled": true}turns it back on. - Receivers must be
https://on port 443 at a public address. Lumeta never follows redirects and reads at most 4 KB of your answer.
Errors
Every error has the same shape, and every response carries an X-Request-Id header. Quote the request id when you write to us and we can find exactly what happened.
{
"error": {
"code": "insufficient_credits",
"message": "Not enough credits for this request. Top up at https://lumeta.ai/subscription/",
"request_id": "01J9…"
}
}The code is stable; build on it, not on the message. param names the input at fault when there is one.
| Status | Code | What to do |
|---|---|---|
| 400 | invalid_json | The body isn't valid JSON. |
| 401 | unauthorized, invalid_token | No key, or a key that was revoked or expired. |
| 402 | insufficient_credits | Not enough credits. Top up, or turn on automatic top-up for this key. |
| 402 | token_cap_exceeded | This key's daily or monthly limit would be passed. |
| 403 | plan_required, insufficient_scope, project_read_only | The plan, the key's permissions, or a view-only project doesn't allow it. |
| 404 | not_found | No such thing on your account. |
| 409 | idempotency_conflict, idempotency_in_progress, already_exists, not_cancelable | The same Idempotency-Key with a different body, or one still running; a duplicate; a run that can't be stopped. |
| 422 | invalid_input, moderation_blocked | An input is wrong (see param), or the request breaks the content rules. |
| 429 | rate_limited, concurrency_limit | Slow down: wait the number of seconds in Retry-After. |
| 502 | provider_error | The model's provider failed. You were not charged; try again. |
| 503 | model_unavailable, service_unavailable | Down for a moment. Wait a few seconds and retry. |
| 500 | internal_error | Our fault. It's logged; send us the request id. |
Limits
| Limit | Per key or connection |
|---|---|
| Requests | 300 a minute |
| New generations | 20 a minute |
| Estimates | 60 a minute |
| Waiting on one request | 45 seconds |
| Items per list page | up to 100 |
| Webhook endpoints | 5 per account |
Over a limit, you get 429 with Retry-After. Some models also allow only a few runs at a time per account, as on the web (concurrency_limit).
Coming from Higgsfield
If you use Higgsfield's API, you can switch by changing one line. https://api.lumeta.ai/higgsfield speaks Higgsfield's dialect (their paths, request and status shapes, and errors) for the models Lumeta also runs, with your Lumeta key and credits.
import higgsfield_client
# HF_KEY = your Lumeta key, exactly as shown when you made it
client = higgsfield_client.SyncClient(base_url="https://api.lumeta.ai/higgsfield")import { config, higgsfield } from "@higgsfield/client/v2";
// Your Lumeta key with its last "_" written as ":" (lm_live_<id>:<secret>)
config({ credentials: "lm_live_<id>:<secret>", baseURL: "https://api.lumeta.ai/higgsfield" });One trap in the Python SDK: its helper functions status(), result() and cancel() that take a request id always call Higgsfield's own servers, whatever base_url says, and would send your Lumeta key there. Poll with the object submit() returns, or use subscribe().
The 19 Higgsfield paths Lumeta answers today (an unlisted path answers 404; each model's own paths are also in compat.higgsfield on GET /v1/models):
Show the paths
/bytedance/seedance-2.0/text-to-video/bytedance/seedance-2.0/image-to-video/bytedance/seedance-2.0/reference-to-video/bytedance/seedance-2.5/text-to-video/bytedance/seedance-2.5/image-to-video/bytedance/seedance-2.5/reference-to-video/bytedance/seedance-2.5/video-extend/minimax/h3/text-to-video/minimax/h3/image-to-video/minimax/h3/reference-to-video/alibaba/happy-horse/v1.1/text-to-video/alibaba/happy-horse/v1.1/image-to-video/alibaba/happy-horse/v1.1/reference-to-video/kling-video/o3/first-last-frame/kling-video/o3/image-reference/kling-video/v3/motion-control/std/kling-video/v3/motion-control/pro/xai/grok-imagine-video/v1.5/reference-to-video/xai/grok-imagine-image-2.0
Add ?hf_webhook=<https url> to a request for a Higgsfield-style callback. New to both? Use the native API above: it adds free estimates, credit limits per key, your Shelf by name and signed webhooks.
Full reference
Every endpoint, field and response is in the API reference, generated from openapi.json (OpenAPI 3.1, webhooks included). Point any OpenAPI tool or code generator at that address.
Questions, or something that should work and doesn't? Write to [email protected] with the request id. How keys, limits and your data are handled: API terms and data.