The Lumeta API

Every Lumeta model over plain HTTPS: the same account, credits, Shelf and projects you use on the web. Ask what a run will cost before you spend, start it, and get the result by waiting, polling or a webhook. Everything you make lands in your Lumeta library with a small API tag.

Quick start

Make a key under Account → Developer (keys are part of the Pro and Ultimate plans), then:

Shell
export LUMETA_API_KEY=lm_live_...

# 1. What will it cost? (free, spends nothing)
curl https://api.lumeta.ai/v1/estimate \
  -H "Authorization: Bearer $LUMETA_API_KEY" -H "Content-Type: application/json" \
  -d '{"model": "nano-banana-2", "input": {"prompt": "a red panda astronaut, studio light"}}'

# 2. Make it
curl https://api.lumeta.ai/v1/generations \
  -H "Authorization: Bearer $LUMETA_API_KEY" -H "Content-Type: application/json" \
  -H "Idempotency-Key: red-panda-001" \
  -d '{"model": "nano-banana-2", "input": {"prompt": "a red panda astronaut, studio light"}}'

# 3. Wait for the result (the id came back from step 2)
curl "https://api.lumeta.ai/v1/generations/gen_123?wait=45" -H "Authorization: Bearer $LUMETA_API_KEY"

Step 2 answers 202 with the generation's id while it runs (or 200 when the tool answers at once). Step 3 holds the request until it finishes or 45 seconds pass, whichever comes first; call it again if it's still running. A finished generation looks like this (shortened):

Response
{
    "id": "gen_123",
    "status": "succeeded",
    "model": {
        "id": 140,
        "name": "Nano Banana 2"
    },
    "outputs": [
        {
            "index": 0,
            "type": "image",
            "url": "https://…/result.jpg",
            "width": 1024,
            "height": 1024
        }
    ],
    "credits_charged": 20,
    "origin": "api",
    "completed_in_seconds": 6.2
}

The same in Python and JavaScript

Python
import os, requests

API = "https://api.lumeta.ai/v1"
H = {"Authorization": "Bearer " + os.environ["LUMETA_API_KEY"]}
body = {"model": "nano-banana-2", "input": {"prompt": "a red panda astronaut, studio light"}}

print(requests.post(f"{API}/estimate", json=body, headers=H).json()["credits"], "credits")
gen = requests.post(f"{API}/generations", json=body, headers={**H, "Idempotency-Key": "red-panda-001"}).json()
while gen["status"] not in ("succeeded", "failed"):
    gen = requests.get(f"{API}/generations/{gen['id']}", params={"wait": 45}, headers=H).json()
print(gen["outputs"][0]["url"] if gen["status"] == "succeeded" else gen["error"])
JavaScript (Node 18+)
const API = "https://api.lumeta.ai/v1";
const H = { Authorization: `Bearer ${process.env.LUMETA_API_KEY}`, "Content-Type": "application/json" };
const body = JSON.stringify({ model: "nano-banana-2", input: { prompt: "a red panda astronaut, studio light" } });

let gen = await (await fetch(`${API}/generations`, { method: "POST", headers: { ...H, "Idempotency-Key": "red-panda-001" }, body })).json();
while (!["succeeded", "failed"].includes(gen.status)) {
  gen = await (await fetch(`${API}/generations/${gen.id}?wait=45`, { headers: H })).json();
}
console.log(gen.status === "succeeded" ? gen.outputs[0].url : gen.error);

Keys and sign-in

Send Authorization: Bearer lm_live_... on every request. Keys are made under Account → Developer, and each one carries its own settings:

  • Permissions (scopes): read (look at models, balance, generations, Shelf and projects), generate (start generations, spends credits) and library (upload files, change your Shelf and projects, manage webhooks).
  • Credit limits: an optional daily and monthly cap. A request that would go over it is refused with 402 token_cap_exceeded, and nothing is spent.
  • Automatic top-up: off by default. When it's off, a key that runs out of credits stops and tells you, instead of charging your card.
  • Expiry: 30, 90, 365 days, or never.

A key is shown once, when you make it. We keep only a fingerprint of it, so treat it like a password: keep it on your server, never in a web page or an app you ship. Revoke it on the same page and it stops working at once.

Building an app other people sign into

Use OAuth 2.1 instead of asking people for keys: the authorization code flow with PKCE, for public clients. Clients can register themselves (dynamic client registration) or use a client ID metadata document. Everything a library needs is at /.well-known/oauth-authorization-server. People see a Lumeta consent page where they choose permissions, a daily limit and top-up, just like for a key. Access tokens (lma_...) last 1 hour; refresh tokens last 90 days and change with every use.

Models and inputs

GET /v1/models lists every model you can run with its starting price (base_credits). Filter with type=image|video|audio or q=, or describe the job with task= and get a recommendation. It also tells you which model your account opens with for each kind (defaults).

Each model names its inputs its own way, so read GET /v1/models/{slug} first: it includes the model's input_schema (JSON Schema 2020-12) with every field, its choices and defaults. Starting prices are a guide; POST /v1/estimate with your exact input is the price you'll pay.

Media inputs

Fields that take an image, video or audio take a reference, never a raw filename:

WriteMeans
file_123A file you uploaded or imported (POST /v1/files, POST /v1/files/import)
shelf_45 or shelf:Blue HatAn item on your Shelf, by id or by the name you gave it
gen_123#0Output 0 of one of your generations (counting from 0)
voice_7One of your own voices (GET /v1/voices)
https://…A public link, imported for you

In a prompt, @Blue Hat (or @[Blue Hat] for a name with spaces) attaches the Shelf item with that name where the model takes it. A word after @ that is no Shelf name is sent as text and noted in warnings, so a handle never breaks a script.

Generations

POST /v1/generations with {"model", "input", "project"} starts one and spends credits. A generation is queued, processing, succeeded or failed. A failed one is refunded by itself (credits_charged reads 0).

  • Waiting. GET /v1/generations/{id}?wait=45 holds the request until it's done or the time is up. Ask again for long videos; they keep running whether you wait or not.
  • Safe retries. Send an Idempotency-Key header on a create. Repeating the same key with the same body within 24 hours returns the first run instead of starting (and paying for) another.
  • Canceling. POST /v1/generations/{id}/cancel stops a run and refunds it where the provider can really stop it (today, models that run on Replicate). Elsewhere it answers 409 not_cancelable and the run finishes normally. A canceled run reads failed with error.code canceled.
  • Your history. GET /v1/generations lists your runs, newest first, with filters for status, model, project and origin (web, api, mcp, cli). Lists take limit (default 20, up to 100) and page with cursor.

Files, Shelf and projects

  • POST /v1/files uploads a file (multipart, part file) with the same checks and limits as the web uploader; POST /v1/files/import takes a public https:// link. Both answer with a file_… id to use as an input.
  • GET /v1/shelves lists your Shelves (proj_0 is All Projects); GET /v1/shelves/{project}/items?q= finds items by name. POST pins an output or file under a name, PATCH renames, DELETE removes the pin (the file stays).
  • GET /v1/projects and POST /v1/projects. A new project becomes your current one, as on the web.
  • GET /v1/account: your plan, credits, and this key's permissions, limits and spend.

Webhooks

Rather than wait, let Lumeta tell you. Register a receiver once:

curl https://api.lumeta.ai/v1/webhook-endpoints \
  -H "Authorization: Bearer $LUMETA_API_KEY" -H "Content-Type: application/json" \
  -d '{"url": "https://example.com/lumeta-webhook", "events": ["*"]}'

The answer includes the endpoint's signing secret (whsec_…), shown this once. From then on, every generation on your account that finishes, whichever way it was started (web, API, assistant or CLI; data.origin says which), is POSTed to your URL a few seconds later as generation.succeeded or generation.failed. data is the generation exactly as GET /v1/generations/{id} shows it. You can have up to 5 endpoints; POST /v1/webhook-endpoints/{id}/test sends a test event right away.

Check the signature

Every request carries Lumeta-Signature: t=<unix seconds>,v1=<hex>, where v1 is the HMAC-SHA256 of "<t>.<raw body>" keyed with your secret. Compute it over the raw bytes (before any JSON parsing) and refuse a t more than 5 minutes old:

Node
import crypto from "node:crypto";

export function verifyLumeta(rawBody, header, secret) {
  const p = Object.fromEntries(header.split(",").map((kv) => kv.split("=")));
  if (!p.t || Math.abs(Date.now() / 1000 - Number(p.t)) > 300) return false;
  const want = crypto.createHmac("sha256", secret).update(`${p.t}.${rawBody}`).digest("hex");
  return typeof p.v1 === "string" && p.v1.length === want.length
    && crypto.timingSafeEqual(Buffer.from(p.v1), Buffer.from(want));
}
Python
import hashlib, hmac, time

def verify_lumeta(raw_body: bytes, header: str, secret: str) -> bool:
    p = dict(kv.split("=", 1) for kv in header.split(","))
    if abs(time.time() - int(p.get("t", 0))) > 300:
        return False
    want = hmac.new(secret.encode(), p["t"].encode() + b"." + raw_body, hashlib.sha256).hexdigest()
    return hmac.compare_digest(want, p.get("v1", ""))
PHP
function verify_lumeta(string $rawBody, string $header, string $secret): bool
{
    parse_str(str_replace(',', '&', $header), $p);
    if (abs(time() - (int)($p['t'] ?? 0)) > 300) return false;
    return hash_equals(hash_hmac('sha256', $p['t'] . '.' . $rawBody, $secret), (string)($p['v1'] ?? ''));
}

Answering, retries and order

  • Answer with any 2xx within 5 seconds. Do the slow work after you answer.
  • Anything else (a timeout, a redirect, a 4xx or 5xx) is tried again after 1 minute, 5 minutes, 30 minutes, 2 hours and 12 hours, with the same event id (Lumeta-Event-Id, also body.id). Use it to skip duplicates.
  • Events can arrive out of order, and a retry carries the generation as it is at that moment, so act on data.status.
  • If an event fails all 6 attempts and your endpoint has not answered a 2xx once in the meantime, the endpoint is switched off (enabled: false). PATCH {"enabled": true} turns it back on.
  • Receivers must be https:// on port 443 at a public address. Lumeta never follows redirects and reads at most 4 KB of your answer.

Errors

Every error has the same shape, and every response carries an X-Request-Id header. Quote the request id when you write to us and we can find exactly what happened.

{
    "error": {
        "code": "insufficient_credits",
        "message": "Not enough credits for this request. Top up at https://lumeta.ai/subscription/",
        "request_id": "01J9…"
    }
}

The code is stable; build on it, not on the message. param names the input at fault when there is one.

StatusCodeWhat to do
400invalid_jsonThe body isn't valid JSON.
401unauthorized, invalid_tokenNo key, or a key that was revoked or expired.
402insufficient_creditsNot enough credits. Top up, or turn on automatic top-up for this key.
402token_cap_exceededThis key's daily or monthly limit would be passed.
403plan_required, insufficient_scope, project_read_onlyThe plan, the key's permissions, or a view-only project doesn't allow it.
404not_foundNo such thing on your account.
409idempotency_conflict, idempotency_in_progress, already_exists, not_cancelableThe same Idempotency-Key with a different body, or one still running; a duplicate; a run that can't be stopped.
422invalid_input, moderation_blockedAn input is wrong (see param), or the request breaks the content rules.
429rate_limited, concurrency_limitSlow down: wait the number of seconds in Retry-After.
502provider_errorThe model's provider failed. You were not charged; try again.
503model_unavailable, service_unavailableDown for a moment. Wait a few seconds and retry.
500internal_errorOur fault. It's logged; send us the request id.

Limits

LimitPer key or connection
Requests300 a minute
New generations20 a minute
Estimates60 a minute
Waiting on one request45 seconds
Items per list pageup to 100
Webhook endpoints5 per account

Over a limit, you get 429 with Retry-After. Some models also allow only a few runs at a time per account, as on the web (concurrency_limit).

Coming from Higgsfield

If you use Higgsfield's API, you can switch by changing one line. https://api.lumeta.ai/higgsfield speaks Higgsfield's dialect (their paths, request and status shapes, and errors) for the models Lumeta also runs, with your Lumeta key and credits.

Python SDK
import higgsfield_client

# HF_KEY = your Lumeta key, exactly as shown when you made it
client = higgsfield_client.SyncClient(base_url="https://api.lumeta.ai/higgsfield")
JavaScript SDK
import { config, higgsfield } from "@higgsfield/client/v2";

// Your Lumeta key with its last "_" written as ":" (lm_live_<id>:<secret>)
config({ credentials: "lm_live_<id>:<secret>", baseURL: "https://api.lumeta.ai/higgsfield" });

One trap in the Python SDK: its helper functions status(), result() and cancel() that take a request id always call Higgsfield's own servers, whatever base_url says, and would send your Lumeta key there. Poll with the object submit() returns, or use subscribe().

The 19 Higgsfield paths Lumeta answers today (an unlisted path answers 404; each model's own paths are also in compat.higgsfield on GET /v1/models):

Show the paths
  • /bytedance/seedance-2.0/text-to-video
  • /bytedance/seedance-2.0/image-to-video
  • /bytedance/seedance-2.0/reference-to-video
  • /bytedance/seedance-2.5/text-to-video
  • /bytedance/seedance-2.5/image-to-video
  • /bytedance/seedance-2.5/reference-to-video
  • /bytedance/seedance-2.5/video-extend
  • /minimax/h3/text-to-video
  • /minimax/h3/image-to-video
  • /minimax/h3/reference-to-video
  • /alibaba/happy-horse/v1.1/text-to-video
  • /alibaba/happy-horse/v1.1/image-to-video
  • /alibaba/happy-horse/v1.1/reference-to-video
  • /kling-video/o3/first-last-frame
  • /kling-video/o3/image-reference
  • /kling-video/v3/motion-control/std
  • /kling-video/v3/motion-control/pro
  • /xai/grok-imagine-video/v1.5/reference-to-video
  • /xai/grok-imagine-image-2.0

Add ?hf_webhook=<https url> to a request for a Higgsfield-style callback. New to both? Use the native API above: it adds free estimates, credit limits per key, your Shelf by name and signed webhooks.

Full reference

Every endpoint, field and response is in the API reference, generated from openapi.json (OpenAPI 3.1, webhooks included). Point any OpenAPI tool or code generator at that address.

Questions, or something that should work and doesn't? Write to [email protected] with the request id. How keys, limits and your data are handled: API terms and data.