Create an image
Submit an image generation. Optionally anchor identity with a character.
POST /v1/images submits an image generation request. Credits are deducted immediately from your team balance; the generation is processed asynchronously. Poll GET /v1/images/{id} for status, or register a webhook endpoint for a push callback when the generation completes or fails.
For a step-by-step walkthrough, see the Quickstart. The full request shape, including all generation parameters, is in the playground below.
Using a style
Styles come fromGET /v1/loras; pass a style’s id (opaque lora_* or slug) as lora_id. The field is tri-state:
style: { id, name } (or null), and style.id round-trips — pass it back as lora_id to reuse the style.
Styles compose with composition acts and subjects — pin a style and an action_id together and both apply. One special case: some catalog entries are composition acts. Sending an act’s id as lora_id pins the act itself, so combining it with a different action_id returns 400 parameter_invalid_combination. The same applies to subjects: the few styles that pick their own model still return 400 parameter_invalid_combination when combined with subjects. lora_id remains incompatible with context_images.
Retired styles
Retired style ids keep working — how depends on the style:- Aliased — the id applies its designated successor style; the response
styleechoes the successor. Update your stored id at your convenience. - Plain — the request succeeds but generates without a style, and the response carries a
warnings[]entry with codestyle_retired_plain. - Discontinued — a small set of styles no longer generate at all:
400with codestyle_retired. Pick a current style fromGET /v1/loras.
Batching with count
count (1–4, whole number) generates that many images in parallel and bills per image. If some images in the batch fail, you receive the ones that succeeded and the difference is refunded automatically — the response image_count reflects the number actually delivered, and output_urls contains one URL per delivered image.
Warnings
The 201 body (and the estimate response) may carrywarnings[] — non-fatal adjustments the platform made to your request:
style_retired_plain (see above) and parameter_ignored (a parameter you sent has no effect on the generation path your request selected — for example seed on a styled generation). The key is omitted when there is nothing to report, appears only on the create and estimate responses (never on GETs, lists, or webhooks), and the code set is open — ignore codes you don’t recognize. Idempotent replays return the original warnings verbatim.
Response fields echo the request parameters as sent, not as used: an ignored parameter (flagged in warnings[]) is echoed back with exactly the value you sent, not the value that was actually applied.
Using a character
Whencharacter_id is set, the platform attaches the character’s saved reference images to the generation as visual anchors for identity consistency. The dispatch path is image-to-image, so denoise_strength becomes effective and influences how closely the output follows the refs vs the prompt.
The character must be in status: ready. Use a synthesizing / reviewing / failed character, or a soft-deleted one, and the request returns 400 character_not_ready.
character_id and reference_image_urls are mutually exclusive. Sending
both returns 400 mutually_exclusive_input. Pick one path per generation.Size
Specify image dimensions one of two ways — never both: Named preset viasize:
<tier>_<ratio> form, e.g. 2k_16_9, 4k_1_1, 2k_2_3.
Custom dimensions via width and height:
- Both required when used.
- Range
[1024, 4096]per side. - Snapped server-side to the nearest multiple of 32 — the response
width/heightreflect the post-snap value.
size and width/height returns 400 parameter_invalid_combination. Sending only one of width/height returns 400 missing_field.
Idempotency
PassIdempotency-Key (any opaque value, 1–256 chars; UUID v4 recommended). Same key + same body within 24h replays the cached response with Aurous-Idempotent-Replayed: true. Same key + different body returns 409 idempotency_key_in_use. The 24h window and 1–256 char bound are documented in Idempotency.
Webhooks
Register a webhook endpoint subscribed toimage.completed / image.failed (POST /v1/webhook_endpoints) to receive a POST callback when the generation reaches a terminal state. The event payload is { event: "image.completed" | "image.failed", data: {...} } where data matches the GET /v1/images/{id} response. See Webhooks for signature verification.Authorizations
Your team API key (starts with al_live_).
Headers
Stripe-style idempotency key (1-256 chars). Same key + same canonical-JSON body returns the cached response with Aurous-Idempotent-Replayed: true. Same key + different body returns 409 invalid_request / idempotency_key_in_use. UUID v4 recommended. Replay window is 24 hours. Absent header is treated as non-idempotent (each call processes anew).
Optional API version pin (YYYY-MM-DD). Defaults to your team's pinned version, or the system default 2026-07-16 for unauthenticated requests.
^\d{4}-\d{2}-\d{2}$"2026-07-16"
Body
The text prompt describing the image to generate. 1-4000 characters; whitespace-only is rejected.
1 - 4000"A golden sunset over mountains, cinematic lighting, 8k resolution"
Optional. Style identifier (lora_*) or slug, from GET /v1/loras. Tri-state: omit and matching runs automatically when your prompt names a look; send null to disable style matching for this request; send an id to pin that style. Styles now compose with action_id and subjects — a composition-act id in this field cannot be combined with a different action_id. Retired style ids keep working: aliased ids apply their successor style (echoed in response.style); other retired ids generate without a style and add a warnings[] entry.
"lora_01HXMQ7Z3K8Y2VNABCDEFGHJKM"
Optional character ID (char_<ulid> from POST /v1/characters; UUID also accepted for legacy back-compat). When set, the character's reference images are sent to the model as visual anchors for identity consistency. The character must be in status: ready — referencing a synthesizing / reviewing / failed character returns 400 character_not_ready. Cross-team character_ids return 404 (existence is never leaked). Mutually exclusive with reference_image_urls: sending both returns 400 mutually_exclusive_input. The output follows your prompt. Superseded by subjects[]; fully supported — successful responses carry an advisory Deprecation: true header when this field is used.
"char_01HXMQ7Z3K8Y2VNABCDEFGHJKM"
Custom output image width in pixels. Use with height OR use size (preset), not both. Range [1024, 4096]; snapped server-side to the nearest multiple of 32. Sending both size and custom dimensions returns 400 with code parameter_invalid_combination. Sending only one of width/height returns 400 with code missing_field.
1024 <= x <= 40962048
Custom output image height in pixels. Use with width OR use size (preset), not both. Range [1024, 4096]; snapped server-side to the nearest multiple of 32. Sending both size and custom dimensions returns 400 with code parameter_invalid_combination. Sending only one of width/height returns 400 with code missing_field.
1024 <= x <= 40962048
Image size as a named preset. Use this OR custom width/height, not both. Format is <tier>_<ratio> where tier is 2k or 4k and ratio matches the supported aspect-ratio set. Sending both size and custom dimensions returns 400 with code parameter_invalid_combination.
2k_1_1, 2k_3_2, 2k_2_3, 2k_4_3, 2k_3_4, 2k_16_9, 2k_9_16, 2k_21_9, 4k_1_1, 4k_3_2, 4k_2_3, 4k_4_3, 4k_3_4, 4k_16_9, 4k_9_16, 4k_21_9 "2k_1_1"
Number of images to generate in this request (1-4, whole number). Images generate in parallel and you are billed per image; if some images in the batch fail, you receive the ones that succeeded and the difference is refunded automatically — the response image_count reflects the number actually delivered.
1 <= x <= 41
When true, an LLM rewrites your prompt before generation to a more detailed, model-friendly form; the rewritten prompt is what reaches the model. This is the only customer-facing prompt-shaping toggle in the public API, and the only one that changes the price: enhanced generations cost a configurable multiplier of the base rate. Styled generations (a pinned or auto-matched style) always shape the prompt around the style — that built-in pass is not the enhancer, never fails a request, and never bills the enhancer multiplier unless you set this flag yourself.
false
Up to 6 reference images. Each entry can be either:
- an opaque file ID
file_<ulid>returned byPOST /v1/files, or - an
https://URL pointing at a public host (max 2048 chars). URLs are server-side fetched through an SSRF-pinned client (rejects private IPs / loopback / link-local / cloud metadata) and materialized as a 24h-TTL file under your team. Image files only — afile_<ulid>uploaded with purposereference_video/reference_audiois rejected (400invalid_format). Pricing matches the reference-image rate (see Pricing). Empty array or omitted is treated as "no references". Mutually exclusive withcharacter_id— sending both returns 400mutually_exclusive_input. Superseded bysubjects[]; fully supported — successful responses carry an advisoryDeprecation: trueheader when this field is used.
6Ordered subjects composed into one image (Image 1, Image 2, …). Max 10 subjects and 10 input images total. Mutually exclusive with character_id/reference_image_urls (both → 400 mutually_exclusive_input). Composes with lora_id when it names a style; the few styles that pick their own model still return 400 parameter_invalid_combination. Omit or [] for text-to-image/style.
10Up to 10 loose reference images for multi-image composition, interpreted from your prompt — no identity grouping or per-subject framing is applied, and no identity consistency is guaranteed. Positions follow array order and can be addressed in the prompt as "Image 1" … "Image 10" (e.g. "the outfit in Image 3"). Each entry is a file_<ulid> ID from POST /v1/files or an https URL to a public host (max 2048 chars). Empty array or omitted is treated as "no context images". Mutually exclusive with subjects, character_id, and reference_image_urls (400 mutually_exclusive_input) and with lora_id (400 parameter_invalid_combination — multi-image composition picks its own model; styles cannot be stacked). The response echoes only a count (context_images: { image_count }), never the image URLs.
10Pin a composition act from GET /v1/actions. Treat the id as opaque — it comes from the catalog and nowhere else. With subjects, the act must support your subject count (see supported_character_counts) — an unsupported count returns 400 action_not_available. Without subjects, the act renders with a new person described by your prompt. Omit to let act detection run automatically; send null to disable detection for this request. Combinable with lora_id when that id names a prompt style (the act and the style compose); a composition-act id sent in lora_id already acts as the pin, so it cannot be combined with a DIFFERENT action_id (400 parameter_invalid_combination). Mutually exclusive with context_images (400 mutually_exclusive_input). An id that is unknown or not visible to your team returns 404 resource_not_found — the same uniform 404 as GET /v1/actions/{id}.
"3f2b6c1e-8a4d-4e2b-9c7a-1d5e8f0a2b3c"
Output format. png yields a transparent background where the composition supports it. Default jpeg.
jpeg, png "jpeg"
Response
Generation created and pending processing
Discriminator — always inference. Mirrors OpenAI's object-field convention so SDK clients can branch on the resource type without inspecting the ID prefix. A single canonical value (inference) covers both image and video generations; use media_type to distinguish the rendering kind.
inference "inference"
Opaque generation ID
"img_01HXMQ7Z3K8Y2VNABCDEFGHJKM"
Current generation status. Lifecycle: pending (created, awaiting dispatch) → processing (running) → one of the terminal values succeeded / failed / cancelled. Additional terminal values may be introduced in future API versions and will be announced via the changelog before they appear on the wire.
pending, processing, succeeded, failed, cancelled "succeeded"
The text prompt used for generation
"A golden sunset over mountains, cinematic lighting"
Creation timestamp (ISO 8601)
"2026-05-04T10:00:00Z"
Distinguishes image vs video generation. May be null for older rows minted before this column existed.
image, video "image"
Generated image proxy URLs. Each URL is anonymous-read (no auth header required) and edge-cached for 24 hours. Available for ~24 hours after generation. Save what you want to keep — long-term storage is intentionally not part of the platform. URLs return 410 Gone after expiry.
Generated video proxy URL (only present on media_type: video). Same 24h TTL as image output_urls.
"https://api.aurous-labs.com/v1/videos/vid_01HXMQ7Z3K8Y2VNABCDEFGHJKM/output?token=..."
The reference image URLs you supplied as visual anchors for this generation, echoed back (snapshotted at inference time). Present only when the generation was driven by your own reference images (reference_image_urls or reference subjects). Omitted entirely for character-driven generations — a character's reference images are managed platform assets and are never echoed.
Human-readable error message if the generation failed. Non-contractual prose — do not parse or match on this value. Switch on error_code instead.
"Content policy violation"
Machine-readable failure reason when status is failed, for the cases where you need to branch in code. null on every successful generation, on failures recorded before this field existed, and on failure paths that have no stable code — error_message is human-facing copy that is re-tuned over time, so never pattern-match it. Currently emitted: generation_interrupted, reference_preparation_failed, content_filtered, generation_failed, and first_frame_too_small (an https:// first_frame_url attached to a video model whose dimensions were below the minimum — a file_<ulid> frame is rejected with the same code as a 400 before any credits are held; see Errors). The set is additive: new codes may appear, so treat an unrecognized value as a generic failure, keep a default branch in your switch statement, and fall back to error_message.
"generation_interrupted"
Processing duration in milliseconds (set on terminal status)
14820
Per-generation cost breakdown — same shape as the estimated_cost returned by POST /v1/{images,videos}/estimate. May be null for older rows from before this field existed; populated for all new generations. The amount reflects the committed charge for terminal-status rows — 0 (with refunded: true) on a failed generation, since the reserved hold was released, never charged.
Resolved output image width in pixels (image generations only). Reflects the post-snap dimension actually generated; may differ from a custom-requested width by up to 31 px due to multiple-of-32 snapping.
2048
Resolved output image height in pixels (image generations only). Reflects the post-snap dimension actually generated; may differ from a custom-requested height by up to 31 px due to multiple-of-32 snapping.
2048
Number of images in the batch. Reflects the requested count while the generation is running; on a terminal status it reflects the number actually DELIVERED — when part of a batch fails you receive the successful images, this count re-stamps to match, and the difference is refunded automatically.
1
Named size preset applied to this generation. null when the request used custom width/height instead of a preset.
2k_1_1, 2k_3_2, 2k_2_3, 2k_4_3, 2k_3_4, 2k_16_9, 2k_9_16, 2k_21_9, 4k_1_1, 4k_3_2, 4k_2_3, 4k_4_3, 4k_3_4, 4k_16_9, 4k_9_16, 4k_21_9 "2k_1_1"
Inference mode dispatched. Images: t2i (text-to-image) — reference images and characters are supplementary inputs to the t2i flow, not a separate mode. Videos (this list also returns vid_* rows): t2v (text-to-video), i2v (image-to-video — frame-driven, subject-driven, or a pinned video model on its own), or r2v (reference mode — the generation was driven by a motion reference). r2v does NOT imply you supplied that reference: it covers BOTH your own clip sent as reference_video_url AND a platform-built reference, produced when a first frame rides with a video model (pinned via video_lora_id, or auto-matched from the image). On the platform-built variant reference_video_url is null — branch on that field, not on inference_type, to tell the two apart. i2i is reserved for a future image-edit endpoint and is not currently emitted.
t2i, t2v, i2v, r2v, i2i "t2i"
CFG rescale factor the customer supplied on the request body, echoed back here. Range 0.0-1.0. Omitted when the customer did not supply a per-request value (the platform applied a precedence-chain default — LoRA, character override, or the global 0.7 — which is not exposed on the response).
0 <= x <= 10.7
Denoising strength the customer supplied on the request body, echoed back here. Range 0.0-1.0. Omitted when the customer did not supply a value or when the generation was a bare text-to-image request (denoise is only applied when reference images or a character are attached).
0 <= x <= 10.6
The random seed the model actually used for this image generation. Populated even when you omit seed on the request — the platform requests a random seed and records the concrete value the provider rolled, so you can reproduce the result by passing it back as seed. Available once status is succeeded; null before then and for failed/cancelled generations. For multi-image batches (image_count > 1) this is the seed of the first image (output_urls[0]); per-image seeds are not yet exposed. Image generations only.
819572108
Video duration in seconds (video generations only)
5
Video resolution (video generations only)
480p, 720p, 1080p "480p"
Video aspect ratio (video generations only)
16:9, 4:3, 1:1, 3:4, 9:16, 21:9, adaptive "16:9"
Whether the generated video includes synchronized audio (video generations only). Echoes the request generate_audio (default true).
true
Resolved reference-video task (video generations only). reference: the generation borrows the clip's motion for a new scene. extend: the generation continues the clip itself. null on every generation that did not supply reference_video_url (including all pre-existing rows). Additional task types may be introduced in future API versions — treat an unrecognized value as opaque.
reference, extend "reference"
Resolved extend direction (video generations only). Non-null only when video_task is extend; null otherwise (including every reference-mode and non-reference generation).
forward, backward "forward"
The reference video URL you submitted, echoed back VERBATIM as you sent it — never re-signed, never a storage path. null for generations that did not supply reference_video_url (including every dashboard-originated row, and including an inference_type: "r2v" generation whose motion reference the platform built from a first frame plus a pinned or matched video model — this field, not inference_type, is what distinguishes your own clip from a platform-built reference).
"https://cdn.example.com/clips/dance-loop.mp4"
The reference audio URL you submitted, echoed back VERBATIM as you sent it. null for generations that did not supply reference_audio_url.
"https://cdn.example.com/audio/voiceover.mp3"
The video model (an id or slug from GET /v1/video_loras) behind this generation (video generations only) — either the video_lora_id you pinned on the request, or the model the platform auto-matched to your prompt when none was pinned. null for a plain (no video model) generation, and always null on image generations.
"lora_01HXMQ7Z3K8Y2VNABCDEFGHJKM"
Display name of the video model in video_lora_id, if any. Omitted for plain generations (no pinned or matched video model).
"Cinematic Pan"
Character ID supplied on the request (char_<ulid> or legacy UUID), echoed back. null when no character was attached to this generation.
"char_01HXMQ7Z3K8Y2VNABCDEFGHJKM"
Ordered subjects composed into this generation (positional — entry N echoes entry N of the identity inputs sent on create; legacy character_id / reference_image_urls inputs are normalized into the same projection). Character entries carry the opaque char_<ulid>; reference entries echo only their image count, never URLs. Video generations echo their cast the same way (up to 2 entries), matching the subjects[] array accepted by POST /v1/videos. Character entries on generations created before cast snapshotting may carry character_id: null; the id is never substituted with an internal identifier. null for generations that did not compose subjects (plain prompt-only or style generations, and context-image generations). The engine is not exposed.
Count-only echo of the context_images[] sent on create — the number of loose reference images supplied, never the image URLs themselves. null for every generation that did not use context_images (plain prompt-only, style, subject, and video generations).
The composition act applied to this generation, echoed as { id, name }. An act is applied either because you pinned it with action_id or because act detection matched one automatically. id is the act identifier from GET /v1/actions; name is its display name (or null when the name was not recorded). null for every generation that did not apply an act — plain prompt-only, style, context-image, and video generations. A new, always-present, nullable key: existing integrations that do not read it are unaffected.
The prompt style applied to this generation, echoed as { id, name }. A style is applied either because you pinned it with lora_id or because style matching resolved one from your prompt — the echo always reflects the style that actually ran (a retired id that aliases to a successor echoes the successor). id is the lora_* identifier from GET /v1/loras and round-trips into lora_id. null for every generation without a style — including lora_id: null requests, retired styles that generate plain, video generations, and all rows minted before styles shipped. A new, always-present, nullable key: existing integrations that do not read it are unaffected.
Non-fatal request adjustments, present ONLY on the POST /v1/images 201 body (and the estimate response) — never on GET reads, list rows, or webhook payloads. Omitted entirely when empty. Current codes: parameter_ignored, style_retired_plain; new codes may be added without a version bump — ignore unknown codes. Idempotent replays return the original warnings verbatim.
Outcome of the prompt-optimization step, echoed for observability. Current values: no_request (the request did not invoke it), succeeded, identical, refused, transport_error, length_overflow; null on rows minted before this field existed. INFORMATIONAL and deliberately an OPEN set (no schema enum — codegen clients must not mint a closed union): do not branch control flow on it; new values may be added without a version bump.
"succeeded"
LoRAs applied to this generation. null for prompt-only and pure-reference generations.
API contract version applied at the time this row was minted (D25 — frozen for replay across future version bumps).
"2026-07-16"
Aurous-Request-Id of the POST that created this row. Quote in support tickets to trace the original create request.
"req_01HXMQ7Z3K8Y2VNABCDEFGHJKM"
Terminal-status timestamp (ISO 8601). NULL until the generation reaches a terminal state.
"2026-05-04T10:00:14Z"

