Skip to main content
POST
Create a new image generation
POST /v1/images submits an image generation request. Credits are deducted immediately from your team balance; the generation is processed asynchronously. Poll GET /v1/images/{id} for status, or register a webhook endpoint for a push callback when the generation completes or fails. For a step-by-step walkthrough, see the Quickstart. The full request shape, including all generation parameters, is in the playground below.

Using a style

Styles come from GET /v1/loras; pass a style’s id (opaque lora_* or slug) as lora_id. The field is tri-state: Whatever happens, the outcome is observable: every generation response echoes style: { id, name } (or null), and style.id round-trips — pass it back as lora_id to reuse the style. Styles compose with composition acts and subjects — pin a style and an action_id together and both apply. One special case: some catalog entries are composition acts. Sending an act’s id as lora_id pins the act itself, so combining it with a different action_id returns 400 parameter_invalid_combination. The same applies to subjects: the few styles that pick their own model still return 400 parameter_invalid_combination when combined with subjects. lora_id remains incompatible with context_images.

Retired styles

Retired style ids keep working — how depends on the style:
  • Aliased — the id applies its designated successor style; the response style echoes the successor. Update your stored id at your convenience.
  • Plain — the request succeeds but generates without a style, and the response carries a warnings[] entry with code style_retired_plain.
  • Discontinued — a small set of styles no longer generate at all: 400 with code style_retired. Pick a current style from GET /v1/loras.

Batching with count

count (1–4, whole number) generates that many images in parallel and bills per image. If some images in the batch fail, you receive the ones that succeeded and the difference is refunded automatically — the response image_count reflects the number actually delivered, and output_urls contains one URL per delivered image.

Warnings

The 201 body (and the estimate response) may carry warnings[] — non-fatal adjustments the platform made to your request:
Current codes are style_retired_plain (see above) and parameter_ignored (a parameter you sent has no effect on the generation path your request selected — for example seed on a styled generation). The key is omitted when there is nothing to report, appears only on the create and estimate responses (never on GETs, lists, or webhooks), and the code set is open — ignore codes you don’t recognize. Idempotent replays return the original warnings verbatim. Response fields echo the request parameters as sent, not as used: an ignored parameter (flagged in warnings[]) is echoed back with exactly the value you sent, not the value that was actually applied.

Using a character

When character_id is set, the platform attaches the character’s saved reference images to the generation as visual anchors for identity consistency. The dispatch path is image-to-image, so denoise_strength becomes effective and influences how closely the output follows the refs vs the prompt. The character must be in status: ready. Use a synthesizing / reviewing / failed character, or a soft-deleted one, and the request returns 400 character_not_ready.
character_id and reference_image_urls are mutually exclusive. Sending both returns 400 mutually_exclusive_input. Pick one path per generation.
cURL
If the character has multiple ref poses, the dispatcher consumes all of them as anchors. There is no current way to limit attachment to a subset of poses — the whole ref set goes in.

Size

Specify image dimensions one of two ways — never both: Named preset via size: Combine into a preset string in <tier>_<ratio> form, e.g. 2k_16_9, 4k_1_1, 2k_2_3. Custom dimensions via width and height:
  • Both required when used.
  • Range [1024, 4096] per side.
  • Snapped server-side to the nearest multiple of 32 — the response width/height reflect the post-snap value.
Sending both size and width/height returns 400 parameter_invalid_combination. Sending only one of width/height returns 400 missing_field.

Idempotency

Pass Idempotency-Key (any opaque value, 1–256 chars; UUID v4 recommended). Same key + same body within 24h replays the cached response with Aurous-Idempotent-Replayed: true. Same key + different body returns 409 idempotency_key_in_use. The 24h window and 1–256 char bound are documented in Idempotency.

Webhooks

Register a webhook endpoint subscribed to image.completed / image.failed (POST /v1/webhook_endpoints) to receive a POST callback when the generation reaches a terminal state. The event payload is { event: "image.completed" | "image.failed", data: {...} } where data matches the GET /v1/images/{id} response. See Webhooks for signature verification.

Authorizations

X-Api-Key
string
header
required

Your team API key (starts with al_live_).

Headers

Idempotency-Key
string

Stripe-style idempotency key (1-256 chars). Same key + same canonical-JSON body returns the cached response with Aurous-Idempotent-Replayed: true. Same key + different body returns 409 invalid_request / idempotency_key_in_use. UUID v4 recommended. Replay window is 24 hours. Absent header is treated as non-idempotent (each call processes anew).

Aurous-Version
string

Optional API version pin (YYYY-MM-DD). Defaults to your team's pinned version, or the system default 2026-07-16 for unauthenticated requests.

Pattern: ^\d{4}-\d{2}-\d{2}$
Example:

"2026-07-16"

Body

application/json
prompt
string
required

The text prompt describing the image to generate. 1-4000 characters; whitespace-only is rejected.

Required string length: 1 - 4000
Example:

"A golden sunset over mountains, cinematic lighting, 8k resolution"

lora_id
string | null

Optional. Style identifier (lora_*) or slug, from GET /v1/loras. Tri-state: omit and matching runs automatically when your prompt names a look; send null to disable style matching for this request; send an id to pin that style. Styles now compose with action_id and subjects — a composition-act id in this field cannot be combined with a different action_id. Retired style ids keep working: aliased ids apply their successor style (echoed in response.style); other retired ids generate without a style and add a warnings[] entry.

Example:

"lora_01HXMQ7Z3K8Y2VNABCDEFGHJKM"

character_id
string
deprecated

Optional character ID (char_<ulid> from POST /v1/characters; UUID also accepted for legacy back-compat). When set, the character's reference images are sent to the model as visual anchors for identity consistency. The character must be in status: ready — referencing a synthesizing / reviewing / failed character returns 400 character_not_ready. Cross-team character_ids return 404 (existence is never leaked). Mutually exclusive with reference_image_urls: sending both returns 400 mutually_exclusive_input. The output follows your prompt. Superseded by subjects[]; fully supported — successful responses carry an advisory Deprecation: true header when this field is used.

Example:

"char_01HXMQ7Z3K8Y2VNABCDEFGHJKM"

width
number

Custom output image width in pixels. Use with height OR use size (preset), not both. Range [1024, 4096]; snapped server-side to the nearest multiple of 32. Sending both size and custom dimensions returns 400 with code parameter_invalid_combination. Sending only one of width/height returns 400 with code missing_field.

Required range: 1024 <= x <= 4096
Example:

2048

height
number

Custom output image height in pixels. Use with width OR use size (preset), not both. Range [1024, 4096]; snapped server-side to the nearest multiple of 32. Sending both size and custom dimensions returns 400 with code parameter_invalid_combination. Sending only one of width/height returns 400 with code missing_field.

Required range: 1024 <= x <= 4096
Example:

2048

size
enum<string>

Image size as a named preset. Use this OR custom width/height, not both. Format is <tier>_<ratio> where tier is 2k or 4k and ratio matches the supported aspect-ratio set. Sending both size and custom dimensions returns 400 with code parameter_invalid_combination.

Available options:
2k_1_1,
2k_3_2,
2k_2_3,
2k_4_3,
2k_3_4,
2k_16_9,
2k_9_16,
2k_21_9,
4k_1_1,
4k_3_2,
4k_2_3,
4k_4_3,
4k_3_4,
4k_16_9,
4k_9_16,
4k_21_9
Example:

"2k_1_1"

count
number

Number of images to generate in this request (1-4, whole number). Images generate in parallel and you are billed per image; if some images in the batch fail, you receive the ones that succeeded and the difference is refunded automatically — the response image_count reflects the number actually delivered.

Required range: 1 <= x <= 4
Example:

1

enhance_prompt
boolean

When true, an LLM rewrites your prompt before generation to a more detailed, model-friendly form; the rewritten prompt is what reaches the model. This is the only customer-facing prompt-shaping toggle in the public API, and the only one that changes the price: enhanced generations cost a configurable multiplier of the base rate. Styled generations (a pinned or auto-matched style) always shape the prompt around the style — that built-in pass is not the enhancer, never fails a request, and never bills the enhancer multiplier unless you set this flag yourself.

Example:

false

reference_image_urls
string[]
deprecated

Up to 6 reference images. Each entry can be either:

  • an opaque file ID file_<ulid> returned by POST /v1/files, or
  • an https:// URL pointing at a public host (max 2048 chars). URLs are server-side fetched through an SSRF-pinned client (rejects private IPs / loopback / link-local / cloud metadata) and materialized as a 24h-TTL file under your team. Image files only — a file_<ulid> uploaded with purpose reference_video/reference_audio is rejected (400 invalid_format). Pricing matches the reference-image rate (see Pricing). Empty array or omitted is treated as "no references". Mutually exclusive with character_id — sending both returns 400 mutually_exclusive_input. Superseded by subjects[]; fully supported — successful responses carry an advisory Deprecation: true header when this field is used.
Maximum array length: 6
Example:
subjects
object[]

Ordered subjects composed into one image (Image 1, Image 2, …). Max 10 subjects and 10 input images total. Mutually exclusive with character_id/reference_image_urls (both → 400 mutually_exclusive_input). Composes with lora_id when it names a style; the few styles that pick their own model still return 400 parameter_invalid_combination. Omit or [] for text-to-image/style.

Maximum array length: 10
context_images
string[]

Up to 10 loose reference images for multi-image composition, interpreted from your prompt — no identity grouping or per-subject framing is applied, and no identity consistency is guaranteed. Positions follow array order and can be addressed in the prompt as "Image 1" … "Image 10" (e.g. "the outfit in Image 3"). Each entry is a file_<ulid> ID from POST /v1/files or an https URL to a public host (max 2048 chars). Empty array or omitted is treated as "no context images". Mutually exclusive with subjects, character_id, and reference_image_urls (400 mutually_exclusive_input) and with lora_id (400 parameter_invalid_combination — multi-image composition picks its own model; styles cannot be stacked). The response echoes only a count (context_images: { image_count }), never the image URLs.

Maximum array length: 10
Example:
action_id
string | null

Pin a composition act from GET /v1/actions. Treat the id as opaque — it comes from the catalog and nowhere else. With subjects, the act must support your subject count (see supported_character_counts) — an unsupported count returns 400 action_not_available. Without subjects, the act renders with a new person described by your prompt. Omit to let act detection run automatically; send null to disable detection for this request. Combinable with lora_id when that id names a prompt style (the act and the style compose); a composition-act id sent in lora_id already acts as the pin, so it cannot be combined with a DIFFERENT action_id (400 parameter_invalid_combination). Mutually exclusive with context_images (400 mutually_exclusive_input). An id that is unknown or not visible to your team returns 404 resource_not_found — the same uniform 404 as GET /v1/actions/{id}.

Example:

"3f2b6c1e-8a4d-4e2b-9c7a-1d5e8f0a2b3c"

output_format
enum<string>

Output format. png yields a transparent background where the composition supports it. Default jpeg.

Available options:
jpeg,
png
Example:

"jpeg"

Response

Generation created and pending processing

object
enum<string>
default:inference
required

Discriminator — always inference. Mirrors OpenAI's object-field convention so SDK clients can branch on the resource type without inspecting the ID prefix. A single canonical value (inference) covers both image and video generations; use media_type to distinguish the rendering kind.

Available options:
inference
Example:

"inference"

id
string
required

Opaque generation ID

Example:

"img_01HXMQ7Z3K8Y2VNABCDEFGHJKM"

status
enum<string>
required

Current generation status. Lifecycle: pending (created, awaiting dispatch) → processing (running) → one of the terminal values succeeded / failed / cancelled. Additional terminal values may be introduced in future API versions and will be announced via the changelog before they appear on the wire.

Available options:
pending,
processing,
succeeded,
failed,
cancelled
Example:

"succeeded"

prompt
string
required

The text prompt used for generation

Example:

"A golden sunset over mountains, cinematic lighting"

created_at
string
required

Creation timestamp (ISO 8601)

Example:

"2026-05-04T10:00:00Z"

media_type
enum<string> | null

Distinguishes image vs video generation. May be null for older rows minted before this column existed.

Available options:
image,
video
Example:

"image"

output_urls
string[] | null

Generated image proxy URLs. Each URL is anonymous-read (no auth header required) and edge-cached for 24 hours. Available for ~24 hours after generation. Save what you want to keep — long-term storage is intentionally not part of the platform. URLs return 410 Gone after expiry.

Example:
video_url
string | null

Generated video proxy URL (only present on media_type: video). Same 24h TTL as image output_urls.

Example:

"https://api.aurous-labs.com/v1/videos/vid_01HXMQ7Z3K8Y2VNABCDEFGHJKM/output?token=..."

reference_image_urls
string[] | null

The reference image URLs you supplied as visual anchors for this generation, echoed back (snapshotted at inference time). Present only when the generation was driven by your own reference images (reference_image_urls or reference subjects). Omitted entirely for character-driven generations — a character's reference images are managed platform assets and are never echoed.

Example:
error_message
string | null

Human-readable error message if the generation failed. Non-contractual prose — do not parse or match on this value. Switch on error_code instead.

Example:

"Content policy violation"

error_code
string | null

Machine-readable failure reason when status is failed, for the cases where you need to branch in code. null on every successful generation, on failures recorded before this field existed, and on failure paths that have no stable code — error_message is human-facing copy that is re-tuned over time, so never pattern-match it. Currently emitted: generation_interrupted, reference_preparation_failed, content_filtered, generation_failed, and first_frame_too_small (an https:// first_frame_url attached to a video model whose dimensions were below the minimum — a file_<ulid> frame is rejected with the same code as a 400 before any credits are held; see Errors). The set is additive: new codes may appear, so treat an unrecognized value as a generic failure, keep a default branch in your switch statement, and fall back to error_message.

Example:

"generation_interrupted"

duration_ms
integer | null

Processing duration in milliseconds (set on terminal status)

Example:

14820

cost
object | null

Per-generation cost breakdown — same shape as the estimated_cost returned by POST /v1/{images,videos}/estimate. May be null for older rows from before this field existed; populated for all new generations. The amount reflects the committed charge for terminal-status rows — 0 (with refunded: true) on a failed generation, since the reserved hold was released, never charged.

Example:
width
number

Resolved output image width in pixels (image generations only). Reflects the post-snap dimension actually generated; may differ from a custom-requested width by up to 31 px due to multiple-of-32 snapping.

Example:

2048

height
number

Resolved output image height in pixels (image generations only). Reflects the post-snap dimension actually generated; may differ from a custom-requested height by up to 31 px due to multiple-of-32 snapping.

Example:

2048

image_count
number

Number of images in the batch. Reflects the requested count while the generation is running; on a terminal status it reflects the number actually DELIVERED — when part of a batch fails you receive the successful images, this count re-stamps to match, and the difference is refunded automatically.

Example:

1

size_preset
enum<string> | null

Named size preset applied to this generation. null when the request used custom width/height instead of a preset.

Available options:
2k_1_1,
2k_3_2,
2k_2_3,
2k_4_3,
2k_3_4,
2k_16_9,
2k_9_16,
2k_21_9,
4k_1_1,
4k_3_2,
4k_2_3,
4k_4_3,
4k_3_4,
4k_16_9,
4k_9_16,
4k_21_9
Example:

"2k_1_1"

inference_type
enum<string> | null

Inference mode dispatched. Images: t2i (text-to-image) — reference images and characters are supplementary inputs to the t2i flow, not a separate mode. Videos (this list also returns vid_* rows): t2v (text-to-video), i2v (image-to-video — frame-driven, subject-driven, or a pinned video model on its own), or r2v (reference mode — the generation was driven by a motion reference). r2v does NOT imply you supplied that reference: it covers BOTH your own clip sent as reference_video_url AND a platform-built reference, produced when a first frame rides with a video model (pinned via video_lora_id, or auto-matched from the image). On the platform-built variant reference_video_url is null — branch on that field, not on inference_type, to tell the two apart. i2i is reserved for a future image-edit endpoint and is not currently emitted.

Available options:
t2i,
t2v,
i2v,
r2v,
i2i
Example:

"t2i"

cfg_rescale
number | null

CFG rescale factor the customer supplied on the request body, echoed back here. Range 0.0-1.0. Omitted when the customer did not supply a per-request value (the platform applied a precedence-chain default — LoRA, character override, or the global 0.7 — which is not exposed on the response).

Required range: 0 <= x <= 1
Example:

0.7

denoise_strength
number | null

Denoising strength the customer supplied on the request body, echoed back here. Range 0.0-1.0. Omitted when the customer did not supply a value or when the generation was a bare text-to-image request (denoise is only applied when reference images or a character are attached).

Required range: 0 <= x <= 1
Example:

0.6

seed
integer | null

The random seed the model actually used for this image generation. Populated even when you omit seed on the request — the platform requests a random seed and records the concrete value the provider rolled, so you can reproduce the result by passing it back as seed. Available once status is succeeded; null before then and for failed/cancelled generations. For multi-image batches (image_count > 1) this is the seed of the first image (output_urls[0]); per-image seeds are not yet exposed. Image generations only.

Example:

819572108

video_duration
integer | null

Video duration in seconds (video generations only)

Example:

5

video_resolution
enum<string> | null

Video resolution (video generations only)

Available options:
480p,
720p,
1080p
Example:

"480p"

video_ratio
enum<string> | null

Video aspect ratio (video generations only)

Available options:
16:9,
4:3,
1:1,
3:4,
9:16,
21:9,
adaptive
Example:

"16:9"

video_generate_audio
boolean | null

Whether the generated video includes synchronized audio (video generations only). Echoes the request generate_audio (default true).

Example:

true

video_task
enum<string> | null

Resolved reference-video task (video generations only). reference: the generation borrows the clip's motion for a new scene. extend: the generation continues the clip itself. null on every generation that did not supply reference_video_url (including all pre-existing rows). Additional task types may be introduced in future API versions — treat an unrecognized value as opaque.

Available options:
reference,
extend
Example:

"reference"

extend_direction
enum<string> | null

Resolved extend direction (video generations only). Non-null only when video_task is extend; null otherwise (including every reference-mode and non-reference generation).

Available options:
forward,
backward
Example:

"forward"

reference_video_url
string | null

The reference video URL you submitted, echoed back VERBATIM as you sent it — never re-signed, never a storage path. null for generations that did not supply reference_video_url (including every dashboard-originated row, and including an inference_type: "r2v" generation whose motion reference the platform built from a first frame plus a pinned or matched video model — this field, not inference_type, is what distinguishes your own clip from a platform-built reference).

Example:

"https://cdn.example.com/clips/dance-loop.mp4"

reference_audio_url
string | null

The reference audio URL you submitted, echoed back VERBATIM as you sent it. null for generations that did not supply reference_audio_url.

Example:

"https://cdn.example.com/audio/voiceover.mp3"

video_lora_id
string | null

The video model (an id or slug from GET /v1/video_loras) behind this generation (video generations only) — either the video_lora_id you pinned on the request, or the model the platform auto-matched to your prompt when none was pinned. null for a plain (no video model) generation, and always null on image generations.

Example:

"lora_01HXMQ7Z3K8Y2VNABCDEFGHJKM"

video_lora_name
string | null

Display name of the video model in video_lora_id, if any. Omitted for plain generations (no pinned or matched video model).

Example:

"Cinematic Pan"

character_id
string | null

Character ID supplied on the request (char_<ulid> or legacy UUID), echoed back. null when no character was attached to this generation.

Example:

"char_01HXMQ7Z3K8Y2VNABCDEFGHJKM"

subjects
object[] | null

Ordered subjects composed into this generation (positional — entry N echoes entry N of the identity inputs sent on create; legacy character_id / reference_image_urls inputs are normalized into the same projection). Character entries carry the opaque char_<ulid>; reference entries echo only their image count, never URLs. Video generations echo their cast the same way (up to 2 entries), matching the subjects[] array accepted by POST /v1/videos. Character entries on generations created before cast snapshotting may carry character_id: null; the id is never substituted with an internal identifier. null for generations that did not compose subjects (plain prompt-only or style generations, and context-image generations). The engine is not exposed.

Example:
context_images
object | null

Count-only echo of the context_images[] sent on create — the number of loose reference images supplied, never the image URLs themselves. null for every generation that did not use context_images (plain prompt-only, style, subject, and video generations).

Example:
action
object | null

The composition act applied to this generation, echoed as { id, name }. An act is applied either because you pinned it with action_id or because act detection matched one automatically. id is the act identifier from GET /v1/actions; name is its display name (or null when the name was not recorded). null for every generation that did not apply an act — plain prompt-only, style, context-image, and video generations. A new, always-present, nullable key: existing integrations that do not read it are unaffected.

Example:
style
object | null

The prompt style applied to this generation, echoed as { id, name }. A style is applied either because you pinned it with lora_id or because style matching resolved one from your prompt — the echo always reflects the style that actually ran (a retired id that aliases to a successor echoes the successor). id is the lora_* identifier from GET /v1/loras and round-trips into lora_id. null for every generation without a style — including lora_id: null requests, retired styles that generate plain, video generations, and all rows minted before styles shipped. A new, always-present, nullable key: existing integrations that do not read it are unaffected.

Example:
warnings
object[]

Non-fatal request adjustments, present ONLY on the POST /v1/images 201 body (and the estimate response) — never on GET reads, list rows, or webhook payloads. Omitted entirely when empty. Current codes: parameter_ignored, style_retired_plain; new codes may be added without a version bump — ignore unknown codes. Idempotent replays return the original warnings verbatim.

Example:
enhancer_outcome
string | null

Outcome of the prompt-optimization step, echoed for observability. Current values: no_request (the request did not invoke it), succeeded, identical, refused, transport_error, length_overflow; null on rows minted before this field existed. INFORMATIONAL and deliberately an OPEN set (no schema enum — codegen clients must not mint a closed union): do not branch control flow on it; new values may be added without a version bump.

Example:

"succeeded"

loras
object[] | null

LoRAs applied to this generation. null for prompt-only and pure-reference generations.

aurous_version
string

API contract version applied at the time this row was minted (D25 — frozen for replay across future version bumps).

Example:

"2026-07-16"

creation_request_id
string

Aurous-Request-Id of the POST that created this row. Quote in support tickets to trace the original create request.

Example:

"req_01HXMQ7Z3K8Y2VNABCDEFGHJKM"

completed_at
string | null

Terminal-status timestamp (ISO 8601). NULL until the generation reaches a terminal state.

Example:

"2026-05-04T10:00:14Z"