Skip to main content
POST
Create a new video generation
POST /v1/videos submits a video generation request. Credits are deducted based on duration × resolution. There are three ways to drive the generation: text-to-video (prompt only), image-to-video (a first frame, or first + last frame interpolation), and subject-driven (a saved character or up to 6 reference images, anchored across the video). A first frame can additionally carry a video model — pinned via video_lora_id, or auto-matched when you omit it. The model supplies the motion, your image supplies the person and the scene. See Start from a still with a video model.
Want the byte-exact Seedance video API instead — SDK-pointable, forwarded verbatim? See the raw Seedance access surface. This POST /v1/videos endpoint is the managed path (characters, reference handling, first/last-frame helpers); the raw API is a direct passthrough.
Polling video status: video generations and image generations share the status endpoint. After POST /v1/videos, poll GET /v1/images/{id} — the ID returned for videos is vid_<ulid> but the polling path is /v1/images/{id}. This unification is part of the v1.0 contract.

Text-to-video

Send prompt alone. This is the default mode — no frame, no character, no reference images.

Image-to-video

Pass first_frame_url to animate from a single starting frame, or both first_frame_url and last_frame_url to interpolate between two frames. Each accepts an opaque file_<ulid> (from POST /v1/files) or an https:// URL (server-side fetched through an SSRF-pinned client, max 2048 chars). last_frame_url requires first_frame_url — sending it alone returns 400 parameter_invalid_combination.

Start from a still with a video model

Send first_frame_url together with a video model — pinned via video_lora_id, or auto-matched by omitting it — and the generation runs in reference mode: the model supplies the motion while your image supplies the person, the scene, and the orientation. Your video starts near your image — same person, same scene, same orientation — but frame 0 is recomposed, not copied. It is not a pixel-exact continuation of the still you sent. If you need the output to open on your exact frame, don’t attach a video model: send first_frame_url with video_lora_id: null, which is guaranteed plain image-to-video.
video_lora_id: null costs more — usually a lot more. It is a fidelity choice with a price attached, not a free one. A video model supplies default_duration, default_resolution and default_ratio, and those defaults are what get priced. Opt out and there is nothing to supply them, so the request falls back to adaptive duration (billed against the 15-second ceiling) at 1080p with no ratio pinned — the largest frame we might deliver.Same request, same still, only the tri-state differing:That is ~2.8×, and it is duration, resolution and ratio compounding, not duration alone. If you want the exact-frame guarantee at a predictable price, send video_lora_id: null with an explicit duration and resolution rather than letting both fall back.
cURL
inference_type: "r2v" on this path does not mean you supplied a reference video. reference_video_url stays null — the motion reference is built by the platform from the video model you pinned or matched. Read r2v as “reference mode”, and branch on reference_video_url if you need to tell the two apart.
A last frame cannot combine with a video model: first_frame_url + last_frame_url + video_lora_id returns 400 parameter_invalid_combination with param: "last_frame_url". Use a first frame alone with the model, or drop the model for first + last interpolation — a first + last pair never auto-matches, so interpolation always runs as plain image-to-video. A subject (character_id, reference_image_urls, subjects[]) still cannot combine with any frame — on this path the still itself is what anchors identity.

Minimum frame size

A first frame you attach to a video model becomes a short seed clip, and we scale it up to get there — capped, so a frame below the floor cannot be animated at all. It must be at least 64 pixels on its shortest side and at least 6,400 pixels in total: 80×80 is the smallest square accepted, and at a 64-pixel short side the long side needs to be at least 100. Larger is better. When you pass a file_<ulid> (from POST /v1/files), we already know its dimensions, so an undersized frame is rejected up front with 400 first_frame_too_small and nothing is held. When you pass an https:// URL the dimensions aren’t known until the bytes are fetched, so the same rule is enforced during preparation: the generation is accepted, transitions to failed with error_code: "first_frame_too_small", and the full hold is refunded. Either way the fix is the same — send a larger frame. The floor applies only when a video model is in play (pinned, or left to auto-match). It fires on the possibility of a model, not on one actually being used — an un-pinned frame is rejected even if no model would have matched your image. With video_lora_id: null no seed clip is built, so no upscaling happens and the rule doesn’t apply. Nor does it apply when a last_frame_url rides alongside — a first + last pair never auto-matches and builds no seed clip either, so no first_frame_too_small ever fires on an interpolation request.

Subject-driven video

Anchor the video’s subject across every frame with either:
  • character_id — a saved character (see Create a character). Must be status: ready — otherwise 400 character_not_ready.
  • reference_image_urls — up to 6 ad-hoc reference images, each a file_<ulid> or an https:// URL. Same shape as image generation’s reference_image_urls.
A subject-driven request still requires a prompt describing the action — the subject supplies identity, not motion; it can’t substitute for a prompt on its own. A subject-driven generation reports inference_type: "i2v" on the response.
character_id and reference_image_urls are mutually exclusive — sending both returns 400 mutually_exclusive_input. A subject (either form) combined with first_frame_url / last_frame_url is also rejected, with 400 parameter_invalid_combination. Pick one of the three modes per request.
cURL

Choosing a video model

GET /v1/video_loras lists the video models available to pin via video_lora_id. The field is tri-state:
  • Pin a model by passing its catalog id or slug as video_lora_id. A pinned model supplies three generation defaults, not just length: default_duration, default_resolution and default_ratio. Anything you send explicitly (duration, resolution, ratio) wins over the model’s default; anything you omit is filled from the model. All three feed the price — see Duration and cost.
  • Omit video_lora_id and the platform picks a suitable model for you. With no first frame, matching reads your prompt; when a first frame is attached, matching is based on what the image shows. If nothing fits, your request generates as plain video — with a first frame, as plain image-to-video; with a subject attached, as a subject-driven video. Matching never fails a request: there is no “no model matched” error.
  • Send video_lora_id: null to opt out of image-based matching. This is meaningful only when first_frame_url is attached, where it guarantees plain image-to-video. Without a first frame, null behaves exactly like omitting the field — your prompt is still auto-matched. There is no way to disable prompt-based matching; pin a model if you need a specific one. null also opts out of the three defaults above, so an otherwise-bare request falls back to adaptive duration at 1080p — measured at ~2.8× a pinned generation. Send duration and resolution alongside it, or see the cost warning.
Treat catalog ids as opaque — don’t assume a fixed prefix. Whichever model ends up behind the generation, pinned or auto-matched, is echoed back on the response: video_lora_id (null for a plain generation) and video_lora_name (omitted for a plain generation).
Estimates don’t run matching, and the quote can miss in either direction. POST /v1/videos/estimate prices exactly the request you sent. If you omit video_lora_id and a model then matches at create, the generation is priced as a video-input generation — and which way the charge moves depends on what you pinned:
  • You left duration and resolution to us — the usual shape for a bare first-frame request. The estimate has no model, so it prices adaptive duration (the 15-second ceiling) at 1080p; the matched model then supplies a shorter, smaller default. The quote is an upper bound that can substantially over-state. Measured on the first-frame path: amount 827.26, actual settled charge 279.91.
  • You pinned duration and resolution yourself. Your values win over the model’s defaults, so a match can only add the motion reference to the price. Here the charge genuinely exceeds the quote.
On an adaptive request, plan against amount_minamount_max, not the amount headline — that band is much wider than the headline and it contained the real charge in the measurement above. Pin video_lora_id whenever you need the estimate to be exact — that is the only value that removes matching from the create path entirely. video_lora_id: null suppresses matching only when a first frame is attached, so it does not make a prompt-only quote exact. This has always been true of auto-matching; a first frame just gives matching one more thing to read.

Duration and cost

Video cost is base_per_second × resolution_factor × duration. POST /v1/videos/estimate prices the same request body with no side effects — no character resolution, no reference materialization, no credit hold.

Adaptive duration (the default)

Pass duration: -1, or omit it when no video model is pinned or matched — the model chooses the natural length for your prompt, a whole number of seconds between 4 and 15. You’re billed for the delivered length, not the ceiling: the 15s ceiling is reserved from your balance up front, and the difference is released once the actual length is known. A team with less than the ceiling available gets 402 balance_too_low even if the eventual (shorter) length would have fit. POST /v1/videos/estimate reflects this as a range:
amount equals amount_max — the ceiling actually held. The real charge, once the generation settles, lands between amount_min and amount_max.

Fixed duration

Pass a whole number of seconds, 4–15, to pin an exact length at an exact price:
If you omit duration while pinning (or auto-matching) a video model that defines its own default_duration, that model’s default length applies — a fixed price, not adaptive. Omitting duration with no model pinned or matched falls back to adaptive. The same substitution happens for resolution and ratio: a pinned or matched model supplies default_resolution and default_ratio for whichever of the two you left out, and both are priced. This is why the same still can settle at very different prices depending only on video_lora_id — a model that defaults to 5s / 720p prices roughly a third of the adaptive-15s / 1080p fallback you get with no model at all. Send duration, resolution and ratio explicitly whenever you want the price to be a property of your request rather than of whichever model ends up behind it.

Audio

generate_audio defaults to true and is echoed back on the response as video_generate_audio. Audio can be declined by a content check independently of the video itself. If a generation fails for this reason, retry with generate_audio: false to skip audio generation, or adjust the prompt.

Output

When the generation reaches status: succeeded, fetch the rendered file via GET /v1/videos/{id}/output. The output URL is valid for ~24 hours; save what you want to keep. A failed generation never produces output and settles at cost: {"amount": 0, "refunded": true} — the reserved hold is released, never charged.

Webhooks

Register a webhook endpoint subscribed to video.completed / video.failed / video.cancelled (POST /v1/webhook_endpoints) to receive a POST callback when the video reaches a terminal state. See Webhooks for signature verification.

Errors

See Errors for the complete catalog, retry policy, and the request_id support workflow.

Authorizations

X-Api-Key
string
header
required

Your team API key (starts with al_live_).

Headers

Idempotency-Key
string

Optional key to safely retry a request. Replaying the same key within 24h returns the original response with Aurous-Idempotent-Replayed: true and never creates a second generation; reusing a key with a different body returns 409 idempotency_key_in_use. A request that failed validation is not cached — retry it with the same key. Same contract as POST /v1/images.

Aurous-Version
string

Optional API version pin (YYYY-MM-DD). Defaults to your team's pinned version, or the system default 2026-07-16 for unauthenticated requests.

Pattern: ^\d{4}-\d{2}-\d{2}$
Example:

"2026-07-16"

Body

application/json
prompt
string

Text prompt describing the video. Required unless first_frame_url is provided (image-to-video). Subject-driven generation (character_id or reference_image_urls) still requires a prompt describing the action — a subject alone cannot substitute for one.

Example:

"A golden sunset over the ocean with gentle waves"

video_lora_id
string | null

Optional, and TRI-STATE. An id (or slug, where one is set) from GET /v1/video_loras pins a specific video model. If OMITTED, the platform picks a suitable model for you: matching reads your prompt, or — when first_frame_url is attached — what the IMAGE shows. If none matches, your request is generated as plain video, plain image-to-video, or (with a subject attached) a subject-driven video; matching never fails the request. Send NULL to opt out of IMAGE-based matching — meaningful only when first_frame_url is attached, where it guarantees plain image-to-video. WITHOUT a first frame, null behaves exactly like omitting the field and your prompt is still auto-matched; pin an id if you need a specific model. A first frame plus a pinned or matched model reports inference_type: "r2v" and starts NEAR your image rather than exactly on it. PRICING: a pinned or matched model also supplies default_resolution and default_ratio alongside default_duration for whatever you left out (anything you send explicitly wins), and all three are priced — so null on an otherwise-bare first-frame request falls back to adaptive duration at 1080p and costs ~2.8x the same still with a model pinned. Send duration and resolution alongside null if you want the exact-frame guarantee at a predictable price.

Example:

"action_01HXMQ7Z3K8Y2VNABCDEFGHJKM"

character_id
string
deprecated

Drive the video from a saved character (reference-to-video). Must be status: ready — otherwise 400 character_not_ready. Cross-team ids 404 (existence never leaked). Mutually exclusive with reference_image_urls (→ 400 mutually_exclusive_input) and with first_frame_url/last_frame_url (→ 400 parameter_invalid_combination). Omit both a subject and a frame for plain text-to-video. Superseded by subjects[] (which also supports multi-person casts); fully supported — successful responses carry an advisory Deprecation: true header when this field is used.

Example:

"char_01HXMQ7Z3K8Y2VNABCDEFGHJKM"

reference_image_urls
string[]
deprecated

Up to 6 subject reference images — each a file_<ulid> (POST /v1/files) or an https:// URL (SSRF-pinned server-side fetch). Image files only: a file_<ulid> uploaded with purpose reference_video/reference_audio is rejected (400 invalid_format). The subject is anchored across the video. Mutually exclusive with character_id and with first_frame_url/last_frame_url. Superseded by subjects[] (which also supports multi-person casts); fully supported — successful responses carry an advisory Deprecation: true header when this field is used.

Maximum array length: 6
Example:
subjects
object[]

Ordered cast for a multi-person video. Up to 2 subjects per generation (current engine limit — may increase; exceeding it returns 400 parameter_invalid_combination). Each subject is a saved character or its own group of 1–4 reference images; character and reference subjects can be mixed. Supersedes the single-subject character_id/reference_image_urls fields — sending both shapes returns 400 mutually_exclusive_input. A cast may accompany reference_video_url (the clip supplies motion, the cast supplies identity) but not video_task: "extend" and not first_frame_url/last_frame_url. Duplicate characters or a reference image reused across subjects return 400. Empty array is treated as omitted.

Maximum array length: 2
first_frame_url
string

First frame for image-to-video / first+last frame interpolation. Either an opaque file_<ulid> ID returned by POST /v1/files, or an https:// URL pointing at a public host (max 2048 chars). Image files only — a file_<ulid> uploaded with purpose reference_video/reference_audio is rejected (400 invalid_format). URLs are server-side fetched through an SSRF-pinned client (rejects private IPs / cloud metadata). MINIMUM SIZE — when a video model is pinned or left to auto-match, this frame is upscaled into a short seed clip, so it must be at least 64 px on its shortest side AND at least 6,400 px in total (80x80 is the smallest square accepted; at a 64 px short side the long side needs 100). A file_<ulid> below that returns 400 first_frame_too_small before any credit hold (dimensions were measured at upload); an https:// URL is accepted and settles at status: failed with error_code: "first_frame_too_small" and a full refund, because validation never fetches the bytes. The check runs on the POSSIBILITY of a model, not on one actually being used — an un-pinned frame is rejected even if no model would have matched your image. No minimum applies with video_lora_id: null, nor when a last_frame_url rides alongside — a first+last pair never auto-matches and builds no seed clip either, so no first_frame_too_small ever fires on an interpolation request.

Maximum string length: 2048
Example:

"file_01HXMQ7Z3K8Y2NABCDEFGHJKMN"

last_frame_url
string

Last frame for first+last frame interpolation. Same rules as first_frame_url. Requires first_frame_url — sending last_frame_url alone returns 400 with code parameter_invalid_combination.

Maximum string length: 2048
Example:

"https://example.com/last.jpg"

resolution
enum<string>

Output video resolution

Available options:
480p,
720p,
1080p
ratio
enum<string>

Output video aspect ratio

Available options:
16:9,
4:3,
1:1,
3:4,
9:16,
21:9,
adaptive
duration
enum<number>
default:-1

Video length in seconds. Pass -1 for ADAPTIVE (default): the model selects the optimal whole-second length within its supported range (currently 4–15 s) and you are billed for the delivered length — a VARIABLE price. Call POST /v1/videos/estimate to see the min–max range before submitting. Pass a whole number 4–15 to pin a fixed length (fixed price). Omitting uses the video model's default length when your request pins or is matched to one; otherwise adaptive (-1). Adaptive holds the 15 s ceiling from your balance until it settles.

Available options:
-1,
4,
5,
6,
7,
8,
9,
10,
11,
12,
13,
14,
15
Example:

5

generate_audio
boolean

Whether the generated video includes synchronized audio. Defaults to the selected video model's audio setting (currently true for all models unless noted on the model).

watermark
boolean

Whether to add watermark

enhance_prompt
boolean

Whether to enhance the prompt with AI

reference_video_url
string

Your reference video — the motion source for a reference-mode generation, or the clip to continue when video_task is extend. Either an opaque file_<ulid> returned by POST /v1/files (uploaded with purpose: reference_video — format/duration/resolution were already validated at upload), or an HTTPS URL. Requirements: MP4 or MOV (H.264/H.265), 2–15 seconds, up to 50 MB, frame area between 409,600 px (≈640×640) and 2,073,600 px (1920×1080, ≈1080p). URLs must be publicly fetchable or pre-signed, serve HTTPS directly (redirects are refused — pass the final URL), and respond within the fetch budget. Cannot be combined with first_frame_url/last_frame_url. While the reference is prepared the generation stays in processing — see the endpoint description for timing.

Maximum string length: 2048
Example:

"https://cdn.example.com/clips/dance-loop.mp4"

reference_audio_url
string

An audio reference (voice, music, or ambience) to guide the generated soundtrack — either a file_<ulid> from POST /v1/files (purpose: reference_audio) or an HTTPS URL. WAV or MP3, 2–15 seconds, up to 15 MB. Needs a companion — a reference video, a subject (subjects[] or the legacy fields), or a pinned video_lora_id — and requires audio output (generate_audio must not be false). Same URL rules as reference_video_url (public or signed HTTPS, no redirects).

Maximum string length: 2048
Example:

"https://cdn.example.com/audio/voiceover.mp3"

video_task
enum<string>
default:reference

What to do with reference_video_url. reference (default): generate a new scene that borrows the clip's motion. extend: continue the clip itself in extend_direction; extend takes the clip as-is, so it cannot be combined with video_lora_id. Only valid when reference_video_url is present.

Available options:
reference,
extend
extend_direction
enum<string>
default:forward

Direction to continue the clip when video_task is extend. Defaults to forward. Only valid with video_task: "extend".

Available options:
forward,
backward

Response

Video generation created

object
enum<string>
default:inference
required

Discriminator — always inference. Mirrors OpenAI's object-field convention so SDK clients can branch on the resource type without inspecting the ID prefix. A single canonical value (inference) covers both image and video generations; use media_type to distinguish the rendering kind.

Available options:
inference
Example:

"inference"

id
string
required

Opaque generation ID

Example:

"img_01HXMQ7Z3K8Y2VNABCDEFGHJKM"

status
enum<string>
required

Current generation status. Lifecycle: pending (created, awaiting dispatch) → processing (running) → one of the terminal values succeeded / failed / cancelled. Additional terminal values may be introduced in future API versions and will be announced via the changelog before they appear on the wire.

Available options:
pending,
processing,
succeeded,
failed,
cancelled
Example:

"succeeded"

prompt
string
required

The text prompt used for generation

Example:

"A golden sunset over mountains, cinematic lighting"

created_at
string
required

Creation timestamp (ISO 8601)

Example:

"2026-05-04T10:00:00Z"

media_type
enum<string> | null

Distinguishes image vs video generation. May be null for older rows minted before this column existed.

Available options:
image,
video
Example:

"image"

output_urls
string[] | null

Generated image proxy URLs. Each URL is anonymous-read (no auth header required) and edge-cached for 24 hours. Available for ~24 hours after generation. Save what you want to keep — long-term storage is intentionally not part of the platform. URLs return 410 Gone after expiry.

Example:
video_url
string | null

Generated video proxy URL (only present on media_type: video). Same 24h TTL as image output_urls.

Example:

"https://api.aurous-labs.com/v1/videos/vid_01HXMQ7Z3K8Y2VNABCDEFGHJKM/output?token=..."

reference_image_urls
string[] | null

The reference image URLs you supplied as visual anchors for this generation, echoed back (snapshotted at inference time). Present only when the generation was driven by your own reference images (reference_image_urls or reference subjects). Omitted entirely for character-driven generations — a character's reference images are managed platform assets and are never echoed.

Example:
error_message
string | null

Human-readable error message if the generation failed. Non-contractual prose — do not parse or match on this value. Switch on error_code instead.

Example:

"Content policy violation"

error_code
string | null

Machine-readable failure reason when status is failed, for the cases where you need to branch in code. null on every successful generation, on failures recorded before this field existed, and on failure paths that have no stable code — error_message is human-facing copy that is re-tuned over time, so never pattern-match it. Currently emitted: generation_interrupted, reference_preparation_failed, content_filtered, generation_failed, and first_frame_too_small (an https:// first_frame_url attached to a video model whose dimensions were below the minimum — a file_<ulid> frame is rejected with the same code as a 400 before any credits are held; see Errors). The set is additive: new codes may appear, so treat an unrecognized value as a generic failure, keep a default branch in your switch statement, and fall back to error_message.

Example:

"generation_interrupted"

duration_ms
integer | null

Processing duration in milliseconds (set on terminal status)

Example:

14820

cost
object | null

Per-generation cost breakdown — same shape as the estimated_cost returned by POST /v1/{images,videos}/estimate. May be null for older rows from before this field existed; populated for all new generations. The amount reflects the committed charge for terminal-status rows — 0 (with refunded: true) on a failed generation, since the reserved hold was released, never charged.

Example:
width
number

Resolved output image width in pixels (image generations only). Reflects the post-snap dimension actually generated; may differ from a custom-requested width by up to 31 px due to multiple-of-32 snapping.

Example:

2048

height
number

Resolved output image height in pixels (image generations only). Reflects the post-snap dimension actually generated; may differ from a custom-requested height by up to 31 px due to multiple-of-32 snapping.

Example:

2048

image_count
number

Number of images in the batch. Reflects the requested count while the generation is running; on a terminal status it reflects the number actually DELIVERED — when part of a batch fails you receive the successful images, this count re-stamps to match, and the difference is refunded automatically.

Example:

1

size_preset
enum<string> | null

Named size preset applied to this generation. null when the request used custom width/height instead of a preset.

Available options:
2k_1_1,
2k_3_2,
2k_2_3,
2k_4_3,
2k_3_4,
2k_16_9,
2k_9_16,
2k_21_9,
4k_1_1,
4k_3_2,
4k_2_3,
4k_4_3,
4k_3_4,
4k_16_9,
4k_9_16,
4k_21_9
Example:

"2k_1_1"

inference_type
enum<string> | null

Inference mode dispatched. Images: t2i (text-to-image) — reference images and characters are supplementary inputs to the t2i flow, not a separate mode. Videos (this list also returns vid_* rows): t2v (text-to-video), i2v (image-to-video — frame-driven, subject-driven, or a pinned video model on its own), or r2v (reference mode — the generation was driven by a motion reference). r2v does NOT imply you supplied that reference: it covers BOTH your own clip sent as reference_video_url AND a platform-built reference, produced when a first frame rides with a video model (pinned via video_lora_id, or auto-matched from the image). On the platform-built variant reference_video_url is null — branch on that field, not on inference_type, to tell the two apart. i2i is reserved for a future image-edit endpoint and is not currently emitted.

Available options:
t2i,
t2v,
i2v,
r2v,
i2i
Example:

"t2i"

cfg_rescale
number | null

CFG rescale factor the customer supplied on the request body, echoed back here. Range 0.0-1.0. Omitted when the customer did not supply a per-request value (the platform applied a precedence-chain default — LoRA, character override, or the global 0.7 — which is not exposed on the response).

Required range: 0 <= x <= 1
Example:

0.7

denoise_strength
number | null

Denoising strength the customer supplied on the request body, echoed back here. Range 0.0-1.0. Omitted when the customer did not supply a value or when the generation was a bare text-to-image request (denoise is only applied when reference images or a character are attached).

Required range: 0 <= x <= 1
Example:

0.6

seed
integer | null

The random seed the model actually used for this image generation. Populated even when you omit seed on the request — the platform requests a random seed and records the concrete value the provider rolled, so you can reproduce the result by passing it back as seed. Available once status is succeeded; null before then and for failed/cancelled generations. For multi-image batches (image_count > 1) this is the seed of the first image (output_urls[0]); per-image seeds are not yet exposed. Image generations only.

Example:

819572108

video_duration
integer | null

Video duration in seconds (video generations only)

Example:

5

video_resolution
enum<string> | null

Video resolution (video generations only)

Available options:
480p,
720p,
1080p
Example:

"480p"

video_ratio
enum<string> | null

Video aspect ratio (video generations only)

Available options:
16:9,
4:3,
1:1,
3:4,
9:16,
21:9,
adaptive
Example:

"16:9"

video_generate_audio
boolean | null

Whether the generated video includes synchronized audio (video generations only). Echoes the request generate_audio (default true).

Example:

true

video_task
enum<string> | null

Resolved reference-video task (video generations only). reference: the generation borrows the clip's motion for a new scene. extend: the generation continues the clip itself. null on every generation that did not supply reference_video_url (including all pre-existing rows). Additional task types may be introduced in future API versions — treat an unrecognized value as opaque.

Available options:
reference,
extend
Example:

"reference"

extend_direction
enum<string> | null

Resolved extend direction (video generations only). Non-null only when video_task is extend; null otherwise (including every reference-mode and non-reference generation).

Available options:
forward,
backward
Example:

"forward"

reference_video_url
string | null

The reference video URL you submitted, echoed back VERBATIM as you sent it — never re-signed, never a storage path. null for generations that did not supply reference_video_url (including every dashboard-originated row, and including an inference_type: "r2v" generation whose motion reference the platform built from a first frame plus a pinned or matched video model — this field, not inference_type, is what distinguishes your own clip from a platform-built reference).

Example:

"https://cdn.example.com/clips/dance-loop.mp4"

reference_audio_url
string | null

The reference audio URL you submitted, echoed back VERBATIM as you sent it. null for generations that did not supply reference_audio_url.

Example:

"https://cdn.example.com/audio/voiceover.mp3"

video_lora_id
string | null

The video model (an id or slug from GET /v1/video_loras) behind this generation (video generations only) — either the video_lora_id you pinned on the request, or the model the platform auto-matched to your prompt when none was pinned. null for a plain (no video model) generation, and always null on image generations.

Example:

"lora_01HXMQ7Z3K8Y2VNABCDEFGHJKM"

video_lora_name
string | null

Display name of the video model in video_lora_id, if any. Omitted for plain generations (no pinned or matched video model).

Example:

"Cinematic Pan"

character_id
string | null

Character ID supplied on the request (char_<ulid> or legacy UUID), echoed back. null when no character was attached to this generation.

Example:

"char_01HXMQ7Z3K8Y2VNABCDEFGHJKM"

subjects
object[] | null

Ordered subjects composed into this generation (positional — entry N echoes entry N of the identity inputs sent on create; legacy character_id / reference_image_urls inputs are normalized into the same projection). Character entries carry the opaque char_<ulid>; reference entries echo only their image count, never URLs. Video generations echo their cast the same way (up to 2 entries), matching the subjects[] array accepted by POST /v1/videos. Character entries on generations created before cast snapshotting may carry character_id: null; the id is never substituted with an internal identifier. null for generations that did not compose subjects (plain prompt-only or style generations, and context-image generations). The engine is not exposed.

Example:
context_images
object | null

Count-only echo of the context_images[] sent on create — the number of loose reference images supplied, never the image URLs themselves. null for every generation that did not use context_images (plain prompt-only, style, subject, and video generations).

Example:
action
object | null

The composition act applied to this generation, echoed as { id, name }. An act is applied either because you pinned it with action_id or because act detection matched one automatically. id is the act identifier from GET /v1/actions; name is its display name (or null when the name was not recorded). null for every generation that did not apply an act — plain prompt-only, style, context-image, and video generations. A new, always-present, nullable key: existing integrations that do not read it are unaffected.

Example:
style
object | null

The prompt style applied to this generation, echoed as { id, name }. A style is applied either because you pinned it with lora_id or because style matching resolved one from your prompt — the echo always reflects the style that actually ran (a retired id that aliases to a successor echoes the successor). id is the lora_* identifier from GET /v1/loras and round-trips into lora_id. null for every generation without a style — including lora_id: null requests, retired styles that generate plain, video generations, and all rows minted before styles shipped. A new, always-present, nullable key: existing integrations that do not read it are unaffected.

Example:
warnings
object[]

Non-fatal request adjustments, present ONLY on the POST /v1/images 201 body (and the estimate response) — never on GET reads, list rows, or webhook payloads. Omitted entirely when empty. Current codes: parameter_ignored, style_retired_plain; new codes may be added without a version bump — ignore unknown codes. Idempotent replays return the original warnings verbatim.

Example:
enhancer_outcome
string | null

Outcome of the prompt-optimization step, echoed for observability. Current values: no_request (the request did not invoke it), succeeded, identical, refused, transport_error, length_overflow; null on rows minted before this field existed. INFORMATIONAL and deliberately an OPEN set (no schema enum — codegen clients must not mint a closed union): do not branch control flow on it; new values may be added without a version bump.

Example:

"succeeded"

loras
object[] | null

LoRAs applied to this generation. null for prompt-only and pure-reference generations.

aurous_version
string

API contract version applied at the time this row was minted (D25 — frozen for replay across future version bumps).

Example:

"2026-07-16"

creation_request_id
string

Aurous-Request-Id of the POST that created this row. Quote in support tickets to trace the original create request.

Example:

"req_01HXMQ7Z3K8Y2VNABCDEFGHJKM"

completed_at
string | null

Terminal-status timestamp (ISO 8601). NULL until the generation reaches a terminal state.

Example:

"2026-05-04T10:00:14Z"