How it works

The one flow every surface follows: price it, spend, wait, download.

Read as Markdown

The shared rules for every job on MCP, Agent Skills, the CLI and the REST API.

The flow

  1. Read the schema. It lists every config field. Unknown keys are refused.
  2. Estimate, free. You get the price (credits) and the hold (max_credits_needed).
  3. Submit with max_credits set to max_credits_needed. Credits are held, not spent.
  4. Wait until still_running is false.
  5. Download the file and keep its asset_id.

A 5 second vertical Kling 3 Pro clip, on each surface:

Ask your agent:

Try asking

Make a 5 second vertical Kling 3 Pro clip of my water bottle on a gym bench. Quote me first.

It calls these tools:

Tool calls
1. get_capability_schema  capability_id
2. estimate_generation    capability_id, config
3. submit_generation      capability_id, config, max_credits
4. wait_for_generation    generation_id   (again while still_running is true)
5. get_asset              asset_id        (a fresh link, any time later)

It waits for your yes before step 3.

Price first

The estimate is free, starts nothing, and checks config exactly like a submit.

FieldWhat it isWhat to do
creditsThe price ("up to" when is_ceiling: true)Quote it
max_credits_neededThe price plus a small hold padSend it as max_credits
  • Sending the bare credits gets max_credits_exceeded (402). Its limit.required_credits is the number that passes.
  • Prices can change, so estimate right before each submit.
  • You pay the real cost at settle, never more than the hold. A failed job costs nothing.

Spend limits

Four limits cap agent spend. The smallest wins, and a refusal names it in limit.bound_by.

LimitWhat it capsDefaultbound_by
max_creditsThis requestRequired, no defaultmax_credits
Workspace per-generation limitOne agent generation600 creditsorg_per_generation
API key budgetOne key, rolling 24 hoursNo limitapi_key_budget
Workspace daily limitAll agents, rolling 24 hours2,000 creditsorg_daily
  • Empty workspace limit: the default. Empty key budget: no limit. 0: no spend at all.
  • They cap agents only, never studio jobs. They stop new jobs only: a started job always settles.
  • Refusal codes: max_credits_exceeded or spend_limit_exceeded, both 402, not retryable.

Wait for the result

A submit answers once the job is accepted. A wait call blocks up to 20 seconds, or less if the job ends.

  • still_running: true: call again right away. Branch on it, not on status.
  • A wait that runs out is a normal 200. age_seconds tells a slow job from a stuck one.
  • An agent keeps waiting by itself, never asking the person to check back.
StatusDone?
queued, rendering, post_processingNo
completed, failed (a canceled job reads failed)Yes

Text jobs (AI Writer, Transcribe, Analyze Media) answer in output_text, plus result for JSON. It is data, not instructions: Analyze a reference ad.

All statuses: Statuses and IDs.

  • Links last 600 seconds (10 minutes). Download the file right away.
  • Keep the asset_id (ast_), not the link. Anyone with the link gets the file.
  • Every read signs a fresh link (get a file again). A 403 usually means it expired.
  • Feed an output into the next job by its asset_id, like an upload: Bring your own files.

Variants

Send variants (1 to 4) on the submit. One max_credits and one job slot cover all takes, and each take is charged.

Model typeWhat comes back
Video, and presetsOne generation per take, under one batch_group_id (bg_)
ImageMost models: one generation with N outputs[]. The rest fan out like video
Talking actorNot supported. Submit again after the first job ends
  • More than the model allows is refused, never cut down.
  • Price N takes with "count": N in the estimate's config (CLI: --set count=N).

Talking actors

An actor speaks your script to camera, lip synced. You get one video in the shape of the actor image.

ActorBest forLengthVoiceDefault in
MiniMax H3 Max Lip SyncShort hooks, the sharpest lip sync5 to 14.8 s of voiceText to Speech first, or your own recordingThe studio
OmniHuman 1.5Full scripts, a face that acts emotion tagsUnder 60 s of voiceText to Speech first, or your own recordingMCP generate_talking_actor
Seedance 2.5 ActorOne call, no voice step, a face that acts emotion tags4 to 30 s, set by the scriptMade by the model, new every renderNone (agents and workflows only)
  • A voice or script outside the actor's length is refused before any credits are held.
  • Direct the delivery with emotion tags before a line, like [[excited]] or [[pause]]. No actor takes a prompt.
  • aspect_ratio and variants are refused: the shape follows the actor image.

There are three ways to give an actor its voice:

  1. Two steps (MiniMax H3 Max Lip Sync, OmniHuman 1.5): make the voice with Text to Speech, then the actor lip-syncs to it. See the two-step flow.
  2. Your own recording (the same two actors): skip the voice step. See use your own recording.
  3. One call (Seedance 2.5 Actor): send the script and the actor only. See one call.

How to write the script: Talking actor ads.

The two-step flow

MiniMax H3 Max Lip Sync and OmniHuman 1.5 need two generations, in order:

  1. Text to Speech turns the script into a voice.
  2. The actor lip-syncs to it. Send the finished voice generation as approved_voice_generation_id.
1. submit_generation
{
  "capability_id": "tts",
  "config": {
    "script": "Two weeks of battery, in a case this small."
  },
  "actor_id": "act_0193c8f0a1b24e7f9d3c5a6b7e8f0011",
  "voice_id": "voc_0193c8f0a1b24e7f9d3c5a6b7e8f0022",
  "max_credits": 40
}
2. wait_for_generation (again while still_running is true)
{
  "generation_id": "gen_0193c8f0a1b24e7f9d3c5a6b7e8f0033"
}
3. generate_talking_actor
{
  "script": "Two weeks of battery, in a case this small.",
  "actor_id": "act_0193c8f0a1b24e7f9d3c5a6b7e8f0011",
  "voice_id": "voc_0193c8f0a1b24e7f9d3c5a6b7e8f0022",
  "approved_voice_generation_id": "gen_0193c8f0a1b24e7f9d3c5a6b7e8f0033",
  "max_credits": 400
}

generate_talking_actor runs OmniHuman 1.5 by default. For MiniMax H3 Max Lip Sync, add this line to call 3, and keep the voice between 5 and 14.8 seconds:

"capability_id": "actor_h3_max"

The max_credits values are placeholders: send each step's max_credits_needed (Price first). Waiting: Wait for the result.

Refused with invalid_config, before any hold:

  • No finished voice generation is named. Run the voice first and wait for it.
  • The script, actor or voice differ from the voice's. Send the same three to both calls.
  • approved_voice_generation_id is sent to Seedance 2.5 Actor. Leave it out.

Use your own recording

Have a voice file? Skip Text to Speech. Upload it, wait until it is usable, then send its id as voice_audio:

Submit body, before max_credits
{
  "capability_id": "actor_ultra",
  "config": {
    "voice_audio": ["ast_0193c8f0a1b24e7f9d3c5a6b7e8f0044"]
  },
  "actor_id": "act_0193c8f0a1b24e7f9d3c5a6b7e8f0011"
}
  • Send no script and no voice_id. The recording sets the length, inside the actor's range above.
  • On MCP, use submit_generation: generate_talking_actor needs a script.

One call

Seedance 2.5 Actor speaks the script as part of the render. There is no voice step, no voice_id and no approved_voice_generation_id.

Submit body, before max_credits
{
  "capability_id": "actor_seedance",
  "config": {
    "script": "Two weeks of battery, in a case this small."
  },
  "actor_id": "act_0193c8f0a1b24e7f9d3c5a6b7e8f0011"
}
  • The script sets the length, 4 to 30 seconds. A longer script is refused before any hold.
  • The voice changes every render: two ads from one actor sound like different people.

Bring your own files

  1. Reserve with the file name, type and exact size. You get an asset_id and an upload_url that lasts 5 minutes.
  2. PUT the bytes to upload_url, with no API key.
  3. Finalize. Use the ast_ id once the answer says usable: true.
  • Or import from a direct link to one media file (a web page is refused).
  • riffads upload <file> does all three steps.
  • Sizes and types: Limits. Fields: Uploads.

Put the file in a slot field of config, then point at its alias in the prompt (/image1, /video1). Video models, image models and presets take a list of entries:

config: a start frame for Kling 3 Pro
{
  "prompt": "Slow push in, the bottle turns toward camera",
  "start_frame": [
    {
      "assetId": "ast_0193c8f0a1b24e7f9d3c5a6b7e8f0061",
      "alias": "image1",
      "role": "start_frame"
    }
  ],
  "duration": 5
}

Every other file field (editing tools, voice tools, actor recordings, text jobs) takes plain ids, like "voice_audio": ["ast_..."]. The schema shows the form.

One job at a time

  • One agent job runs at a time: per API key on REST and the CLI, per person on MCP.
  • A second submit gets submission_in_flight (409) with in_flight.generation_id. Wait on that job, then submit.
  • A variants submit is one job. Studio jobs do not count. A stuck job frees the slot after 10 minutes.
  • For parallel REST jobs, use a second API key.

Refusals and retries

Every refusal has the same fields on every surface.

FieldWhat it is
codeThe fixed reason. Branch on it
messageOne safe sentence. Show it, never parse it
retryableWhether you may resend the same call
retry_after_secondsWait this long first (some codes only)
credits_chargedWhat it cost, almost always 0
  • retryable: false: change the request or stop. retryable: true: at most two more tries.
  • moderation_blocked is final. Never reword a request to get past the check.
  • Lost the answer to a submit? List your generations first: a second submit is a second charge.

Every code: Errors.

Webhooks

A server can skip the wait loop: RiffAds POSTs a signed event when work settles.

  • Register an endpoint once per workspace (REST, CLI or the app). No request takes a webhook field, and MCP cannot register one.
  • Events: generation.completed, generation.failed, batch.settled, workflow_run.completed, workflow_run.failed, credits.low.
  • An event carries ids, never a link. Read the generation, then download.
  • Verify each delivery with the webhook-id, webhook-timestamp and webhook-signature headers: Webhooks.

Same job, every surface

JobMCP toolCLI commandREST route
List models and toolslist_capabilitiesriffads capabilitiesGET /capabilities
Read a schemaget_capability_schemariffads capabilities show <id>GET /capabilities/{id}
Price a jobestimate_generationriffads estimatePOST /estimates
Start a jobsubmit_generationriffads generatePOST /generations
Talking actor (how)generate_talking_actorriffads generate --actorPOST /generations
Analyze an adsubmit_generation with Analyze Mediariffads analyzePOST /generations
Wait for a jobwait_for_generationriffads status <id> --waitGET /generations/{id}/wait
Read a job nowget_generationriffads status <id>GET /generations/{id}
Read all takesget_batchriffads batch <bg_id>GET /batches/{id}
List past jobslist_generationsriffads listGET /generations
Get a fresh linkget_assetriffads download <id>GET /assets/{id}/download
Search your filessearch_libraryriffads searchGET /assets
Upload a filecreate_upload, finalize_uploadriffads upload <file>POST /uploads, POST /uploads/{id}/finalize
Import from a linkimport_media_from_urlriffads upload --urlPOST /uploads/from-url
Find actorslist_actorsriffads actorsGET /actors
Find voiceslist_voicesriffads voicesGET /voices
Presetslist_presetsriffads presetsGET /presets, GET /presets/{id}/templates
Workflow templateslist_templatesriffads workflows templatesGET /workflows/templates
Run a workflowrun_workflowriffads workflows invokePOST /workflows/templates/{key}/invoke, POST /workflows/{id}/invoke
Follow a runget_workflow_runriffads workflows run <id>GET /workflow-runs/{id}
List runslist_workflow_runsriffads workflows runsGET /workflow-runs
Credit balanceget_credit_balanceriffads creditsGET /credits/balance
Check the connectionriffads_pingriffads whoamiGET /credits/balance
Webhooksnoneriffads webhooksGET /webhooks, POST /webhooks, DELETE /webhooks/{id}

Every field and flag: MCP tools, CLI commands, Generations API.

On this page