# Editing tools (/editing-tools)



Editing tools change a file you already have: an upload or an earlier result. You run them like a model: [check the price first](/how-it-works#price-first), then start the job.

Building with the API? [List capabilities](/api/capabilities#list-capabilities) shows each tool's id next to its name.

Each tool is a capability: send its id as `capability_id`. Ids: `auto_caption`, `text_overlay`, `trim_video`, `stitch`, `split_scenes`, `merge_layers`, `extract_frame`, `extract_audio`, `change_speed`, `change_voice`, `resize`, `upscale`, `remove_background`, `skin_enhance`, `camera_angle`.

## All tools [#all-tools]

| Tool                               | You send                                                                | You get back                                       | Longest input    |
| ---------------------------------- | ----------------------------------------------------------------------- | -------------------------------------------------- | ---------------- |
| Add Captions (`auto_caption`)      | `source_video`                                                          | Captions burned in                                 | 600 s            |
| Text Overlay (`text_overlay`)      | `source_media` (photo or video), `text`                                 | Your text on the file                              | Video 600 s      |
| Trim Video (`trim_video`)        | `source_video`                                                          | One section                                        | 600 s            |
| Stitch Videos (`stitch`)            | `videos`: 2 to 10 clips                                                 | One video, in list order                           | 120 s in total   |
| Split Into Scenes (`split_scenes`)      | `source_video`                                                          | Up to 10 clips, one per scene (min 0.5 s)          | 600 s            |
| Merge Layers (`merge_layers`)      | `background` (video or photo), `layers` (1 to 6 photos, clips or audio) | One composited video                               | Background 120 s |
| Extract Frame (`extract_frame`)     | `source_video`                                                          | One PNG                                            | 600 s            |
| Extract Audio (`extract_audio`)     | `source_video`                                                          | One 192 kbps MP3. A video with no sound is refused | 600 s            |
| Change Speed (`change_speed`)      | `source_media` (video or audio)                                         | Faster or slower, pitch kept                       | 600 s            |
| Change Voice (`change_voice`)      | `source_video`, plus `voice_id`                                         | Same words, new voice                              | 120 s            |
| Resize (`resize`)            | `source_media` (photo or video)                                         | A new shape                                        | Video 120 s      |
| Upscale (`upscale`)           | `source_media` (photo or video)                                         | A sharper, larger copy                             | Video 30 s       |
| Remove Background (`remove_background`) | `source_media` (photo or video)                                         | The subject, no background                         | No limit         |
| Skin Enhancer (`skin_enhance`)      | `source_image` (a portrait)                                             | A retouched copy                                   | One photo        |
| Camera Angle (`camera_angle`)      | `source_image`                                                          | A new camera angle                                 | One photo        |

* A file slot takes plain asset ids: `"source_video": ["ast_..."]`.
* A longer file is refused before the tool runs. Trim it first. Upload sizes: [Limits](/reference/limits).
* Change Voice (`change_voice`) takes a voice id (`voc_`) as the top-level `voice_id`, the same as Text to Speech (`tts`). To revoice a recording with no picture, use Speech to Speech (`speech_to_speech`).

Voice tools that start from audio, with their fields: [Voice and music](/models#voice-and-music).

## Settings [#settings]

Defaults are in bold. The live schema wins: [Read one schema](/api/capabilities#read-one-schema).

| Tool                           | Field               | Values                                                                                               |
| ------------------------------ | ------------------- | ---------------------------------------------------------------------------------------------------- |
| Add Captions (`auto_caption`)  | `caption_style`     | **`punch`**, `impact`, `boxed`, `word_pop`, `clean`, `marker`, `handwritten`, `terminal`             |
|                                | `caption_position`  | **`auto`**, `top`, `center`, `bottom`                                                                |
| Text Overlay (`text_overlay`)  | `text`              | 1 to 300 characters. Required                                                                        |
|                                | `position`          | **`top`**, `center`, `bottom`                                                                        |
|                                | `text_align`        | `left&#x60;, &#x2A;*`center`**, `right`                                                              |
|                                | `start`, `duration` | Seconds, video only: start 0 to 600 (**0**), duration 0.1 to 600 (**3**)                             |
|                                | `font`              | **`insta`**, `tiktok`                                                                                |
|                                | `background`        | **`none`**, `dark`, `white`                                                                          |
|                                | `text_size`         | **`default`**, or `custom` with `text_size_px` 8 to 400                                              |
| Trim Video (`trim_video`)    | `start`, `length`   | Seconds: start 0 to 600 (**0**), length 0.1 to 600 (**5**). Stops at the clip's end                  |
| Split Into Scenes (`split_scenes`)  | `mode`              | **`auto`** finds the cuts, `manual` makes equal parts                                                |
|                                | `threshold`         | 1 to 90 (**17**). Lower finds more scenes. `auto` only                                               |
|                                | `scene_count`       | 2 to 10 (**6**): the most in `auto`, the exact number in `manual`                                    |
| Extract Frame (`extract_frame`) | `timestamp`         | Seconds, 0 to 600 (**0**). Past the end gives the last frame                                         |
| Change Speed (`change_speed`)  | `speed`             | 0.25 to 4 (**1**). Below 1 is slower                                                                 |
| Resize (`resize`)        | `aspect_ratio`      | **`9:16`** (1080x1920), `4:5` (1080x1350), `1:1` (1080x1080), `16:9` (1920x1080)                     |
|                                | `mode`              | **`fit`** (black bars), `crop` (fills the frame), `smart` (a model extends the picture, photos only) |
| Camera Angle (`camera_angle`)  | `horizontal`        | **`front`**, `left_45`, `right_45`, `left_profile`, `right_profile`, `behind`                        |
|                                | `vertical`          | **`eye_level`**, `low`, `high`, `overhead`                                                           |
|                                | `zoom`              | **`keep`**, `wide`, `medium`, `close_up`                                                             |

The other tools take only their files.

## Chain tools [#chain-tools]

Put one job's output (its `asset_id`) in the next tool's file slot. Nothing is downloaded in between.

Each step is its own job, so [wait for the result](/how-it-works#wait-for-the-result) before you start the next. Split Into Scenes (`split_scenes`) gives back several clips: chain each one on its own.

Example: trim a video to 15 seconds, make it vertical, then add captions.

**Trim** with Trim Video (`trim_video`).

```json title="Job 1"
{
  "capability_id": "trim_video",
  "config": {
    "source_video": ["ast_0193c8f0a1b24e7f9d3c5a6b7e8f0061"],
    "length": 15
  }
}
```

**Make it vertical** with Resize (`resize`). The file is job 1's output.

```json title="Job 2"
{
  "capability_id": "resize",
  "config": {
    "source_media": ["ast_0193c8f0a1b24e7f9d3c5a6b7e8f0062"],
    "aspect_ratio": "9:16",
    "mode": "crop"
  }
}
```

**Add captions** with Add Captions (`auto_caption`). The file is job 2's output.

```json title="Job 3"
{
  "capability_id": "auto_caption",
  "config": {
    "source_video": ["ast_0193c8f0a1b24e7f9d3c5a6b7e8f0063"],
    "caption_style": "boxed"
  }
}
```

### Exact words with Text Overlay [#exact-words-with-text-overlay]

Image and video models still misspell words. For an exact headline, price or brand name, make the picture without the words. Then add them with Text Overlay (`text_overlay`).

* One text block per job, typed text only (not a logo).
* For a second line, run it again on the file the first job gave back.

```json title="A line at the bottom, on a dark background"
{
  "capability_id": "text_overlay",
  "config": {
    "source_media": ["ast_0193c8f0a1b24e7f9d3c5a6b7e8f0064"],
    "text": "Free shipping this week",
    "position": "bottom",
    "background": "dark"
  }
}
```
