Use the Generate rail to create new images, video clips, and audio using AI provider models.
How Generate Works
Generation in Studio is straightforward: you choose what kind of media you want, pick a provider and model, describe what you want in a prompt, and click Generate. Studio sends your prompt and any reference media to the chosen provider, then imports the result as a take in your project.
Every take is saved in the asset shelf with a link back to the provider job that created it — so you can always see what prompt, model, and settings produced it.
Generation never sends your project files, credentials, or other assets to a provider unless you explicitly include them in the request (as a prompt, a reference, or context). Your provider keys stay in Keychain. See Provider Credentials and Privacy for details.
Choose a Mode
The Generate rail has three modes, selectable at the top:
- Image — Generate a single still image.
- Video — Generate a video clip, with optional audio and motion controls.
- Audio — Generate a voiceover, dialogue, sound effect, or ambience.
Switch between modes freely. Studio remembers your last settings for each mode and reapplies your saved Provider Defaults when you switch — so Image always returns to your preferred image model, and Audio always returns to your preferred audio provider.
Providers and Models
What Providers Are
A provider is an AI service that does the actual generation. Studio includes adapters for several providers. You bring your own API keys for the providers you want to use.
When you select a provider, Studio shows the models that provider offers for the current mode. Not every provider supports every mode.
Studio supports generation through OpenRouter and ElevenLabs. The available models depend on the provider account and the selected generation mode.
Choosing a Model
Models vary in what they can produce. Studio checks what the selected model supports and disables controls that do not apply. If a setting is grayed out, it means the current model cannot use it — try a different model.
When in doubt, start with the default model for your chosen mode. It is a safe, capable starting point.
Refreshing Model Availability
Providers occasionally add or remove models. Click the refresh button next to the model picker to update the list. This reads from the provider's current offerings using your stored credentials.
Writing Prompts
The prompt is the main way you describe what you want to generate. It appears as a text field in the Generate rail.
Good Prompts
A good prompt is specific but not cluttered. Describe:
- The subject — what or who is in the scene.
- The action or mood — what is happening, or what feeling the image should convey.
- The setting — where it takes place.
- Visual or audio qualities — lighting, composition, pacing, tone of voice.
Example image prompt:
A wooden workbench in a sunlit workshop, tools organized on a pegboard, warm afternoon light through a window, shallow depth of field.
Example video prompt:
A barista pouring latte art in a bright café, slow smooth motion, overhead shot, warm natural light, 5 seconds.
Example audio prompt:
A calm, clear voiceover introducing a product. Warm and professional tone, unhurried pacing.
Prompts With Context
Studio can add context to your prompt when you select it. Generate includes explicit controls for production context such as character context, location context, and scene context. Selecting a storyboard scene can load its prompt, bindings, and references into Generate.
Use context when you want consistency across multiple generations — for example, keeping the same character description and setting across several takes.
Reference Media
The Generate rail includes wells for attaching images as references:
- Edit Source — An image that the model should use as the starting point or main visual guide. Think of this as "generate something based on this."
- Visual Reference — A secondary image that provides style, color, composition, or subject guidance without being the main input.
You can drag images from the asset shelf into these wells, or use the Use in Generate Image option when right-clicking an asset.
Not every model supports image references. Studio lets you know if the selected model cannot use them.
Use the reference wells with the video Frame Lock controls below when you need continuity between takes.
Advanced Options
The Advanced section in the Generate rail exposes controls that are less commonly changed but useful for specific workflows.
Video Advanced Options
- Duration — How long the generated video should be.
- Aspect Ratio — The shape of the frame (16:9, 1:1, 9:16, etc.).
- Motion Presets — Pre-built motion behaviors for video models that support them.
- Frame Lock — Lock the first or last frame to a reference image for shot-to-shot continuity.
- Generate From Prior Scene — Use the last frame of a timeline clip or previous take as the starting frame for the next generation.
- Native Audio — For video models that can also generate synchronized audio.
Audio Advanced Options
- Audio Subtype — The kind of audio: voiceover, dialogue, sound effect, or ambience. Available choices depend on the selected provider and model.
- Voice Profile — A saved voice selection from your provider.
- Target Duration — How long the generated audio should be.
JSON and Provider Overrides
The Advanced JSON and Provider Override fields are for provider-specific options that do not yet have dedicated controls in Studio. Leave these at their defaults unless you know the provider expects a specific value. The inline ? buttons explain what each field is for.
What Happens When You Click Generate
- Studio checks that your provider key is available and the model supports what you have asked for.
- The job appears in the Activity panel as "preparing."
- Your prompt, reference media, and settings are sent to the provider.
- The Activity panel shows progress.
- When the provider returns a result, Studio imports it as a take in the asset shelf.
- The take is linked back to the provider job record, so you can always review what created it.
If generation fails, Studio shows the failure reason in the Activity panel. See Troubleshooting for common issues.
Tips
- Generate several takes before choosing. The same prompt can produce different results. Try two or three takes, then pick the best one.
- Start with shorter durations for video. A 3-second clip generates faster than a 10-second one and gives you quicker feedback on whether the prompt and model combination is working.
- Use the
?buttons. Every Generate field with an inline?has a short explanation of what it controls. - Save reusable prompt patterns in Scene Cards. If you generate consistently for a character or setting, create a Scene Card with that context rather than retyping prompts. See Scene Cards in Elements.
- Check model capabilities before troubleshooting. Many "bugs" are simply a model not supporting a requested feature. The Generate rail shows you what is available.
Watch Out For
- Provider Defaults reapply when you switch modes. If you have different defaults for Image, Video, and Audio, switching modes will change the selected provider and model. This is intentional — check the provider picker after switching to confirm.
- Video generation can take a minute or more. This is normal. The Activity panel shows status.
- Not every model supports image references. If the reference wells are disabled, try a model that lists reference support.
- Context from Elements is additive. If you have a long character description, a detailed location, and a scene card all adding context, the total prompt sent to the provider may be quite long. If results seem unfocused, try reducing the context or writing a tighter prompt.
- Output quality varies by model. If a take misses the goal, refine the prompt, shorten the request, or try another model that supports the same mode.