Free AI Video Generation with LTX 2.5: Prompts and Settings
6 min read

Once LTX 2.5 is installed, three things decide the result: the structure of your prompt, the resolution and frame-count rules, and which workflow you choose. This post is the last part of the series. I introduced the model in What Is LTX 2.5, and wrote about the installation and hardware requirements separately. Here we look at producing your first video and getting the most out of free AI video generation.
This post is based on the official documentation; I have not tried generating video myself. The rules and settings come from the model card, LTX's README, and the ComfyUI page. Sources are at the end.
Which workflow should you choose?
ComfyUI has three ready-made templates:
| Goal | Template | Note |
|---|---|---|
| I only have an idea | Text-to-Video | Two-stage |
| I have an image and want motion | Image-to-Video | Two-stage; the image becomes the first frame |
| I know my start and end frames | First-Frame / Last-Frame | Single-stage; generates what's in between |
On the Python side there are additional pipelines too: video from an audio input (audio-to-video), keyframe interpolation, region regeneration (Retake), lip-synced dubbing (Dub-It), and SDR-to-HDR conversion. These are advanced; in the first few days the text and image workflows are enough.
Practical advice: If you have an image, start with Image-to-Video. With text-to-video, composition is left to chance; with an image, you keep control. The two-stage templates use a spatial upscaler for high resolution; the single-stage First/Last Frame template has no upscaler.
How to write a prompt
The guidelines in LTX's README:
- Be detailed and chronological. Write events in the order they happen.
- Start with the direct action, and keep descriptions concrete and plain.
- Stay within 200 words.
- Follow this order: main action, movements, appearances, background, camera, lighting, changes over time.
- State movement, appearance, camera angle, and environment details explicitly.
An example prompt that follows this structure (not official; I wrote it to illustrate the guideline):
A golden retriever runs through a sunny meadow, ears flapping. The grass
bends as it passes; wildflowers sway. Behind it, a wooden fence and rolling
hills. The camera tracks alongside at ground level, shallow depth of field.
Warm afternoon light, long shadows.The official examples and the prompt guide are all in English. I have no data in my sources on how well prompts in other languages work, so I make no claim; writing in English matches the official examples exactly.
Prompt enhancer
By default, the templates keep the prompt enhancer on, which expands a short prompt into a more detailed cinematic description. If you want your own sentence used exactly as written, turn off the Prompt Enhance setting.
Audio
The model generates audio at the same time. Describing sound and speech details in the prompt (ambient sound, a speaking character, music) steers the result. One independent hands-on post reported problems with Japanese audio; I found no data for other languages, so test it yourself.
Multi-shot generation
The new feature in LTX 2.5: several connected shots in one generation. Character, lighting, and sound are kept consistent across shots. How to write multi-shot prompts is covered in LTX's official "How to prompt LTX-2" guide; I haven't read and tested it yet, so I won't go into detail.
Resolution and frame-count rules
If you don't follow these, you get errors or a corrupted output:
| Rule | Value |
|---|---|
| Frame count | frames % 8 == 1 (9, 17, 25, 33, 41, 49, 57, 65, ..., 121, ...) |
| Width and height, single-stage pipeline | Divisible by 32 |
| Width and height, two-stage pipeline | Divisible by 64 |
| Stage 2 resolution | Twice that of stage 1 |
That is why broadcast labels like 720 or 1080 are not used directly; values like 704 and 768 are used instead. The Diffusers example uses 544×960 and 121 frames (24 fps) for stage 1; the second stage runs at twice that.
If you don't want to choose the duration yourself, the duration predictor node, or leaving the --num-frames flag empty on the command line, lets the model pick the length from the prompt.
Settings for your first video
From the "workflow tips" section of the ComfyUI page and LTX's low-VRAM guide:
- Run a ready-made template as-is first. Confirm that the tool works.
- Test with low resolution and few frames. The ComfyUI page suggests 480×720 and 41 to 81 frames.
- Raise quality once you're happy. Increase the resolution first, then the frame count, one at a time.
- Compare with the same seed. Change only one setting at a time.
- If you get a memory error, follow the order in the requirements post: FP8, lower resolution, distilled model, tiled VAE.
- Keep your workspace clean. Close unused workflows and clear the cache between large generations.
A one-line image-to-video example on the command line:
--image path/to/first_frame.jpg 0 1.0 \
--prompt "The camera slowly dollies out as wind moves through the grass"The --image flag takes the path, the frame number (0 is the first frame), and the strength.
Things to watch for with free generation
- License: Commercial use is free if annual revenue is under $10 million. Above that, a paid license is required. Details are in the what-is post.
- Electricity and time: Free means you pay no money; it doesn't cover your card's running time.
- Accuracy: The model is not designed to provide factual information. Don't publish content that looks like a real person or event as-is.
- Acceptable use: The prohibited areas in the license (child sexual abuse, non-consensual sexual content, and so on) apply.
- Images of real people: If you will generate video from someone else's photo, make sure you have permission.
For editing, captions, and publishing steps after generation, see the complete guide to making video with AI.
Fine-tuning and LoRA
The "dev" transformer is fully trainable, and LoRA and IC-LoRA can be trained with the LTX-2 Trainer. According to the model card, the large majority of LoRAs trained on LTX-2.3 work on LTX-2.5 without changes; there are a few exceptions, so verify before using them in production.
Frequently Asked Questions
How do I generate my first video with LTX 2.5?
In ComfyUI, open the Text-to-Video or Image-to-Video template from the Templates menu, click Download all, write a prompt, and press Queue Prompt. The output lands in the ComfyUI/output/ folder.
Can I write prompts in languages other than English?
My sources say nothing about other languages. The official examples and guides are in English; do your first tests in English and compare other languages yourself.
Why am I getting an error, is it the frame count or resolution?
Most likely. The frame count must satisfy % 8 == 1, and width and height must be divisible by 32 in single-stage and 64 in two-stage.
How do I control the audio?
Describe sound and speech details in the prompt. The audio comes from a separate audio VAE.
What is the maximum video length?
Sources mention 6 to 20 seconds. The practical limit is set by your hardware.
Sources
- Lightricks, LTX-2.5 model card (Hugging Face): restrictions, image-to-video example, duration predictor, fine-tuning note, limitations.
- Lightricks, LTX-2 repository (GitHub): prompt guidelines, pipelines.
- LTX, Using ComfyUI with LTX: templates, workflow tips.
- LTX, How to Run a Video Generation Model Locally: frame and resolution rules, step-by-step scaling.
- TrueNorthAI, RTX 5090 hands-on: note on the Japanese audio issue (third party).
- The Rundown, LTX-2.5 Review: acceptable use policy.
Related Posts
What Is LTX 2.5? Free Open-Source AI Video Generator
LTX 2.5 is a 22-billion-parameter open-weight model that generates video with sound. Here is what it does, its license, why it counts as free, and who it suits.
How to Install LTX 2.5 in ComfyUI: Step-by-Step Guide
Install LTX 2.5 with ComfyUI templates, by adding nodes manually, or from the Python command line. File list, folders, and fixes for the most common errors.
LTX 2.5 Requirements and Speed: How Much VRAM?
The official minimum for LTX 2.5 is 32 GB of VRAM. What happens on 24 GB and 16 GB cards, how FP8 and INT8 save memory, and the generation times people report.