İçeriğe geç / Skip to content / Zum Inhalt

LTX 2.5 Requirements and Speed: How Much VRAM?

Ahmet Balaman

8 min read

AI VideoLTX 2.5VRAMGraphics CardLow VRAMRTX 5090
LTX 2.5 Requirements and Speed: How Much VRAM?

The official minimum for LTX 2.5 is an NVIDIA graphics card with 32 GB of VRAM. The community also runs it on smaller cards, but that is not an officially supported route, and it is slow. The model gets called the "fastest local AI video generator," so how fast is it really, and how long will you wait on which card? In this post I gather the answer with numbers and their sources.

This is the third part of the series. You may want to read What Is LTX 2.5 first, then how to install it.

This post is a compilation; I did not measure the times myself. The numbers come from LTX's documentation, NVIDIA's news post, and community reports, and I wrote the source next to each one. Sources are at the end.

Official requirements

According to LTX's ComfyUI page:

Item Value
Graphics card CUDA-compatible, 32 GB or more of VRAM
Free disk Over 100 GB (models and cache)
Python 3.12 or later

LTX's low-VRAM guide additionally lists 32 GB of system memory, CUDA 13.2 or later, Python 3.10 or later, and PyTorch 2.7. The model card, on the other hand, recommends CUDA 12.7 or later. So CUDA and Python versions are not consistent across sources; check the README in the repository before installing.

The recommended configuration in the guide is an A100 (80 GB) or H100, 64 GB of system memory, and a 200 GB SSD. The full-precision model runs comfortably in this class; on consumer cards, quantized and distilled versions are used.

What happens on which card?

The guide's tiers:

  • 80 GB and above (A100, H100): The "dev" checkpoint at full precision and the two-stage pipeline.
  • 32 GB (RTX 5090 class): The distilled checkpoint with FP8 quantization.
  • Under 32 GB: Below the documented minimum. LTX says this is not supported, though some settings may work with low resolution and aggressive memory optimization.

Community experience from independent sources:

Card Reported result Source
RTX 5090 (32 GB) A 4-second 720p clip in about 25 seconds NVIDIA guide (via RunAIHome)
RTX 5090 (32 GB) An 8-second clip in about 3 minutes once weight streaming kicks in RunAIHome
RTX 5090 Generated an 8-second 720p image-to-video in about 80 seconds (BF16) TrueNorthAI (Note)
RTX 3090/4090 (24 GB) 5-second clips complete; about 2.5 minutes per clip on a 3090 with the distilled model; 10-second clips may run out of memory in the final decode step RunAIHome
RTX 5060 Ti (16 GB) 1344×768, 24 fps in about 170 seconds with INT8/NVFP4 distilled weights; over 92% of 32 GB system memory was used AI Creative Log (Note)
RTX 3060 (12 GB) About 0.5 megapixels, a 10-second video in about 180 seconds, with 96 GB of system memory AI Creative Log (Note)

Every one of these numbers is a single person's community report. There is no official measurement showing the speed or quality difference from quantization; the AI Creative Log author also notes this explicitly. Read the times not as a range but as "something close to this may happen."

The key detail: memory is not just sampling

Three traps come up in the sources I keep running into:

  1. The VAE decode step. After sampling finishes, the step that decodes the image needs extra memory; on 24 GB and even 16 GB cards, memory can run out at exactly this point. The suggested fix is to switch to the lighter Conv video VAE and split decoding into tiles.
  2. System memory. Weight streaming and CPU offload use system RAM. On 16 GB cards, the practical minimum is 32 GB of RAM.
  3. Clip length. Memory multiplies as length grows; start with a short clip.

Ways to reduce memory

The methods in LTX's low-VRAM guide, ordered by effect:

  • Distilled checkpoint. A fixed 8-step pipeline, much lighter and faster than the "dev" model. This is the starting point on any card under 32 GB.
  • FP8 quantization. Cuts memory use by about 40%. There are two backends: fp8-cast (works on most cards with FP8 support) and fp8-scaled-mm (faster on Hopper cards). On the command line, --quantization fp8-cast.
  • INT8 ConvRot files for ComfyUI. File names are in the install post. One source puts this file family at about 21.5 GB.
  • NVFP4. An even smaller version for NVIDIA Blackwell cards with FP4 support.
  • Tiled VAE decoding. Reduces peak memory and slows decoding slightly.
  • Leaving text encoding to the API. The Gemma 4 encoder and the prompt enhancer hold VRAM; with an API key you can run them remotely.
  • Environment variable: Before running any pipeline, set PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True.

Scaling up step by step

When memory is tight, LTX's advice is to start from the smallest setting and raise one thing at a time:

  1. Distilled checkpoint, 512 pixels, 49 frames, fp8-cast.
  2. Increase resolution step by step: 512, 576, 640, 704, 768. In two-stage pipelines, width and height must be divisible by 64, and in single-stage by 32; 720 or 1080 don't follow this rule.
  3. Then increase the frame count. The rule for frame count is that (frames - 1) is divisible by 8: 9, 17, 25, 33, 41, 49, 57, 65, and up.
  4. If you are adding IC-LoRA, add one group at a time.

If you get a memory error, go in this order: check the environment variable, turn on FP8, lower the resolution or frame count, switch to the distilled model, increase the VAE tile count, drop to a single IC-LoRA group, and close other applications that use GPU memory.

How solid is the "fastest" claim?

This section answers the question in the title. Here is what the sources say:

  • According to Pexo's compilation, a 10-second 720p clip was generated in about 6.8 seconds on two NVIDIA GB200 graphics cards (a vendor demonstration). The same post relays that VentureBeat found the model "several times faster than its closest closed-model competitors," but offers no independent verification.
  • NVIDIA announced up to 2x performance gains and 40% memory savings for LTX 2.5 on RTX graphics cards (with optimizations such as NVFP4 and FastVideo).
  • On the official pipeline side, the distilled pipeline runs a fixed 8 steps (8 in the first stage, 4 in the second); this is the pipeline LTX calls "fastest."

So saying "one of the fastest among local video models" is defensible; to say "definitely the fastest," you would need an independent comparison under the same conditions, and I could not find one in the sources.

Which hardware should you buy?

This is my own interpretation, drawn from the numbers in the sources:

  • If you already have a 32 GB card (RTX 5090 class), you will be comfortable with short clips.
  • On a 24 GB card, 5-second clips are possible, but the decode step is borderline.
  • For a card with 16 GB or less, don't buy one specifically for this; the RunAIHome author also recommends renting a cloud GPU for 8 to 12 GB (citing about $0.99 per hour).
  • Don't neglect system memory: 32 GB of RAM is the practical minimum on the low-VRAM route.

Frequently Asked Questions

How much VRAM does LTX 2.5 need?

The official minimum is 32 GB. It is reported to run on 24 GB and 16 GB cards with quantized and distilled files, but this is not a supported route.

Does it work on a 16 GB graphics card?

According to community reports, yes, very slowly and using nearly all of 32 GB of system memory. There is no official support.

Does it work with an AMD graphics card?

My sources only mention NVIDIA CUDA. The LTX guide says it "targets NVIDIA."

How many minutes does it take to generate a video?

It depends on the card and clip length. On an RTX 5090 the reported time for a 4-second 720p clip is about 25 seconds; on 24 GB cards, a 5-second clip takes minutes.

Does it work on a Mac?

My sources have no information about Mac, so I won't say anything.

Sources

Comments