What Are InSpatio World 1.5, Sol Refiner and PixelUMM?
7 min read
Three research projects stood out on the image and video side this week: InSpatio World 1.5, Nvidia's SoL-Refiner (called "Sol Refiner" in the video) and PixelUMM. One turns a single image into video that roams the scene, one upscales low-resolution video to 4K in one step, and the third questions the basic architecture of image generation.
This article summarizes the relevant sections of the AI Search channel's weekly AI news video on YouTube and extends them with information from the three projects' own pages; the video link is at the end. This is a news roundup, and I have not tested the tools. The numbers are the researchers' own measurements.
What is InSpatio World 1.5?
InSpatio World takes an image and a camera path and generates video that roams the scene along that path. Version 1.0 came out with the idea of building a world beyond the original camera view from a single video. 1.5 widens the input options and the range of viewpoints:
- Input: A single image, an image set, a panorama or a video. Multiple images increase consistency.
- Real-time scene roaming: Starting from a single image you can explore the scene from a wide range of viewpoints; areas hidden in the original view are revealed, and newly revealed areas stay consistent with the realism and style of the source image.
- Next-view prediction in dynamic worlds: With video input you can move the viewpoint freely through a moving scene and choose when to observe the action. According to the video, the motion of the source video is carried into the scene too.
- Bullet time (the effect of orbiting the camera around frozen time): Starting from several synchronized images, you can design a camera path around the subject to build a smooth bullet-time sequence.
The page states that subject identity, spatial structure and scene details are preserved even as the camera moves dramatically. A live demo, code, Hugging Face and ModelScope links are available; the code is open source. According to the video the model is based on the 1.3-billion-parameter version of Wan 2.1 and is under 6 GB, so it can fit most consumer graphics cards. The presenter says he was surprised this base was chosen over MiniMax H3. Project page: inspatio.github.io/inspatio-world-1.5.
Who is it for? Cinematography and visual effects, immersive experiences, and research in spatial and embodied intelligence (dynamic environment modeling for robots) are the use cases the page lists.
What is SoL-Refiner?
SoL-Refiner stands for "Speed-of-Light One-Step Refinement for High-Resolution Video." It is work by Nvidia Research's Efficient AI team and Singapore lab. The idea: high-resolution video generation is expensive because compute grows quickly with the number of spatiotemporal tokens. Generating a low-resolution draft first and then applying a refiner is a practical route, but conventional multi-step refinement introduces a second sampling bottleneck. SoL-Refiner turns the low-resolution output of different generator models into 4K video with a single target-resolution step.
The training recipe has three parts: high-resolution continual training, frame-based reinforcement learning post-training and one-step distribution-matching distillation. According to the video it is derived from LTX Refiner.
SoL-Refiner measurements
Speed measurements on the page (the researchers' own):
| Generator | Direct generation | Draft + one-step refinement | Speedup |
|---|---|---|---|
| MiniMax H3 (5 s, 1344×768, 24 fps, GB200) | 152.3 s | 5.64 s | 27.03x |
| Wan (1280×720, 81 frames, H100) | 272.69 s | 78.71 s | 3.46x |
| Cosmos-Nano (35 steps, 189 frames, H100) | (35-step direct) | With one-step refinement | 2.81x |
| SANA-Video (81 frames, H100) | 20.04 s | 12.54 s | 1.60x |
The MiniMax H3 example uses a two-GPU pipeline: stage one takes 4.06 s at 896×512 and the refiner 1.56 s. The page also states that it is competitive at 2K and improves two reported metrics over LTX-2.3 at 4K.
The presenter's main criticism in the video is fidelity: SoL-Refiner gives much sharper results but changes details noticeably in the examples shown; the details of a barn differ after upscaling. The page's "refiner off/on" comparisons show gains in details such as faces, hair, fabric and texture. Where detail matters, such as a product, a face or text, you need to compare the output with the original. The code is released. Project page: nvlabs.github.io/Sana/Sol-Refiner.
What is PixelUMM?
PixelUMM is a unified model that both understands and generates images and video; it is work by researchers at Nvidia and the University of Waterloo. What sets it apart is the architecture: no VAE and no vision encoder. A single decoder-only Transformer reads and writes raw pixels. Images are split into 16×16 patches and videos into 4-frame tubes. The backbone is Qwen3-8B; there are separate "expert" layers for understanding and generation, and all share one self-attention across text, clean pixels and noisy pixels.
Most image and video models produce content in a "latent space" compressed by a VAE and then decode back to pixels. PixelUMM skips that step. The presenter in the video describes this as "no encoder"; according to the page, the accurate statement is that it has neither a VAE nor a vision encoder.
A limitation the page states openly: the linear pixel output head can leave faint grid-aligned patch artifacts in smooth regions (such as sky), which become more pronounced at high classifier-free guidance (around CFG 6). Convolutional heads reduce this, but the released model, all demos and the videos on the page were produced with the default linear heads; switching to a convolutional head needs further training. The page also has 1.7-billion and 8-billion parameter versions and a comparison of pixel-space and VAE-space training.
The video's verdict is clear: quality and consistency do not reach frontier models. But this is a research preview of a new architecture; code and model are released, and most source files are under the permissive Apache 2 license. Project page: nv-tlabs.github.io/PixelUMM.
All three together
| InSpatio World 1.5 | SoL-Refiner | PixelUMM | |
|---|---|---|---|
| What it does | Roaming a scene along a camera path from an image or video | Upscales low-resolution video to 4K in one step | Image and video understanding and generation in pixel space |
| Basis | Wan 2.1 1.3B based (per the video) | Refiner that works with different generators | Transformer with a Qwen3-8B backbone |
| Strength | Multi-input, consistent roaming | 27x end-to-end speedup on MiniMax H3 | VAE-free architecture |
| Weakness | Not stated on the page | Changes details | Behind frontier models, patch artifacts |
For the speed-up side, see the same week's DMAD and PDMD article; for installing an open-source video model, there is the How to install LTX 2.5 guide.
Frequently Asked Questions
What inputs does InSpatio World 1.5 accept?
A single image, an image set, a panorama or a video.
Does SoL-Refiner upscale video faithfully?
No. It gives much sharper results but changes details; the difference is visible in the video's examples. The speed measurements on the page show up to a 27x end-to-end speedup on MiniMax H3.
What is different about PixelUMM?
It works directly in pixel space with a single Transformer, without a VAE or a vision encoder. Its quality does not yet match frontier models.
Are all three open?
According to the video, code is released for all three. Most of PixelUMM's source files are under Apache 2; see the project pages for the others' licenses.
Source
- AI Search, Gemini 4, GPT 6.1, Dots, Claude Sonnet 5.5, Ideogram 4.5, Flux 3: AI NEWS (YouTube, October 4, 2026): the InSpatio World 1.5, Sol Refiner and PixelUMM sections.
- InSpatio, InSpatio-World 1.5: project page.
- Nvidia Research, SoL-Refiner: project page and measurements.
- Nvidia and University of Waterloo, PixelUMM: project page.
Related Posts
What Is Comfy Agent? An Agent That Builds ComfyUI Nodes
Comfy Agent is an agent built into ComfyUI: describe what you want and it plans, builds and runs the workflow. Comfy Cloud only for now, free to try.
DMAD and PDMD: 4-Step AI Video Generation Explained
ByteDance's DMAD and PDMD run MiniMax H3 in 4 steps. On VideoGen-Eval, PDMD's total score beats the 50-step teacher model.
What Is LTX 2.5? Free Open-Source AI Video Generator
LTX 2.5 is a 22-billion-parameter open-weight model that generates video with sound. Here is what it does, its license, why it counts as free, and who it suits.