LTX-2.5 vs MiniMax H3: Two New Local AI Video Models in ComfyUI
Aug 21, 2026
LTX-2.5 and MiniMax H3 arrived within days of each other, and both can generate video and audio through official ComfyUI templates. In this first look, I open the basic local workflows, run the included starting points, play both demo results, and explain what I want to test next.
This is not a scientific benchmark and I am not declaring a winner. LTX-2.5 interests me for its improved distilled model, diffusion fidelity rendering, and connected multishot generation. MiniMax H3 interests me for native stereo audio, multimodal references, first-and-last-frame control, and its reference-driven workflows.
The important question is not simply which model is better. It is which one is better for a specific production job.
Which deeper test should come first: subject consistency, dialogue and lip sync, image-to-video, multishot generation, or reference workflows?
Chapters
00:00 Opening demos
00:10 LTX-2.5 and MiniMax H3 introduction
00:50 Why these two models matter
01:46 LTX-2.5 first look
04:59 MiniMax H3 first look
07:44 Different strengths, no winner yet
08:33 What we test next
09:23 Final thoughts
AI Disclosure
This video contains machine-generated demonstration footage created with LTX-2.5 and MiniMax H3. The host narration, testing, editorial choices, and final edit are original Baynum Tech Works production work.
Show More Show Less #Science

