Imbutus

News

04 Aug 2026, 15:00

MiniMaxDirector: a timeline for H3

I wrote a ComfyUI node for MiniMax H3 and released it as open source: MiniMaxDirector. Instead of packing an entire film into one prompt box, you lay it out on a timeline — what happens on screen, how the camera moves, what is heard — and it compiles that into the single structured prompt H3 actually reads.

What it does for you:

  • Three tracks. Shots, camera and audio, each segment with its own text and its own span.
  • Cut times are computed, so the model is told exactly when each shot begins.
  • Legal lengths only. The clip is snapped to a duration H3 accepts, so a run cannot fail on arithmetic.
  • Attach a file to a segment and it becomes a numbered reference in the prompt automatically — no tokens to type by hand.

It is already installed on your GPU. Start the MiniMax H3 bundle and open Docs → Media → MiniMax H3 → the MiniMaxDirector workflow; the graph is ready to run, nothing to install.

On your own machine, it is on the ComfyUI registry as minimax-director — searchable in the ComfyUI Manager, or comfy node install minimax-director. Source and issues: github.com/imbutus/ComfyUI-MiniMaxDirector. MIT licensed.

03 Aug 2026, 18:00

MiniMax H3 walkthrough

The MiniMax H3 bundle now has a walkthrough on the Docs page, in English, Russian and Chinese, with subtitles.

It goes through a first clip end to end:

  • Where the workflows are and which one to open first.
  • Writing sound into the prompt — what to describe, and what the model does with it.
  • Legal clip lengths — why an arbitrary duration is snapped, and which numbers land on whole seconds.
  • Attaching references so a subject stays the same from one shot to the next.

Watch it on YouTube, or open Docs → Media → MiniMax H3.

03 Aug 2026, 15:00

MiniMax H3 is available

A new video bundle is on the Media page: MiniMax H3, an omni-modal model that generates picture and sound in one pass — the audio is not dubbed on afterwards, it comes out of the same generation as the video.

What is worth knowing before your first clip:

  • Sound is part of the model. Describe what is heard the same way you describe what is seen, and it is generated with the shot. Stereo, no separate voice step.
  • Reference images, audio and video. Attach them and point at them from the prompt; the model uses them to keep a subject or a look consistent across shots.
  • Clip length is not free. H3 accepts lengths on a fixed lattice at 24 fps, so the workflow snaps whatever you ask for to the nearest legal value. 8, 25 and 42 seconds are the whole-second ones.
  • One prompt, many shots. The model reads a shot list with cut times, so a single generation can contain several shots rather than one continuous take.

Open Media → Video → MiniMax H3 to start it.

28 Jul 2026, 11:28

Voice bundles reorganized

The four voice bundles now have one clear job each, one shared set of workflow names, and they start faster.

  • Fish Audio S2 — the dubbing pick. Text-to-speech and voice cloning across 80+ languages, with voice-to-SRT and SRT-to-voice that keep the original timing.
  • CosyVoice 3 — the change-voice pick. Native voice conversion keeps the original words, pauses and delivery and swaps only the timbre, with no transcription step in between.
  • Qwen3-TTS — the all-Qwen bundle. Three-second voice cloning plus voice design, and the only bundle that ships Qwen3-ASR as a second transcription engine.
  • WhisperX is now the transcription engine everywhere. It is the more accurate one in real use, and dropping the second engine from the other three bundles removed about 3.7 GB from each pod, so voice pods are ready sooner.
  • Two workflows you did not have before. common-align-script-to-srt is in every voice bundle now — paste a script with a blank line between sections and it aligns them to your narration as an SRT. CosyVoice 3 also gained a proper cosyvoice3-voice-clone.

Some workflow names changed in the _examples folder, so look for the new ones: *-tts-clone is now *-voice-clone, *-audio-to-srt is now common-audio-to-srt, and qwen3tts-change-voice is now qwen3tts-redub-voice — it transcribes and re-speaks the words rather than converting the voice, so use CosyVoice 3 when you want the delivery kept.

The Docs page now lists every bundle with its models and every workflow it ships, each with the same notes you see on the ComfyUI canvas.

26 Jul 2026, 14:49

Sulphur-2 video tutorial

The Sulphur-2 bundle now has a full walkthrough on the Docs page — just under eight minutes, in English, Russian and Chinese, with subtitles.

It covers the LTX Director workflow end to end:

  • Image → video — make a person in a photo say the text you give it.
  • Image + audio → video — feed a voice clip you prepared in Fish Audio S2 and have the person in the picture speak that exact recording.
  • Extending a clip — put an existing video at the start or the end of the timeline and let the model generate the missing part. Match the framerate of your source, or the result drifts.
  • fp8 vs bf16 — these are different models, not quality presets. fp8 is the compressed one: lighter on resources, weaker results.

Also covered: why to test on a short time span before committing to a full render, why splitting prompts into short intervals generates more accurately, and how the wrong duration makes the model hallucinate. The voice itself can be replaced afterwards with CosyVoice3.

Open Docs → Media → Video tutorial to watch it.

26 Jul 2026, 08:32

SCAIL-2 is ready to use

SCAIL-2 has been through a full round of testing and is out of beta. Both of its jobs work end to end, and the Beta chip is gone from the bundle card.

  • Character animation — drive your own character with the motion from any clip. The background is generated fresh; the driving video's scene is not kept.
  • Character replacement — swap one person in an existing video and keep the original scene, background and everyone else untouched.
  • Two characters at once — one reference photo holding both of them, animated together from a clip with two moving subjects.

Every workflow now opens with step-by-step notes inside ComfyUI: which inputs to fill, how to point the tracker at one person when the clip has several, what to check in the mask previews before committing to a long render, and how to go past the default five seconds.

Two things worth knowing before you start. The render is silent — SCAIL-2 generates picture only, so add sound afterwards in a video editor. And the reference image has to be a real photograph; for two characters that means one picture of both of them together, not two photos joined side by side.

A video tutorial is coming later.

24 Jul 2026, 06:25

Set your own auto-stop timeout

Your GPU stops after a stretch of inactivity, so you are never billed for a machine you forgot about. That timeout used to be the same for everyone. Now it is yours to set.

  • Set LLM and Media timeouts independently on the Settings page.
  • Choose anything from 5 minutes up — or switch auto-stop off entirely for long jobs.
  • Media treats a busy ComfyUI queue as activity, so a render that runs for hours keeps the pod alive. The countdown only starts once the queue is empty.
  • LLM counts your own session, so your timeout never cuts anyone else off.

Running out of balance always stops a GPU, whatever you choose here.

22 Jul 2026, 08:34

See your usage and spending

You can now see exactly how much GPU time you have used and what it cost — right on the Monitor page.

  • Every LLM and media (image / video / voice) session is listed with its length and price.
  • A totals line shows your lifetime spend, split between LLM and media.

Open Monitor from the top menu to review your usage anytime.

18 Jul 2026, 15:00

Prepared OSINT Workflows

You can now trigger ready-made OSINT lookups directly in chat — no need to know which tools to run or in what order.

Triggers:

  • osint:email <address> — checks which platforms an email is registered on, plus domain/MX lookup
  • osint:person <name> — public profile and username enumeration across platforms
  • osint:company <name> — resolves a company name to its official domain, then maps infrastructure and org info
  • osint:domain <domain> — whois, subdomain enumeration, DNS records
  • Send osint on its own to see the full list again

Each workflow runs real tools (whois, dig, holehe, sherlock, theHarvester, crt.sh) on your own rented Kali Linux VPS — rent one in the Virtual Machines section if you don't have one yet. If you don't have a machine, the model will tell you so directly instead of guessing or making anything up.

18 Jul 2026, 15:00

Introducing the News Page

This page — the one you're reading right now. Short announcements about new features and changes, in English, Russian, and Chinese.

Nothing complicated: new posts show up here as they're published, newest first.

Page 1