Imbutus

Документация

Регистрация и активация

После регистрации пополните баланс до порогового значения. Как только баланс достигнет нужной суммы, аккаунт активируется автоматически. Порог может вырасти со временем — активируйтесь раньше.

Минимальный баланс для активации: $18 (будет расти, не опоздайте)

LLM

Видеоинструкция

Полное руководство:

Тарификация

Оплата только за фактическое использование GPU по рыночной цене — примерно столько же, сколько прямая аренда. Если несколько пользователей работают на одном GPU одновременно, стоимость делится. Зовите друзей — платите меньше.

Модели

Список моделей ограничен намеренно. Четыре причины:

  • 4B и 9B стоят примерно одинаково, но 9B значительно лучше — дублирующие модели исключены.
  • Некоторые модели нереализуемы в масштабе. Kimi-K2.6 требует 8 топовых GPU одновременно — стабильно обеспечить это практически невозможно.
  • Некоторые модели просто слишком большие, чтобы стартовать быстро. Чекпоинт в несколько сотен гигабайт может скачиваться часами на свежий GPU, прежде чем ответит хоть что-то.
  • Каждая модель требует индивидуальной настройки железа и ПО. Добавление модели — это реальная работа.

Доступные модели:

Дешёвые

  • huihui-ai/Huihui-Qwythos-9B-Claude-Mythos-5-1M-abliterated — Улучшенная 9B — новейшая дистилляция Claude (Mythos-5), контекст до 1 млн токенов, ввод изображений, без цензуры. Доступная и способная.
    256K ctxрассужденияприём изображений

Универсальные

Для программирования

  • imbutus/YuYu1015-Ornith-1.0-35B-abliterated — Альтернативная 35B — Ornith 1.0, аблитерированная и мультимодальная. Полная точность на одном GPU.
    256K ctxрассужденияприём изображений
  • huihui-ai/Huihui-Qwen3-Coder-Next-abliterated — Самая мощная модель для программирования. Используйте для написания кода и генерации программ.
    256K ctxбез рассуждений

Модели добавляются со временем. Каждую можно найти на Hugging Face по тому же имени.

Агентный API

Помимо модели, API включает базу CVE и эксплойтов с обновлением каждые 6 часов, а также встроенные знания о подключённых Kali Linux VPS. Дополнительные функции можно запросить через систему поддержки.

OSINT-сценарии

Отправьте "osint" в чат — модель ответит списком этих сценариев и их точным синтаксисом, запоминать ничего не нужно.

  • osint:email <address>

    Проверяет, на каких платформах зарегистрирован email-адрес, плюс поиск домена и MX-записей.

    whoisdigholehe
    Как это работает
    1. Проверяет WHOIS-данные домена адреса.
    2. Проверяет MX-записи домена (маршрутизацию почты).
    3. Запускает holehe, чтобы проверить адрес на 120+ платформах на наличие аккаунта.
    4. Сообщает, на каких платформах адрес зарегистрирован — неоднозначные результаты помечаются, а не додумываются.
  • osint:person <name>

    Ищет публичные профили и вероятные имена пользователей на разных платформах по имени.

    sherlocktheHarvester
    Как это работает
    1. Формирует вероятные имена пользователей на основе имени (и транслитерацию, если имя не на латинице).
    2. Запускает sherlock, чтобы проверить эти имена пользователей на 300+ сайтах.
    3. Если указана организация, дополнительно запускает theHarvester.
    4. Сообщает найденные профили с оговоркой о совпадении имён — соответствие никогда не утверждается как точное.
  • osint:company <name>

    Определяет официальный домен компании по названию, затем анализирует её инфраструктуру и организацию.

    theHarvesterwhoiscrt.shdig
    Как это работает
    1. Ищет и подтверждает официальный домен компании.
    2. Запускает полный сценарий Domain (WHOIS, поддомены, DNS) для этого домена.
    3. Запускает theHarvester для поиска сотрудников/email и вероятного шаблона email.
    4. Сообщает подтверждённый домен, структуру организации и данные инфраструктуры вместе.
  • osint:domain <domain>

    Выполняет whois, поиск поддоменов и DNS-запросы для домена.

    whoiscrt.shdigtheHarvester
    Как это работает
    1. WHOIS-запрос для регистратора и даты создания.
    2. Поиск поддоменов через логи сертификатов (crt.sh).
    3. Запрос DNS-записей — A, MX и NS.
    4. Проход theHarvester для поиска дополнительных email/хостов.

Каждому сценарию нужна активная машина Kali Linux — весь инструментарий (theHarvester, sherlock, whois, holehe и другое) уже установлен на ней, поэтому модель может выполнять всё вживую. Арендуйте машину в разделе Virtual Machines.

Рабочий процесс

Я лично использую Imbutus с PI (pi.dev) — но вы можете подключить любой поддерживаемый клиент и использовать так, как удобно вам.

Чат в веб-интерфейсе работает, но не является основным способом использования. Если при запросе GPU офлайн — клиент покажет прогресс загрузки в реальном времени.

Выделенная VM — задайте имя машине при создании (например, kalinux01). AI-модель знает её по этому имени и получает прямой доступ к оболочке: запускает nmap, metasploit, sqlmap и любые инструменты. Просто упомяните имя в промпте — модель подключится и будет работать с ней. Оплата посуточно — завершить можно в любой момент.

Списание идёт пока сессия активна. Чтобы остановить — нажмите кнопку Stop на главной странице (появляется при выбранной модели) или напишите модели: «останови нашу сессию». Если в этот момент никто другой не использует GPU — он выключится и списание прекратится немедленно.

Поддерживаемые агенты

Почти каждый AI-инструмент, расширение для IDE и агентный фреймворк работает с одним из двух форматов API — Messages API от Anthropic (/v1/messages) или Chat Completions от OpenAI (/v1/chat/completions). Поддерживаются оба, поэтому всё, что подключается к Claude или ChatGPT, заработает здесь без дополнительных настроек. Выберите клиент ниже для инструкции.

Anthropic API · /v1/messagesOpenAI API · /v1/chat/completions

Любой агент или инструмент с поддержкой Anthropic или OpenAI API также работает здесь — не только те, что перечислены выше.

Медиа

Генерация медиа особенно хороша для задач социальной инженерии: клонирование голоса для учений по вишингу, изображения и видео для фишинговых предлогов и обучения распознаванию дипфейков, синтетические лица для sock-puppet OSINT-персон. Всё работает через ComfyUI — визуальный редактор воркфлоу — на выделенном GPU, по независимым наборам: Изображения, Видео и Голос. Каждый набор работает на своём GPU, их можно запускать одновременно; оплата посекундная, пока GPU активен. В каждом наборе есть готовые примеры воркфлоу в панели Templates ComfyUI.

Эти модели не ограничены безопасностью — используйте их для чего угодно.

Видеоинструкция

Общий обзор того, что общего у всех медиа-бандлов:

Видео для остальных бандлов добавляются постепенно — работа уже идёт.

Тарификация

Генерация медиа (видео, голос, изображения) работает иначе: каждой сессии выделяется отдельный GPU, который обрабатывает только один запрос одновременно. Поскольку GPU не делится между пользователями, стоимость тоже не делится — вы оплачиваете полную стоимость GPU на время его работы.

Бандлы и их модели

Каждый бандл — это один GPU-под со своими моделями и готовыми к запуску воркфлоу. Вы арендуете бандл, а не отдельную модель.

Видео

MiniMax H3БандлЧастично снята

Генерирует видео с родным стереозвуком — реплики, звуковые эффекты и музыка создаются вместе с картинкой за один проход, а не накладываются потом. Работает от текста, изображения или референсов: зафиксируйте персонажа, стиль, движение, движение камеры или голос по 9 изображениям, 3 видео и 3 аудиоклипам.

Модерация MiniMax работает на их облачном API и не входит в открытые веса, поэтому здесь ваши запросы ничем не фильтруются — но что именно базовая модель обучена отклонять, нигде не описано и нами не проверено.

Воркфлоу ComfyUI: MiniMaxDirector

Готовые воркфлоу

minimaxh3-director
MiniMaxDirector Lay out shots on the Director node's timeline; it compiles them into the single structured prompt MiniMax H3 reads, and snaps the clip to a length H3 accepts (length % 17 == 5 at 24 fps). • prompt and report outputs show exactly what was built and what the linter thinks. • Drop an image on a shot with Add Image; it becomes <Picture 1> automatically. • Models: the H3 bundle (ref2va unet, Qwen3-VL text encoder, video + audio VAEs). The three prompt buttons
ButtonTrackMakes
Add Video PromptMAINwhat happens on screen
Add Sound PromptAUDIOwhat is heard -- H3 generates it, no file
Add Camera PromptCAMERAhow the camera moves
Add Image / Add Audio / Add Video attach a real file instead, and the prose is given a <Picture n> / <Audio n> / <Video n> token pointing at it. With a block selected that carries no file yet, the file lands on that block; otherwise it gets a block of its own. Presets, copying, and the keyboard Clear empties the timeline: every block, the global prompt and the music. The preset dropdown replaces it with a worked example -- a talking avatar, a three-shot scene, a two-hander, a reference shot. Each one compiles and runs as it stands, so it is a starting point rather than a form. One undo step puts your own work back.
KeyDoes
Deleteremove the selected blocks
Ssplit them at the playhead
Cmd/Ctrl+C · Cmd/Ctrl+Vcopy the selection, paste it at the playhead keeping its spacing
Cmd/Ctrl+Dduplicate in place at the playhead
Cmd/Ctrl+Zundo
The playhead The red line is where every Add button puts its block -- click the empty part of a track to move it, or drag the scrubber. If it is standing inside a block there is no room, so the new one goes on the end instead. • Dragging a block or its edge snaps to the playhead and to the edge of every other block, on any track, within a few pixels -- so a cue can start exactly where a shot does. • S cuts the selected blocks in two at the playhead. The second half keeps the prose and drops any attached file, so the same picture is never in the prompt twice. • Zooming with + / - recentres the view on it. The clip settings, left to right
FieldWhat it does
durationLength of the whole piece, in frames. Type it here; the arrows step 17 at a time so you stay on H3's lattice. Zero or empty means the clip follows its content.
= ... sThe same length in seconds. Read-only -- see below.
frame rateAlways 24. H3 has no other rate, so this is shown, never chosen.
width / heightOutput resolution, in multiples of 32. Mirrors of the node's own widgets.
resizeHow reference images are fitted. match scales them to the output size; max keeps them larger, which holds a face or a logo together better and costs more time.
renders N framesWhat will actually be generated, after rounding up to 17n+5. If it differs from what you typed, the rounding is shown.
The three tabs The panel under the toolbar shows one of three things, and remembers which one across a reload: TIMELINE (the tracks and the selected block's fields), CAST (one card per person, with the count on the tab) and GLOBAL (the two clip-wide prompt boxes). The segment panel, under the timeline Select a block first -- with nothing selected there is nothing to edit and the fields are not on screen. Everything here edits that block.
FieldWhat it does
SEGMENT PROMPTWhat happens in this block. On MAIN it becomes the shot's sentence; on AUDIO the sound; on CAMERA a note added to the move.
start / end / lengthThe block's span in frames. Editing end moves the right edge and leaves the start alone -- the same edit as dragging the right grip.
line / faces / how / languageMAIN blocks only. One row per spoken line, + line for another -- see below.
cameraCAMERA blocks only. The move, from the vocabulary below.
describes / used as / keep fileBlocks carrying a file only. See the next section.
detach mediaRemoves the file, keeps the block and its prose.
The GLOBAL tab holds the two fields set once for the whole piece:
FieldWhat it does
GLOBAL PROMPTStyle and scene constants for the whole clip. Compiles into the opening of Shot 1.
GLOBAL MUSICScore only the audience hears, as instrumentation, tempo and dynamics -- not mood words. Empty compiles to non_diegetic_music: N/A.
Why seconds are read-only. Every greyed = ... s box is a reading, not an input. A second is 24 frames wide, so a block typed as 1.08 came back as 26 frames and was shown as 1.08 again -- the number actually set was never on screen. The clip is written in seconds and cut in frames, and only one of those can be the field you edit. Dialogue H3 makes the voice and the picture in one pass, and the guide's form for it is exact: The young woman with a quiet, breathy voice (S1) says: <d>[English] I get off at the next station.</d> The editor writes it for you, and splits it the way the work actually splits: **who the people are is written once in the CAST tab, and a block only says who talks and what they say**. CAST is one card per person:
FieldWhat it does
name themA short name, yours, so the faces on a dialogue row are readable.
fromWhich file on the timeline this person is drawn from. That binding is what makes a face and a voice one person; the card then shows the <Subject n> badge the prompt will use.
keep themHow much of the person survives, compiled as subject_retention. Not the block's keep file: the photo may be fully_preserved while the face taken out of it is an attribute_transfer onto somebody else.
what they look likeFor a card with a file. Becomes their line in subject_definitions.
how they soundAge, gender, pitch, timbre, accent, on screen or off. H3 fixes the voice from this, so an empty one is a voice nobody chose and the linter says so.
+ character adds a card; they speak switches dialogue off for the whole clip -- every row and every <d> at once, cards kept. Describing the same speaker two different ways in two shots used to be possible; to the model that reads as two people wearing one label. On the block:
FieldWhat it does
lineThe words themselves, sent verbatim -- never translated, punctuation kept.
facesWho says it: click a face from the cast. Two lit on one row is the guide's (S1,S2) -- the same words spoken by both at the same instant.
howHow it is performed. Becomes the verb: says, whispers, shouts, answers -- free text, used as written.
languageNames the language of the words; it does not translate them.
+ line adds another row, so one block can hold a conversation: a line each, spoken in turn, compiled as one <d> apiece. × removes a row. Three readings, kept apart on purpose: a chorus is one row with two faces, a conversation is two rows, and an argument -- overlapping speech with no agreed words -- is neither. Write that one in the segment prompt and put the sound in an AUDIO cue; there is nothing for H3 to quote. Attached files: used as, describes and keep file used as -- what the file is for. It decides the task type the summary opens with, and the guide wants every relationship named:
Used asTask type it produces
referencereference generation -- guidance for a character, scene, style or camera move
first frame / keyframe / last framekeyframe completion -- the image is a concrete frame of the target video, and retention_analysis says which
continue fromvideo continuation
editvideo editing
An attached audio adds audio reuse or audio reference depending on its keep. Several at once combine: [keyframe completion + video continuation + audio reuse]. A segment holding a real file gets two extra fields, and they are what switch the prompt into H3's full-reference format -- six sections instead of three. describes -- what the file is, and what must stay the same. One noun phrase, written for someone who cannot see it. It becomes that file's line in subject_definitions: <Picture 1> is the raccoon, grey and black with a striped tail and a masked face. Leave it empty and the linter warns: an unnamed reference is one H3 has to guess at. keep file -- how much of the file survives into the video. One per file, always; it also sits on the block itself, bottom right. A cast card drawn from this file carries its own keep them for the person, which may differ. The sentence it produces lands in retention_analysis, as <Picture 1> (appears in [Shot 2]): fully_preserved - the raccoon, ...
ValueMeans
fully_preservedcopy it -- same subject, same look, unchanged
partially_preservedkeep the subject, let pose, angle or lighting change
attribute_transfertake one trait -- a face, a colour, a texture -- onto something else
weak_referenceloose inspiration only: style, palette, grade, nothing literal
These are H3's own words, not ours. It reads them as instructions, so a wrong one is worse than a vague describes: fully_preserved on a style reference asks the model to reproduce the whole frame. Subjects live on cast cards. Point a card's from at a file and the person becomes a <Subject n> of their own, tracked apart from the picture they came from: <Subject 1> is the man's face, from <Picture 2>. That separation is what a face swap needs. The picture stays a weak_reference -- you do not want the whole frame back -- while the card's keep them is attribute_transfer onto the person in another shot. The shot then mentions <Subject 1> rather than <Picture 2>, because naming both asks for two different things at once. Camera modes Each mode contributes one sentence to the shot it covers. H3 reads prose, not enum values, so this is the exact wording it receives:
ModeSentence sent to the model
staticThe camera holds a static shot.
dolly_inThe camera pushes in with small amplitude at slow speed.
dolly_outThe camera pulls out with small amplitude at slow speed.
pan_leftThe camera pans left.
pan_rightThe camera pans right.
tilt_upThe camera tilts up.
tilt_downThe camera tilts down.
orbitThe camera moves in an arc around the subject.
handheldThe camera shakes slightly.
crash_zoomThe camera zooms in with large amplitude at fast speed.
A note typed into a camera segment is appended after the sentence, so dolly_in + "closing on the apple" becomes *The camera pushes in with small amplitude at slow speed. closing on the apple* -- so write the note as a continuation, not as a sentence of its own. Camera work is its own block rather than folded into the shot line, because a move can straddle a cut -- and merging them would silently pick a side.
minimaxh3-director-turbo
MiniMaxDirector — Turbo The director workflow with larryvrh's Turbo LoRA between the model loader and the sampler: 6 steps instead of 20, so a draft costs about a fifth of the GPU time. Same timeline, same prompt, same output. • Stay in 4-8 steps. 6-8 looks best; past 8 it stops helping and starts over-sharpening. • Leave LoRA strength at 1.0 and the scheduler on simple. • Preview weights: audio and fast motion are the weak spots. For a final take, use minimaxh3-director. --- Lay out shots on the Director node's timeline; it compiles them into the single structured prompt MiniMax H3 reads, and snaps the clip to a length H3 accepts (length % 17 == 5 at 24 fps). • prompt and report outputs show exactly what was built and what the linter thinks. • Drop an image on a shot with Add Image; it becomes <Picture 1> automatically. • Models: the H3 bundle (ref2va unet, Qwen3-VL text encoder, video + audio VAEs). The three prompt buttons
ButtonTrackMakes
Add Video PromptMAINwhat happens on screen
Add Sound PromptAUDIOwhat is heard -- H3 generates it, no file
Add Camera PromptCAMERAhow the camera moves
Add Image / Add Audio / Add Video attach a real file instead, and the prose is given a <Picture n> / <Audio n> / <Video n> token pointing at it. With a block selected that carries no file yet, the file lands on that block; otherwise it gets a block of its own. Presets, copying, and the keyboard Clear empties the timeline: every block, the global prompt and the music. The preset dropdown replaces it with a worked example -- a talking avatar, a three-shot scene, a two-hander, a reference shot. Each one compiles and runs as it stands, so it is a starting point rather than a form. One undo step puts your own work back.
KeyDoes
Deleteremove the selected blocks
Ssplit them at the playhead
Cmd/Ctrl+C · Cmd/Ctrl+Vcopy the selection, paste it at the playhead keeping its spacing
Cmd/Ctrl+Dduplicate in place at the playhead
Cmd/Ctrl+Zundo
The playhead The red line is where every Add button puts its block -- click the empty part of a track to move it, or drag the scrubber. If it is standing inside a block there is no room, so the new one goes on the end instead. • Dragging a block or its edge snaps to the playhead and to the edge of every other block, on any track, within a few pixels -- so a cue can start exactly where a shot does. • S cuts the selected blocks in two at the playhead. The second half keeps the prose and drops any attached file, so the same picture is never in the prompt twice. • Zooming with + / - recentres the view on it. The clip settings, left to right
FieldWhat it does
durationLength of the whole piece, in frames. Type it here; the arrows step 17 at a time so you stay on H3's lattice. Zero or empty means the clip follows its content.
= ... sThe same length in seconds. Read-only -- see below.
frame rateAlways 24. H3 has no other rate, so this is shown, never chosen.
width / heightOutput resolution, in multiples of 32. Mirrors of the node's own widgets.
resizeHow reference images are fitted. match scales them to the output size; max keeps them larger, which holds a face or a logo together better and costs more time.
renders N framesWhat will actually be generated, after rounding up to 17n+5. If it differs from what you typed, the rounding is shown.
The three tabs The panel under the toolbar shows one of three things, and remembers which one across a reload: TIMELINE (the tracks and the selected block's fields), CAST (one card per person, with the count on the tab) and GLOBAL (the two clip-wide prompt boxes). The segment panel, under the timeline Select a block first -- with nothing selected there is nothing to edit and the fields are not on screen. Everything here edits that block.
FieldWhat it does
SEGMENT PROMPTWhat happens in this block. On MAIN it becomes the shot's sentence; on AUDIO the sound; on CAMERA a note added to the move.
start / end / lengthThe block's span in frames. Editing end moves the right edge and leaves the start alone -- the same edit as dragging the right grip.
line / faces / how / languageMAIN blocks only. One row per spoken line, + line for another -- see below.
cameraCAMERA blocks only. The move, from the vocabulary below.
describes / used as / keep fileBlocks carrying a file only. See the next section.
detach mediaRemoves the file, keeps the block and its prose.
The GLOBAL tab holds the two fields set once for the whole piece:
FieldWhat it does
GLOBAL PROMPTStyle and scene constants for the whole clip. Compiles into the opening of Shot 1.
GLOBAL MUSICScore only the audience hears, as instrumentation, tempo and dynamics -- not mood words. Empty compiles to non_diegetic_music: N/A.
Why seconds are read-only. Every greyed = ... s box is a reading, not an input. A second is 24 frames wide, so a block typed as 1.08 came back as 26 frames and was shown as 1.08 again -- the number actually set was never on screen. The clip is written in seconds and cut in frames, and only one of those can be the field you edit. Dialogue H3 makes the voice and the picture in one pass, and the guide's form for it is exact: The young woman with a quiet, breathy voice (S1) says: <d>[English] I get off at the next station.</d> The editor writes it for you, and splits it the way the work actually splits: **who the people are is written once in the CAST tab, and a block only says who talks and what they say**. CAST is one card per person:
FieldWhat it does
name themA short name, yours, so the faces on a dialogue row are readable.
fromWhich file on the timeline this person is drawn from. That binding is what makes a face and a voice one person; the card then shows the <Subject n> badge the prompt will use.
keep themHow much of the person survives, compiled as subject_retention. Not the block's keep file: the photo may be fully_preserved while the face taken out of it is an attribute_transfer onto somebody else.
what they look likeFor a card with a file. Becomes their line in subject_definitions.
how they soundAge, gender, pitch, timbre, accent, on screen or off. H3 fixes the voice from this, so an empty one is a voice nobody chose and the linter says so.
+ character adds a card; they speak switches dialogue off for the whole clip -- every row and every <d> at once, cards kept. Describing the same speaker two different ways in two shots used to be possible; to the model that reads as two people wearing one label. On the block:
FieldWhat it does
lineThe words themselves, sent verbatim -- never translated, punctuation kept.
facesWho says it: click a face from the cast. Two lit on one row is the guide's (S1,S2) -- the same words spoken by both at the same instant.
howHow it is performed. Becomes the verb: says, whispers, shouts, answers -- free text, used as written.
languageNames the language of the words; it does not translate them.
+ line adds another row, so one block can hold a conversation: a line each, spoken in turn, compiled as one <d> apiece. × removes a row. Three readings, kept apart on purpose: a chorus is one row with two faces, a conversation is two rows, and an argument -- overlapping speech with no agreed words -- is neither. Write that one in the segment prompt and put the sound in an AUDIO cue; there is nothing for H3 to quote. Attached files: used as, describes and keep file used as -- what the file is for. It decides the task type the summary opens with, and the guide wants every relationship named:
Used asTask type it produces
referencereference generation -- guidance for a character, scene, style or camera move
first frame / keyframe / last framekeyframe completion -- the image is a concrete frame of the target video, and retention_analysis says which
continue fromvideo continuation
editvideo editing
An attached audio adds audio reuse or audio reference depending on its keep. Several at once combine: [keyframe completion + video continuation + audio reuse]. A segment holding a real file gets two extra fields, and they are what switch the prompt into H3's full-reference format -- six sections instead of three. describes -- what the file is, and what must stay the same. One noun phrase, written for someone who cannot see it. It becomes that file's line in subject_definitions: <Picture 1> is the raccoon, grey and black with a striped tail and a masked face. Leave it empty and the linter warns: an unnamed reference is one H3 has to guess at. keep file -- how much of the file survives into the video. One per file, always; it also sits on the block itself, bottom right. A cast card drawn from this file carries its own keep them for the person, which may differ. The sentence it produces lands in retention_analysis, as <Picture 1> (appears in [Shot 2]): fully_preserved - the raccoon, ...
ValueMeans
fully_preservedcopy it -- same subject, same look, unchanged
partially_preservedkeep the subject, let pose, angle or lighting change
attribute_transfertake one trait -- a face, a colour, a texture -- onto something else
weak_referenceloose inspiration only: style, palette, grade, nothing literal
These are H3's own words, not ours. It reads them as instructions, so a wrong one is worse than a vague describes: fully_preserved on a style reference asks the model to reproduce the whole frame. Subjects live on cast cards. Point a card's from at a file and the person becomes a <Subject n> of their own, tracked apart from the picture they came from: <Subject 1> is the man's face, from <Picture 2>. That separation is what a face swap needs. The picture stays a weak_reference -- you do not want the whole frame back -- while the card's keep them is attribute_transfer onto the person in another shot. The shot then mentions <Subject 1> rather than <Picture 2>, because naming both asks for two different things at once. Camera modes Each mode contributes one sentence to the shot it covers. H3 reads prose, not enum values, so this is the exact wording it receives:
ModeSentence sent to the model
staticThe camera holds a static shot.
dolly_inThe camera pushes in with small amplitude at slow speed.
dolly_outThe camera pulls out with small amplitude at slow speed.
pan_leftThe camera pans left.
pan_rightThe camera pans right.
tilt_upThe camera tilts up.
tilt_downThe camera tilts down.
orbitThe camera moves in an arc around the subject.
handheldThe camera shakes slightly.
crash_zoomThe camera zooms in with large amplitude at fast speed.
A note typed into a camera segment is appended after the sentence, so dolly_in + "closing on the apple" becomes *The camera pushes in with small amplitude at slow speed. closing on the apple* -- so write the note as a continuation, not as a sentence of its own. Camera work is its own block rather than folded into the shot line, because a move can straddle a cut -- and merging them would silently pick a side.
minimaxh3-i2v
MiniMax H3 MiniMax H3 is MiniMax's general-purpose, omni-modal generation model. It jointly understands text, image, video, and audio, and generates video with native stereo audio: voice, sound effects, and music are modeled jointly in a single forward pass, not layered on afterward. Output is up to 2K resolution, 24fps, and up to about 15 seconds. About this workflow This template runs the Image to Video task (MiniMaxH3ImageToVideo node), which covers both: • t2va (text-to-video), when no images are connected • fl2va (first/last-frame image-to-video), when first_frame and/or last_frame are connected Key inputsfirst_frame / last_frame: optional keyframes; the model generates the motion between them • prompt: describe the shots, motion, and the accompanying audio (dialogue, SFX, music) in one block • width / height: set via Resolution Selector. H3's native canvas is a 768px short edge, capped at 768x1344 pixels, rounded to a multiple of 32 • duration (seconds): converted to a valid frame length by the Math Expression node, snapping up to the model's 17-frame-per-block (17k+5) grid at 24fps
minimaxh3-r2v
MiniMax H3 MiniMax H3 is MiniMax's general-purpose, omni-modal generation model. It jointly understands text, image, video, and audio, and generates video with native stereo audio: voice, sound effects, and music are modeled jointly in a single forward pass, not layered on afterward. Output is up to 2K resolution, 24fps, and up to about 15 seconds. ComfyUI links • ComfyUI#15224🤗 Comfy-Org/MiniMax-H3 About this workflow This template runs the reference-to-video (ref2va) task using the MiniMaxH3ReferenceToVideo node. It takes any mix of reference images, videos, and standalone audio, and weaves them into the generation to lock in a character's identity, a style, a motion, a camera move, or a voice. Key inputsref_images / ref_videos / ref_video_audios / ref_audios: up to 9 reference images, 3 reference videos (each may carry its own paired soundtrack), and 3 standalone reference audio clips • prompt: reference the inputs by tag, in the exact order they were connected, for example <Picture 1>, <Video 1>, <Audio 1>, then describe the target scene, motion, and audio • ref_image_size: match scales references down to the generation's resolution (faster); max keeps up to a 2048px short edge for stronger identity fidelity, at the cost of speed since reference tokens ride along every sampling step • width / height: set via Resolution Selector. • duration (seconds): converted to a valid frame length by the Math Expression node Sampling and decode • Sampler: res_multistep. beta or normal scheduler tends to outperform simple for reference-heavy prompts like this one • The sampler's joint audio+video LATENT output feeds directly into both VAEDecode (video, minimax_h3_video_vae_fp16) and VAEDecodeAudio (audio, minimax_h3_audio_vae_fp32); each decode node automatically pulls its own half out of the packed latent. CreateVideo then muxes the two into a single MP4 with synced sound • The diffusion model here is minimax_h3_ref2va_pruned_int8_convrot.safetensors, a different set of weights from the fl2va model used by the t2v/i2v templates Ref2va's output is very sensitive to prompt wording; matching the reference tags precisely and being explicit about which reference drives which part of the shot tends to work best.
minimaxh3-t2v
MiniMax H3 MiniMax H3 is MiniMax's general-purpose, omni-modal generation model. It jointly understands text, image, video, and audio, and generates video with native stereo audio: voice, sound effects, and music are modeled jointly in a single forward pass, not layered on afterward. Output is up to 2K resolution, 24fps, and up to about 15 seconds. ComfyUI links • ComfyUI#15224🤗 Comfy-Org/MiniMax-H3 About this workflow Key inputsprompt: describe the shots, camera moves, and the accompanying audio (dialogue, SFX, music) in one block • width / height: set via Resolution Selector. H3's native canvas is a 768px short edge, capped at 768x1344 pixels, rounded to a multiple of 32 • duration (seconds): converted to a valid frame length by the Math Expression node, snapping up to the model's 17-frame-per-block (17k+5) grid at 24fps

Видеоинструкция

Обзор бандла MiniMax H3:

Модели в этом бандле

Sulphur-2БандлРасцензурирована

Генерирует видеоклипы из текстового запроса, изображения, аудио или видео с помощью LTX Director 2.0, плюс липсинк-дубляж — дайте готовый клип и свою речь, и он перегенерирует кадр так, чтобы губы совпадали со словами. Sulphur-2 — нецензурированная версия LTX 2.3.

Воркфлоу ComfyUI: LTXDirector

Готовые воркфлоу

sulphur2-lipdub
How to use Lip-sync an existing clip to speech you supply. 1. Load Video (red) — upload the clip whose mouth should move. 2. Positive Prompt (red) — describe the shot and write what the person says. 3. IC-LoRA (orange) — ltx-2.3-22b-ic-lora-dubit-0.9, strength 1.0 is the tested value. 4. Press Run — stage 1 generates at low resolution, stage 2 refines it. The audio comes from the loaded video. To dub a different voice, mux your voiceover onto the clip before uploading it.
sulphur2-ltx-director-2

Слишком большой граф, чтобы описать его здесь — смотрите разбор: LTXDirector

Видеоинструкция

Обзор бандла Sulphur-2:

Модели в этом бандле

  • Sulphur 2

    Генерирует видеоклипы из текстового запроса, изображения, аудио или видео с помощью LTX Director 2.0. Нецензурированная версия LTX 2.3.

  • LTX DubIt IC-LoRA

    Дополнение для липсинка к LTX 2.3 / Sulphur-2. Перегенерирует клип так, чтобы губы совпадали с вашей речью — дубляж на другой язык или замена сказанного.

SCAIL-2БандлИзначально без цензуры

Анимация и замена персонажа — управляйте референсным персонажем видео движения; люди маскируются автоматически (SAM 3.1), ручной риггинг не нужен. Построено на Wan2.1 14B. Результат без звука — добавьте его потом в видеоредакторе.

Готовые воркфлоу

scail2-animation
SCAIL-2 — Character Animation Take the motion out of one video and put your own character into it. The driving video's background is not kept — the scene is generated fresh around your character. Fill these in 1. Load Video — your driving clip. Only the movement is used, never the appearance. 2. Load Image — the character to animate (person, mascot, drawing). One character only. The output video is sized to this image, so a portrait photo gives a portrait video. 3. Run SAM3 Video Track — upper node, fed by Load Video. Its text box names the subject to copy motion from. One person in the clip: leave human. Several people: pick one, e.g. man, woman in a red dress — otherwise SCAIL-2 gets two motion tracks and one character, and the result breaks. 4. Run SAM3 Video Track — lower node, fed by Load Image. Leave at human for a person. For a non-human character use its noun, e.g. dog, robot. 5. CLIP Text Encode (Positive Prompt) — describe your character and the action, e.g. a short-haired man in a striped shirt, hands on his hips, full body. Add full body if you want legs in frame. 6. Press Run. Check the masks first The two Preview Image nodes show the tracking masks. Exactly one subject should be coloured in each. If extra subjects light up, make the prompt in step 3 or 4 more specific, or raise detection_thres above 0.50. Do this before any long render — a bad mask wastes the whole run. Length Default is 81 frames at 16 fps — about 5 seconds, taken from the start of your clip. To go longer, raise these two together and keep them equal: • Load Video → frame_load_capWan SCAIL To Video → length 161 ≈ 10 s, 321 ≈ 20 s. SCAIL-2 is trained at 81 frames, so longer runs cost more VRAM and the character may drift. To start somewhere other than the beginning, set Load Video → skip_first_frames (in frames, at 16 fps — 160 skips 10 s). Sound This workflow does not do sound. The render is always silent — SCAIL-2 generates picture only, and nothing in the graph carries audio through to the output. Add the soundtrack afterwards in a video editor, using your original clip as the audio source. Trying to attach it here is not worth it: the render is a short slice of your clip, so the audio would not line up anyway. Video format The upload button accepts .mp4, .webm, .mkv and .gif. H.264 MP4 is the safe choice. If a clip is rejected with "Invalid video file", re-encode it before uploading. A common cause is a movie-rip audio track (AC-3) that the pod's ffmpeg cannot decode: ``` ffmpeg -i input.mp4 -c:v libx264 -pix_fmt yuv420p -an clean.mp4 ``` -an drops the audio, which this workflow does not use anyway. Keep the clip's own resolution — it is resized internally, so a huge 4K source only costs upload time. Leave alone unless you know why • Negative Prompt — a fixed quality filter, not something to describe your video with. • KSamplersteps 6, cfg 1.0, euler / simple. These are tuned for the distilled LoRA; raising steps or cfg makes it worse, not better. • Create SCAIL-2 Colored Mask → replacement_modefalse here on purpose. Setting it true switches to the replacement behaviour (keeps the original background), which is what the scail2-replacement workflow already does. • Every model this workflow needs is already installed on the pod.
scail2-animation-multi-char
SCAIL-2 — Character Animation (two characters) Same graph as the single-character animation, driven by a clip with two moving subjects. Only the inputs and the prompts differ. There is only ONE Load Image — and that is correct There is no second image node, and you should not add one. Both characters come from a single reference image that already contains both of them. SAM3 finds both figures inside that one picture and hands SCAIL-2 two separate coloured regions. It has to be one real photograph of both subjects together — one background, one camera, one light. The output video is built from this frame, so whatever you hand over becomes the scene. Do not glue two separate photos side by side. A collage keeps both backgrounds and the seam between them, and the render comes out looking like two videos in one frame. If all you have is a separate photo of each character, this workflow cannot merge them — run the single-character scail2-animation workflow on each one instead. Fill these in 1. Load Image — one picture holding both characters. The output video is sized to this image, so a wide image gives a wide video. Both characters should be clearly separated and, if you want limbs animated, fully in frame. 2. Load Video — a driving clip with two moving subjects. Their motion is copied; their appearance is not. 3. Run SAM3 Video Track — upper node, fed by Load Video. Leave at human when the clip has exactly two people and you want both. Use a narrower word only if there are extra people to exclude. 4. Run SAM3 Video Track — lower node, fed by Load Image. Leave at human for two people. For non-human characters use a word that matches both, e.g. mascot or character — a word matching only one of them will leave the other untracked. 5. CLIP Text Encode (Positive Prompt) — describe both characters and what they do together, e.g. a black dog mascot character and a green-and-cream bird mascot character holding hands and dancing on a white stage. 6. Press Run. Check the masks first — this matters most here The two Preview Image nodes show the tracking masks. You need two differently coloured regions in each: two in the driving mask, two in the reference mask. If either shows one region, or three, fix the prompt in step 3 or 4 before rendering. Who maps to whom is decided by Create SCAIL-2 Colored Mask → sort_by (area by default — biggest region first in both masks). If the wrong character gets the wrong motion, that pairing is the reason. Length Default is 81 frames at 16 fps — about 5 seconds, from the start of the clip. To go longer, raise these together and keep them equal: • Load Video → frame_load_capWan SCAIL To Video → length 161 ≈ 10 s, 321 ≈ 20 s. SCAIL-2 is trained at 81 frames — longer runs cost more VRAM and drift more, and two characters drift faster than one. To start later in the clip, use Load Video → skip_first_frames (frames at 16 fps — 160 skips 10 s). Sound This workflow does not do sound. The render is always silent — SCAIL-2 generates picture only, and nothing in the graph carries audio through to the output. Add the soundtrack afterwards in a video editor, using your original clip as the audio source. Trying to attach it here is not worth it: the render is a short slice of your clip, so the audio would not line up anyway. Video format The upload button accepts .mp4, .webm, .mkv and .gif. H.264 MP4 is the safe choice. If a clip is rejected with "Invalid video file", re-encode it before uploading. A common cause is a movie-rip audio track (AC-3) that the pod's ffmpeg cannot decode: ``` ffmpeg -i input.mp4 -c:v libx264 -pix_fmt yuv420p -an clean.mp4 ``` -an drops the audio, which this workflow does not use anyway. Keep the clip's own resolution — it is resized internally, so a huge 4K source only costs upload time. Leave alone unless you know why • Negative Prompt — a fixed quality filter, not a place to describe your video. • KSamplersteps 6, cfg 1.0, euler / simple, tuned for the distilled LoRA. Raising steps or cfg makes it worse. • replacement_modefalse here on purpose; the background is meant to be generated fresh. To keep an original background instead, use the scail2-replacement workflow. • Every model this workflow needs is already installed on the pod.
scail2-replacement
SCAIL-2 — Character Replacement Swap one person in your video for your own character. The original scene, background and everyone else stay exactly as they are. Steps 1. Load Video — upload your clip. The output is sized to this video, not to your image. 2. Load Image — upload the character who takes their place. A clear, full-body photo works best. 3. Run SAM3 Video Track (the upper one, fed by Load Video) — its text box says who gets replaced. One person in the clip: leave human. Several people: name the one you want, e.g. man or woman in a red dress. 4. Run SAM3 Video Track (the lower one, fed by Load Image) — leave it at human. 5. Positive Prompt — describe your new character inside the video's setting, e.g. bearded man in a grey suit sitting at the desk. 6. Press Run. Check the masks before a long render The two Preview Image nodes show the tracking masks. In the driving mask, only the person being replaced should be coloured. If extra people light up, make the prompt in step 3 more specific, or raise detection_thres above 0.50. Length A default run is 81 frames at 16 fps — about 5 seconds, taken from the start of your clip. For longer output raise these two together and keep them equal: • Load Video → frame_load_capWan SCAIL To Video → length 161 ≈ 10 seconds, 321 ≈ 20 seconds. SCAIL-2 is trained at 81 frames, so longer runs cost more VRAM and the character may drift. To start later in the clip, set Load Video → skip_first_frames (frames at 16 fps — 160 skips 10 s). Running the same clip in 81-frame slices at 0, 81, 162, 243 and joining the files is the alternative to one long render; expect a visible seam at each join. Sound This workflow does not do sound. The render is always silent — SCAIL-2 generates picture only, and nothing in the graph carries audio through to the output. Add the soundtrack afterwards in a video editor, using your original clip as the audio source. Trying to attach it here is not worth it: the render is a short slice of your clip, so the audio would not line up anyway. Video format The upload button accepts .mp4, .webm, .mkv and .gif. H.264 MP4 is the safe choice. If a clip is rejected with "Invalid video file", re-encode it before uploading. A common cause is a movie-rip audio track (AC-3) that the pod's ffmpeg cannot decode: ``` ffmpeg -i input.mp4 -c:v libx264 -pix_fmt yuv420p -an clean.mp4 ``` -an drops the audio. Since this workflow keeps the original scene, re-encode from the highest-quality source you have — the output resolution is taken from this clip. Leave alone unless you know why • Replace mode is already on — the toggles on Create SCAIL-2 Colored Mask and Wan SCAIL To Video are true. Turning them off gives the animation behaviour instead (original background discarded). • Negative Prompt — a fixed quality filter, not a place to describe your video. • KSamplersteps 6, cfg 1.0, euler / simple, tuned for the distilled LoRA. Raising steps or cfg makes results worse, not better. • Output size comes from the video via Get Image from Batch, not from your reference image — so a portrait clip stays portrait no matter what you upload. • Every model this workflow needs is already installed on the pod.

Модели в этом бандле

Изображения

FLUX.2 kleinБандлРасцензурирована

FLUX.2 klein 9B (без цензуры) — быстрый текст → изображение + многореференсное редактирование, плюс RefControl контроль глубины/структуры. Эстетический тюн «True V3» (Q8 GGUF) с абилитерированным текстовым энкодером. Личный фаворит: не дайте низкой цене себя обмануть — многореференсное редактирование и RefControl не уступают более дорогим наборам, всё на одной GPU с 24GB.

Готовые воркфлоу

flux2-klein-controlnet
ControlNet (depth) — klein 1. Load Image depth (red) — drop the photo whose structure you want to copy. A depth map is auto-extracted (Depth Anything V2). 2. Load Image reference (red) — drop the subject/style reference to place into that structure. 3. Prompt — describe the result (default refcontrol). 4. RefControl strength (yellow) — the Lora node. Higher = follow the depth structure more strictly. 5. Press Run — result in Save Image. Powered by the RefControl depth LoRA for klein 9B. The first depth run downloads the preprocessor weights (~1.3GB) once.
flux2-klein-darkbeast-faceswap
Face Swap — DarkBeast (klein 9B) 1. Target photo (red, top-left) — the picture whose face gets replaced. Ships with example.png so the graph runs out of the box. 2. Face to swap in (red, bottom-left) — the source face/identity to paste on. 3. Prompt — inside the Face Swap node; keep it simple, e.g. swap the face of the person with the reference face, keep pose, expression and lighting. 4. Steps = 5, CFG = 1 — DarkBeast is a distilled BFS model tuned for 5 steps / CFG 1. Do not raise them — higher values make it worse, not better. 5. Color Match (orange) regrades the result to the target's lighting so the swap blends in. Lower strength (or 0) to disable. 6. Press Run — the result is saved in Save Image. DarkBeast Klein 9b V2 BFS — face-swap-specialized klein 9B (safetensors, loaded via UNETLoader).
flux2-klein-edit
Image Edit — klein (multi-reference) 1. Load Image (red, group image 1) — drop the reference picture. It ships with example.png so the graph runs out of the box. 2. Prompt — describe the edit inside the Image Edit node. 3. Reference images toggle — switch image 2image 10 on to use more reference pictures. Each toggle enables one Load Image group on the left. 4. Color Match (orange) — regrades the result to image 1's lighting/colors so edits blend in. Lower strength (or 0) to disable. 5. Press Run — the result is saved in Save Image. FLUX.2 klein 9B (uncensored) — fast text → image + up to 10 reference images.
flux2-klein-text-to-image
How to use 1. Prompt (orange) — type what you want to generate. 2. Two variants: Standard and Distilled (faster). Enable one and bypass the other with Ctrl-B. 3. Press Run — the result is saved in Save Image. FLUX.2 klein 9B (uncensored) — fast text → image.

Видеоинструкция

Обзор бандла FLUX.2 klein:

Модели в этом бандле

  • FLUX.2 klein 9B (True V3, uncensored)

    Быстрая 9B FLUX.2 klein (без цензуры) — эстетический тюн «True V3» с абилитированным текстовым энкодером. Текст-в-изображение, многоракурсное редактирование и RefControl (контроль глубины/структуры), помещается на одну видеокарту 24GB.

  • DarkBeast Klein 9b V2 BFS (face-swap, uncensored)

    DarkBeast Klein 9b V2 BFS — специализированный на замене лиц тюн FLUX.2 klein 9B, без цензуры. Дистиллирован под 5 шагов при CFG 1 (технология Best Face Swap): подайте целевое фото и референс лица — модель переносит личность, сохраняя позу, выражение и освещение. Идёт рядом с тюном «True V3» — просто выберите его в загрузчике. fp8 на видеокарте 24GB, bf16 на 32GB.

Ideogram 4БандлЧастично освобождена

Текст-в-изображение с сильной типографикой — отлично для читаемого текста и дизайна. Визуально размещайте текст и элементы в точных позициях на холсте (контроль регионального макета).

Абилитерированный энкодер убирает отказы по запросу, но NSFW-контент был отфильтрован из обучающих данных модели, поэтому результаты в этой области всё ещё нестабильны.

Готовые воркфлоу

ideogram4-text-to-image
How to use 1. Prompt Builder (red) — type your prompt in the Description field (plain language is fine). For layout control, open the Ideogram 4 editor and drag boxes to place objects/text in regions. 2. Resolution Selector (orange) — choose aspect ratio / size. 3. Press Run — the image appears in Save Image. "Image blocked by safety filter" comes from the model's own safety training, not ComfyUI.

Модели в этом бандле

Qwen-Image-2512-Edit-2511БандлПолностью без цензуры

Содержит две модели: Qwen-Image-2512 генерирует изображения из текста, а Qwen-Image-Edit-2511 редактирует существующие изображения по запросу (до 3 входных изображений). Включает воркфлоу Multi-angle Camera — потяните за 3D-рукоятку, чтобы изменить ракурс камеры на любом фото.

Готовые воркфлоу

qwen-image-edit
How to use 1. Load Image — upload image 1 (required). Type the edit instruction in the Image Edit prompt field. 2. To combine pictures, enable Load Image 2 / 3 (right-click → Set Mode → Always) and upload. 3. Press Run — result in Save Image. Predefined example — reset every GPU start; use Workflows → Save As to keep your own copy.
qwen-image-edit-multiangle-camera
How to use 1. Load Image (red) — drop the photo whose camera angle you want to change. 2. Qwen Multiangle Camera (red) — drag the 3D handle to set the angle, or pick a preset. The prompt is built for you. 3. Press Run — the re-angled image appears in Save Image. Powered by Qwen-Image-Edit-2511 + the multi-angle camera LoRA (4-step Lightning).
qwen-style-transfer
The quality of the style transfer depends largely on the quality of the RF inversion. These settings work well, but feel free to try other values.
qwen-text-to-image
How to use 1. Text to Image (red) — type your prompt in the text field, set width / height (and seed if you want). 2. Press Run — the image appears in Save Image. Sizes: 1:1 1328×1328 · 16:9 1664×928 · 9:16 928×1664 · 4:3 1472×1104 · 3:4 1104×1472 Predefined example — reset on every GPU start. Use Workflows → Save As to keep your own copy.
qwen-upscale-4k
Upscale to 4K 1. Load Image (red) — drop any image (e.g. one you made with the Text-to-Image workflow). 2. Target size (yellow) — Scale to Total Pixels sets the working resolution. 4 MP ≈ 4K; raise/lower for your GPU. 3. Refine (yellow) — the KSampler re-renders detail at the new size. denoise ~0.35–0.45: higher = more new detail, lower = closer to the original. 4. Press Run — the upscaled image lands in Save Image. This is a single refine pass on an existing image. Generate first in the Text-to-Image workflow, then upscale here.

Модели в этом бандле

  • Qwen-Image-2512

    Генерирует изображения из текста с высокой точностью следования запросу и качественным отображением текста на изображении.

  • Qwen-Image-Edit-2511

    Редактирует существующие изображения по запросу — замена фона, добавление/удаление объектов, изменение стиля (до 3 входных изображений).

Boogu-ImageБандлВ основном освобождена

Две модели в одном бандле: Boogu Turbo для быстрого текст-в-изображение и Boogu Edit для редактирования изображений по инструкции. Сильный билингвальный рендеринг текста.

Абилитерированный энкодер убирает отказы по запросу, но у модели мягкая защита, и результаты в этой области всё ещё могут быть нестабильны.

Готовые воркфлоу

boogu-edit
How to use 1. Load Image (red) — upload the image you want to edit. 2. Instruction — double-click the red Image Edit (Boogu) subgraph and type what to change in the prompt box. 3. Size (yellow) — output matches the input by default. Bypass Resize Image/Mask to keep the original size, or raise its megapixels (e.g. 4) for higher resolution — depends on your GPU. 4. Press Run — compare input vs result in Image Compare; the result is saved in Save Image. Boogu Edit is instruction-based image editing with strong bilingual (English / 中文) text rendering.
boogu-turbo-t2i
How to use 1. Prompt — double-click the red Text to Image (Boogu Turbo) subgraph and type your description in the prompt box. 2. Resolution (yellow) — pick aspect ratio / size in Resolution Selector. 3. Press Run — the image appears in Save Image. Boogu Turbo is a fast text-to-image model with strong bilingual (English / 中文) text rendering. For instruction-based image editing, open the Boogu Edit workflow.

Модели в этом бандле

  • Boogu-Image Turbo

    Быстрая генерация текст-в-изображение за 4 шага с сильным фотореализмом и двуязычным (англ./кит.) отображением текста.

  • Boogu-Image Edit

    Редактирование изображений по инструкции — опишите изменение текстом, чтобы добавить, заменить или перерисовать объекты на изображении.

Krea-2БандлПолностью без цензуры

Быстрый фотореалистичный текст-в-изображение с разрешением до 2K, с 9 выбираемыми стилевыми LoRA для разных образов.

Готовые воркфлоу

krea2-text-to-image
How to use 1. Prompt — double-click the red Text to Image (Krea-2 Turbo) subgraph and type your description in Text String (User Prompt). 2. Resolution (orange) — pick aspect ratio / size in Resolution Selector. 3. Press Run — the image appears in Save Image. Prompt enhancement is on by default; it expands your prompt using the model's own text encoder (no extra model needed). Toggle prompt_enhance inside the subgraph to turn it off. Style LoRAs — set enable_lora? to true inside the subgraph, pick a krea2_* file in LoraLoaderModelOnly; the matching trigger word is added automatically. All 9 LoRAs are pre-installed.
LoRATrigger WordStrength
krea2_darkbrushmonochrome ink wash style1.0
krea2_dotmatrixmonochrome stippling style1.0
krea2_kidsdrawingnaive expressive sketch style1.0
krea2_neondriptextured abstract style1.0
krea2_rainywindowrainy window style1.0
krea2_retroanimepurple retro anime style1.0
krea2_softwatercolorart deco watercolor style1.0
krea2_sunsetblurethereal motion blur style1.0
krea2_vintagetarotvintage tarot style1.0

Модели в этом бандле

Голос

Общие модели

  • WhisperX

    Движок распознавания речи для голос-в-SRT (по умолчанию). Ядро Whisper плюс фонемное выравнивание для очень точного пословного тайминга субтитров и разделения по говорящим.

    Используется в: Fish Audio S2 · CosyVoice 3 · Qwen3-TTS · Chatterbox Multilingual

Fish Audio S2БандлБез фильтра контента

Выбор для дубляжа — синтез речи и клонирование голоса на 80+ языках (Fish Audio S2 Pro) с лучшими в классе показателями точности. Голос-в-SRT и SRT-в-голос сохраняют исходный тайминг, распознавание через WhisperX.

Готовые воркфлоу

common-align-script-to-srt
Script → per-section SRT 1. Load audio (red, left) — upload the recording of your narration. 2. script (in the Align node) — paste your script; a blank line starts a new section. Each section becomes one SRT cue. 3. language — leave auto, or set ru / en / zh. 4. Press Run — WhisperX aligns your exact text to the speech and writes an SRT with one cue per section (exact start/end). A ⬇ Download SRT button appears when it finishes. Sections whose words aren't found in the audio (e.g. a line you skipped while reading) are dropped and logged. Runs best on a GPU tier.
common-audio-to-srt
AUDIO -> SRT (transcribe, keep timing). 1. Upload your source audio to 'Load source audio'. 2. On 'Audio -> SRT' pick the spoken language (auto / en / ru / zh). ENGINE is whisperx — Whisper core + phoneme forced-alignment; the tightest word-level timing and the best accuracy of the two engines. The Qwen3-TTS bundle additionally offers 'qwen3-asr' as an alternative engine; every other voice bundle is WhisperX-only. 3. Run. The timestamped subtitles are saved to output/subs/transcript.srt and shown in the node. Then translate that SRT text (keep the timestamps unchanged) and feed it to the 'SRT -> Audio' workflow to get a dubbed track with the SAME timing.
fishs2-srt-to-audio
SRT -> dubbed AUDIO, voice cloned, original timing preserved (Fish Audio S2 Pro). 1. Paste your TRANSLATED subtitles into 'Fish S2 SRT Dub' (keep the same timestamps as the source SRT, only the text is translated). 2. Upload a clean 10-30s clip of the speaker to 'Reference voice'. 3. Run. The output language is detected from the SRT text itself. Optional: press '🎙 Transcribe' on the dub node to preview the WhisperX transcription of your reference clip in 'ref_text' and fix it before Run. Left empty, the transcript is derived automatically. Knobs: fit_to_timing (on = lock each line into its SRT slot), max_stretch (cap before audio sounds sped-up).
fishs2-voice-clone
TTS + VOICE CLONE (Fish Audio S2 Pro, 80+ languages). 1. Upload a clean 10-30s clip of the target speaker to 'Reference voice'. 2. 'ref_text' — the exact words spoken in that clip. LEAVE EMPTY to have it transcribed automatically (WhisperX); type it manually for maximum accuracy. 3. Type the text to speak into 'text' — the language is detected from the text itself. 4. Run, listen in 'Save audio'. Tip: leave 'Reference voice' unconnected to let the model pick a random voice. The S2 server starts at boot; the first request after boot may wait a bit while it warms up.

Модели в этом бандле

CosyVoice 3БандлБез фильтра контента

Выбор для смены голоса — нативное преобразование сохраняет исходные слова, паузы и подачу и меняет только тембр, без промежуточной транскрипции. Также клонирование голоса, синтез речи, голос-в-SRT и SRT-в-голос, распознавание через WhisperX.

Готовые воркфлоу

common-align-script-to-srt
Script → per-section SRT 1. Load audio (red, left) — upload the recording of your narration. 2. script (in the Align node) — paste your script; a blank line starts a new section. Each section becomes one SRT cue. 3. language — leave auto, or set ru / en / zh. 4. Press Run — WhisperX aligns your exact text to the speech and writes an SRT with one cue per section (exact start/end). A ⬇ Download SRT button appears when it finishes. Sections whose words aren't found in the audio (e.g. a line you skipped while reading) are dropped and logged. Runs best on a GPU tier.
common-audio-to-srt
AUDIO -> SRT (transcribe, keep timing). 1. Upload your source audio to 'Load source audio'. 2. On 'Audio -> SRT' pick the spoken language (auto / en / ru / zh). ENGINE is whisperx — Whisper core + phoneme forced-alignment; the tightest word-level timing and the best accuracy of the two engines. The Qwen3-TTS bundle additionally offers 'qwen3-asr' as an alternative engine; every other voice bundle is WhisperX-only. 3. Run. The timestamped subtitles are saved to output/subs/transcript.srt and shown in the node. Then translate that SRT text (keep the timestamps unchanged) and feed it to the 'SRT -> Audio' workflow to get a dubbed track with the SAME timing.
cosyvoice3-change-voice
CHANGE VOICE (CosyVoice 3 native voice conversion) — keeps the words and the delivery, swaps only the timbre. 1. Upload the speech you want re-voiced (any length) to 'Source speech'. It is auto-split into <=25s chunks, re-voiced, and stitched back, so length is unlimited. 2. Upload a SHORT clean clip (3-30s) of the target speaker to 'Target voice'. Keep it SHORT (<=30s). 3. Run, listen in 'Save audio'. This is REAL voice conversion — no transcription step, so pauses, emphasis and pacing survive intact. To generate NEW speech from typed text instead, use the 'Voice Clone' workflow. First run downloads the CosyVoice model (~GB) — give it a few minutes. ⚠️ Predefined example — it resets to the original every GPU start; edits here are lost. Save under a NEW name (Save As) to keep your own copy; custom workflows persist between sessions.
cosyvoice3-srt-to-audio
SRT -> dubbed AUDIO, voice cloned, original timing preserved. 1. Paste your TRANSLATED subtitles into 'CosyVoice SRT Dub' (keep the same timestamps as the source SRT, only the text is translated). 2. Upload a SHORT clean clip (3-15s, no music) of the speaker to 'Reference voice' — CosyVoice clones this timbre. 3. Pick language (auto works; en/zh are reliable). 4. Run. Each line is synthesized and time-stretched to fit its slot, so the output lines up with your video. Knobs: fit_to_timing (on = lock to SRT timing), max_stretch (cap before audio sounds sped-up), speed. Languages: EN and ZH are solid. Other officially supported languages can be less reliable in this model — prefer the Qwen3-TTS bundle for those. First run downloads the CosyVoice model (~GB).
cosyvoice3-voice-clone
VOICE CLONE (CosyVoice 3 zero-shot). 1. Upload a SHORT clean clip (3-30s, no music) of the target speaker to 'Reference voice'. A reference clip is REQUIRED. 2. Type the text you want spoken into 'text' on the Voice Clone node. 3. Run, listen in 'Save audio'. First run downloads the CosyVoice model (~GB) — give it a few minutes. To re-voice EXISTING speech instead of generating new speech, use the 'Change Voice' workflow — it keeps the original words and delivery and only swaps the timbre. ⚠️ Predefined example — it resets to the original every GPU start; edits here are lost. Save under a NEW name (Save As) to keep your own copy; custom workflows persist between sessions.

Модели в этом бандле

Qwen3-TTSБандлБез фильтра контента

Полностью Qwen-бандл — синтез речи с клонированием голоса по 3-секундному образцу плюс дизайн голоса: опишите голос словами, и он заговорит (Qwen3-TTS 1.7B). Также голос-в-SRT и SRT-в-голос, и это единственный бандл с Qwen3-ASR в качестве второго движка распознавания рядом с WhisperX.

Готовые воркфлоу

common-align-script-to-srt
Script → per-section SRT 1. Load audio (red, left) — upload the recording of your narration. 2. script (in the Align node) — paste your script; a blank line starts a new section. Each section becomes one SRT cue. 3. language — leave auto, or set ru / en / zh. 4. Press Run — WhisperX aligns your exact text to the speech and writes an SRT with one cue per section (exact start/end). A ⬇ Download SRT button appears when it finishes. Sections whose words aren't found in the audio (e.g. a line you skipped while reading) are dropped and logged. Runs best on a GPU tier.
common-audio-to-srt
AUDIO -> SRT (transcribe, keep timing). 1. Upload your source audio to 'Load source audio'. 2. On 'Audio -> SRT' pick the spoken language (auto / en / ru / zh). ENGINE is whisperx — Whisper core + phoneme forced-alignment; the tightest word-level timing and the best accuracy of the two engines. The Qwen3-TTS bundle additionally offers 'qwen3-asr' as an alternative engine; every other voice bundle is WhisperX-only. 3. Run. The timestamped subtitles are saved to output/subs/transcript.srt and shown in the node. Then translate that SRT text (keep the timestamps unchanged) and feed it to the 'SRT -> Audio' workflow to get a dubbed track with the SAME timing.
qwen3tts-redub-voice
RE-DUB VOICE (Qwen3-TTS pipeline: transcribe -> re-speak, timing preserved). This is NOT voice conversion. Qwen3-TTS has no native VC, so this workflow chains two steps: the source audio is transcribed with timestamps (Audio -> SRT), then every line is re-spoken from scratch by the cloned TARGET voice at its original timestamp (SRT Dub). 1. Upload the recording you want re-dubbed to 'Source audio'. 2. Upload a SHORT clean clip (3-15s, no music) of the TARGET voice to 'Target voice'. Leave 'ref_text' empty — it is transcribed automatically. 3. On the dub node pick the LANGUAGE of the source speech and run. Line timing is preserved, but the words are re-generated, so intonation, pauses and emphasis inside each line are the model's, not the original speaker's. Two side effects worth knowing: a transcription error becomes a wrong word in the output, and any non-speech audio is dropped. For REAL voice conversion — same words, same delivery, only the timbre swapped — use the CosyVoice 3 bundle (the change-voice pick) or Chatterbox Multilingual. Both do it natively with no transcription step.
qwen3tts-srt-to-audio
SRT -> dubbed AUDIO, voice cloned, original timing preserved (Qwen3-TTS). 1. Paste your TRANSLATED subtitles into 'Qwen3-TTS SRT Dub' (keep the same timestamps as the source SRT, only the text is translated). 2. Upload a SHORT clean clip (3-15s, no music) of the speaker to 'Reference voice'. 3. 'ref_text' — the exact words spoken in that clip. LEAVE EMPTY to have it transcribed automatically (WhisperX); type it manually for maximum accuracy. 4. Pick the language of the TRANSLATED text and run. Knobs: fit_to_timing (on = lock each line into its SRT slot), max_stretch (cap before audio sounds sped-up). Each line is synthesized with the cloned voice and placed at its SRT timestamp, so the output lines up with your video.
qwen3tts-voice-clone
VOICE CLONE (Qwen3-TTS 1.7B Base). 1. Upload a SHORT clean clip (3-15s, no music) of the target speaker to 'Reference voice'. 2. 'ref_text' — the exact words spoken in that clip. LEAVE EMPTY to have it transcribed automatically (WhisperX); type it manually for maximum accuracy. 3. Type the text you want spoken into 'text' and pick its language. 4. Run, listen in 'Save audio'. First run loads the model (pre-baked at boot, a few seconds).
qwen3tts-voice-design
VOICE DESIGN (Qwen3-TTS 1.7B VoiceDesign). No reference audio needed — describe the voice you want in plain language. 1. Type the text to speak into 'text'. 2. Describe the voice in 'instruct' (gender, age, mood, pace, accent — e.g. 'A raspy old pirate, slow and theatrical'). 3. Pick the language and run. Tip: to REUSE a designed voice, save its output and feed it into the Voice Clone workflow as the reference sample.

Модели в этом бандле

  • Qwen3-TTS-1.7B-VoiceDesign

    Модель дизайна голоса — опишите нужный голос простыми словами (пол, возраст, настроение, акцент), и она озвучит ваш текст этим голосом. 10 языков.

  • Qwen3-TTS-1.7B-Base

    Модель клонирования голоса — 3-секундный образец задаёт итоговый голос. 10 языков.

  • Qwen3-ASR

    Альтернативный движок распознавания речи для голос-в-SRT, входит только в бандл Qwen3-TTS. Встроенные пословные метки времени.

Chatterbox MultilingualБандлБез фильтра контента

Синтез речи, клонирование и нативное преобразование голоса на 23 языках (Chatterbox Multilingual v3) — модель, обошедшая ElevenLabs в слепых тестах. Также голос-в-SRT и SRT-в-голос, распознавание через WhisperX.

Готовые воркфлоу

common-align-script-to-srt
Script → per-section SRT 1. Load audio (red, left) — upload the recording of your narration. 2. script (in the Align node) — paste your script; a blank line starts a new section. Each section becomes one SRT cue. 3. language — leave auto, or set ru / en / zh. 4. Press Run — WhisperX aligns your exact text to the speech and writes an SRT with one cue per section (exact start/end). A ⬇ Download SRT button appears when it finishes. Sections whose words aren't found in the audio (e.g. a line you skipped while reading) are dropped and logged. Runs best on a GPU tier.
common-audio-to-srt
AUDIO -> SRT (transcribe, keep timing). 1. Upload your source audio to 'Load source audio'. 2. On 'Audio -> SRT' pick the spoken language (auto / en / ru / zh). ENGINE is whisperx — Whisper core + phoneme forced-alignment; the tightest word-level timing and the best accuracy of the two engines. The Qwen3-TTS bundle additionally offers 'qwen3-asr' as an alternative engine; every other voice bundle is WhisperX-only. 3. Run. The timestamped subtitles are saved to output/subs/transcript.srt and shown in the node. Then translate that SRT text (keep the timestamps unchanged) and feed it to the 'SRT -> Audio' workflow to get a dubbed track with the SAME timing.
chatterbox-change-voice
VOICE CONVERSION (Chatterbox VC). Replaces the VOICE in a recording while keeping the words, pacing and intonation of the original. 1. Upload the recording you want to convert to 'Source audio'. 2. Upload a SHORT clean clip (3-15s, no music) of the TARGET voice to 'Target voice'. 3. Run, listen in 'Save audio'. First run loads the VC model (pre-baked at boot, a few seconds).
chatterbox-srt-to-audio
SRT -> dubbed AUDIO, voice cloned, original timing preserved (Chatterbox Multilingual v3). 1. Paste your TRANSLATED subtitles into 'Chatterbox SRT Dub' (keep the same timestamps as the source SRT, only the text is translated). 2. Upload a SHORT clean clip (3-15s, no music) of the speaker to 'Reference voice'. No transcript needed. 3. Pick the language code of the TRANSLATED text (en / ru / zh / ...) and run. Knobs: fit_to_timing (on = lock each line into its SRT slot), max_stretch (cap before audio sounds sped-up). Each line is synthesized with the cloned voice and placed at its SRT timestamp, so the output lines up with your video.
chatterbox-voice-clone
TTS + VOICE CLONE (Chatterbox Multilingual v3, 23 languages). 1. Upload a SHORT clean clip (3-15s, no music) of the target speaker to 'Reference voice'. No transcript needed. 2. Type the text to speak into 'text' and pick its language code (en / ru / zh / ...). 3. Run, listen in 'Save audio'. Tip: leave 'Reference voice' unconnected to use the model's default voice. First run loads the model (pre-baked at boot, a few seconds).

Модели в этом бандле

Воркфлоу и ноды

Каждый бандл открывается в ComfyUI с рабочими воркфлоу, которые можно запускать как есть — откройте панель Workflows и загляните в папку «_examples». В каждом воркфлоу на холсте есть заметка «How to use»: она объясняет, какие поля заполнить и в каком порядке, так что запоминать ничего не нужно. Результаты появляются во вкладке Assets. Воркфлоу — это граф нод; трогать нужно только те, что требуют ввода. Ноды выделены цветом:

  • Обязательный — нужно задать перед запуском — загрузить файл или ввести промпт.
  • Важный — стоит проверить или часто меняется — соотношение сторон, режим или ключевые параметры.

Авто-стоп GPU

GPU, простаивающий 60 минут, выключается автоматически — вы не платите за машину, о которой забыли. Что считается простоем, для LLM и Media отличается.

LLM — отсчёт идёт по вашей сессии и сбрасывается при каждом запросе. Сессии индивидуальны, поэтому завершение вашей не затрагивает других; сама машина выключается, когда закрывается последняя сессия.

Media — под следит за своей очередью ComfyUI. Всё, что рендерится или ждёт в очереди, считается активностью, поэтому многочасовой рендер удерживает под. Отсчёт начинается только когда очередь пуста.

Это можно изменить. LLM и Media настраиваются отдельно на странице настроек, и любой из них можно перевести в режим «никогда не останавливать».

При нулевом балансе GPU останавливается в любом случае, независимо от таймаута.

Логи и данные

Я запускаю модели на выделенных GPU-серверах — ваши запросы не отправляются никакому стороннему AI-провайдеру.

Переписка: чаты сохраняются, только если вы используете веб-чат, чтобы вы могли к ним вернуться. Запросы через API (/v1/messages, /v1/chat/completions) не сохраняются — их содержимое не записывается в базу данных.

Отладочные логи: для диагностики я храню короткоживущие технические логи с ротацией в 3 дня (ничего старше 3 дней не остаётся). В них — только диагностика: модель, тайминги, число токенов, названия инструментов — но не переписка. Логи GPU-сервера хранятся на самом инстансе и исчезают, когда он выключается.

Партнёрская программа

На странице Referrals получите промокод и поделитесь им. Каждый пользователь кода получит скидку 10%. Вы зарабатываете 15% с каждой их траты.

Философия

Это философия проекта. Прежде всего, я создал эти инструменты для собственной повседневной работы и поделился ими со всеми. Я также работаю по запросам клиентов — если что-то в принципе можно сделать, значит, я могу это реализовать. Просто напишите в систему поддержки. Если одну и ту же функцию или инструмент просят несколько пользователей, я её создам. Можете считать это бутиком ИИ-инструментов под управлением вашего друга.

Верификация

Чтобы каждый мог убедиться, что этот проект подлинный и находится под контролем его создателя, я публикую свой открытый PGP-ключ ниже. Если вы когда-либо столкнётесь со спором или имитацией этого проекта, любое сообщение, подписанное ключом с этим отпечатком, можно считать полученным от меня.

Отпечаток: 64B0 A423 385A 6518 49FD F920 E491 5625 837E 5AD9

Показать открытый ключ
-----BEGIN PGP PUBLIC KEY BLOCK-----

mQINBGpSurIBEACj0dS+3/A2xbCdnwWOrBNOnbzBoDxSxKdl8F7397HsuRaSuPv9
251Dm0WlgR7uAnvjhmMuuquCrrVNih2c5BDy8TUkKegTIccg7Iqkn6E5Xa5AbJxT
O2NA935H3WOX5LG1TrJrS8TW4sFnk8LGNdrkKfSYwdvRrvDtAYGI/bratwweenJw
ZLqZIP80UK2+Y8rS012k519KqOMaPx0w+5Ltz9WFqFq/cipQjTUNuQ05j60a8VOJ
Migki8UjCf4pp04kZ2903WNdIkfvKh7IU09V08y+LlQlGbB4qETZACNt7B7DZ3n+
d/4Io+bK2W7BLTvz2zriLfaFhcvNx5BpNyDqdLmAFWCCouCe+4LBBTxQfL95UW0P
Tx0Jz/KXfJnP5h/mQf6fO+tqbM5CRPhAItFWOYlDe+uSkkXsDKdOb2vqznUO39Cw
mARDjE40oSjrf+9PMHe/j1Y0hYTiQ+U/OQOY+zjXTL+rhZJrophLkctLbfwOnSVb
74L3f+0SMXiTFEnyNuMXolPjsg4115ghdi5NMQOzlb2EylM+nms8wjBdVDxhIGCv
daHXy4aosDBsrtSdWEwy2Pw35GdQQ9l21lYE4u9vJE7eeiaKwGMNZp8hpDFbU0Si
ri0E6c0Mve6GNlwAujK5pkCL0D5D67ESdZ7k4guyVsKHZEMotos0/2+ZYQARAQAB
tANrZXmJAm4EEwEIAFgWIQRksKQjOFplGEn9+SDkkVYlg35a2QUCalK6shsUgAAA
AAAEAA5tYW51MiwyLjUrMS4xMiwwLDMDGy8EBQsJCAcCAiICBhUKCQgLAgQWAgMB
Ah4HAheAAAoJEOSRViWDflrZregP/1jpyy8wEyUC4Ulq4qTNE3WZ63TWn6LNSuBb
y1VmDQ2HVeWlZO94uOZSqvpyACp9Z8EDCuypHT9BHlTjaXrHMHCsa2ZQBBn64dzZ
us75QAfviccKCgHOFZ6sJvik7m+tBW7JyrPPEdhcAp03yQ9CY5c8riutFLq+EHPL
0yKQJoVWDRQidx54m3b6OPp4ZLmzJWMFX1NSSs3FD4Gj53+xWTxWvw8YrxJ5o3AQ
yHl/IsHmF7FSmO9HRzbj0OX0yv3lIl0q1umh5Jbep9R7Jsh909JsKBAGjy6liVeC
vOvdINFaNof7+u6ln6zF55aq7Q9ALyDwgYgVdxi/p9g7HJdEBipRcEP4jM3indd8
ZMVRDYxDxQz3at8jFCpmSW5NwyDyKSsOYR7vqOVHvZLIP46Ci7Ir8/+2Q7Xm4sZA
yTkwKOF+zjYlQPHs0ulH8VmF/UFcuZp+aMQVUfadISCMe9Gl+UBjMuFPwygnkTdp
RJlTPlO8L8t2ISFW7oy5P5WNk33IKdtLHqoMu3eT/0oISSpqL1SovuV6K4Mn8Je1
5f8afvalMzbWorUS9cyZQfoHok6wxNAdipBAKt3PNiGcE/1AHd3DcNOS6RQmxQMP
YQv3Lr+gV0c89US2Fd2vzMuZ6f/huGMEN8PSdweLrAQahQRUzc56joxd60KucIdR
eEbBD3eE
=hE65
-----END PGP PUBLIC KEY BLOCK-----