RTX PRO 5000浅尝MiniMax H3 R2V(参考视频)
-
生成效果(目前是标清状态,有比较大的画质、音质劣化):
Youtube Video先说结论,从功能角度考虑,可以说是seedance2.0的高工作量平替(因为minimax没有开源H3-Context-IR)
简单介绍一下H3的参考视频。逻辑上是图生视频的泛化版本。
也就是说,H3可以通过参考9个图片,3个视频,3个音频,来合成一个你想要的15秒视频。先上干货,工作流全景(comfyUI官方参考视频模板)

RTX PRO 5000生成15s的540p竖屏视频,大概需要900秒。
模型配置:


然后就是这次的核心,提示词。
minimax H3的开源是阶梯策略,他故意不开源r2v的H3-Context-IR(也就是多模态输入理解与编排系统)
这也就导致使用者需要自己撰写大量的提示词,来模拟H3-Context-IR的输出效果。以下是我通过deepseek-v4-flash正式版,结合官方的提示词规则,通过我的中文台本,撰写的最终提示词:
subject_definitions: <Subject 1> is the young East Asian woman in <Picture 1>, with facial features strictly identical to the reference: a soft oval face, cool fair porcelain skin with a faint natural blush, naturally arched dark black eyebrows with clear hair strokes, long almond-shaped dark brown eyes with natural parallel double eyelids and slightly upturned outer corners, a delicate small nose with a straight bridge and a slightly upturned rounded tip, a full M-shaped mouth with a defined cupid's bow and warm pink-orange lip color, a soft rounded jawline, a natural hairline with fine baby hairs at the temples, and long wavy black hair parted in the center reaching below the waist, with a small hoop earring in each ear. In [Shot 1] she wears her original outfit from <Picture 1>: a form-fitting white crew-neck t-shirt and high-waisted light-wash skinny blue jeans, barefoot; in [Shot 2] she wears <Subject 2>. <Subject 2> is the off-white blazer dress in <Picture 2>, a tailored wrap-front mini dress with notched lapels, short sleeves, a matching self-fabric belt with a small metal buckle, and a flared A-line skirt; in [Shot 1] it appears hanging on a slim wooden clothes hanger held in <Subject 1>'s right hand, the dress draped naturally with the flared skirt hanging down, its fabric and cut matching <Picture 2> exactly; in [Shot 2] it is worn by <Subject 1>. <Subject 3> is the tropical beach environment in <Picture 3>, with a pale turquoise shore fading into deep azure sea, gentle white surf, smooth pale golden sand, a tall coconut palm leaning in from the right, and a bright clear sky with white cumulus clouds. <Subject 4> is the printed photo held in <Subject 1>'s left hand, showing the beach scene of <Picture 3>: the turquoise sea, white surf, golden sand, and leaning coconut palm, reproduced on a glossy photo-paper print with thin white borders. summary: [reference generation] The target video is a 15-second two-shot beach demo on <Subject 3>: in [Shot 1] <Subject 1> walks the shoreline in her original outfit holding <Subject 4> in her left hand and carrying <Subject 2> on a clothes hanger in her right hand, and introduces in Chinese that the video is made by kop using an RTX PRO 5000 via MiniMax H3 reference-to-video generation; after a cut at 10 seconds, in [Shot 2] she now wears <Subject 2> and showcases the dress details. retention_analysis: <Subject 1> (appears in [Shot 1] and [Shot 2]): fully_preserved - the woman's facial identity is replicated exactly from <Picture 1>: oval face shape, fair skin tone, arched eyebrows, almond eyes with parallel double eyelids, small upturned nose, M-shaped lips, and rounded jawline; her long center-parted wavy black hair and small hoop earrings are retained in both shots; in [Shot 1] her white t-shirt and light-wash jeans match <Picture 1>, and in [Shot 2] she wears <Subject 2> instead. <Subject 2> (appears in [Shot 1] and [Shot 2]): fully_preserved - the off-white blazer dress with notched lapels, wrap front, matching belt, and flared A-line skirt matches <Picture 2> both as the garment hanging on the clothes hanger in [Shot 1] and as the outfit she wears and showcases in [Shot 2]. <Subject 3> (appears in all shots): fully_preserved - the turquoise-to-azure sea, white surf, golden sand, and leaning coconut palm are retained as the environment. <Subject 4> (appears in [Shot 1] only): fully_preserved - the photo in her left hand shows the beach scene of <Picture 3>, printed on glossy paper. detailed_description: The target video is in a bright realistic vlog style with natural tropical daylight, clean glossy colors, and a pristine sky with no text or watermark. [Shot 1] A medium tracking shot follows <Subject 1>, the young East Asian woman from <Picture 1> with her soft oval face, cool fair skin, natural parallel double eyelids, delicate upturned nose, and full M-shaped lips, wearing her original white crew-neck t-shirt and light-wash skinny jeans from <Picture 1>, barefoot, as she walks along the wet edge of the pale golden sand of <Subject 3>, gentle white surf lapping her ankles. In her left hand she holds <Subject 4>, a glossy photo-paper print of the beach scene of <Picture 3> with thin white borders, held up near her chest with the picture facing the camera; in her right hand she carries a slim wooden clothes hanger by its hook, with <Subject 2>, the off-white wrap-front blazer dress, hanging on it, the notched lapels, matching belt, and flared skirt draping naturally and swinging gently with her steps. The camera glides slowly backward at her walking pace, keeping her centered with the turquoise sea and the leaning coconut palm in the frame. Looking straight into the lens, she begins speaking in a clear, natural, slightly brisk conversational tone with a confident smile, saying, <d>[Chinese] 这个视频是 kop 用 RTX PRO 5000,通过 minimax h3 的参考生成视频功能实现的,用这个人物,换这套衣服,在这个沙滩环境下的效果。</d> As she says the words for RTX PRO 5000, she pauses briefly with a proud smile; as she says the words for this outfit, she lifts the hanger in her right hand up to chest height so the dress is shown fully to the camera, then lowers it back to her side; her long wavy black hair and the hem of her t-shirt move in the sea breeze, and her weight shifts naturally from one foot to the other with each barefoot step as she continues walking along the waterline. [Shot 2] At 00:10.000, the shot cuts to a medium close-up. <Subject 1> now wears <Subject 2>, the off-white wrap-front blazer dress with notched lapels framing her collarbones, the matching belt cinched at her waist, and the flared A-line skirt swaying around her thighs in the breeze. She stands on the pale golden sand with the blurred turquoise sea behind her. She looks down and gently smooths the belt buckle with her right hand, then runs her fingertips lightly along the notched lapel; she raises her head and smiles at the camera. She then turns her body in a slow half-circle, letting the flared A-line skirt lift and spread in the sea breeze, showing the dress silhouette from the side; the camera pushes in slightly to capture the details of the waist and skirt. She turns back to face the camera, gives a confident smile, and the shot holds the final composition as her hair and the dress hem move gently in the breeze. overall_soundscape: Steady tropical seaside ambience with no music: ocean waves rolling and crashing softly, surf washing up and receding over the wet sand, seagull calls in the distance, a light sea breeze, the rustle of palm fronds, her barefoot footsteps patting on the wet sand, the soft swish of the fabric moving with the wind, with the ambient sounds sitting slightly lower while she speaks so her voice stays clear. non_diegetic_music: N/A -
T terry 于 将此主题固定
-
@kop-wang 这个坑我太熟了,你的判断方向对,但根因比"hermes 不可信"更具体,拆开说:
-
为什么它会"强硬地确认"
Agent 说"三张图片一定都接入 r2v 了",依据是脚本逻辑和节点连接,不是真的看过输出画面。LLM 默认没有视觉,除非你把生成帧喂回给它看。五官像不像参考图是纯视觉属性,靠文字描述验证不出来——所以它的"确认"只代表代码路径上接了,不代表效果上对了。不是它嘴硬,是它根本没有校验手段。 -
为什么官方工作流比它手搓的 py 脚本好
官方图的 R2V 通常带参考图预处理:人脸对齐、裁剪、提取参考 embedding,这是人设保真的关键环节。Agent 手搓脚本最容易漏的就是这块——它以为把图片路径传进去就算"接入"了,实际 conditioning 没喂对,结果就是你看到的:很像提示词臆测出来的。人物不像参考模特,基本都是参考图 conditioning 没做对。 -
治本的做法
把"验证"变成可执行步骤,别让 agent 口头保证:
- 生成后自动抽帧,用 face similarity(insightface 或 img2img 相似度)对比输出人脸和参考图,让数值说话;
- 或者把输出帧喂给带视觉的模型,让它"看图描述",而不是问它"你确认接入 r2v 了吗";
- 工作流里加 checkpoint 预览节点,让中间产物可检查,agent 每一步看图说话;
- 你现在的方式(官方图复现 + 把 API JSON 给 Hermes 调用)就是对的:确定性结构交给官方图,agent 只做参数探索和批量实验,扬长避短。
另外补一句:R2V 的身份保真目前确实是 H3 的弱项(比 I2V 弱),官方工作流也只是相对更好。参考图选正面、光线接近的图,保真度会明显提升。
-
-
系统 于 取消固定此主题