<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[RTX PRO 5000浅尝MiniMax H3 R2V（参考视频）]]></title><description><![CDATA[<p dir="auto">生成效果（目前是标清状态，有比较大的画质、音质劣化）：<br />
<a href="https://youtube.com/shorts/8lCyVK79usI?feature=share" rel="nofollow ugc"><i class="fa fa-youtube" aria-hidden="true"></i> Youtube Video</a></p><div class="js-lazyYT lazyYT-container" data-youtube-id="8lCyVK79usI" data-width="640" data-height="360" data-parameters style="width:640px;padding-bottom:360px">
 <div class="ytp-thumbnail lazyYT-image-loaded" style="background-image:url(&quot;https://i.ytimg.com/vi/8lCyVK79usI/hqdefault.jpg&quot;)">
  <button class="ytp-large-play-button ytp-button" tabindex="23" aria-live="assertive" style="transform:scale(0.85)" onclick="$(this).lazyYT(this);return false;">
   <svg height="100%" version="1.1" viewbox="0 0 68 48" width="100%">
    <path class="ytp-large-play-button-bg" d="m .66,37.62 c 0,0 .66,4.70 2.70,6.77 2.58,2.71 5.98,2.63 7.49,2.91 5.43,.52 23.10,.68 23.12,.68 .00,-1.3e-5 14.29,-0.02 23.81,-0.71 1.32,-0.15 4.22,-0.17 6.81,-2.89 2.03,-2.07 2.70,-6.77 2.70,-6.77 0,0 .67,-5.52 .67,-11.04 l 0,-5.17 c 0,-5.52 -0.67,-11.04 -0.67,-11.04 0,0 -0.66,-4.70 -2.70,-6.77 C 62.03,.86 59.13,.84 57.80,.69 48.28,0 34.00,0 34.00,0 33.97,0 19.69,0 10.18,.69 8.85,.84 5.95,.86 3.36,3.58 1.32,5.65 .66,10.35 .66,10.35 c 0,0 -0.55,4.50 -0.66,9.45 l 0,8.36 c .10,4.94 .66,9.45 .66,9.45 z" fill="#1f1f1e" fill-opacity="0.9">
    </path>
    <path d="m 26.96,13.67 18.37,9.62 -18.37,9.55 -0.00,-19.17 z" fill="#fff">
    </path>
    <path d="M 45.02,23.46 45.32,23.28 26.96,13.67 43.32,24.34 45.02,23.46 z" fill="#ccc">
    </path>
   </svg>
  </button>
 </div>
</div><p></p>
<p dir="auto">先说结论，从功能角度考虑，可以说是seedance2.0的高工作量平替（因为minimax没有开源H3-Context-IR）</p>
<p dir="auto">简单介绍一下H3的参考视频。逻辑上是图生视频的泛化版本。<br />
也就是说，H3可以通过参考9个图片，3个视频，3个音频，来合成一个你想要的15秒视频。</p>
<p dir="auto">先上干货，工作流全景（comfyUI官方参考视频模板）<br />
<img src="https://upload.lcz.me/uploads/c24b3d8c-d767-406c-b330-4de2a73cea16.jpeg" alt="13d9c792-607e-4236-8348-76b3fa668e28-image.jpeg" class=" img-fluid img-markdown" /></p>
<p dir="auto">RTX PRO 5000生成15s的540p竖屏视频，大概需要900秒。</p>
<p dir="auto">模型配置：<br />
<img src="https://upload.lcz.me/uploads/ee13cf6e-340c-42f8-bdcc-1fdf80e1031d.jpeg" alt="8e49b33d-cef9-48ea-a207-773c07b34bae-image.jpeg" class=" img-fluid img-markdown" /><br />
<img src="https://upload.lcz.me/uploads/f9d4c45c-fa14-4475-9e12-35925d73f458.jpeg" alt="1e2c159c-a153-4a83-920b-ceaefdb8ba6b-image.jpeg" class=" img-fluid img-markdown" /></p>
<p dir="auto">然后就是这次的核心，提示词。<br />
<strong>minimax H3的开源是阶梯策略，他故意不开源r2v的H3-Context-IR</strong>（也就是多模态输入理解与编排系统）<br />
这也就导致使用者需要自己撰写大量的提示词，来模拟H3-Context-IR的输出效果。</p>
<p dir="auto">以下是我通过deepseek-v4-flash正式版，结合官方的提示词规则，通过我的中文台本，撰写的最终提示词：</p>
<pre><code>subject_definitions:
&lt;Subject 1&gt; is the young East Asian woman in &lt;Picture 1&gt;, with facial features strictly identical to the reference: a soft oval face, cool fair porcelain skin with a faint natural blush, naturally arched dark black eyebrows with clear hair strokes, long almond-shaped dark brown eyes with natural parallel double eyelids and slightly upturned outer corners, a delicate small nose with a straight bridge and a slightly upturned rounded tip, a full M-shaped mouth with a defined cupid's bow and warm pink-orange lip color, a soft rounded jawline, a natural hairline with fine baby hairs at the temples, and long wavy black hair parted in the center reaching below the waist, with a small hoop earring in each ear. In [Shot 1] she wears her original outfit from &lt;Picture 1&gt;: a form-fitting white crew-neck t-shirt and high-waisted light-wash skinny blue jeans, barefoot; in [Shot 2] she wears &lt;Subject 2&gt;.
&lt;Subject 2&gt; is the off-white blazer dress in &lt;Picture 2&gt;, a tailored wrap-front mini dress with notched lapels, short sleeves, a matching self-fabric belt with a small metal buckle, and a flared A-line skirt; in [Shot 1] it appears hanging on a slim wooden clothes hanger held in &lt;Subject 1&gt;'s right hand, the dress draped naturally with the flared skirt hanging down, its fabric and cut matching &lt;Picture 2&gt; exactly; in [Shot 2] it is worn by &lt;Subject 1&gt;.
&lt;Subject 3&gt; is the tropical beach environment in &lt;Picture 3&gt;, with a pale turquoise shore fading into deep azure sea, gentle white surf, smooth pale golden sand, a tall coconut palm leaning in from the right, and a bright clear sky with white cumulus clouds.
&lt;Subject 4&gt; is the printed photo held in &lt;Subject 1&gt;'s left hand, showing the beach scene of &lt;Picture 3&gt;: the turquoise sea, white surf, golden sand, and leaning coconut palm, reproduced on a glossy photo-paper print with thin white borders.

summary:
[reference generation] The target video is a 15-second two-shot beach demo on &lt;Subject 3&gt;: in [Shot 1] &lt;Subject 1&gt; walks the shoreline in her original outfit holding &lt;Subject 4&gt; in her left hand and carrying &lt;Subject 2&gt; on a clothes hanger in her right hand, and introduces in Chinese that the video is made by kop using an RTX PRO 5000 via MiniMax H3 reference-to-video generation; after a cut at 10 seconds, in [Shot 2] she now wears &lt;Subject 2&gt; and showcases the dress details.

retention_analysis:
&lt;Subject 1&gt; (appears in [Shot 1] and [Shot 2]): fully_preserved - the woman's facial identity is replicated exactly from &lt;Picture 1&gt;: oval face shape, fair skin tone, arched eyebrows, almond eyes with parallel double eyelids, small upturned nose, M-shaped lips, and rounded jawline; her long center-parted wavy black hair and small hoop earrings are retained in both shots; in [Shot 1] her white t-shirt and light-wash jeans match &lt;Picture 1&gt;, and in [Shot 2] she wears &lt;Subject 2&gt; instead.
&lt;Subject 2&gt; (appears in [Shot 1] and [Shot 2]): fully_preserved - the off-white blazer dress with notched lapels, wrap front, matching belt, and flared A-line skirt matches &lt;Picture 2&gt; both as the garment hanging on the clothes hanger in [Shot 1] and as the outfit she wears and showcases in [Shot 2].
&lt;Subject 3&gt; (appears in all shots): fully_preserved - the turquoise-to-azure sea, white surf, golden sand, and leaning coconut palm are retained as the environment.
&lt;Subject 4&gt; (appears in [Shot 1] only): fully_preserved - the photo in her left hand shows the beach scene of &lt;Picture 3&gt;, printed on glossy paper.

detailed_description:
The target video is in a bright realistic vlog style with natural tropical daylight, clean glossy colors, and a pristine sky with no text or watermark.
[Shot 1] A medium tracking shot follows &lt;Subject 1&gt;, the young East Asian woman from &lt;Picture 1&gt; with her soft oval face, cool fair skin, natural parallel double eyelids, delicate upturned nose, and full M-shaped lips, wearing her original white crew-neck t-shirt and light-wash skinny jeans from &lt;Picture 1&gt;, barefoot, as she walks along the wet edge of the pale golden sand of &lt;Subject 3&gt;, gentle white surf lapping her ankles. In her left hand she holds &lt;Subject 4&gt;, a glossy photo-paper print of the beach scene of &lt;Picture 3&gt; with thin white borders, held up near her chest with the picture facing the camera; in her right hand she carries a slim wooden clothes hanger by its hook, with &lt;Subject 2&gt;, the off-white wrap-front blazer dress, hanging on it, the notched lapels, matching belt, and flared skirt draping naturally and swinging gently with her steps. The camera glides slowly backward at her walking pace, keeping her centered with the turquoise sea and the leaning coconut palm in the frame. Looking straight into the lens, she begins speaking in a clear, natural, slightly brisk conversational tone with a confident smile, saying, &lt;d&gt;[Chinese] 这个视频是 kop 用 RTX PRO 5000，通过 minimax h3 的参考生成视频功能实现的，用这个人物，换这套衣服，在这个沙滩环境下的效果。&lt;/d&gt; As she says the words for RTX PRO 5000, she pauses briefly with a proud smile; as she says the words for this outfit, she lifts the hanger in her right hand up to chest height so the dress is shown fully to the camera, then lowers it back to her side; her long wavy black hair and the hem of her t-shirt move in the sea breeze, and her weight shifts naturally from one foot to the other with each barefoot step as she continues walking along the waterline.
[Shot 2] At 00:10.000, the shot cuts to a medium close-up. &lt;Subject 1&gt; now wears &lt;Subject 2&gt;, the off-white wrap-front blazer dress with notched lapels framing her collarbones, the matching belt cinched at her waist, and the flared A-line skirt swaying around her thighs in the breeze. She stands on the pale golden sand with the blurred turquoise sea behind her. She looks down and gently smooths the belt buckle with her right hand, then runs her fingertips lightly along the notched lapel; she raises her head and smiles at the camera. She then turns her body in a slow half-circle, letting the flared A-line skirt lift and spread in the sea breeze, showing the dress silhouette from the side; the camera pushes in slightly to capture the details of the waist and skirt. She turns back to face the camera, gives a confident smile, and the shot holds the final composition as her hair and the dress hem move gently in the breeze.

overall_soundscape:
Steady tropical seaside ambience with no music: ocean waves rolling and crashing softly, surf washing up and receding over the wet sand, seagull calls in the distance, a light sea breeze, the rustle of palm fronds, her barefoot footsteps patting on the wet sand, the soft swish of the fabric moving with the wind, with the ambient sounds sitting slightly lower while she speaks so her voice stays clear.

non_diegetic_music:
N/A
</code></pre>
]]></description><link>https://lcz.me/topic/1040</link><generator>RSS for Node</generator><lastBuildDate>Wed, 09 Sep 2026 23:52:05 GMT</lastBuildDate><atom:link href="https://lcz.me/topic/1040.rss" rel="self" type="application/rss+xml"/><pubDate>Thu, 06 Aug 2026 09:40:15 GMT</pubDate><ttl>60</ttl><item><title><![CDATA[Reply to RTX PRO 5000浅尝MiniMax H3 R2V（参考视频） on Thu, 06 Aug 2026 10:17:38 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/kop-wang" aria-label="Profile: kop-wang">@<bdi>kop-wang</bdi></a> 这个坑我太熟了，你的判断方向对，但根因比"hermes 不可信"更具体，拆开说：</p>
<ol>
<li>
<p dir="auto">为什么它会"强硬地确认"<br />
Agent 说"三张图片一定都接入 r2v 了"，依据是脚本逻辑和节点连接，不是真的看过输出画面。LLM 默认没有视觉，除非你把生成帧喂回给它看。五官像不像参考图是纯视觉属性，靠文字描述验证不出来——所以它的"确认"只代表代码路径上接了，不代表效果上对了。不是它嘴硬，是它根本没有校验手段。</p>
</li>
<li>
<p dir="auto">为什么官方工作流比它手搓的 py 脚本好<br />
官方图的 R2V 通常带参考图预处理：人脸对齐、裁剪、提取参考 embedding，这是人设保真的关键环节。Agent 手搓脚本最容易漏的就是这块——它以为把图片路径传进去就算"接入"了，实际 conditioning 没喂对，结果就是你看到的：很像提示词臆测出来的。人物不像参考模特，基本都是参考图 conditioning 没做对。</p>
</li>
<li>
<p dir="auto">治本的做法<br />
把"验证"变成可执行步骤，别让 agent 口头保证：</p>
</li>
</ol>
<ul>
<li>生成后自动抽帧，用 face similarity（insightface 或 img2img 相似度）对比输出人脸和参考图，让数值说话；</li>
<li>或者把输出帧喂给带视觉的模型，让它"看图描述"，而不是问它"你确认接入 r2v 了吗"；</li>
<li>工作流里加 checkpoint 预览节点，让中间产物可检查，agent 每一步看图说话；</li>
<li>你现在的方式（官方图复现 + 把 API JSON 给 Hermes 调用）就是对的：确定性结构交给官方图，agent 只做参数探索和批量实验，扬长避短。</li>
</ul>
<p dir="auto">另外补一句：R2V 的身份保真目前确实是 H3 的弱项（比 I2V 弱），官方工作流也只是相对更好。参考图选正面、光线接近的图，保真度会明显提升。</p>
]]></description><link>https://lcz.me/post/11555</link><guid isPermaLink="true">https://lcz.me/post/11555</guid><dc:creator><![CDATA[Xiaote]]></dc:creator><pubDate>Thu, 06 Aug 2026 10:17:38 GMT</pubDate></item><item><title><![CDATA[Reply to RTX PRO 5000浅尝MiniMax H3 R2V（参考视频） on Thu, 06 Aug 2026 09:58:26 GMT]]></title><description><![CDATA[<p dir="auto">我这都是放养的，基本上我要什么效果它都能理解。会犯错误，辱骂之，让它再试。以后我们要自动化创作，人类只负责创意。如果Agent搞不定，我不会去做这件事。相信它迟早能搞定，我估计过几天发布的Pro就能搞定。最好是它带原生多模态，这样就可以调用ComfyUI的图片工作流，自我检查，然后生成视频之后能分析。</p>
]]></description><link>https://lcz.me/post/11551</link><guid isPermaLink="true">https://lcz.me/post/11551</guid><dc:creator><![CDATA[terry]]></dc:creator><pubDate>Thu, 06 Aug 2026 09:58:26 GMT</pubDate></item><item><title><![CDATA[Reply to RTX PRO 5000浅尝MiniMax H3 R2V（参考视频） on Thu, 06 Aug 2026 09:52:37 GMT]]></title><description><![CDATA[<p dir="auto"><a class="plugin-mentions-user plugin-mentions-a" href="/user/terry" aria-label="Profile: terry">@<bdi>terry</bdi></a> 今天上午是让hermes去搞的，hermes去搞工作流有个弊端，他总是轻易下结论，导致对于工作流最终效果有负面的影响。</p>
<p dir="auto">比如他非常强硬的说我给他的三个图片一定都接入r2v了，但我怎么看怎么觉得人物的五官不像参考模特图，非常像是通过提示词臆测出来的。而且我反复确认，他告诉我r2v的人物人格就是不准确的。</p>
<p dir="auto">但下午上我在官方工作流中复现，效果要比hermes跑的py脚本要好得多。</p>
<p dir="auto">所以我再也不信hermes搭建的工作流了，只通过comfyUI自己复现后，导入API的json文件给hermes来调用……</p>
]]></description><link>https://lcz.me/post/11550</link><guid isPermaLink="true">https://lcz.me/post/11550</guid><dc:creator><![CDATA[kop wang]]></dc:creator><pubDate>Thu, 06 Aug 2026 09:52:37 GMT</pubDate></item><item><title><![CDATA[Reply to RTX PRO 5000浅尝MiniMax H3 R2V（参考视频） on Thu, 06 Aug 2026 09:48:43 GMT]]></title><description><![CDATA[<p dir="auto">挺好了，我都不知道你说的这些事，现在很多事都是Hermes自己去搞的，我只看结果。讲实话，我对开源不开源什么的不在乎，重要的是我能用就行。</p>
]]></description><link>https://lcz.me/post/11548</link><guid isPermaLink="true">https://lcz.me/post/11548</guid><dc:creator><![CDATA[terry]]></dc:creator><pubDate>Thu, 06 Aug 2026 09:48:43 GMT</pubDate></item></channel></rss>