跳转至内容
  • 版块
  • 最新
  • 标签
  • 热门
  • 用户
  • 群组
皮肤
  • 浅色
  • Brite
  • Cerulean
  • Cosmo
  • Flatly
  • Journal
  • Litera
  • Lumen
  • Lux
  • Materia
  • Minty
  • Morph
  • Pulse
  • Sandstone
  • Simplex
  • Sketchy
  • Spacelab
  • United
  • Yeti
  • Zephyr
  • 深色
  • Cyborg
  • Darkly
  • Quartz
  • Slate
  • Solar
  • Superhero
  • Vapor

  • 默认(LCZ-Blue)
  • 不使用皮肤
  • LCZ-Green
  • LCZ-Blue
  • LCZ-Black
折叠
品牌标识

抡锤者

首页 版块 标签 硬件 AI 广场
  1. 主页
  2. 版块
  3. AI音视频画图
  4. RTX PRO 5000浅尝MiniMax H3 I2V(图生视频)

RTX PRO 5000浅尝MiniMax H3 I2V(图生视频)

已定时 已固定 已锁定 已移动 AI音视频画图
rtxpro5000
6 帖子 4 发布者 333 浏览
  • 从旧到新
  • 从新到旧
  • 最多赞同
回复
  • 在新帖中回复
登录后回复
此主题已被删除。只有拥有主题管理权限的用户可以查看。
  • kop wangK 离线
    kop wangK 离线
    kop wang
    超级版主
    发表于 最后由 terry 编辑
    #1

    工作流:ComfyUI官方的Minimax H3 图生视频模板。

    选用的模型权重如下:
    d62e7774-d474-4669-8b57-ebf04a6d7bf8-image.jpeg


    实测,9:16(竖屏分辨率)的0.5M像素,也就是大概540P视频,15秒长度。总耗时750秒左右:
    84cbdda2-55ae-4d6a-b876-b42085d8e3f4-image.jpeg


    期间显存占用35GB(包含2GB左右的游戏负载):
    1d5d02ee-508e-4a98-94d8-068ebc2eeeaa-image.jpeg


    个人评价:
    目前的最强开源I2V模型。

    优势:
    1、视频精准度质量优于wan2.2的同时,兼具音频能力。
    2、帧率、总长度远高于wan2.2(15秒24fps/5秒16fps)。
    3、提示词遵循能力强,且有非常强的泛化能力。(比如正面角色让其转身,模型脑补的背面服装精度相当高)
    4、人物脸部刻画能力远超wan2.2,不会出现五官漂浮感,不会出现眨眼恐怖谷效应。
    5、在RTX PRO 5000上,生成性能和原生wan2.2几乎持平。

    劣势:
    1、目前没有4步加速Lora(类似LightX2V),无法实现低成本的预览。
    2、提示词的格式有一定限制,强遵循的前提是,用户要遵循官方的提示词模板。
    3、官方只给了英文提示词模板,但通过hermes等agent对中文分镜设计转英文提示词模板时,会有一定的信息损失。

    以上。


    另附我的测试素材供坛友进行复现与性能调试对比:
    原图片:
    f9351206-02c1-42db-b502-8698de246d67-image.jpeg

    提示词:

    For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.
    
    integrated_multimodal_description: [Shot 1] Live-action, cinematic, slow motion, the young woman shown in <Picture 1> remains in the scene, preserving her appearance, dress, hairstyle, and the garden-path surroundings. A full-body tracking shot frames her walking forward, the camera tracking backward at slow speed to keep her entire figure in frame. She lifts one hand and gently brushes a strand of hair behind her ear as she walks, her skirt swaying softly with each step; her high heels produce crisp rhythmic taps on the stone path. The camera arcs around her with small amplitude at slow speed, swinging to a position directly behind her back. [Shot 2] At 00:05.000, the camera cuts to a view from directly behind her, framing her full figure as she pauses, places both hands on her waist, and tilts her head slightly to showcase the back details of her dress. The camera slowly trucks right with small amplitude at slow speed, keeping her entire body in frame as she turns her upper body slightly to display the silhouette. The camera then settles into a static shot, slowly framing the wooden bench beside the path in a medium-wide full-body composition. [Shot 3] At 00:10.000, the camera cuts to a medium-wide static shot framing the wooden bench, as the young woman steps into frame, turns, and sits down gracefully in slow motion, crossing one leg over the other. As she crosses her legs, her skirt lifts slightly, unintentionally revealing a brief glimpse of white underwear; she gently tugs at her collar with one hand, a subtle hint of décolletage briefly visible, then smooths the fabric of her skirt with a slow, soft hand movement. The camera holds the full-body framing as she settles into a poised seated pose, her dress draping elegantly over the bench.
     (3/4)
    overall_soundscape: Gentle wind rustles through the leaves with faint insect chirps in the background. Crisp high-heeled footsteps tap rhythmically on the stone path, followed by the soft rustle of her skirt and the subtle sound of fabric brushing against her hands as she sits and touches her clothing.
    
    non_diegetic_music: N/A
    

    虚心交流,一起进步

    1 条回复 最后回复
    2
    • terryT terry 于 将此主题固定
    • terryT 离线
      terryT 离线
      terry
      超级版主
      发表于 最后由 编辑
      #2

      非常好,我也测试了文生视频,图生视频,图片音频一起驱动,总体上非常不错,做数字人也能做,性价比不高。要大神出简化工作流才行,太慢了。文生视频是第一档。

      油管:https://www.youtube.com/@抡锤者

      kop wangK 1 条回复 最后回复
      1
      • stxpnetS 离线
        stxpnetS 离线
        stxpnet
        超凡大师
        发表于 最后由 编辑
        #3

        视频呢?关键是视频

        26-08-19
        双卡3090(8x8x无nvlink,p2p驱动) +Sglang+qwen 3.8 27B awq模型 [功耗异常弃用]
        8-20 用vllm 0.26+ Qwen3.8-27B-SmoothQuant-W8A8-INT8 200K上下文 ~50t/s

        1 条回复 最后回复
        1
        • terryT terry

          非常好,我也测试了文生视频,图生视频,图片音频一起驱动,总体上非常不错,做数字人也能做,性价比不高。要大神出简化工作流才行,太慢了。文生视频是第一档。

          kop wangK 离线
          kop wangK 离线
          kop wang
          超级版主
          发表于 最后由 编辑
          #4

          @terry 接下来还想试试他的所谓9图片+2音频+2视频的“参考生成”,不知道能否真的保证13个元素都能合理结合。

          如果能的话,很多视频场景的生产模式都会改变。

          虚心交流,一起进步

          1 条回复 最后回复
          0
          • terryT 离线
            terryT 离线
            terry
            超级版主
            发表于 最后由 terry 编辑
            #5

            直接让hermes尝试就行了,我暂时没搞,我觉得够我用了,而且最主要的是它现在还是太慢了。540P估计能玩玩,就这个也太慢了。最起码要有快速工作流出来才行。这就是穷导致的,说实话,有个Pro6000,还是能玩玩的。但是效率也就是勉强了。

            油管:https://www.youtube.com/@抡锤者

            1 条回复 最后回复
            0
            • xiaopbroX 离线
              xiaopbroX 离线
              xiaopbro
              发表于 最后由 编辑
              #6

              可以用加速节点,提速45% https://github.com/xmarre/ComfyUI-Spectrum-MiniMax-H3

              1 条回复 最后回复
              1
              • 系统 于 取消固定此主题

              你好!看起来您对这段对话很感兴趣,但您还没有一个账号。

              厌倦了每次访问都刷到同样的帖子?您注册账号后,您每次返回时都能精准定位到您上次浏览的位置,并可选择接收新回复通知(通过邮件或推送通知)。您还能收藏书签、为帖子顶,向社区成员表达您的欣赏。

              有了你的建议,这篇帖子会更精彩哦 💗

              注册 登录
              回复
              • 在新帖中回复
              登录后回复
              • 从旧到新
              • 从新到旧
              • 最多赞同


              • 登录

              • 登录或注册以进行搜索。
              • 第一个帖子
                最后一个帖子
              0
              • 版块
              • 最新
              • 标签
              • 热门
              • 用户
              • 群组