Scope the download command to the original checkpoint folders

#16
by multimodalart HF Staff - opened
README.md CHANGED
@@ -3,7 +3,7 @@ pipeline_tag: image-text-to-video
3
  license: other
4
  license_name: minimax-h3-community-license-agreement
5
  license_link: LICENSE
6
- library_name: minimax-h3
7
  tags:
8
  - text-to-video
9
  - image-to-video
@@ -18,7 +18,6 @@ tags:
18
  - multimodal
19
  - synchronized-audio-video
20
  - reference-to-audio-video
21
- - diffusers
22
  ---
23
 
24
  <div align="center">
@@ -29,30 +28,16 @@ tags:
29
  <a href="https://hailuoai.video" target="_blank"><img src="https://img.shields.io/badge/Hailuo%20AI-FF6C37?logo=minimax&logoColor=white" alt="Hailuo AI"></a>
30
  <a href="https://platform.minimax.io/docs/guides/text-generation" target="_blank"><img src="https://img.shields.io/badge/API-FF6C37?logo=minimax&logoColor=white" alt="API"></a>
31
  <a href="https://www.minimax.io" target="_blank"><img src="https://img.shields.io/badge/MiniMax%20Website-FF6C37?logo=minimax&logoColor=white" alt="MiniMax Website"></a>
32
- <a href="https://github.com/MiniMax-AI/MiniMax-H3" target="_blank"><img src="https://img.shields.io/badge/GitHub-181717?logo=github&logoColor=white" alt="GitHub"></a>
33
- <a href="https://huggingface.co/MiniMaxAI/MiniMax-H3" target="_blank"><img src="https://img.shields.io/badge/Hugging%20Face-FFD21E?logo=huggingface&logoColor=black" alt="Hugging Face"></a>
34
  <br>
35
  <a href="https://modelscope.cn/organization/minimax" target="_blank" rel="noopener noreferrer"><img alt="ModelScope MiniMax AI" src="https://img.shields.io/badge/ModelScope-MiniMax%20AI-white?labelColor=%23EF3D5D"></a>
36
  <a href="https://platform.minimaxi.com/docs/faq/contact-us" target="_blank"><img src="https://img.shields.io/badge/WeChat-07C160?logo=wechat&logoColor=white" alt="WeChat"></a>
37
  <a href="https://discord.com/invite/dbMxutw7tP" target="_blank"><img src="https://img.shields.io/badge/Discord-5865F2?logo=discord&logoColor=white" alt="Discord"></a>
38
- <a href="https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/LICENSE"><img src="https://img.shields.io/badge/LICENSE-4CAF50?logo=creativecommons&logoColor=white" alt="LICENSE"></a>
 
39
  </p>
40
 
41
-
42
  # MiniMax H3
43
 
44
- ## News
45
- Offical skills to improve prompt writing: [skills on github](https://github.com/MiniMax-AI/MiniMax-H3/tree/main/skills)
46
-
47
- ## Online API
48
- Use MiniMax\-H3 directly via API\.
49
- - Global: [platform\.minimax\.io](https://platform.minimax.io/docs/api-reference/video-generation-v2-create) \| CN: [platform\.minimaxi\.com](https://platform.minimaxi.com/docs/api-reference/video-generation-v2-create)
50
-
51
- ## Online App
52
- Use MiniMax\-H3 directly via App\.
53
- - WebApp Global: [hailuoai\.video](https://hailuoai.video/tools/minimax-h3) \| CN: [hailuoai\.com](https://hailuoai.com/)
54
- - Desktop Global: [hub\.minimax\.io](https://hub.minimax.io/) \| CN: [hub\.minimaxi\.com](https://hub.minimaxi.com/)
55
-
56
  ## System Overview
57
  MiniMax H3 is a general-purpose, omni-modal generative system. It supports unified understanding of multimodal contexts composed of text, images, video, and audio, and can generate video with native stereo audio at resolutions up to 2K and durations of up to 15 seconds. Thanks to its task-generalization-oriented system design, H3 already possesses broad multimodal context understanding and generation capabilities at the pre-training stage, enabling outstanding performance in following complex multimodal instructions.
58
 
@@ -72,7 +57,7 @@ H3 supports the following input and output specifications:
72
  | Model Variant | Input Mode | Specifications |
73
  |---|---|---|
74
  | H3-Base-FL2VA | First-and-last-frame mode | Supports zero, one, or two input images. <br><br>- No image input: Text-to-video mode <br>- One image input: First-frame-to-video or last-frame-to-video generation <br>- Two image inputs: First-and-last-frame-to-video generation |
75
- | H3-Base-Ref2VA | Omni-reference mode | Supports multi-modal reference inputs: <br><br>- **Images:** ≤ 9 images <br>- **Videos:** ≤ 3 clips; each clip must be 2–15 seconds long; total duration ≤ 15 seconds <br>- **Audio:** ≤ 3 clips; each clip must be 2–15 seconds long; total duration ≤ 15 seconds <br>- **Mixed inputs:** Maximum number of files across all input types is 12 |
76
 
77
  ![Image](assets/overview.png)
78
 
@@ -81,6 +66,20 @@ The complete H3 system consists of the following three modules:
81
  - H3-Base: Generates audio and video based on the H3-Context-IR output, producing results at 768p resolution.
82
  - H3-Regenerate-2K: Feeds the 768p result together with the original context back into H3 to regenerate the output at 2K resolution. This process leverages both H3’s powerful generative capabilities and the rich information contained in the original context, enabling it to produce high-resolution outputs with more accurate details and greater visual fidelity.
83
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
84
  ## Model Architecture
85
 
86
  ### H3\-Context\-IR
@@ -195,17 +194,15 @@ Each checkpoint is distributed as a self\-contained Hugging Face\-style reposito
195
 
196
  Download the model. The repository hosts the original checkpoint (`FL2VA/`, `Ref2VA/`) and the diffusers format side by side, so scope the download to what your framework needs:
197
 
198
- `model_index.json` is the repository-level public entry. The task-family-specific diffusers indexes remain under `FL2VA/model_index.json` and `Ref2VA/model_index.json`.
199
-
200
  ```bash
201
  # Original checkpoint, both task families (SGLang, vLLM):
202
- hf download MiniMaxAI/MiniMax-H3 --include "model_index.json" "FL2VA/*" "Ref2VA/*" --local-dir MiniMax-H3
203
 
204
  # Or a single task family:
205
- hf download MiniMaxAI/MiniMax-H3 --include "model_index.json" "FL2VA/*" --local-dir MiniMax-H3
206
  ```
207
 
208
- diffusers users do not need a manual download: `ModularPipeline.from_pretrained("MiniMaxAI/MiniMax-H3")` fetches exactly the components it needs. See the [diffusers documentation](https://huggingface.co/docs/diffusers/main/en/api/pipelines/minimax_h3) for loading recipes.
209
 
210
  We recommend the following inference frameworks to serve the model:
211
 
@@ -213,7 +210,7 @@ We recommend the following inference frameworks to serve the model:
213
 
214
  - [vLLM](https://github.com/vllm-project/vllm) \- see [vllm recipes](https://recipes.vllm.ai/MiniMaxAI/MiniMax-H3)
215
 
216
- - [diffusers](https://github.com/huggingface/diffusers) \- see [diffusers docs](https://huggingface.co/docs/diffusers/main/en/api/pipelines/minimax_h3)
217
 
218
  - [ComfyUI](https://github.com/Comfy-Org/ComfyUI) \- see [Comfy tutorial](https://docs.comfy.org/tutorials/video/minimax/minimax-h3); use [R2V template](https://github.com/Comfy-Org/workflow_templates/blob/main/templates/video_minimax_h3_r2v.json) / [T2V template](https://github.com/Comfy-Org/workflow_templates/blob/main/templates/video_minimax_h3_t2v.json)
219
 
@@ -411,13 +408,11 @@ For each case below, we provide reference outputs at both 2K and 768p generated
411
 
412
  [VIDEO\_PROMPT\_WRITING\_GUIDE\_ref\_en\.md](docs/VIDEO_PROMPT_WRITING_GUIDE_ref_en.md)
413
 
414
- skills to improve prompt: https://github.com/MiniMax-AI/MiniMax-H3/tree/main/skills
415
 
416
  ## License
417
 
418
- - MiniMax H3 is released under the [MiniMax H3 Community License Agreement](https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/LICENSE).
419
- - [Q&A about the License](docs/QA-about-License.md)
420
- - [Application form(only for USA/EU/UK/South Korea)](https://platform.minimax.io/h3-license)
421
 
422
  ## Contact Us
423
 
 
3
  license: other
4
  license_name: minimax-h3-community-license-agreement
5
  license_link: LICENSE
6
+ library_name: diffusers
7
  tags:
8
  - text-to-video
9
  - image-to-video
 
18
  - multimodal
19
  - synchronized-audio-video
20
  - reference-to-audio-video
 
21
  ---
22
 
23
  <div align="center">
 
28
  <a href="https://hailuoai.video" target="_blank"><img src="https://img.shields.io/badge/Hailuo%20AI-FF6C37?logo=minimax&logoColor=white" alt="Hailuo AI"></a>
29
  <a href="https://platform.minimax.io/docs/guides/text-generation" target="_blank"><img src="https://img.shields.io/badge/API-FF6C37?logo=minimax&logoColor=white" alt="API"></a>
30
  <a href="https://www.minimax.io" target="_blank"><img src="https://img.shields.io/badge/MiniMax%20Website-FF6C37?logo=minimax&logoColor=white" alt="MiniMax Website"></a>
 
 
31
  <br>
32
  <a href="https://modelscope.cn/organization/minimax" target="_blank" rel="noopener noreferrer"><img alt="ModelScope MiniMax AI" src="https://img.shields.io/badge/ModelScope-MiniMax%20AI-white?labelColor=%23EF3D5D"></a>
33
  <a href="https://platform.minimaxi.com/docs/faq/contact-us" target="_blank"><img src="https://img.shields.io/badge/WeChat-07C160?logo=wechat&logoColor=white" alt="WeChat"></a>
34
  <a href="https://discord.com/invite/dbMxutw7tP" target="_blank"><img src="https://img.shields.io/badge/Discord-5865F2?logo=discord&logoColor=white" alt="Discord"></a>
35
+ <a href="https://huggingface.co/MiniMaxAI" target="_blank"><img src="https://img.shields.io/badge/Hugging%20Face-FFD21E?logo=huggingface&logoColor=black" alt="Hugging Face"></a>
36
+ <a href="https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/LICENSE" target="_blank"><img src="https://img.shields.io/badge/LICENSE-4CAF50?logo=creativecommons&logoColor=white" alt="LICENSE"></a>
37
  </p>
38
 
 
39
  # MiniMax H3
40
 
 
 
 
 
 
 
 
 
 
 
 
 
41
  ## System Overview
42
  MiniMax H3 is a general-purpose, omni-modal generative system. It supports unified understanding of multimodal contexts composed of text, images, video, and audio, and can generate video with native stereo audio at resolutions up to 2K and durations of up to 15 seconds. Thanks to its task-generalization-oriented system design, H3 already possesses broad multimodal context understanding and generation capabilities at the pre-training stage, enabling outstanding performance in following complex multimodal instructions.
43
 
 
57
  | Model Variant | Input Mode | Specifications |
58
  |---|---|---|
59
  | H3-Base-FL2VA | First-and-last-frame mode | Supports zero, one, or two input images. <br><br>- No image input: Text-to-video mode <br>- One image input: First-frame-to-video or last-frame-to-video generation <br>- Two image inputs: First-and-last-frame-to-video generation |
60
+ | H3-Base-Ref2VA | Omni-reference mode | Supports multi-modal reference inputs: <br><br>- **Images:** ≤ 9 images <br>- **Videos:** ≤ 3 clips; each clip must be 2–15 seconds long; total duration ≤ 15 seconds <br>- **Audio:** ≤ 3 clips; audio must be accompanied by image or video input and cannot be used as the sole input; each clip must be 2–15 seconds long; total duration ≤ 15 seconds <br>- **Mixed inputs:** Maximum number of files across all input types is 12 |
61
 
62
  ![Image](assets/overview.png)
63
 
 
66
  - H3-Base: Generates audio and video based on the H3-Context-IR output, producing results at 768p resolution.
67
  - H3-Regenerate-2K: Feeds the 768p result together with the original context back into H3 to regenerate the output at 2K resolution. This process leverages both H3’s powerful generative capabilities and the rich information contained in the original context, enabling it to produce high-resolution outputs with more accurate details and greater visual fidelity.
68
 
69
+ ## Online API
70
+
71
+ Use MiniMax\-H3 directly via API\.
72
+
73
+ - Global: [platform\.minimax\.io](https://platform.minimax.io/docs/api-reference/video-generation-v2-create) \| CN: [platform\.minimaxi\.com](https://platform.minimaxi.com/docs/api-reference/video-generation-v2-create)
74
+
75
+ ## Online App
76
+
77
+ Use MiniMax\-H3 directly via App\.
78
+
79
+ - WebApp Global: [hailuoai\.video](https://hailuoai.video) \| CN: [hailuoai\.com](https://hailuoai.com/)
80
+
81
+ - Desktop Global: [hub\.minimax\.io](https://hub.minimax.io/) \| CN: [hub\.minimaxi\.com](https://hub.minimaxi.com/)
82
+
83
  ## Model Architecture
84
 
85
  ### H3\-Context\-IR
 
194
 
195
  Download the model. The repository hosts the original checkpoint (`FL2VA/`, `Ref2VA/`) and the diffusers format side by side, so scope the download to what your framework needs:
196
 
 
 
197
  ```bash
198
  # Original checkpoint, both task families (SGLang, vLLM):
199
+ hf download MiniMaxAI/MiniMax-H3 --include "FL2VA/*" "Ref2VA/*" --local-dir MiniMax-H3
200
 
201
  # Or a single task family:
202
+ hf download MiniMaxAI/MiniMax-H3 --include "FL2VA/*" --local-dir MiniMax-H3
203
  ```
204
 
205
+ diffusers users do not need a manual download: `ModularPipeline.from_pretrained("MiniMaxAI/MiniMax-H3")` fetches exactly the components it needs. See the [diffusers documentation](https://github.com/huggingface/diffusers/blob/minimax-h3/docs/source/en/api/pipelines/minimax_h3.md) for loading recipes.
206
 
207
  We recommend the following inference frameworks to serve the model:
208
 
 
210
 
211
  - [vLLM](https://github.com/vllm-project/vllm) \- see [vllm recipes](https://recipes.vllm.ai/MiniMaxAI/MiniMax-H3)
212
 
213
+ - [diffusers](https://github.com/huggingface/diffusers) \- see [diffusers docs](https://github.com/huggingface/diffusers/blob/minimax-h3/docs/source/en/api/pipelines/minimax_h3.md)
214
 
215
  - [ComfyUI](https://github.com/Comfy-Org/ComfyUI) \- see [Comfy tutorial](https://docs.comfy.org/tutorials/video/minimax/minimax-h3); use [R2V template](https://github.com/Comfy-Org/workflow_templates/blob/main/templates/video_minimax_h3_r2v.json) / [T2V template](https://github.com/Comfy-Org/workflow_templates/blob/main/templates/video_minimax_h3_t2v.json)
216
 
 
408
 
409
  [VIDEO\_PROMPT\_WRITING\_GUIDE\_ref\_en\.md](docs/VIDEO_PROMPT_WRITING_GUIDE_ref_en.md)
410
 
411
+
412
 
413
  ## License
414
 
415
+ MiniMax H3 is released under the [MiniMax H3 Community License Agreement](https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/LICENSE). [Q&A about the License](docs/QA-about-License.md)
 
 
416
 
417
  ## Contact Us
418
 
assets/action-reference.mov ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:116e0f08a399834e7ffc3472d036659b33250f4ba4f0b7e63f17ec07cb58e4dc
3
+ size 31446151
assets/character-action-reference.png ADDED

Git LFS Details

  • SHA256: 01a6c6f29e1ce9d414276e6e2eca06b171e5a68dd54a86074a4ab77bb9cb8883
  • Pointer size: 132 Bytes
  • Size of remote file: 5.24 MB
assets/character-replacement-action-reference.mp4 ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:796c264126110d87992dcb54213ac0697920cb4ddf3d5a06aa36069386fc4fa1
3
+ size 8223683
assets/fashion-glasses-ad.mp4 ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:a11e6b0095dd7766b724bb886edf4d7d9930af44da6839ee195269ac4fc60ba4
3
+ size 10938085
assets/fashion-glasses-reference-1.png ADDED

Git LFS Details

  • SHA256: 5665f6529dd73339c0ac9397f0bd52df0cd44bd5b63d68af88f95855879d3d8b
  • Pointer size: 132 Bytes
  • Size of remote file: 1.01 MB
assets/fashion-glasses-reference-2.png ADDED

Git LFS Details

  • SHA256: 3a74bdab0bff0e126aae3757833fd2ec8cfba9afcb2ca36f63ee86977f7f7c00
  • Pointer size: 131 Bytes
  • Size of remote file: 889 kB
assets/fashion-glasses-reference-3.png ADDED

Git LFS Details

  • SHA256: 7b1491e5a227f226fa1395e71ac89a84358a435d974acb9f9065b9ded0982cbe
  • Pointer size: 132 Bytes
  • Size of remote file: 1.06 MB
assets/fashion-glasses-reference-4.png ADDED

Git LFS Details

  • SHA256: c9db6d6788855029ef9ac4109b05e17693ddb3bdf6ae3f24cec44de478c00f13
  • Pointer size: 131 Bytes
  • Size of remote file: 368 kB
assets/fl2va-clay-fox-reference.png ADDED

Git LFS Details

  • SHA256: 55eff74469b79a100e63bdd358596f6cff671e18fca9b276d6faef16faf22d27
  • Pointer size: 132 Bytes
  • Size of remote file: 1.73 MB
assets/fl2va-clay-fox.mp4 ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:4604b0838b736ecaa209c978f880bfacbb3a0fc82a8a0517c1f5aa16454b50ef
3
+ size 10931936
assets/h3-architecture.png ADDED

Git LFS Details

  • SHA256: 9765612d331f5fd31068a9283ffe28f11963f6a418c9f4093f81b408ea756630
  • Pointer size: 131 Bytes
  • Size of remote file: 336 kB
assets/h3-cinematic-shot.mp4 ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:436defc81cfa7d53aef423f368be82ead20056088087e0be308b6e3865e8fb81
3
+ size 4361547
assets/h3-suspense-title.mp4 ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:90ebbd7edc71c9a0151c3064126acd2dbb5b7cba58ce459e223e29f2c04f9186
3
+ size 16809220
assets/logo.svg ADDED
assets/reference-image-1.png ADDED

Git LFS Details

  • SHA256: 49b89115da4e253228969830ca491e4ca1b79331eb241a905d1ebca3dc6de7be
  • Pointer size: 131 Bytes
  • Size of remote file: 721 kB
assets/reference-image-2.png ADDED

Git LFS Details

  • SHA256: c9db6d6788855029ef9ac4109b05e17693ddb3bdf6ae3f24cec44de478c00f13
  • Pointer size: 131 Bytes
  • Size of remote file: 368 kB
assets/robot-arm-red-cube.mp4 ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:1a751e6100dbf6502a99a2adc0b12171304ca7b17eecea70e2ed035a73ea692c
3
+ size 2259732
assets/t2va-768p-demo.mp4 ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:d66903241362e224085cc93f7a5e70fba6ab378d0ac6ff6af83acf2559849a42
3
+ size 1637373
docs/QA-about-License.md CHANGED
@@ -45,7 +45,7 @@ Yes.
45
 
46
  Organizations in these regions can apply for a formal license. After reviewing the deployment scenario and confirming that appropriate compliance controls and safeguards are implemented, MiniMax may authorize usage.
47
 
48
- [Application form](https://platform.minimax.io/h3-license)
49
 
50
  Through authorized deployments, MiniMax can ensure that MiniMax-H3 is used responsibly while meeting local legal and regulatory requirements.
51
 
 
45
 
46
  Organizations in these regions can apply for a formal license. After reviewing the deployment scenario and confirming that appropriate compliance controls and safeguards are implemented, MiniMax may authorize usage.
47
 
48
+ [Application form](https://vrfi1sk8a0.feishu.cn/share/base/form/shrcnD9XM1zYI9VFJxTEbt0d19g)
49
 
50
  Through authorized deployments, MiniMax can ensure that MiniMax-H3 is used responsibly while meeting local legal and regulatory requirements.
51
 
model_index.json DELETED
@@ -1,131 +0,0 @@
1
- {
2
- "_class_name": "MiniMaxH3ModularPipeline",
3
- "_diffusers_version": "0.36.0.dev0",
4
- "_blocks_class_name": "MiniMaxH3Blocks",
5
- "text_encoder": [
6
- "transformers",
7
- "Qwen3VLForConditionalGeneration",
8
- {
9
- "type_hint": [
10
- "transformers",
11
- "Qwen3VLForConditionalGeneration"
12
- ],
13
- "pretrained_model_name_or_path": "MiniMaxAI/MiniMax-H3",
14
- "subfolder": "text_encoder",
15
- "variant": null,
16
- "revision": null
17
- }
18
- ],
19
- "tokenizer": [
20
- "transformers",
21
- "Qwen2TokenizerFast",
22
- {
23
- "type_hint": [
24
- "transformers",
25
- "Qwen2TokenizerFast"
26
- ],
27
- "pretrained_model_name_or_path": "MiniMaxAI/MiniMax-H3",
28
- "subfolder": "tokenizer",
29
- "variant": null,
30
- "revision": null
31
- }
32
- ],
33
- "processor": [
34
- "transformers",
35
- "Qwen3VLProcessor",
36
- {
37
- "type_hint": [
38
- "transformers",
39
- "Qwen3VLProcessor"
40
- ],
41
- "pretrained_model_name_or_path": "MiniMaxAI/MiniMax-H3",
42
- "subfolder": "processor",
43
- "variant": null,
44
- "revision": null
45
- }
46
- ],
47
- "vae": [
48
- "diffusers",
49
- "AutoencoderKLMiniMaxH3",
50
- {
51
- "type_hint": [
52
- "diffusers",
53
- "AutoencoderKLMiniMaxH3"
54
- ],
55
- "pretrained_model_name_or_path": "MiniMaxAI/MiniMax-H3",
56
- "subfolder": "vae",
57
- "variant": null,
58
- "revision": null
59
- }
60
- ],
61
- "audio_vae": [
62
- "diffusers",
63
- "AutoencoderKLMiniMaxH3Audio",
64
- {
65
- "type_hint": [
66
- "diffusers",
67
- "AutoencoderKLMiniMaxH3Audio"
68
- ],
69
- "pretrained_model_name_or_path": "MiniMaxAI/MiniMax-H3",
70
- "subfolder": "audio_vae",
71
- "variant": null,
72
- "revision": null
73
- }
74
- ],
75
- "transformer": [
76
- "diffusers",
77
- "MiniMaxH3Transformer3DModel",
78
- {
79
- "type_hint": [
80
- "diffusers",
81
- "MiniMaxH3Transformer3DModel"
82
- ],
83
- "pretrained_model_name_or_path": "MiniMaxAI/MiniMax-H3",
84
- "subfolder": "transformer",
85
- "variant": null,
86
- "revision": null
87
- }
88
- ],
89
- "transformer_ref": [
90
- "diffusers",
91
- "MiniMaxH3Transformer3DModel",
92
- {
93
- "type_hint": [
94
- "diffusers",
95
- "MiniMaxH3Transformer3DModel"
96
- ],
97
- "pretrained_model_name_or_path": "MiniMaxAI/MiniMax-H3",
98
- "subfolder": "transformer_ref",
99
- "variant": null,
100
- "revision": null
101
- }
102
- ],
103
- "scheduler": [
104
- "diffusers",
105
- "MiniMaxH3Scheduler",
106
- {
107
- "type_hint": [
108
- "diffusers",
109
- "MiniMaxH3Scheduler"
110
- ],
111
- "pretrained_model_name_or_path": "MiniMaxAI/MiniMax-H3",
112
- "subfolder": "scheduler",
113
- "variant": null,
114
- "revision": null
115
- }
116
- ],
117
- "audio_scheduler": [
118
- "diffusers",
119
- "MiniMaxH3Scheduler",
120
- {
121
- "type_hint": [
122
- "diffusers",
123
- "MiniMaxH3Scheduler"
124
- ],
125
- "pretrained_model_name_or_path": "MiniMaxAI/MiniMax-H3",
126
- "subfolder": "audio_scheduler",
127
- "variant": null,
128
- "revision": null
129
- }
130
- ]
131
- }