MiniMax-H3 ORB360 CardSpin (step 50)

A rank-32 LoRA for MiniMax-H3 Ref2VA that does two things from one file, picked by the prompt:

  • ORB360_CARDSPIN: your photo turns in space like a thin physical photo card. The person or animal inside it turns with it: profile on the edge-on sliver (nose sticking out past the edge), the back of their head on the back of the card, the other profile on the way round, and a clean landing back on your photo after exactly 5.125 s.
  • ORB360_CW: the ordinary thing it was built for, a single smooth clockwise 360-degree camera orbit around a frozen subject that ends on the starting view.

Powered by MiniMax H3.

The card spin

One reference image (examples/whitecat.png), the prompt in prompts/cardspin_caption.txt, 1024 x 768, 124 frames, 20 steps, seed 20260926.

Same file, normal orbit

Same image, same seed, same LoRA file; only the prompt changed (prompts/orbit_realscene_caption.txt). The 50 card-spin steps did not break the orbit.

How it happened

We were testing an orbit LoRA (the ORB360 project) on real photos instead of the grey Blender renders it was trained on, and fed it an 1867 portrait of Sir John Herschel by Julia Margaret Cameron. Instead of orbiting a man, it decided the photograph was the object. This is the clip that started it all:

It was too good to leave as a one-off. So we took that single generated video, wrote a caption describing what happens in it with timestamps, and trained 50 more steps on top of the 750-step orbit LoRA using only that one clip. After 50 steps (about half an hour on one GPU) the effect transferred to photos it had never seen: other Cameron portraits, and a colour photo of a cat in a hat. We also checked 100 and 150 steps; they work too, with slightly crisper card edges, but 50 has the most charm and is the lightest touch on the orbit behaviour, so that is the one released here.

Using it

ComfyUI: load it with the standard Load LoRA node (model only, strength 1.0) on a MiniMax-H3 Ref2VA model. It is a kohya-style LoRA (lora_unet_*, lora_down/lora_up/alpha); all 200 modules map onto ComfyUI's MiniMax-H3 weights. Give it one reference image and paste one of the prompts from prompts/ as the text.

musubi-tuner (kohya-ss/musubi-tuner, v0.3.5 or later):

python src/musubi_tuner/minimax_h3_generate_video.py --task ref2va \
  --dit minimax_h3_ref2va_bf16.safetensors --prune_adaln \
  --video_vae minimax_h3_video_vae_fp16.safetensors --audio_vae minimax_h3_audio_vae_fp32.safetensors \
  --text_encoder qwen3vl_32b_minimax_h3_int8_convrot.safetensors --text_encoder_attn_mode sdpa \
  --lora_weight minimax_h3_orb360_cardspin_step50.safetensors --lora_multiplier 1.0 --lora_runtime_attach \
  --video_size 832 672 --video_length 124 --infer_steps 20 --attn_mode sdpa --seed 20260926 \
  --prompt "$(cat prompts/cardspin_caption.txt)" --ref your_photo.png --save_path out/
  • One reference image (--ref in musubi). It becomes <Picture 1> in the prompt and is the first and last frame.
  • Keep 124 frames at 24 fps: the prompt's timestamps (edge-on at 1.0 s, back of the card from 1.7 s, other profile at 3.8 s, back on the photo at 5.125 s) assume that length.
  • Resolutions we have seen work: 672 x 832 and 832 x 1024 portraits, 1024 x 768 and 1152 x 768 landscapes.
  • Portraits of people and animals are the sweet spot. It works on colour photos as well as old sepia prints.
  • For the normal orbit, use prompts/orbit_realscene_caption.txt. It keeps the photo's real surroundings; earlier wording like "clean studio" pulled some scenes towards a grey studio backdrop.
  • Roll a few seeds. This was taught from one clip, so some seeds commit to the card harder than others.
  • The prompts follow MiniMax's official Ref2VA prompt layout (subject_definitions, summary, retention_analysis, detailed_description, overall_soundscape, non_diegetic_music). Both audio sections are N/A, which asks for silence; without them H3 tends to invent a soundtrack.

Training

Both stages used musubi-tuner (v0.3.5, --task ref2va; one local patch that only affects step counting on --resume) with the same settings:

Base MiniMax-H3 Ref2VA, BF16 transformer (minimax_h3_ref2va_bf16 repack), loaded with --prune_adaln
Training adapter ostris/minimax_h3_training_adapter minimax_h3_ref2va_training_adapter_v1 via --base_weights (training only; not used at inference)
LoRA rank 32, alpha 32 (networks.lora_minimax_h3)
Optimizer AdamW, lr 1e-4 constant, max grad norm 1.0
Precision bf16 mixed, gradient checkpointing, SDPA
Batch 1, no accumulation, video only (no audio loss), seed 20260926
Hardware one NVIDIA RTX PRO 6000 (96 GB)

Stage 1: ORB360 orbit, 750 steps, from scratch. Four assets rendered in Blender, each as one 124-frame, 24 fps, 512 x 512 clockwise orbit with exactly one full turn. Three reference images per clip, rendered with the exact cameras of frames 0, 41 and 82 (0, 120 and 240 degrees). The captions give timestamped milestones (for example "At 1.708333 seconds (zero-based frame 41), the camera has advanced 120 degrees clockwise ... corresponds to <Picture 2>") and ask for a completely frozen subject. About 18 s per step.

Stage 2: CardSpin, +50 steps. Initialised from the stage-1 weights (fresh optimizer). One training clip: the Herschel glitch above (generated at 832 x 1024, trained at 672 x 832, 124 frames), with the original Herschel photo as the single reference and the card-spin caption in prompts/cardspin_caption.txt. About 40 s per step, 73 GB peak VRAM.

Limitations

  • Learned from a single example. The card always turns the same way with roughly the same timing, and some subjects or seeds may commit to the card less than others.
  • The first stage saw only four rendered assets at 512 x 512, so orbit quality on complex real scenes varies.
  • Evaluated by eye on a handful of images and seeds, not with a benchmark.

What's next

This is a side quest. The main ORB360 work is training a much steadier 360 LoRA on a larger rendered set: more assets, several orbit lengths and speeds, higher and lower camera heights, closer and wider framing, and captions that describe the exact camera for every clip. The aim is to see whether consistent, controllable orbits can lead to some kind of stable geometry adapter. More on that when it works.

Files

  • minimax_h3_orb360_cardspin_step50.safetensors: the LoRA (fp32, 597 MB)
  • prompts/cardspin_caption.txt, prompts/orbit_realscene_caption.txt: the two prompts behind the cat videos
  • examples/whitecat.png: reference image for the cat clips
  • examples/herschel_1867_cameron_met263166.png: the Herschel portrait (Julia Margaret Cameron, 1867; The Metropolitan Museum of Art, Open Access, CC0), resized
  • assets/: the three videos above and their poster frames
  • LICENSE (MiniMax H3 Community License Agreement), NOTICE

Base model and licence

This is a LoRA for MiniMax-H3 by MiniMax; it does nothing without their base weights, and its second stage was trained on a clip generated with them. It is distributed under the MiniMax H3 Community License Agreement, including its acceptable-use, distribution, commercial and territorial provisions. It is not MIT. Read the complete upstream terms; this repository does not expand them. See NOTICE for attribution and the modification notice. Thanks to MiniMax for releasing H3, to Ostris for the training adapter, and to kohya-ss for musubi-tuner.

Downloads last month
5
Inference Providers NEW

Model tree for MATLOWAI/MiniMax-H3-ORB360-CardSpin

Adapter
(98)
this model