I am currently training a LoRA for the MiniMax-H3 FL2VA model. However, after training on 10k videos, both the video clarity and temporal consistency have noticeably degraded. I am unsure if this issue stems from the prompts, the dataset, or certain training parameters.
I noticed that DiffSynth open-sourced a LoRA checkpoint (https://modelscope.ai/models/DiffSynth-Studio/MiniMax-H3-LoRA-LineartAnime). Could you share the specific training details for this model? I am particularly interested in the prompt design (how to align with the original MiniMax prompt templates), training steps, batch size, timestep shift, learning rate, and LoRA rank. I would love to reference these details to improve my training quality and consistency. Thank you!
I am currently training a LoRA for the MiniMax-H3 FL2VA model. However, after training on 10k videos, both the video clarity and temporal consistency have noticeably degraded. I am unsure if this issue stems from the prompts, the dataset, or certain training parameters.
I noticed that DiffSynth open-sourced a LoRA checkpoint (https://modelscope.ai/models/DiffSynth-Studio/MiniMax-H3-LoRA-LineartAnime). Could you share the specific training details for this model? I am particularly interested in the prompt design (how to align with the original MiniMax prompt templates), training steps, batch size, timestep shift, learning rate, and LoRA rank. I would love to reference these details to improve my training quality and consistency. Thank you!