Civitai Description:
beta 4:
rebuilt on beta3 config with some loras changed out for newer versions. Turbo was consensus merged to construct a ref/t2va hybrid turbo lora. That's the key element to the merge, there is no non-turbo version. The non-turbo version is bad. That custom turbo merge will need further improvement. This one preforms video, reference, and motion well in 6-8 steps without the issues from beta3, but audio needs shift configuration.
Use sampling like Euler/simple 8 steps with sampling shift - 12 video/ 7+ audio. LCM/simple or beta with 6-8 steps and no shift can also be better for audio and drawn styles.
Audio is lackluster and it's becoming somewhat apparent that H3's integrated audio is not good, like terrible actually. Any future multimodal models should avoid integrated audio if they intend to open source. I have almost no control over how the audio works inside the model. Don't post about it. I focused on motion and prompt response and of course when I get those working well, the audio ends up bad, go figure. I'll look at what kind of different turbo configs can enable more audio crispness to come back or likely will have to wait for Sulphur to replace MysticXXX which is contributing to the audio quality drop.
