Say hello to DreamX-Creator 1.0, a 7B native audio-video generator with local 2K refinement. Apache 2.0.
🤖
modelscope.ai/models/GD-ML/D…
🏆 Delivers results comparable to much larger SOTA models like MiniMax H3 on Verse Bench, reaching 0.1351 DeSync after RL and 0.6930 VQ with the 2K Refiner.
🎬 Turn a first frame and prompt into synchronized video, dialogue, action sounds, ambience, weather, crowds, and music. No separate dubbing stage.
🔊 Gated Cross-Modal Attention lets sound and visuals shape each other, while modality-aware RL improves video, audio, and their alignment.
✨ The one-step 2K Refiner tops the compared refinement methods on MUSIQ (0.7073) and MANIQA (0.4382) while preserving motion and audio timing.