Research Scientist/Engineer at Alibaba Tongyi Lab. Core authors of Wan2.7/2.6/2.5, Crafter family: VideoCrafter1, DynamiCrafter, ToonCrafter, ViewCrafter

Hong Kong
We are launching AA-Video-T2V v2.0, our new benchmark for evaluating text to video models, alongside AA-Video-T2V-Silent v2.0 for video generation without audio. Built on a new methodology, it judges every model at 1080p on a regularly refreshed prompt set, and ranks them across 10 use cases, 10 capabilities and a wide range of styles. Video models are being adopted across more industries and workflows, from film studios to advertising agencies. Our new benchmark not only ranks models overall, but also shows which model is best for specific use case and video generation capability. Use cases are grounded in how consumers and enterprises use video generation. Capabilities draw on lab and academic research, and on how creators and businesses push video models today. We tag every prompt by use case (such as Live-Action Film and Marketing & Advertising) and by the capability it tests (such as Text Rendering and Audio Synchronization), and the overall benchmark samples evenly across both. Because of this, the overall ranking reflects a model's versatility across use cases and well-roundedness across capabilities. We also tag each prompt by visual style, such as photorealistic, 3D render, cartoon and anime, and hand-drawn illustration. AI video is also moving onto bigger screens and into production, from microdramas to movie theaters, while low barrier to generate is resulting in a proliferation of low quality AI video content. The quality bar keeps rising, so we now judge every clip at 1080p and high bitrate. We are launching AA-Video-T2V v2.0 with more than 68,000 high quality human preference votes from private evaluators based in US/UK over 1,000 prompts, and AA-Video-T2V-Silent v2.0 with more than 47,000 votes over 500 prompts. Initial insights from an in-depth analysis of the 10 highest ranking models on the Artificial Analysis AA-Video-T2V v2.0 Leaderboard: ➤ Wan 3.0 ranks #1 overall and leads 10 of the 20 category boards, including Cartoon and Anime style and Animation & Gaming use case, at $12 per minute of video. ➤ Dreamina Seedance 2.5 ranks #2 and is the human performance specialist, #1 on both Human Anatomy and Dialogue & Lip Sync. At $34.12/min it is the most expensive model in the top 10. ➤ MiniMax H3 (768p) ranks #3, statistically tied with Seedance 2.5 at $4.80/min, about 1/7 of the price. It is also #1 on Text Rendering. ➤ FLUX 3 ranks #4, with its strongest results on Text Rendering (#3) and Dialogue & Lip Sync (#2). ➤ Gemini Omni Flash 1.1 ranks #5 and is the graphic 2D and audio specialist, #1 on UI/UX & Motion Design use case, Flat Design style and Audio Synchronization capability. See below for the use case, capability and style breakdowns 🧵
28
15
214
31,855
Excited to introduce Wan3.0. Simple Input. Smart Creation. • Native 30-Second Video Generation • Reality-Grade Rendering • Omni-Reference: Beyond text, images, audio, and video—now with documents, spreadsheets, slides, webpages, and more.
1
242
Jinbo Xing retweeted
Introducing Wan3.0 — now in Public Beta. Simple Input. Smart Creation. Where Imagination Meets Reality. • Native 30-Second Video Generation • Reality-Grade Rendering • Omni-Reference: Beyond text, images, audio, and video—now with documents, spreadsheets, slides, webpages, and more. Public Beta is now live. Apply now and start creating ↓
352
337
2,784
20,102,458
The Wan-Image (Wan 2.7) technical report is now live! 📄✨ 🤩Curious about how it works under the hood? Check out all the details and our latest findings here: arxiv.org/abs/2604.19858v2 🎯Web service: wan.video/
4
169
Jinbo Xing retweeted
1/8 Meet Wan2.7-Image — Our unified model for image generation and editing. One model that generates, edits, and understands images: Realistic faces with full control over bone structure, eyes, and contour Color Palette with HEX codes and reference image extraction 3K-token text rendering across 12 languages, print-quality Interactive editing with visual instructions on your images Up to 12 consistent images in one generation Details on each below ↓
67
123
1,065
3,449,644
Jinbo Xing retweeted
这是目前我看到的最自然也是最流畅的AI数字人。 只要你有能力,完全可以自己创造一个专属的IP,让它去出镜帮你完成很多事情,比如做博主、拍广告、做评测、讲课…就像真人那样。
Made with AI
50
113
829
114,112
🚩 We’re thrilled to announce our CVPR2026 Workshop on GenAI/AIGC: 2nd 𝐇𝐢𝐆𝐞𝐧: 𝐇𝐮𝐦𝐚𝐧-𝐈𝐧𝐭𝐞𝐫𝐚𝐜𝐭𝐢𝐯𝐞 𝐆𝐞𝐧𝐞𝐫𝐚𝐭𝐢𝐨𝐧 𝐚𝐧𝐝 𝐄𝐝𝐢𝐭𝐢𝐧𝐠. 📷 higen-2025.github.io 📷openreview.net/group?id=thec… 📷Deadline: March 15 2026⏲️
1
2
8
2,228
Jinbo Xing retweeted
NVIDIA just dropped PersonaPlex-7B 🤯 A full-duplex voice model that listens and talks at the same time. No pauses. No turn-taking. Real conversation. 100% open source. Free. Voice AI just leveled up. huggingface.co/nvidia/person…
149
1,024
9,349
2,639,947
Jinbo Xing retweeted
Wan App is now live on iOS & Android! 🚀 Unlock the full spectrum of AI visual creation: - Starring: Cast your characters into any video—perfect identity & voice consistency, no retraining needed. - Video: From T2V to I2V, enjoy cinematic motion, precise control, native A/V sync, and multi-shot storytelling that brings your vision to life. - Image: Craft photorealistic visuals or refine with studio-grade editing tools. Ready to create? Scan the QR code and launch your AI journey today!
175
234
2,767
5,239,395
Jinbo Xing retweeted
Ready to be amazed? 🚀 Our new concept film is out! See how Wan2.2 makes life bloom in every frame. See it for yourself! 👇
71
234
1,369
168,939
🚩 We’re thrilled to announce our ICCV 2025 Workshop on GenAI/AIGC: 𝐇𝐢𝐆𝐞𝐧: 𝐇𝐮𝐦𝐚𝐧-𝐈𝐧𝐭𝐞𝐫𝐚𝐜𝐭𝐢𝐯𝐞 𝐆𝐞𝐧𝐞𝐫𝐚𝐭𝐢𝐨𝐧 𝐚𝐧𝐝 𝐄𝐝𝐢𝐭𝐢𝐧𝐠. 🎯 higen-2025.github.io 🚪Submit: openreview.net/group?id=thec… ⏲️Deadline: June 30 2025 23:59 UTC
2
12
2,988
Jinbo Xing retweeted
Excited to share our #TrajectoryCrafter, a diffusion model for Redirecting Camera Trajectory in Monocular Videos! Try to explore the world underlying your videos~ Page: trajectorycrafter.github.io Demo: huggingface.co/spaces/Doubii… Code: github.com/TrajectoryCrafter…
5
52
265
14,094
Jinbo Xing retweeted
😘This is not a drill—Wan 2.1 OPEN SOURCE is finally here! ⏰Live Broadcast Time: Feb 25th,2025, 11:00 PM(UTC+8)
48
111
884
98,984
Jinbo Xing retweeted
We propose GenProp✨, a generative video propagation framework, which can seamlessly propagate any first frame edit through the video. 🧵(1/n) - Arxiv: arxiv.org/abs/2412.19761 - Project Page: genprop.github.io/ - Video: piped.video/watch?v=GC8qfWzZ…
4
9
21
3,811
Jinbo Xing retweeted
We present UniReal, a universal framework for multiple image generation and editing tasks. Webpage: xavierchen34.github.io/UniRe… Paper: arxiv.org/abs/2412.07774
6
12
62
16,881
Jinbo Xing retweeted
Introducing Hyperstroke, a vector quantized representation to efficiently represent strokes during the incremental drawing process. Potential downstreams including artistic drawing, sketch understanding and processing, etc. dl.acm.org/doi/10.1145/36817… (SIGGRAPH Asia 2024 TC Track)
4
7
61
17,614
Glad to see our 𝐓𝐨𝐨𝐧𝐂𝐫𝐚𝐟𝐭𝐞𝐫 has been selected in the SIGGRAPH Asia 2024 Technical Papers Program Trailer! piped.video/watch?v=hUktDm6W… #SIGGRAPHasia
1
147
Jinbo Xing retweeted
Alibaba presents MIMO Controllable Character Video Synthesis with Spatial Decomposed Modeling Character video synthesis aims to produce realistic videos of animatable characters within lifelike scenes. As a fundamental problem in the computer vision and graphics community, 3D works typically require multi-view captures for per-case training, which severely limits their applicability of modeling arbitrary characters in a short time. Recent 2D methods break this limitation via pre-trained diffusion models, but they struggle for pose generality and scene interaction. To this end, we propose MIMO, a novel framework which can not only synthesize character videos with controllable attributes (i.e., character, motion and scene) provided by simple user inputs, but also simultaneously achieve advanced scalability to arbitrary characters, generality to novel 3D motions, and applicability to interactive real-world scenes in a unified framework. The core idea is to encode the 2D video to compact spatial codes, considering the inherent 3D nature of video occurrence. Concretely, we lift the 2D frame pixels into 3D using monocular depth estimators, and decompose the video clip to three spatial components (i.e., main human, underlying scene, and floating occlusion) in hierarchical layers based on the 3D depth. These components are further encoded to canonical identity code, structured motion code and full scene code, which are utilized as control signals of synthesis process. The design of spatial decomposed modeling enables flexible user control, complex motion expression, as well as 3D-aware synthesis for scene interactions. Experimental results demonstrate effectiveness and robustness of the proposed method.
13
211
1,295
149,284
Jinbo Xing retweeted
Introducing our StereoCrafter, a tool for transforming any video into stereo content, cooperating with depth estimation approaches, like our DepthCrafter. Hope this can benefit the XR community. Project page: stereocrafter.github.io/ Paper: arxiv.org/abs/2409.07447
2
16
100
7,595
Introducing 𝚅𝚒𝚎𝚠𝙲𝚛𝚊𝚏𝚝𝚎𝚛 🥳. 𝚅𝚒𝚎𝚠𝙲𝚛𝚊𝚏𝚝𝚎𝚛 can generate high-fidelity novel views from single or sparse input images with accurate camera pose control! ✨Paper: arxiv.org/abs/2409.02048 🎯Code: github.com/Drexubery/ViewCra… 🥁Demo: huggingface.co/spaces/Doubii…
7
77
377
40,776