Introducing WorldCrafter🌍 Turn an image or text into explorable video worlds. Move, turn, and revisit places. 🎮 Precise camera control 🧠 3D-aware memory & consistent views ⚡ Interactive demo 🤗 Open code & weights github.com/TencentARC/WorldC… #WorldModels #SpatialIntelligence
3
4
338
Wangbo Yu retweeted
🚀 LongVILA is open-sourced: our comprehensive solution for scaling long-context Visual-Language Models (VLMs) to tackle the challenges of long video understanding! 🎥📖 - Paper: arxiv.org/pdf/2408.10188 - Code: github.com/NVlabs/VILA/tree/… - Models: huggingface.co/collections/E… 🔍 What is LongVILA? A full-stack framework co-designed for both algorithm and system optimization, enabling long-context multi-modal learning at an unprecedented scale. 💡 Why it matters: Long-context multi-modal understanding is key for applications like real-world robotics, GPT-like multi-modal agents, and long video summarization. 🔑 Key Contributions: 1️⃣ Long Context Capability: Handling over 1M tokens for tasks like Needle-in-a-Haystack, achieving 99.8% accuracy. 2️⃣ Cutting-Edge System Design: Developed the Multi-Modal Sequence Parallelism (MM-SP) system: Scales to 2M tokens context length on 256 GPUs 🚀 2.1×–5.7× faster than ring-style parallelism ⚡ 3️⃣ Efficient Training Pipeline: Context extension, and long-context fine-tuning for robust video understanding, with long video SFT datasets. 4️⃣ Strong Benchmark Performance on 9 video benchmarks, from ActivityNet-QA to VideoMME. Many thanks for the great team! @XueFz @DachengLi177 @huqinghao @Yao__Lu @songhan_mit Let us know your thoughts and questions! #LLM #LLMs #AI #Video #NVIDIA
3
19
66
11,130
Wangbo Yu retweeted
With ViewCrafter that generates high-quality novel views from a single image, we can get a sense of looking around the Creation Pillars, an enormous structure spanning 70 by 55 light-years and located 6,500 light-years away! Paper: arxiv.org/abs/2409.02048 Code: github.com/Drexubery/ViewCra… Model: huggingface.co/spaces/Doubii…
4
28
2,548
Wangbo Yu retweeted
Excited to share our DepthCrafter, a super consistent video depth model for long open-world videos! Project webpage: depthcrafter.github.io/
DepthCrafter Generating Consistent Long Depth Sequences for Open-world Videos paper page: huggingface.co/papers/2409.0… Despite significant advancements in monocular depth estimation for static images, estimating video depth in the open world remains challenging, since open-world videos are extremely diverse in content, motion, camera movement, and length. We present DepthCrafter, an innovative method for generating temporally consistent long depth sequences with intricate details for open-world videos, without requiring any supplementary information such as camera poses or optical flow. DepthCrafter achieves generalization ability to open-world videos by training a video-to-depth model from a pre-trained image-to-video diffusion model, through our meticulously designed three-stage training strategy with the compiled paired video-depth datasets. Our training approach enables the model to generate depth sequences with variable lengths at one time, up to 110 frames, and harvest both precise depth details and rich content diversity from realistic and synthetic datasets. We also propose an inference strategy that processes extremely long videos through segment-wise estimation and seamless stitching. Comprehensive evaluations on multiple datasets reveal that DepthCrafter achieves state-of-the-art performance in open-world video depth estimation under zero-shot settings. Furthermore, DepthCrafter facilitates various downstream applications, including depth-based visual effects and conditional video generation.
13
178
923
253,621
Wangbo Yu retweeted
Introducing our StereoCrafter, a tool for transforming any video into stereo content, cooperating with depth estimation approaches, like our DepthCrafter. Hope this can benefit the XR community. Project page: stereocrafter.github.io/ Paper: arxiv.org/abs/2409.07447
2
16
100
7,595
Introducing 𝚅𝚒𝚎𝚠𝙲𝚛𝚊𝚏𝚝𝚎𝚛 🥳. 𝚅𝚒𝚎𝚠𝙲𝚛𝚊𝚏𝚝𝚎𝚛 can generate high-fidelity novel views from single or sparse input images with accurate camera pose control! ✨Paper: arxiv.org/abs/2409.02048 🎯Code: github.com/Drexubery/ViewCra… 🥁Demo: huggingface.co/spaces/Doubii…
7
77
377
40,776