During inference, WorldDiT acts for a few steps, observes what changed, and replans. World modeling stays in training, while light-weight deployment remains action only.
1
8
585
WorldDiT learns to sample robot actions and predict a future view of the world in parallel through one diffusion backbone. That richer signal helps it reach strong LIBERO results with far fewer parameters. Table below shows the full benchmark comparison.
1
9
1,844
Many robot policies need a massive pretrained VLM to generate actions, bringing billions of parameters and their deployment cost into the control loop. WorldDiT does not need one. It gives us an architecture we can scale through joint world and action modeling.
1
1
8
901
We are releasing WorldDiT, a unified architecture for robotics world modeling and control. On the LIBERO benchmark, it performs the best among all publicly released methods that do not need a VLM to generate actions. Its size and performance sit on the reported Pareto frontier.
28
54
154
71,076
Today we're releasing Paris 2.0, to our knowledge the first decentralized-trained video generation model. At Bagel Labs, we believe frontier models should not require homogeneous clusters of premium, supply constrained GPUs. Paris 1.0 proved this for image generation. Paris 2.0 extends the recipe to video generation and lays the substrate for global-scale world models. To test the approach, we trained two models head-to-head in an iso-FLOP, iso-data comparison. One was a monolithic model trained conventionally, on a single premium GPU cluster. The other was Paris 2.0, trained across an extreme mix of GPU types, generations, and vendors distributed around the globe. Against the monolithic model under matched data and compute, the results were: FVD: 561.04 → 279.01 (a ~2x improvement) CLIP text-video alignment and aesthetic score both improved. To our knowledge, this is the first distributed training architecture to surpass its monolithic counterpart under matched data and compute. Technical Report: arxiv.org/abs/2605.26064 Model Weights: huggingface.co/bageldotcom/p…
We're releasing Paris 2.0, which, to our knowledge, is the world's first decentralized trained video generation model. We benchmarked it against a monolithic model trained on the same data and compute budget, and Paris 2.0 outperformed the monolithic by ~2x on FVD benchmark.
6
6
35
7,895
In town for NVIDIA GTC? If you're building generative world models or investing in the people who are - we're putting the right people in one room for you tomorrow night in Palo Alto. Co-hosted by Alumni Ventures. Signup link below.
5
6
21
11,573
Diffusion models are becoming the foundation for image, video, and world models. We are hosting a founders and investors gathering on that topic during NVIDIA GTC week, co-hosted by our friends at Alumni Ventures. Mar 16, Menlo Park. Sign up below. luma.com/nvidia-gtc-generati…
3
8
18
6,024
Paris - made with ❤️ by bagel labs
Introducing Paris - world's first decentralized trained open-weight diffusion model. We named it Paris after the city that has always been a refuge for those creating without permission. Paris is open for research and commercial use.
9
6
58
8,758
The results. These images came from 8 experts that never spoke to each other during training. We believe if we can scale this approach, this is the first real step towards open source superintelligence. But that requires solving some more really really hard problems. If you're interested in helping us achieve this while doing the best open-source work of your life, come work with us, jobs.bagel.com
3
2
68
13,198
The numbers. Paris achieved comparable results to SOTA decentralized approaches while using: 14× less training data (11M vs 158M images) 16× less compute (120 A40 GPU-days vs ~1176 A100-days) Paris also wins against monolithic training baselines. Our Top-2 routing on DiT-B/2 hits FID-50K of 22.60, a 7.04 point improvement over single-model training (29.64).
1
2
43
4,837
Here's what we did differently. Distributed training typically uses parallelism techniques like data parallelism, pipeline parallelism, model parallelism etc. All require synchronization between compute nodes. We removed this requirement entirely with Paris through decentralized flow matching. After training, we built a lightweight DiTRouter, also in complete isolation, that learned to select experts at inference based on noisy latents.
1
2
58
5,084
Paris does something that shouldn't work. It's a combination of smaller expert diffusion models pre-trained from scratch, across different continents in complete isolation. Absolutely zero synchronization among each other during training. This zero communication protocol achieves comparable quality to SOTA distributed approaches using 14× less data and 16× less compute. How? See our full technical report and model weights below. Full Technical Report: github.com/bageldotcom/paris… Model Weights: huggingface.co/bageldotcom/p…
2
9
125
22,337
Introducing Paris - world's first decentralized trained open-weight diffusion model. We named it Paris after the city that has always been a refuge for those creating without permission. Paris is open for research and commercial use.
146
156
963
607,809
Octave Convolution
7
4
52
26,371
bagel dot com
54
18
362
110,061
Mixture of Experts
9
7
90
14,027
Forward diffusion
12
9
104
10,093
Thank you all for the incredible turnout at ICML! It was an absolute pleasure connecting with so many of you during our paper presentation, at our exhibitor booth, and at the exclusive distributed ML mixer. See you all next time, with more frontier distributed ML research and more merch. Together, we can just do things!
We’re excited to share that Bagel Labs is one of the official ICML exhibitors! Next week, the Bagel team will be exhibiting novel research, participating in poster sessions, and hosting a special side event at ICML, Vancouver. See below to find out where you can find us!
5
7
45
21,202
Replying to @UnityEagle

ALT You Get A Kid And You Get A Kid GIF

1
4
117
Replying to @UnityEagle

ALT Bagel Excited GIF

1
52
And lastly, we love this prompt for an outro with text! Prompt: an incredibly detailed first-person POV of acceleration, retro futuristic old school pulp magazines BAGEL LABS, bright orange with DVD scanlines and authentic film grain as well as CRT glitching, trails of light escaping through VFX, cinematic lighting and composition, mysterious and eerie atmosphere at the end with hints of gorgeous cinematography, synthwave music
1
1
15
1,260
Here, we added one small change to the prompt and the music felt slightly more liminal. Prompt: a synthwave retro futuristic old school music video, bagels, bright orange with DVD scanlines and authentic film grain as well as CRT glitching, trails of light escaping through VFX, cinematic lighting and composition, mysterious retrowave atmosphere, with hints of gorgeous cinematography, synthwave music in the key of C, vaporwave elements
1
1
17
699
Instead of trying to specify chord progressions, you can also just try keys and BPM! It's not perfect, but it is interesting. Prompt: a synthwave retro futuristic old school music video, bagels, bright orange with DVD scanlines and authentic film grain as well as CRT glitching, trails of light escaping through VFX, cinematic lighting and composition, mysterious retrowave atmosphere, with hints of gorgeous cinematography, synthwave music in the key of F minor, 147 BPM
17
430

ALT Bagel Excited GIF

1
48
We’re excited to share that Bagel Labs is one of the official ICML exhibitors! Next week, the Bagel team will be exhibiting novel research, participating in poster sessions, and hosting a special side event at ICML, Vancouver. See below to find out where you can find us!
8
11
89
17,138
Here’s how the performance of the training of a model using the library looks. The evaluation data is displayed in the TensorBoard dashboard for the Qwen-3 model, demonstrating that the Tiny Tool Use library offers clear and interpretable training and evaluation metrics along with improved model capability for function calling.
1
13
933
Replying to @BrentLynch

ALT You Get A Kid And You Get A Kid GIF

2
27
Replying to @Alterverse_AI
This is amazing! Stellar work, it's incredibly cinematic :).

ALT Wow Amazed GIF

1
1
26
Replying to @FotachuARGUY

ALT You Get A Kid And You Get A Kid GIF

1
2
259
We're excited to share that our zkLoRA paper will be presented at the Open Source Development Workshop at ICML 2025. One of the first zero-knowledge verification works ever at ICML. We believe it's impossible for decentralized, collaborative training systems to work under adversarial attack without a 100% verifiability guarantee. zkLoRA delivers those guarantees for low-rank compression - one of the go-to methods for bandwidth-efficient collaborative model training - with zero-knowledge proofs. At real-time latency.
8
14
56
29,474
Replying to @taherdhanera

ALT Youre The Best Christina Aguilera GIF

17
Replying to @Alterverse_AI

ALT You Rock GIF

1
20
Replying to @weirdai_art @bidhan

ALT Lou Bagel Dance GIF

1
31
Replying to @aDimensionDoor

ALT Happy Chris Pratt GIF by Parks and Recreation

1
15

ALT Hehe GIF

1
1
21
zkLoRA v0.2 is out! zkLoRA is the first ever open-source framework that lets many parties train or serve a model while everyone can verify steps with zero-knowledge proofs in real time. What's new in this update, - The new version adds polynomial commitments for every single activation. That means, the system can patch thousands of LoRA adapters into one run and still prove every contribution, in just one single forward pass. - Each activation is sealed with a proof that prevents tampering and hides the parameters. That means planet scale collaboration without the need for auditing weights. - Replaces the old file-hash with a BLAKE3 Merkle tree that outputs a 32-byte root and a 32-byte random nonce. Allowing verification in logarithmic time so the cost stays low even for thousands of contributors. Give it a try. Paper and github link below.
3
14
56
22,709
Return on Experience (RoE) Can a single number line up every milestone so far and hint at the next leap? We think maybe. We call that number Return on Experience (RoE). It’s an early phase benchmark. Link: blog.bagel.net/p/return-on-e…
1
4
658
With Great Data, Comes Great Responsibility A practitioner’s guide to Privacy-Preserving machine learning. Differential privacy, Federated Learning, FHE, TEE etc. And they compare against each other, Link: blog.bagel.net/p/with-great-…
1
1
97
Missed out on the bagels last week? Follow us @bagelopenAI for updates on where to find more delicious bagels at our next event. Btw, we're hiring : jobs.bagel.net
3
671
“Open access to super intelligence is the most important right of the modern age." - @bidhan shared his vision of a world where intelligence is created by all and for all.
1
1
127
And we’re on! Come by at 225 Richmond St W in Toronto
1
11
835
Builder’s Day with Google for Startups on 2/8 in San Francisco 🥯 RSVP to connect with dozens of builders committed to a future of monetizable open source AI Sign up now: lu.ma/mbtw3s0b
1
3
18
894
Monetizable Open Source AI turned NeurIPS 2024 orange. Timestamp: @NeurIPSConf 2024
4
12
1,851
The Open Source AI Mixer @NeurIPSConf sparked bold ideas. Thanks to all who joined. Conversations like these will shape the future of Monetizable Open Source AI 🥯 More to come.
5
14
1,922
Nearly 200 of the top AI builders and experts attending @NeurIPSConf have registered to attend. Join us. We'll be on the ground with the team from @MorpheusAIs talking all things Monetizable Open Source AI and more. See you tomorrow.
2
4
14
4,099
We're excited for our founder @bidhan to present today at AI Keynotes at Microsoft on how Bagel is monetizing open-source AI for developers everywhere.  If you're in the Bay area we invite you to attend 🥯 Details in comments 👇
1
5
17
722
Available now on Bakery🥯: Qwen2.5-Coder-7B-Instruct This model rivals GPT-4 in coding, excelling in math, reasoning, and 40+ programming languages. It scores high on EvalPlus, LiveCodeBench, Aider (73.7), and McEval (65.9), making it a leader in code generation and repair. As the flagship of the Qwen2.5-Coder family, it supports code assistants, visual artifact creation, and enterprise-grade workflows. Licensed under Apache 2.0, it’s built for debugging, reasoning, and efficient development. Link in comments to start using Qwen2.5-Coder-7B-Instruct on Bakery today.
2
14
78
1,170
This month we’re expanding the private beta launch of our Bakery platform, marking a major step forward towards our vision. For the first time, developers gain access to decentralized fine-tuning and secure AI resource sharing, within a system that rewards contributions and collaboration. Sign up below for access🥯 bakery.bagel.net
1
76
Here's a brief example of how fine tuning works using Bagel: To start the fine tuning process, your dataset is loaded along with a model (which is available for purchase on the Bakery).
1
163
In 2nd place, we have pincast @amacdowell , who used the Bakery to train their LLM on a fascinating part of NYC’s open data. They’re launching a mobile app as their first real-world experiment! See their demo here: piped.video/STgVXIKkK9g
2
98
In the forward diffusion, Gaussian noise is incrementally added to an image, transforming it step-by-step into a noisy version. The reverse diffusion uses a deep neural network to reconstruct the image from from the noisy data, aiming to learn a probability distribution represented by model parameters θ.
3
189
Diffusion Models operate through a dual-process system: forward diffusion adds noise to data, and reverse diffusion reconstructs it. This method is grounded in principles of non-equilibrium thermodynamics, making it a powerful tool for synthetic data generation.
2
3
188
5/10 ZKML training involves validating the correct execution of a neural network on a labeled dataset, iteratively over multiple epochs, without revealing the evolving weights. A compressed proof is generated, verifying the accuracy of the training process. This requires an additional arithmetic circuit to implement the optimization function and minimize the loss. Each training epoch generates a proof, confirming the correctness of the previous epochs.
1
2
224
3/10 In ZKML inference, a pretrained neural network processes an unlabeled dataset, producing labels and ZK proofs. The proofs attest to the existence of specific weights that generated the labels, without revealing the weights. This is accomplished by representing the neural network as an arithmetic circuit, with the weights and dataset acting as private and public inputs, respectively. The output consists of the label and a zero-knowledge argument.
1
2
282
Why AI Needs Web3? The panel will take place on April 19th at TOKEN2049. @bidhan will be alongside @mheinrich and @xiangrenNLP, at the Zeebu Stage, to discuss about the intersection.
3
95
Join us on April 17th at OFR Dubai for the Singularity: AI x Crypto event presented by IOSG Ventures. @IOSGVC Our CEO @bidhan will dive into the Privacy First Decentralized AI Ecosystem during the Data + AI Model panel. Explore how Bagel creates a secure environment for human-AI interaction, trading, and the improvement of LLMs through machine learning datasets. Secure your spot: lu.ma/ofr.dubai
1
5
145