Researcher in deep-learning optimization, LLM training, and reasoning 🧠. Associate professor at UniBasel. Past: EPFL, ETH Zurich (Switzerland) 🇨🇭

Zurich
Aurelien Lucchi retweeted
🎯 Learning is most effective at the frontier of capability: problems that are too easy or too hard teach nothing. For LLM reasoners trained with GRPO this is literal: problems the model always or never solves give zero gradient. We introduce Frontier Learning👇🧵
28
66
1,045
97,279
Aurelien Lucchi retweeted
arXiv has updated our policy on rate limiting for all submitters. This update was made to fairly distribute moderator time & support the arXiv community of staff, volunteers, readers & authors. Please read our announcement to learn more: blog.arxiv.org/2026/10/01/up…
71
483
1,690
695,437
Aurelien Lucchi retweeted
🚨 We turned European train delays into a live league table. Germany finished last. 🇩🇪 Switzerland isn't even #1. 🇨🇭 Over the last 30 days, DelayBahn tracked 900,000+ train stops across 6 countries. Welcome to Europe's Rail Leaderboard 🏆🚆 See where your country ranks 👇 delaybahn.com/rangliste Some of the results are wild 🧵 (1/6)
212
323
3,302
421,138
Aurelien Lucchi retweeted
#AISTATS 2027 Call for Paper is out! This year's edition is taking place in Montreal on May 3-6th, 2027 🇨🇦 Aymeric Dieuleveut and I are chairing virtual.aistats.org/Conferen… New this year: AI review, and more! Abstract deadline: September 29, 2026 Paper deadline: October 6, 2026
4
15
77
4,620
Aurelien Lucchi retweeted
Important points. Note also that users can opt out from use of their (de-identified) data in training. Anthropic has similar policies, though I do not remember such discussions when they announced mathematical results.
Two things to distinguish: Did any human or agent look at user data as part of the Navier Stokes effort? No. Do we use user feedback and de-identified data to improve ChatGPT and Codex in a holistic way? Yes. And so does every LLM company.
21
11
155
75,975
Anthropic: “Claude Team plan for Scientists!” Me: applies Anthropic: “We were unable to verify that you are a scientist.”
2
13
1,602
Aurelien Lucchi retweeted
London’s Hyde Park has transformed into the Dubai desert.
451
1,558
15,642
4,082,385
Nice. Didn't know this was an option.
PSA for anyone preparing a NeurIPS rebuttal (or any OpenReview response). Draft in Google Docs, go to Tools > Preferences > Enable Markdown. When you’re done you can simply select the text, right-click, and “Copy as Markdown”. No more fixing formatting by hand. You’re welcome!
3
1,043
We’re hiring! 🇨🇭 Join the University of Basel as a PhD student or postdoctoral researcher in the field of LLM training. Work on: 🤖 LLM pre- and post-training 🎯 Reinforcement learning and reasoning 📈 Reliable generalization ⚙️ Large-scale experiments 📍 Basel, Switzerland
8
32
235
21,344
This is a joint position with @ilijabogunovic, so you get two advisors instead of one.
1
5
881
Aurelien Lucchi retweeted
Releasing the model weights and technical report of Kimi K3. Kimi K3 is our most capable model: a 2.8T MoE model with native visual understanding and a 1M-token context window. New model architecture: 2.5x the intelligence per unit of compute, not just more params. Alongside Kimi K3, we're opening up more of the stack behind it — high-performance attention kernels, MoE communication library, and infrastructure for running agent environments at scale. Model weights: huggingface.co/moonshotai/Ki… Tech report: github.com/MoonshotAI/Kimi-K… Tech blog: kimi.com/blog/kimi-k3
1,534
7,202
45,924
14,614,736
Aurelien Lucchi retweeted
Slides from my ICML tutorial "Is numerical optimization theory irrelevant to machine learning practice in 2026?": cs.ubc.ca/~schmidtm/Document… (Updated to fix some typos, incorporate feedback, and add some things I did not have time for. Will link to the video here when posted.)
5
85
538
45,367
Aurelien Lucchi retweeted
(1/6) A Warp Drive for Neural PDE Solvers??? Tomorrow in Seoul, Till and Alex are presenting their floral #ICML2026 Spotlight. Don’t miss it if you care about neural surrogates for physics, foundation models, physics world models, …
2
14
69
5,715
Aurelien Lucchi retweeted
Judging optimizer gaps by looking only at language modeling with a fixed batch size is dangerous: one gets only 1/2 of the story. @orientino_ and @ruuustem_10 went beyond. Turns out that for every model, task, and data, there is always a setup where Adam > SGD. 🧵
2
15
74
10,570
Aurelien Lucchi retweeted
Hiring 2 summer ML research interns at the University of Basel 🇨🇭. Research topics: RL/diffusion LLM post-training, reasoning, or LLM orchestration. Possible fully funded PhD offers to follow. I'll be at ICLR this week and happy to chat. Apply: forms.gle/TeeeNU6e7kDH3jX96
12
34
368
23,402
After a lot of hands-on time with Codex and Claude: Claude is significantly better at actually following instructions. Not even close right now.
Codex Compute efficient ✅ Always up, never down ✅ Best at hardcore engineering ✅ Crazy good app, first to escape the terminal ✅
1
730
🚨 PhD position in Reasoning for LLMs at the University of Basel 🇨🇭 Work on: • reasoning in LLMs • diffusion LLMs • theory ↔ real-world applications Top venues (ICML, NeurIPS, ICLR) + strong math/ML focus Joint position with I. Bogunovic @ilijabogunovic
3
15
172
21,129
Why Basel / Switzerland (lifestyle + funding matters) 📍 Basel = top research environment + high quality of life 💰 fully funded PhD with competitive Swiss salary
1
8
1,401