Creating something new. Ex-Meta TBD. On-Leave from Stanford w/ @sanmikoyejo. Prev @ Gemini, MIT, Harvard, Uber, UCL, UC Davis

Mountain View, CA
This week, a new scaling axis was created We can't wait to see how high this new scaling ladder lets us climb ❤️‍🔥
2
5
108
13,851
Tokenizer Myths and Where to Find Them! Slides from talk last night - preprint forthcoming - feedback welcome 🙏 1/11
1
7
111
13,880
One might be wondering: Is a better metric possible? We prove a tokenizer trilemma: out of three reasonable desiderata, a metric can only ever obtain two. Choosing the metric then implicitly chooses what research can be done. 10/11
1
1
6
543
Where does the field go from here? 1. BPB tells us about coding length, nothing else. We have given BPB meaning beyond what it means 2. (tin foil hat) Intelligence is not about compression, but about discriminating correct from plausible-but-incorrect continuations 11/11
1
1
9
504
Gave a fun talk tonight titled "Tokenizer Myths and Where To Find Them"! If anyone wants to give early feedback, please DM for the slides! (Note: This work is totally unrelated to our startup)
3
3
57
3,091
(in full sincerity: congratulations to the Gemini team!)
1
45
2,066
Kinda crazy to get an Oral for your first paper - congratulations @alexxx_zzz6825 !! 🥳🎉🎊
Excited to share our EMNLP oral paper! We find that RL teaches model to traverse parametric knowledge more effectively, thus accessing knowledge previously inaccessible in instruction-tuned models. With @niloofar_mire, @RylanSchaeffer and Manasa Kaniselvan 🧵
3
4
28
7,593
IP Lawyers: Are you using open source software? My cofounder: We don't know what software we're using. Our agents will get back to you.
1
11
1,502
"Tokenization is a fundamental component of every modern language model: before a model sees data, its tokenizer has already determined how the data appears [...] In this work, we identify and challenge three myths in the tokenization research literature." 👀
6
4
77
11,028
are you taking agents seriously?
19
2,079
A new tinfoil hat theory for why RSI might be slower to achieve than math & code capabilities might predict: - Most math on Arxiv is probably mostly correct - Most code on GitHub is probably mostly correct - Most AI/ML research on Arxiv is probably mostly incorrect
30
33
1,151
78,822
- Labs hire based on these probably mostly incorrect publications, meaning that the additional training data they acquire from employees is noisy at best and bad at worst
1
86
8,538