Machine Learning Engineer. I learn by teaching. Opinions are my own. Building bestgpusforai.com/

Toronto, Ontario
Dmitry Noranovich retweeted
Reasoning from scratch, round number 6! An introduction (and implementation) of Reinforcement Learning with Verifiable Rewards (RLVR) and Group Relative Policy Optimization (GRPO). 00:00 Introduction 01:54 What makes a reasoning model different? 04:25 Reasoning traces and model capability 08:29 Accuracy and format rewards 11:34 Aha moments and DeepSeek-R1 training 14:41 Reasoning effort and answer length 18:38 RLHF and RLVR 23:04 GRPO vs. PPO 26:40 GRPO explained with a cooking analogy 31:43 The KL term and simplified GRPO 35:04 Loading the pretrained model 36:07 Loading the MATH training data 39:26 Sampling model responses 46:30 Computing verifiable rewards 49:55 Computing advantages 51:54 Token and sequence log probabilities 55:29 Implementing sequence log probabilities 57:37 Fixing the inference-mode error 1:02:24 Computing the GRPO loss 1:04:37 Putting the GRPO step together 1:09:19 The GRPO training loop 1:12:57 Training settings, logging, and checkpoints 1:17:24 Running training and inspecting outputs 1:19:28 Loading and evaluating checkpoints 1:22:33 MATH-500 results and training stability 1:24:05 Memory requirements and next steps
63
149
1,371
56,885
Dmitry Noranovich retweeted
I am super excited about Vector Databases for Sparse Attention, and I think they'll be an essential component of every LLM and Agentic system. 🧬 Why should an Agent grep text when it could start from a precomputed map of its whole environment? 👾 Here is the math on REFRAG. TL;DR: chunks of tokens are stored at multiple granularities, a compressed vector and the full tokens. A lightweight policy between the LLM and the Vector Database decides which chunks to expand. At a compression rate of 32, each 512 token block is represented with 16 vectors. Attention now scales with vectors instead of tokens. So in principle, a model with a 128K window that expands 1% of chunks would effectively process about 3.2 million tokens. And imagine expansion that stops at intermediate resolutions instead of jumping from 1 vector all the way back to 32 token vectors, or even higher compression rates. Further, for many applications, 1% expansion may be generous... 🚀 I left this discussion with Siddharth super excited about this direction for AI, I hope you find it interesting!
Long Context LLMs and Search! Is RAG dead 💀 or here to stay? Is this a battle of two competing technologies? Or is it an opportunity where the work in Long Context LLMs can improve Search and vice versa? 🧬 I am SUPER EXCITED to share a new episode of the Weaviate Podcast with Siddhart Gollapudi, a Ph.D. student at U.C Berkeley! 🎉 Siddharth has recently lead the work behind, “Can Language Models Actually Retrieve In-Context? Drowning in Documents at Million Token Scale”. Alongside his collaborators at UT Austin and UC Berkeley, they have trained 0.6B parameter in-context retriever, named BlockSearch! 🔥 They show that BlockSearch can compete with single vector dense retrievers up to million token scale, or say 20K, 500 token documents! The podcast begins with all the quick questions: Is RAG dead? How does this compare to generative retrieval? Is this a reranker? Does this subsume all database functionality, as well as multi-hop search style tasks? What about Recursive Language Models? I think the opportunity to use this in reranking is amazingly unexplored!! And the implications of where this technology is headed is also absolutely fascinating! The opportunity for Vector Databases to help with Sparse Attention mechanisms is so exciting for Weaviate, as well as others working in the space! 🚀 This was a super fun conversation, I learned so much, and I hope you find it useful! 💚 YouTube: piped.video/gCGEg7jJXh0 Spotify: spotifycreators-web.app.link…
3
4
13
844
Dmitry Noranovich retweeted
I am super excited about the opportunity for listwise rerankers powered by LLMs that can read longer context windows. 🚀 “Drowning in Documents” from @mat_jacob1002 et al. showed that we get diminishing returns with scaling cross encoder reranking. We have made similar findings at Weaviate in the development of our Query Agent, and our paper on this will hopefully be published soon. The TLDR is that we identify the ceiling on scaling cross encoder reranking, and then from that point, hand this ceiling to a third stage listwise reranker. Similarly to how Cross Encoders are Drowning in Documents, we found that Listwise Rerankers do not scale as far as you would think! 🫠 We found plateauing results just going from 20 to 100 documents, let alone the 500 document scale Siddharth explored in BlockSearch! Once these models reach this scale, given the current capability of cross encoders as a second stage, and the ceiling the third stage inherits, we will have incredibly accurate search systems! Anyways, I hope you find this discussion interesting!
Long Context LLMs and Search! Is RAG dead 💀 or here to stay? Is this a battle of two competing technologies? Or is it an opportunity where the work in Long Context LLMs can improve Search and vice versa? 🧬 I am SUPER EXCITED to share a new episode of the Weaviate Podcast with Siddhart Gollapudi, a Ph.D. student at U.C Berkeley! 🎉 Siddharth has recently lead the work behind, “Can Language Models Actually Retrieve In-Context? Drowning in Documents at Million Token Scale”. Alongside his collaborators at UT Austin and UC Berkeley, they have trained 0.6B parameter in-context retriever, named BlockSearch! 🔥 They show that BlockSearch can compete with single vector dense retrievers up to million token scale, or say 20K, 500 token documents! The podcast begins with all the quick questions: Is RAG dead? How does this compare to generative retrieval? Is this a reranker? Does this subsume all database functionality, as well as multi-hop search style tasks? What about Recursive Language Models? I think the opportunity to use this in reranking is amazingly unexplored!! And the implications of where this technology is headed is also absolutely fascinating! The opportunity for Vector Databases to help with Sparse Attention mechanisms is so exciting for Weaviate, as well as others working in the space! 🚀 This was a super fun conversation, I learned so much, and I hope you find it useful! 💚 YouTube: piped.video/gCGEg7jJXh0 Spotify: spotifycreators-web.app.link…
2
7
19
1,664
Dmitry Noranovich retweeted
We have a Stitch MCP. We have a Stitch SDK. But, we don't have a Stitch CLI. Well, not until now. Introducing the @google/stitch CLI: 🔷 Connect to your local coding agents 🔷 Generate screens and design systems 🔷 Send a local dev server snapshot to Stitch Do it all without leaving the terminal or better yet, ask your favorite harness like @antigravity 😎 Learn more 👇
46
158
1,394
196,805
Dmitry Noranovich retweeted
Announcing 𝙾𝚙𝚎𝚗𝙳𝚘𝚝𝚜 An open source OpenAI Dots that's self-hostable and works with any agent harness. Available on Mobile and Web! Clone it. Build on top of it. Repo → github.com/CopilotKit/OpenDo…
🎉 Introducing 𝙾𝚙𝚎𝚗𝙳𝚘𝚝𝚜 Self-hostable, always-on AI coworkers that works with ANY agent harness. Includes: - Computer use: browser, terminal & files - Bring agents to Slack, Teams etc - Spaces and Pages for projects - Voice calls - Web and Mobile Repo → github.com/CopilotKit/OpenDo… Powered by @CopilotKit and AG-UI. Clone this template and customize it however you want. Enterprise-ready.
14
54
406
34,887
Dmitry Noranovich retweeted
Long Context LLMs and Search! Is RAG dead 💀 or here to stay? Is this a battle of two competing technologies? Or is it an opportunity where the work in Long Context LLMs can improve Search and vice versa? 🧬 I am SUPER EXCITED to share a new episode of the Weaviate Podcast with Siddhart Gollapudi, a Ph.D. student at U.C Berkeley! 🎉 Siddharth has recently lead the work behind, “Can Language Models Actually Retrieve In-Context? Drowning in Documents at Million Token Scale”. Alongside his collaborators at UT Austin and UC Berkeley, they have trained 0.6B parameter in-context retriever, named BlockSearch! 🔥 They show that BlockSearch can compete with single vector dense retrievers up to million token scale, or say 20K, 500 token documents! The podcast begins with all the quick questions: Is RAG dead? How does this compare to generative retrieval? Is this a reranker? Does this subsume all database functionality, as well as multi-hop search style tasks? What about Recursive Language Models? I think the opportunity to use this in reranking is amazingly unexplored!! And the implications of where this technology is headed is also absolutely fascinating! The opportunity for Vector Databases to help with Sparse Attention mechanisms is so exciting for Weaviate, as well as others working in the space! 🚀 This was a super fun conversation, I learned so much, and I hope you find it useful! 💚 YouTube: piped.video/gCGEg7jJXh0 Spotify: spotifycreators-web.app.link…
5
9
18
3,981