Long Context LLMs and Search! Is RAG dead 💀 or here to stay? Is this a battle of two competing technologies? Or is it an opportunity where the work in Long Context LLMs can improve Search and vice versa? 🧬
I am SUPER EXCITED to share a new episode of the Weaviate Podcast with Siddhart Gollapudi, a Ph.D. student at U.C Berkeley! 🎉
Siddharth has recently lead the work behind, “Can Language Models Actually Retrieve In-Context? Drowning in Documents at Million Token Scale”. Alongside his collaborators at UT Austin and UC Berkeley, they have trained 0.6B parameter in-context retriever, named BlockSearch! 🔥
They show that BlockSearch can compete with single vector dense retrievers up to million token scale, or say 20K, 500 token documents!
The podcast begins with all the quick questions: Is RAG dead? How does this compare to generative retrieval? Is this a reranker? Does this subsume all database functionality, as well as multi-hop search style tasks? What about Recursive Language Models?
I think the opportunity to use this in reranking is amazingly unexplored!!
And the implications of where this technology is headed is also absolutely fascinating! The opportunity for Vector Databases to help with Sparse Attention mechanisms is so exciting for Weaviate, as well as others working in the space! 🚀
This was a super fun conversation, I learned so much, and I hope you find it useful! 💚
YouTube:
piped.video/gCGEg7jJXh0
Spotify:
spotifycreators-web.app.link…