The document retrieval layer for AI

London
Just as AlphaGo helped ignite the modern AI wave, self-improving agents may shape its future; hierarchical search is a fundamental building block for the next stage of AI development.
my personal faves are: AutoResearch for Feature Extraction Use LLMs to generate Jev questions's who's probabilities are fed into a classic ML algorithm like logistic regression to predict a label. Hierarchical Classification Traverse a classification taxonomy using Jev and beam sesarch. Jev is great for graph traversal.
1
3
467
Jev 🤝 PageIndex Search long documents like a human 📖
Jev can search 🔍 through a long document just like you would! 🕵️
8
479
An exciting new direction for the future of AI retrieval: organize knowledge into a hierarchy, then use a decision model like Jev to decide level by level to find the relevant part.
Search like a human does - by semantically traversing a hierarchy with jev
7
563
Inspired by @EGafni’s Twitter thread on combining Jev with PageIndex. We show how to build long-document search with @typesafeai Jev + PageIndex. No vector database. No embeddings. Open source: github.com/VectifyAI/jev-doc… 🧵👇
11
24
174
96,799
PageIndex fixes both. It turns the document into a tree of sections, each with a title and a summary. Jev then makes one small choice per level: a section, then a page inside it. Every choice has only a handful of options, however long the document.
1
10
1,735
PageIndex + Jev is cooking right now, keep updated!
Been wanting people to use Jev for this for a while now :) Semantic hierarchical search! just like a human would. check out our hierarchical classification cookbook
1
6
628
Wow, this can be a perfect model for RAG, with string/number/integer types.
Not another Jev. TypeLLM adds type-safe generation to LLMs — more types (string, number, boolean, enum), vision, dependent fields and thinking. Playground is open to all with $5 credit. API is rolling out to early-access users in the coming days. typellm.ai/dashboard/playgro…
1
383
Not too late to scales the retrieval with the model!
if this takes off, then i was 4 years too early 😭 x.lingyaoai.com/jerryjliu0/status/1590… only OGs remember gpt tree index
1
350
Block references are now live for all PageIndex Cloud users. Every citation points to the exact block, not just the page. Docs → docs.pageindex.ai/sdk/chat#c…
2
253
As we scale up PageIndex Cloud, we’re introducing a new, simpler pricing model: • $0.01 per page indexed • $0.001 per active page / month • Unlimited retrieval on active pages Simple usage-based billing. No monthly commitment. Get started with $10 in free credits.
1
1
5
306
An illustration of three ways to ask a document: • Vector DB: Split into chunks and search by vector similarity • PageIndex: Navigate a document tree and retrieve by node relevance • Full context: Put the whole document into the context— through an LLM provider’s file API
4
8
595
LLMs can solve the Navier–Stokes problem, but they’re still very poor at understanding PDFs. GPT-6 Astra scores just 32.2% on the GDP-PDF benchmark: surgehq.ai/benchmarks/gdp-pd…
7
3
12
457
PageIndex is built on this principle: a persistent semantic representation of knowledge, with retrieval accuracy that improves as the LLM scales.
General coding agents now significantly outperform purpose-built data agents. What is left for systems researchers to solve? We break it down in our latest paper.
8
5
10
814
A simple visualization of how PageIndex works. Index: generate a tree index for each document. Retrieve: agentically tree search with LLM Learn more: pageindex.ai/developer
1
7
227