Search @perplexity_ai | prev @JinaAI_ @sap @HPI_DE | Multimodal LLMs, embeddings

Berlin, Germany
Sarah Eslami retweeted
Hey, if you've tried any of @perplexity_ai APIs (Search, Agent, Embeddings, Decision, etc.) and were disappointed by anything, please DM me! Would be glad to learn what we should do better :)
8
4
50
35,819
Sarah Eslami retweeted
New day, new model :)
Introducing the Perplexity Decisions API. It's powered by pplx-decider-v1-27b, our multimodal decision model trained to output a probability distribution over a fixed set of answers instead of text. It costs $0.04/million input tokens and scores 85.71% across benchmarks.
2
2
26
2,343
Sarah Eslami retweeted
in the past months we worked closely with tpuf team, iterating our models beyond web search. During the journey, we built a new way of training contextual embedding models, which connects our previous work, the context pruner. We're sharing how we train the model, and how it performs on @turbopuffer 's private bench. great work from @ESL_Sarah , Markus, @antoine_chaffin , Louis, Max and special thanks to @n0riskn0r3ward , who helps evaluate, iterate again and again 🐐. We'll work closely together to continue improve it, and eventually offer it to all.
We built a new way to train contextual embedding models, which encode each chunk of a document with the whole document in view. pplx-embed-v2-context-9b-preview sets a new state of the art on ConTEB and @turbopuffer's new, privately held context-bench. perplexity.ai/hub/blog/conte…
3
8
64
5,119
Sarah Eslami retweeted
Hey search folks, look what we just shipped! A strong contextual embedding model, already available on Hugging Face: huggingface.co/perplexity-ai…
We built a new way to train contextual embedding models, which encode each chunk of a document with the whole document in view. pplx-embed-v2-context-9b-preview sets a new state of the art on ConTEB and @turbopuffer's new, privately held context-bench. perplexity.ai/hub/blog/conte…
1
8
64
4,494
So glad to have @antoine_chaffin on the team! It’s not often you see someone make so much impact in just a few weeks! 🙌 🐐
Replying to @antoine_chaffin
As usual, this has been a team effort, so a big congratulations to @ESL_Sarah for leading the effort and all the other members (@MaxSchall1, @bo_wangbo, Markus & Louis), I am very happy we managed to push it! Also thanks to @turbopuffer and @n0riskn0r3ward for the effort of building the benchmark, it has been very cool to discuss evaluations and the capacity the contextual models should have!
1
2
11
689
Was quite a pleasure working with @n0riskn0r3ward towards achieving this milestone! 🐐 Stay tuned for the full tech report and more details of our approach!
I’ve been convinced contextual embedding models were dope since voyage-context-3 was released. Started using them and benchmarking them, realized that, esp for long documents, I needed them to rank not just the answer chunk, but the disambiguating context chunks highly. Done well, this could allow you or your search agent to read only the relevant fraction of the document rather than every page of a long document. So I started trying to measure that too. Offered to share my benchmark results with the perplexity team sometime after they released their contextual model. It’s worth mentioning - not all modeling teams want feedback from third parties like me. The perplexity team was eager for it and within a week had a new model for me to benchmark. So I’d run their new checkpoint on my benchmark, stare at the results, get a little frustrated that the benchmark was imperfect and wasn’t measuring everything as well as I wanted it to, iterate on it a bit and give them new results. We ran that back many times over the last few weeks. I’ve re-worked some part of the benchmark at least 9 times now looking over the results. Instead of getting frustrated with my moving target of a benchmark they just kept improving their model and refining their approach until they were clearly on top. Notably, they didn’t achieve this via privileged access to the benchmark, they simply iterated until they had a good model. Was really a pleasure working with @ESL_Sarah, @bo_wangbo, Markus, @antoine_chaffin, Louis, and Max. For those who aren’t going to read the blog I’ll share more about how “evidence recall” works soon.
1
115
Sarah Eslami retweeted
Contextual models are a solution to the chunking problem But how should we train them? Today, we release a new preview version of our contextual embedding model and discuss a new training approach that pushed the SOTA results including on a new benchmark from @turbopuffer
We built a new way to train contextual embedding models, which encode each chunk of a document with the whole document in view. pplx-embed-v2-context-9b-preview sets a new state of the art on ConTEB and @turbopuffer's new, privately held context-bench. perplexity.ai/hub/blog/conte…
7
11
104
13,144
Sarah Eslami retweeted
Introducing Perplexity Research Fellowship: a program for early-career researchers, engineers, and analysts from any technical or quantitative discipline to work on AI research. perplexity.ai/hub/research/f…
50
184
1,909
266,119
Excited to share our latest work: Q2D-Web — a large-scale retrieval benchmark for agentic RAG, with 190M documents, ~70K agent-reformulated queries, and deep relevance judgments. 📄 Preprint: arxiv.org/abs/2609.08887 🏆 Leaderboard: huggingface.co/spaces/perple…
We're introducing Q2D-Web (Query2Doc-Web), a benchmark and public leaderboard for evaluating retrieval in agentic RAG systems. Q2D-Web tests how embedding models perform on large-scale web search using agent-reformulated search queries. Read more: perplexity.ai/hub/blog/q2d-w…
1
2
8
594
Shout out to @MaxSchall1 for driving this work.
55
Sarah Eslami retweeted
We released pplx-embed-v1 early this year, with Q2D benchmark to evaluate how embedding model performs for web search, not looking at nDCG@10, but Recall@1000. Q2D-Web further scale it up 7 times. With 70k agentic reformulated queries and 190 million web corpus. We hope the benchmark becomes the modern “MS MARCO”
We're introducing Q2D-Web (Query2Doc-Web), a benchmark and public leaderboard for evaluating retrieval in agentic RAG systems. Q2D-Web tests how embedding models perform on large-scale web search using agent-reformulated search queries. Read more: perplexity.ai/hub/blog/q2d-w…
5
15
77
8,241
Sarah Eslami retweeted
we are releasing a new benchmark and a public leaderboard for embedding models that measure the performance of large-scale agentic web search. more info can be found here. paper: arxiv.org/pdf/2609.08887 blog: perplexity.ai/hub/blog/q2d-w… leaderboard: huggingface.co/spaces/perple…
We're introducing Q2D-Web (Query2Doc-Web), a benchmark and public leaderboard for evaluating retrieval in agentic RAG systems. Q2D-Web tests how embedding models perform on large-scale web search using agent-reformulated search queries. Read more: perplexity.ai/hub/blog/q2d-w…
5
13
93
9,960
Sarah Eslami retweeted
You can now use the Perplexity Search API in Hermes Agent
Perplexity Search API is now available in Hermes Agent. Search API gives Hermes access to an index of more than 400 billion URLs. It returns real-time results and ranks snippets by relevance. pplx.ai/hermes
49
56
920
190,419
Sarah Eslami retweeted
Perplexity Search inside Hermes. Perplexity's index currently includes 450B+ high-quality URLs. Rapidly advancing to a trillion by EOY with high-quality snippets.
Perplexity Search API is now available in Hermes Agent. Search API gives Hermes access to an index of more than 400 billion URLs. It returns real-time results and ranks snippets by relevance. pplx.ai/hermes
36
47
586
95,381
Sarah Eslami retweeted
it's been great working with @NousResearch to integrate our search api into Hermes Agent. we are making a lot of cool improvements to our search stack to make it the most accurate and cost-efficient.
Perplexity Search API is now available in Hermes Agent. Search API gives Hermes access to an index of more than 400 billion URLs. It returns real-time results and ranks snippets by relevance. pplx.ai/hermes
7
4
30
4,150
Sarah Eslami retweeted
wtf is this way to handle mathematicians work and scientific communication TLDR: Leven and Tristan worked over several months on one of the Millenium Prize Problems with various AIs to reach final interesting results. OpenAI apparently heard about it in the last days and prompted their latest models to work on the direction Leven and Tristan found fruitful. They then tried to push for controlling communication of the result and dropping Leven from authorship with some very bad taste social pressure. Hope this is not a glimpse of the future we’ll get in science research with these dominating players playing marketing games hurtful for the real scientific community.
84
377
2,902
697,970
Sarah Eslami retweeted
We’re open sourcing WANDR. WANDR is an internal benchmark we built and used for building deep and wide research capabilities inside Perplexity Computer. research.perplexity.ai/artic…
64
72
667
251,999
Sarah Eslami retweeted
if you're testing a new retrieval model or long-context LLM, it's a waste of your time (and ours...) to report 0.2% gains on the many saturated and expired benchmarks if you're in that position and looking for way to rescue your great new idea, put it to the test on OBLIQ-Bench
We set out to build a better retriever, so we looked for the hardest IR benchmarks. For each, we asked how much headroom remained by running oracle reranking with a frontier LLM. Most had little room left! So we built OBLIQ-Bench to study much harder search queries than before.
10
11
171
25,344