I’ve been convinced contextual embedding models were dope since voyage-context-3 was released.
Started using them and benchmarking them, realized that, esp for long documents, I needed them to rank not just the answer chunk, but the disambiguating context chunks highly. Done well, this could allow you or your search agent to read only the relevant fraction of the document rather than every page of a long document.
So I started trying to measure that too.
Offered to share my benchmark results with the perplexity team sometime after they released their contextual model.
It’s worth mentioning - not all modeling teams want feedback from third parties like me. The perplexity team was eager for it and within a week had a new model for me to benchmark. So I’d run their new checkpoint on my benchmark, stare at the results, get a little frustrated that the benchmark was imperfect and wasn’t measuring everything as well as I wanted it to, iterate on it a bit and give them new results.
We ran that back many times over the last few weeks. I’ve re-worked some part of the benchmark at least 9 times now looking over the results.
Instead of getting frustrated with my moving target of a benchmark they just kept improving their model and refining their approach until they were clearly on top.
Notably, they didn’t achieve this via privileged access to the benchmark, they simply iterated until they had a good model.
Was really a pleasure working with
@ESL_Sarah,
@bo_wangbo, Markus,
@antoine_chaffin, Louis, and Max.
For those who aren’t going to read the blog I’ll share more about how “evidence recall” works soon.