Nikhil Benesch retweeted
If you want allocation, you gotta join the company - @Sirupsen @nikhilbenesch 🫡
1
2
35
6,790
latency hiding is such a simple but powerful technique the more of your search pipeline that lives in turbopuffer, the more places we have to hide latency
moving embedding into turbopuffer lets us parallelize the embedding call with other work in the query plan whole-system puffin' shaves 100+ ms off your semantic searches turbopuffer.com/blog/native-…
4
1,387
introducing the buff puff
9
1
101
5,397
apparently @dylan522p was very confused about this behavior the first time we met at a dinner party (SemiAnalysis should puff)
everyone at WarpStream sells, but @turbopuffer takes it to the next level the employees will look you in the eyes at a dinner party after just meeting you five seconds ago and ask “does your company puf? should you be puffin’?” with a completely straight face
10
2,670
we are so web scale
fun things you get to see when working at @turbopuffer: today i'm fixing my code because it exceeds `i64::MAX` when summing byte size of searched namespaces in the last minute. that's 9,760,691,815,929,332,437 bytes. per minute. in a single production cluster. 9.8 exabytes :D
1
18
1,826
Nikhil Benesch retweeted
All the S3-native products when a new S3-native product is launched
3
7
124
16,724
never go shimless
the first time I met CTO @nikhilbenesch was for lunch in Dec '24 in NYC, and the table at the Greek restaurant was wobbly. My grandmother Grete always carried a shim in her purse, and to pay homage to her, I always carry one too so I immediately I pulled it out at that lunch and fixed the table. that's when @nikhilbenesch knew this was it. he would join turbopuffer. I gifted him the shim. yesterday I went shimless to a restaurant in NYC. a reckless endeavour for anyone of Grete's linage, especially in NYC. fortunately, @nikhilbenesch brought his shim. he had waited 16 months for this moment. ❤️
23
3,003
what's not stated here: @natevanben built v0 of sharding in <2 weeks simplicity scales 📈
1
1
24
2,177
our goal is "no 500s ever", but if that ever slips we do the next best thing!
Got a note from the @turbopuffer team that they were seeing 500 errors on a small number of our requests and they would investigate. Found the issue, patched it, and shipped it to prod. Continually impressed with their team. Can't think of another vendor has been so proactive
3
3
126
68,587
we are also in the market for anyone interested in pushing the frontier of late interaction. the current LI models are too small to do the technique justice!
tpuf now supports late interaction [beta] use models like ColBERT to represent text as a set of vectors (1 per token) tpuf uses a single-vector ANN index for a fast first pass, then reranks hits using exact late interaction scoring to boost recall docs: turbopuffer.com/docs/query#l…
7
1,725
we have a slack channel called #goodgraphs where we celebrate good graphs. today's is especially good
autoscaling is deceptively hard our indexer fleet scales nodes on job queue time. if a queue suddenly went quiet, we'd scale down too hard & new jobs could queue up waiting for nodes to claim them we tweaked the HPA signal to prevent underprovisioning → ~2x shorter queue time
24
2,941
Nikhil Benesch retweeted
Nemotron 3 Embed is now available on tpuf native embeddings via our partner @baseten contact us for beta access, full model list here: turbopuffer.com/docs/embeddi…
Today we released Nemotron 3 Embed 8B and it reached #1 overall on RTEB 🏆 RTEB benchmarks retrieval accuracy across real-world tasks. Better retrieval gives agents more relevant context, helping improve response accuracy.
1
3
76
7,454
one of my fav tpuf integrations just had its official launch 🐘 → 🐡
we are releasing our first open source project! puffgres is a sync service between Postgres and @turbopuffer. instead of two DB calls everywhere, you write ‘dumb’ JS transforms (i.e. tokenization, embedding) and the core handles logical replication, retries, batching, etc
1
1
27
2,811
we are so ready for your hardest search workloads 🐡
turbopuffer crossed $100M run-rate in March. 19mo after $1M. Profitable & <$1M raised. Cursor・Anthropic・Notion・Cognition・Harvey・Bridgewater・Ramp・Linear・Legora・Superhuman・Atlassian・Granola We’d be nowhere without them. We work like hell to exceed their expectations.
4
64
6,357
we think this is what the future of search looks like
SID-1 is an agentic search model by @SID_AI → 1.9x recall over RAG + rerank → 24x faster, 99% cheaper than GPT-5.1 trained using large-scale RL on turbopuffer at 1k+ QPS bursts over 10M+ document corpora across thousands of steps tpuf.link/sid-1
3
6
84
29,376
Nikhil Benesch retweeted
new: namespace pinning pin namespaces to reserved compute for faster, cheaper, more predictable p99 on high QPS workloads ~50x cheaper than shared compute for a 128GB namespace at 500 QPS docs: turbopuffer.com/docs/pinning
5
6
71
18,203
Nikhil Benesch retweeted
we couldn’t afford the sphere so we got the next best thing costs less, hits harder
34
42
1,076
198,456
Nikhil Benesch retweeted
I don’t think there’s a product on our stack that improves and gets cheaper than tpuf. Absolutely wild
we’ve reduced query prices by up to 94%, thanks to infrastructure improvements we've made under the hood puff harder, for much less
1
1
48
8,751
AI coding data point: models are getting closer to being able to implement a feature of this complexity. For now @morgallant firmly has the edge.
new: ContainsAnyToken filter return documents that match any token in the query string faster than BM25 when you just need a binary match (e.g. "showing 36 of 131,072" in your search UI) as it skips score computation entirely docs: tpuf.link/containsany
1
8
1,190
For anything correctness sensitive we're in the uncanny valley. The models are good enough to produce dangerously plausible code, but often miss a few crucial edge cases.
1
7
473