Developer Relations Engineer @ Qdrant

New York, USA
Dylan Couzon retweeted
Build an assistant that can store and search what it sees, hears, and is told, all without leaving the device. In Building AI Assistants with On-Device Memory, you'll build a local memory system that recalls text, voice, and images, and teach it to recognize something new from a few photos. Built in partnership with @qdrant_engine and taught by @DylanCouzon, Developer Experience Engineer at Qdrant. This course also includes our new AI coding lab, so you can practice building with an AI coding agent from the DeepLearning.AI mobile app. Enroll for free: hubs.la/Q04y3S0F0
18
32
224
13,895
One of the things I love most is to go deep into technical topics and explore them But if there's something I love even more, it's to get my hands dirty and write some good ol' Rust to implement what I learnt🦀 That's what I did with cross-encoder inference at @qdrant_engine: I learnt about it, then made cross-encode-rs, a rust library built on top ONNX and that you can use to run any cross-encoder model as a reranker⚙️ I built a server, too, that you can deploy with Docker and/or on your Kubernetes clusters just in a few commands🐋 I also benchmarked it against sentence-transformers and fastembed (python), but the findings weren't fully what you would expect👀 Read more in the blog post I wrote: qdrant.tech/blog/oxidizing-c… And star the GitHub repo: github.com/AstraBert/cross-e…
2
9
341
Re-embedding a corpus to test a faster query model is a pain. We trained two smaller models to search Stella vectors, so Zero, Nano, and full Stella can use the same index. The tradeoffs are in the diagram. Research preview: qdrant.tech/blog/constella-r…
1
5
147
Dylan Couzon retweeted
We moved 400K+ document vectors from Milvus to Qdrant and wrote up the whole thing: kaivid.com/blogs/migrate-mil… The post covers how we ran the migration, verified that the data arrived intact, and compared the two on latency, throughput, and memory. Latency landed in a similar range, @qdrant_engine used noticeably less memory in this workload, and speed depended on concurrency, so it's worth testing with your own query mix.
1
3
8
377
Dylan Couzon retweeted
🗽In NYC tonight? @DylanCouzon & @qdrant_engine are hosting a debate night, and @deepset_ai is teaming up Hear some hot AI takes, get a chance to win prizes, and learn about @Haystack_AI and @qdrant_engine along the way! luma.com/nyc-meetup-qdrant-0…
4
6
348
Dylan Couzon retweeted
Qdrant 1.19 brings some interesting new features, and I really like the idea of memory tiers for keeping everything aligned. I haven't had a chance to look into the TurboQuant datatype yet, but I'll get to it.
Let’s Talk about Qdrant 1.19! It’s time to discuss what’s new. Join us for our next Qdrant Discord Office Hours as we walk through the latest features and updates, followed by an open Q&A and community discussion. Date: August 20 Time: 5:00 PM CEST Link: discord.gg/PGemY6ugf?event=1… Bring your questions, projects, and ideas. We’d love to hear what you’re working on! More details about 1.19 update: qdrant.tech/blog/qdrant-1.19… See you there!
1
1
376
Dylan Couzon retweeted
How @Bayer built an enterprise search engine for 116,000 employees on Qdrant Hybrid Cloud:1 35M points on 4 nodes. Hybrid search for deep agents. Semantic caching to control agent cost. A measured 20% efficiency gain. Three years in production, from prototype to platform: qdrant.tech/blog/case-study-…
2
1
16
559
How I use Claude Code and Codex on a daily basis
84
Dylan Couzon retweeted
Prompt: Let’s design an agentic chat bot for devsupport and customer support teams, create Tickets on JIRA Customer Support and Dev Support boards, auto assign tickets. Read tickets from the email. Use RAG to query previously fixed tickets and help with the answers, if a new tickets comes. Create context for that ticket, an embedding pipeline for rag in vector DB. Let’s consider Qdrant, also add the Solutions in the embedding to query later. System Design Diagram
1
2
70
Dylan Couzon retweeted
AI has gotten better at reasoning, but it still forgets almost everything. Join @DylanCouzon from Qdrant and James Le from @twelvelabs on July 28 for an evening dedicated to AI memory, video intelligence, and retrieval systems. Hear from engineers building production AI systems as they cover: → Persistent memory for video AI → Collective memory for Edge AI with Qdrant Edge → Practical retrieval architectures and live demos Want to showcase your own project? We’re also hosting community demos, and only 2 demo slots remain! If you’re building with AI, retrieval, video intelligence, or agentic workflows, we’d love to have you present your work. 📅 July 28 📍 Bellevue, WA 🎟️ Register here: luma.com/kyksgkak See you there!
2
2
12
776
My most honest PR
1
1
59
Dylan Couzon retweeted
The qdrant-advisor skill is now easily installed to any of your coding assistants: npx skills add qdrant/skills/meta/qdrant-advisor Installing this one skill keeps your agents primed with the latest skills context for your projects and deployments. Even as we continually improve and add to said skills. Read more here: github.com/qdrant/skills
3
7
345
Dylan Couzon retweeted
Elastic's benchmark claims Qdrant is 7x slower on disk search. They never enabled the two features built for that workload. We turned them on and found 2x throughput, half the latency, 1/3 the compute. Read more: qdrant.tech/blog/benchmark-e…
7x higher vector search throughput at comparable recall. Elasticsearch 9.4.1 DiskBBQ vs Qdrant 1.18.1, tested on network-attached persistent storage. The storage topology most K8s and managed-cloud deployments actually run on. Not local NVMe. The gap is disk access. DiskBBQ searches a compact quantized index and limits full-precision reads. Qdrant rescores against original vectors on disk. On network-attached storage, those random reads get expensive. Elasticsearch latency: 120 to 150ms across recall levels. Qdrant: 315ms to 900ms as recall increases. Benchmark tool, dataset, and configs are all published below.
4
4
19
1,743
"Opus 4.7 is our smartest model yet"
2
85
Been messing with RAG pipelines and this actually helped: Forget finding “the perfect chunk size.” Try indexing multiple sizes (100/200/500), retrieving from all, and fusing with RRF. Gains ranged from 1–37% depending on the benchmark. Shoutout @AI21Labs Code 👇
1
3
80
Dylan Couzon retweeted
Comparing experiments in Phoenix just got a lot easier. The new List View helps you quickly scan results with per-example metrics, while the Metrics View gives you a high-level look at how changes impact cost, latency, tokens, and more. Upgrade to v11.32.1 to start exploring.
3
8
374
Dylan Couzon retweeted
See the new Phoenix evals library in action and bring your questions for our virtual workshop this Thursday! RSVP: luma.com/45eopucf
2
5
408