It doesn't have to be a vector DB - and in fact I'm seeing increasing skepticism...

rcarmo · 2024-07-24T16:23:48 1721838228

I've been using SQLite FTS (which is essentially BM25) and it works so well I haven't really bothered with vector databases, or Postgres, or anything else yet. Maybe when my corpus exceeds 2GB...

niam · 2024-07-24T17:31:01 1721842261

What are the arguments for embedded vector DBs being suboptimal in RAG, out of curiosity?

simonw · 2024-07-24T17:53:35 1721843615

The biggest one is that it's hard to get "zero matches" from an embeddings database. You get back all results ordered by distance from the user's query, but it will really scrape the bottom of the barrel if there aren't any great matches - which can lead to bugs like this one: https://simonwillison.net/2024/Jun/6/accidental-prompt-injec...

The other problem is that embeddings search can miss things that a direct keyword match would have caught. If you have key terms that are specific to your corpus - product names for example - there's a risk that a vector match might not score those as highly as BM25 would have so you may miss the most relevant documents.

Finally, embeddings are much more black box and hard to debug and reason about. We have decades of experience tweaking and debugging and improving BM25-style FTS search - the whole field of "Information Retrieval". Throwing that all away in favour of weird new embedding vectors is suboptimal.

kgeist · 2024-07-25T05:41:27 1721886087

>but because embeddings search orders by similarity score it will ALWAYS return results, really scraping the bottom of the barrel if it has to

Why not have a similarity threshold? Say, if the distance is below 0.7, do not accept the search result.

simonw · 2024-07-25T06:24:22 1721888662

It turns out picking that threshold is extremely difficult - I've tried! The value seems to differ for different searches, so picking eg 0.7 as a fixed value isn't actually as useful as you would expect.

zmccormick7 · 2024-07-25T15:08:31 1721920111

Agreed that thresholds don't work when applied to the cosine similarity of embeddings. But I have found that the similarity score returned by high-quality rerankers, especially Cohere, are consistent and meaningful enough that using a threshold works well there.

kgeist · 2024-07-29T10:04:04 1722247444

I use similarity threshold (to remove absolutely irrelevant results) and then use a reranker to get Top N.

jairuhme · 2024-07-25T15:11:54 1721920314

I'll add to what the other commenter noted, but sometimes the difference between results get very granular (i.e. .65789 vs .65788) so deciding on where that threshold should be is little trickier.

ianbutler · 2024-07-24T17:16:41 1721841401

In 2019 I was using vector search to narrow the search space within 100s of millions of documents and then do full text search on the top 10k or so docs.

That seems like a better stacking of the technologies even now

ramoz · 2024-07-25T10:36:02 1721903762

Interesting. Why did you need to “narrow” the search space using vector space? Did you build custom embeddings and feel confident about retrieval segments?

I did similar in 2019 but typically in reverse, FTS, and a dual tower model to rerank. Vector search was an additional capability but never augmented the FTS.

ianbutler · 2024-07-27T04:25:37 1722054337

It was in consideration of how slow our FTS at the time was over large amount of documents and the window we wanted to keep response times in and you're correct, we had custom embeddings and we had a reasonably high confidence.

So vector search would reduce the space to like 10k documents and then we'd take the document ids and FTS acted as the final authority on the ranking.