other interesting topics on postgres and kubernetes
Have not covered yet, a few other cool topics. Can go into more detail later, but, for now some highlights. Runtime embeddings to save money At one point in a recent postgres pgvector retrieval project, a really cool epiphany was, w.r.t. indexing in pgvector both a concatenated blob of items and granular items, that it is not necessary to also embed the low level items because they can be embedded at runtime and it is not that time consuming. The dataset is a corpus dishes grouped by menus. The initial design was to allow a two step embedding search, first across the concatenated menu embeddings to narrow down, and then on the menu level, to search the dish item level embeddings a second time. But disk space wise, it became clear that the dish items and their embedding hnsw index would take up an enormous amount. There are typically 10 dishes per menu so a proportionally ~10x larger space was needed. But the duhh moment was that maybe we don’t even need to store the dish item level index at all because maybe we can embed the dishes on the fly when we already have done a first pass, ending up with maybe 5 to 10 menus. Now we would only need to embed maybe 50 to 100 dishes and that would not be too slow. And indeed embedding and searching through 50 to 100 dishes did take time but it appeared to be worth the precious postgresql disk space saved. Of course depending on future query latency requirements it is always possible to reconsider this. ...