PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 3, 20250 citationsOpen Access

Toward Efficient and Scalable Design of In-Memory Graph-Based Vector Search

View Full Paper
IAIlias AziziKEKarima EchihabTPThemis Palpanas

Key Points

  • Graph-based methods improve vector search efficiency, particularly using incremental insertion and neighborhood diversification.
  • The analysis covers 12 leading algorithms on data collections with up to 1 billion vectors, highlighting scalability challenges.
  • Key insights indicate that the choice of the base graph affects performance and scalability of vector search algorithms.
  • Future research should focus on adaptive seed selection and diversification strategies to enhance graph-based methods.

Abstract

Vector data is prevalent across business and scientific applications, and its popularity is growing with the proliferation of learned embeddings. Vector data collections often reach billions of vectors with thousands of dimensions, thus, increasing the complexity of their analysis. Vector search is the backbone of many critical analytical tasks, and graph-based methods have become the best choice for analytical tasks that do not require guarantees on the quality of the answers. Although several paradigms (seed selection, incremental insertion, neighborhood propagation, neighborhood diversification, and divide-and-conquer) have been employed to design in-memory graph-based vector search algorithms, a systematic comparison of the key algorithmic advances is still missing. We conduct an exhaustive experimental evaluation of twelve state-of-the-art methods on seven real data collections, with sizes up to 1 billion vectors. We share key insights about the strengths and limitations of these methods; e.g., the best approaches are typically based on incremental insertion and neighborhood diversification, and the choice of the base graph can hurt scalability. Finally, we discuss open research directions, such as the importance of devising more sophisticated data adaptive seed selection and diversification strategies.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Azizi et al. (2025) studied this question.

synapsesocial.com/papers/68e02f3cf0e39f13e7fa24efhttps://doi.org/10.48550/arxiv.2509.05750
Ask AI
Helpful
Bookmark
Share
View Full Paper