PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
September 5, 2025Proceedings of the VLDB Endowment1 citations

BigVectorBench: Heterogeneous Data Embedding and Compound Queries are Essential in Evaluating Vector Databases

View Full Paper
GKGuoxin KangChinese Academy of SciencesZGZhongxin GeChinese Academy of SciencesJHJingpei HuChinese Academy of Sciences

Key Points

  • BigVectorBench improves how vector databases manage heterogeneous data embedding and compound queries.
  • Evaluations confirm that existing benchmarks neglect performance metrics for compound queries and embedding tasks.
  • The proposed benchmark suite highlights performance bottlenecks in leading vector database technologies.
  • BigVectorBench integrates essential evaluations for dual tasks in real-world vector database applications.

Abstract

Vector databases are designed to effectively store, organize, and retrieve high-dimensional vectors, enabling faster and more accurate querying and analysis. This study highlights that the performance of cutting-edge vector databases hinges on their proficiency in managing heterogeneous data embedding and handling compound queries. The former task revolves around converting varied data types into a cohesive vector format, while the latter involves processing multimodal or single-modal queries with precise constraints. The paper advocates for evaluating these dual tasks within an integrated benchmark framework. However, state-of-the-art vector database benchmarks overlook heterogeneous data embedding and compound queries, creating a gap in evaluating vector database performance. To address this gap, we introduce BigVectorBench, a benchmark suite designed to evaluate vector database performance. BigVectorBench contributes by defining and evaluating the embedding performance of heterogeneous data. Additionally, it abstracts compound queries, which are increasingly used in real-world applications, replacing unimodal vector searches. Our rigorous evaluations validate the two design decisions of BigVectorBench and identify performance bottlenecks of mainstream vector databases. Its source code and user manual are available from https://github.com/BenchCouncil/BigVectorBench.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Kang et al. (2025) studied this question.

synapsesocial.com/papers/68bb3a2b2b87ece8dc954a9ehttps://doi.org/10.14778/3718057.3718078
Ask AI
Helpful
Bookmark
Share
View Full Paper