Benchmark evaluation reveals distinct detection mechanics and cluster absorption in text streams, highlighting the need for enrichment metrics over simple detection rates.
Key Points
To establish a standardized, reproducible benchmark for evaluating proactive novel topic detection algorithms against complete ground truth under identical conditions.
Injected large language model-generated articles about fictitious topics into real news and science-technology streams under controlled semantic distance, volume, and timing across 16 scenarios.
Evaluated four detection approaches—micro-cluster stream algorithms (DBSTREAM, DenStream), a windowed topic model (BERTrend), and a cosine novelty detector—paired with two sentence encoders across four injection schedules.
Stream-clustering detections consistently occurred via absorption into pre-existing clusters across all runs, seeds, and encoders rather than forming new clusters, with topical relatedness dictated by the corpus rather than the detector.
Only the windowed topic modeling paradigm produced an emergence signal, and detection methods exhibited complementary performance across difficulty regions, demonstrating that raw detection rates are misleading without enrichment metrics.