Comparative evaluation demonstrates that direct LLM clustering achieves comparable quality to embedding methods while enabling promptable criteria, suggesting viable tradeoffs in cost and...
The standard pipeline for text clustering (embed, then run a geometric clustering algorithm) produces clusters that are accurate but opaque: a practitioner debugging the output sees anonymous centroids, not reasons. We study methods that use a large language model (LLM) as the entire clustering operator, and provide a controlled comparison of the natural inference topologies under one oracle abstraction: three base strategies, FullContext (the whole corpus in one call), ChunkMerge (parallel chunk clusterings unified by a prototype merge call), and Incremental (online, one assign-or-create call per item), plus a derived hybrid, SampleSeed (one call over a random seed sample, then parallel batch assignment against the frozen result). We derive each strategy's cost profile and measure quality, dollar cost, latency, order sensitivity, output reliability, and label interpretability against embedding baselines on four corpora. On noisy financial tweets the best cost-tier LLM configurations obtain mean ARI numerically comparable to the strongest embedding baseline (0.25-0.29 vs. 0.25, not significant on this fixed sample) at USD 0.02-0.65 per run; on a coarse control the outcome depends on who is told the true k, implicating granularity alignment as a central success factor. The clustering criterion is promptable: asked to cluster by sentiment instead of topic, the LLM reaches ARI 0.30 where generic embeddings stay at or below 0.08. FullContext degrades catastrophically with scale (over 2 output-repair events per item at N=1000), strategy rankings change across model tiers, and SampleSeed, sized by a coverage bound independent of corpus size, is the most scale-stable configuration at the cost tier, holding ARI 0.27 at N=1000 with a flat 10-25 second critical path. We quantify what each choice costs.
No takes yet. Share an insight, caveat, or question.
Sirui Ray Li (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: