We present Cascaded Signal Re nement (CSR), a three-layer architecture for analyzing large text corpora using small language models (18B parameters) runningon consumer hardware. CSR combines a zero-cost deterministic layer (lexical detection, semantic clustering, sliding-window co-occurrence scoring) with a fast LLM screening layer (binary classi cation) and a deep LLM analysis layer (structured extraction with domain-injected context). We evaluate CSR on a compliance monitoring task: a 27,020 message multilingual corporate communications corpus processed by a 4B-parameter model with 48K context running locally via llama.cpp. Layer 1 reduces the corpus to 89 candidate blocks covering 11.4% of messages in under 1 second. Layer 2 screening rejects 73.0% of candidates, and the expensive Layer 3 deep analysis ultimately processes only 2.9% of the original corpus (776 messages), achieving a 96.9% token reduction while producing 22 con rmed ndings with 17 additional deterministic safety-net detections in 17 APIcalls totaling 98.7 seconds. The core contribution is demonstrating that intelligent pre-filtering makes small models competitive with frontier models for needle-in-haystack text analysis
del Amor Herrera Miguel (Fri,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: