Current large language models rely on web-scale datasets that lack cultural grounding, resulting in systems biased toward globally dominant distributions at the expense of regional contexts. While high-quality corpora like FineWeb provide excellent global foundations, they are not structured to support the efficient extraction of culturally coherent, national-level datasets under modest computational constraints. In this work, we present a scalable pipeline for constructing sovereign corpora by extending the FineWeb framework with targeted ccTLD extraction and structural refinement. Using the Australian context (.au) as a primary case study, we demonstrate the ability to isolate large-scale national datasets that maintain broad topical coverage and internal coherence. We validate our approach across multiple jurisdictions (.uk, .ca, .nz), processing tens of terabytes of data to yield a combined 1.3 trillion tokens. Our results establish a practical foundation for building regionally anchored datasets for culturally grounded AI development without reliance on hyperscale infrastructure.
Altenburg et al. (Sat,) studied this question.