Fine-tuning a language model on a narrow domain requires paired instruction/response data that usually does not exist for that domain. The common workaround — crawl a site, chunk the text, and prompt a stronger model to synthesize question–answer pairs from each chunk — is easy to prototype and easy to get subtly wrong. We describe DATAFORGE, an open-source pipeline that treats each of those failure modes as a first-class design constraint rather than an afterthought: HTTP failure semantics arehonored (transient vs. permanent status codes, Retry-After, robots.txt Crawl-delay); every artifact (page, chunk, sample) is checkpointed so a multi-hour run is resumable rather than restartable; collection and generation are overlapped through bounded, backpressured queues to avoid idling an LLM behind a slow crawler; unattended LLM spend is bounded by an explicit, concurrency-safe budget rather than only indirect knobs like page count; and—the concern motivating this document—dataset splitting is group-aware, so that samples generated from the same source page, which share its facts, cannot appear on both sides of a train/test boundary and inflate held-out evaluation scores. We describe the system’s architecture and the reasoning behind each design decision, and evaluate it end to end on four websites: 831 samples for $0.12, with page-disjoint splits on every site. The same benchmark exposed six defects that unit tests had missed, including a declared Crawl-delay that was not enforced while scraping; all are fixed and reported. Of fourteen sites considered, ten were excluded after a review of their terms and access policies, although robots.txt alone would have permitted at least six of them. The LLM judge approved 97.6% of samples, but the same model generated and judged them, so we treat that figure as unvalidated until a human audit is complete. We also describe agent access through the Model Context Protocol, situate the resulting synthetic, fine-tuning-oriented dataset relative to retrieval augmented generation as a complementary route to reducing hallucination, discuss the responsible-use considerations inherent to an unattended web-scraping and synthetic-data tool, and situate the system relative to prior work in web-scale data collection and synthetic instruction-tuning data generation.
No takes yet. Share an insight, caveat, or question.
Ian K. Too (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: