The individual job search produces a substantial volume of personal digital trace data, yet this data is typically fragmented across email inboxes, job-board platforms, and informal records, making it difficult for applicants to understand their own search as a coherent empirical process. This paper presents an exploratory quantitative study that treats the job search as a data science problem, documenting both the construction of a human-in-the-loop AI pipeline for reconstructing job-search records and the empirical findings that pipeline produces. Using a personal job-search archive as the empirical case, the study integrates data from two Gmail exports, LinkedIn (554 records), Indeed, and Glassdoor through an iteratively developed processing pipeline. The pipeline development is documented in full, including the specific failures that motivated each design change, the AI models used at each stage, and the human validation decisions that resolved ambiguous cases. The resulting dataset contains 2,841 application-event records spanning 1,616 unique companies. Key findings include: 64% of application records show no recorded employer response; email accounts for 97% of interview-stage evidence while platform exports capture only submission events; 24 of the top 30 most frequent companies are recruitment agencies rather than direct employers; and the raw status field, containing over 40 variants, documents the label drift that emerges when multiple AI models classify the same data without a controlled output schema. The paper contributes both a replicable methodological framework for personal job-search analytics and an honest account of how AI-assisted research pipelines are built in practice, including what breaks, why, and how it gets fixed.
Shivkumar Ram Natarajan (Sat,) studied this question.