Historical analysis traces multiset data structures to seventeenth-century collections of alterity, suggesting new frameworks for evaluating how large language models encode race.
Key Points
To provide a genealogical account of 'bags' or 'multisets'—unsorted computational data collections used by modern large language models—by tracing their origins to early modern representational practices.
Conducted a historical and genealogical analysis of representational practices from 1600 to 1665.
Examined seventeenth-century still-life paintings, botanical specimens, and ethnographic records to contrast disordered multisets with rigid taxonomies.
Traced modern computational 'bag' representations back to early modern techniques that aggregated oddities, automata, and colonial subjects without prior categorization.
Demonstrated that disordered multisets developed historically as a functional counterpart to Linnaean taxonomies, providing an alternative model for how large language models process racial difference.