Key points are not available for this paper at this time.
Abstract Benchmarking large language models (LLMs) is a key practice for evaluating their capabilities and risks. This paper considers the development of “BIG Bench,” a crowdsourced benchmark designed to test LLMs “Beyond the Imitation Game.” Drawing on linguistic anthropological and ethnographic analysis of the project's GitHub repository, we examine how contributors developed tasks based on their lay understandings of language, cognition, and intelligence. By tracing how contributors make implicit judgments about what constitutes a meaningful test of intelligence, we show how widespread language ideologies shape the evaluation of LLMs and the imaginaries that guide their development.
Building similarity graph...
Analyzing shared references across papers
Loading...
Anna Weichselbraun (Mon,) studied this question.
www.synapsesocial.com/papers/69403ba12d562116f290cb91 — DOI: https://doi.org/10.1111/jola.70035
Anna Weichselbraun
Journal of Linguistic Anthropology
University of Vienna
Building similarity graph...
Analyzing shared references across papers
Loading...