Randomized trial evaluates automated species identification in Colombian vertebrates, indicating significant challenges for rare taxa.
Multimodal artificial intelligence (AI) is increasingly adopted by conservation organizations to overcome taxonomic bottlenecks. However, its reliability may systematically fail for sparsely documented and recently described taxa. This study evaluates automated species identification using 10,000 research-grade iNaturalist images of Colombian vertebrates across amphibians, birds, fishes, mammals, and reptiles. Testing OpenAI GPT-5.2 with localized prompts but no taxonomic metadata, we found that species-level accuracy scales non-linearly with global photographic representation and decays severely for recently recognized taxa. At the 10th percentile of data availability, predicted accuracy collapsed to 6–12% for most groups, with mammals performing best at 41%. Furthermore, severe taxonomic hallucinations frequently crossed class boundaries, particularly among fishes and reptiles. These findings demonstrate that AI errors are not random, but structurally inherent to the long-tail distributions of training data. While general-purpose AI offers immense value for rapid sorting, its operational integration requires strict safeguards, including forced abstention protocols, geographic filters, and mandatory expert escalation for rare or recently described species.
No takes yet. Share an insight, caveat, or question.
Gustavo Nicolas Paez Salamanca (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: