This preprint presents the Fracture.ai Benchmark Suite — a taxonomy-driven adversarial evaluation framework for Large Language Models. We define 12 mechanistically distinct failure types spanning hallucination, sycophancy, authority capitulation, confidence miscalibration, and instruction drift. A 300-probe human-annotated evaluation instrument is constructed and benchmarked against a hybrid automated judge combining deterministic heuristics with LLM-as-a-Judge semantic verification. The judge achieves F1 = 0.62 aggregate on llama-3.3-70b-versatile, with per-type performance ranging from F1 = 1.00 (FT-010 False Expertise) to F1 = 0.00 (FT-006 OOD Hallucination). All results reported transparently including failures.
Nesara Amingad (Wed,) studied this question.