Demonstrates a universal geometric differentiation between experiential and factual concepts in language models, suggesting deep cognitive implications.
Key Points
To investigate how human language differentiates between experiential and factual semantic content through geometric representation.
Analyzed 10 large language models from 9 organizations across 4 countries.
Used Grassmann subspace distance to measure geometric distinctions in semantic categories.
Examined model behavior in multiple languages including English, Chinese, French, and Arabic.
Observed neural layer activity, particularly in MLP layers, during processing of different concepts.
A universal two-cluster structure was observed across all models and training paradigms.
Experiential concepts were geometrically separated from factual concepts as predicted.
Identical structural patterns were found in pre-RLHF and post-RLHF model versions.
Self-referential and extrinsic content remained consistently isolated in terms of geometry.