The visual system can extract statistical information from multiple human faces in the form of ensemble representation. Despite extensive research, the automaticity of face ensemble encoding is still under debate. Some studies suggest that face ensemble encoding occurs automatically, while others postulate a capacity-limited processing that requires attentional resources. This inconsistency might be attributed to methodological differences in set size: A group-based size comprises several faces, whereas a crowd-based size exceeds eight faces. Such variation in set size may modulate the role of attention, where the group-based size favors automatic encoding, and the crowd-based size necessitates attentional engagement due to extensive noise. We addressed these questions by recording event-related potentials (ERP) of electroencephalography (EEG) in 35 subjects. Participants underwent three conditions: the implicit condition (ignoring faces while performing a fixation task), the explicit condition (detecting changes in the mean emotion), and the categorization condition (recognition of emotion proportion for each trial). We found that early components showed consistent set size effects across all tasks, indicating automatic sensory processing. However, participants could detect changes in average emotional expressions only in the explicit condition (compared to the implicit one), as was reflected in the P3 and LPP components of ERP. The categorization task results suggest that ensemble extraction alone is insufficient for forming predictions. Moreover, theta-band oscillations supported our findings only in the explicit condition for the group-based set size. These results suggest that face ensemble perception involves both automatic sensory registration and attention-dependent comparison of ensemble statistics.
Koch et al. (Fri,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: