Theoretical analysis proposes breath coordination pedagogy in human vocal training, suggesting a somatic physiological benchmark for artificial general intelligence.
Speech is exhaled, shaped time. A breath comes before every sentence, and the state of the speaker travels in the sound whether he means it to or not. This paper takes that fact as the ground for a criterion of artificial general intelligence, in place of the text benchmarks and the economic definition by which the field currently measures itself. A system meets the criterion if it teaches a person, untrained or a trained singer, the breath coordination that bel canto demands, so that the person feels it, keeps it under real conditions including stage fright, speaks and sings with noticeably less effort, and moves a listener who knows nothing of the process. Bel canto is the road and the singing is the receipt; what is taught is breathing that is efficient, and the awareness to keep it when the singing stops. The premise underneath is a claim about the human being: breathing is the variable in which body, emotion, attention, communication and self-regulation are integrated, communication is a form of organised breathing, and voice, nerves, presence, singing and listening are appearances of one system. Whoever does not understand breathing does not understand the human being, and a system that claims to understand human beings in general cannot be blind to their central regulating mechanism. The difficulty is documented over four centuries. The statements that are physiologically true bring about nothing, and the statements that work are not true. Knowledge was never the missing piece; the last obstacle is the relationship in which honest judgement costs the teacher something. A system that hears the sound itself rather than words about the sound, keeps its own record of every attempt, and gives plain instructions that set off in the learner a feedback only he can perceive, while the recording holds the fact, has no relationship to protect and so removes that obstacle. The criterion turns on the system as well: anything that reorganises human communication is already acting on human breath, and a system that understood the human situation would measure itself by whether people speak more freely after using it. Today’s systems stop short for a nameable reason — the missing integration of audio analysis, voice physiology, psychophysiology and the knowledge of which instruction comes next — and none of it can be designed or judged without the disciplines that have kept the record. The paper closes with the words themselves: German has Empirie for experience held in the first person and Empirismus for the doctrine about it; English has one word, empiricism, for both, and so cannot say what the singer has and the literature lacks. Romanian keeps the distinction and, in its word for soul, keeps the breath.
No takes yet. Share an insight, caveat, or question.
Fatima C. Spisländer (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: