PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 24, 2024113 citationsOpen Access

AudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head

View Full Paper
RHRongjie HuangMLMingze LiDYDongchao Yang

Key Points

Key points are not available for this paper at this time.

Abstract

Large language models (LLMs) have exhibited remarkable capabilities across a variety of domains and tasks, challenging our understanding of learning and cognition. Despite the recent success, current LLMs are not capable of processing complex audio information or conducting spoken conversations (like Siri or Alexa). In this work, we propose a multi-modal AI system named AudioGPT, which complements LLMs (i.e., ChatGPT) with 1) foundation models to process complex audio information and solve numerous understanding and generation tasks; and 2) the input/output interface (ASR, TTS) to support spoken dialogue. With an increasing demand to evaluate multi-modal LLMs of human intention understanding and cooperation with foundation models, we outline the principles and processes and test AudioGPT in terms of consistency, capability, and robustness. Experimental results demonstrate the capabilities of AudioGPT in solving 16 AI tasks with speech, music, sound, and talking head understanding and generation in multi-round dialogues, which empower humans to create rich and diverse audio content with unprecedented ease. Code can be found in https://github.com/AIGC-Audio/AudioGPT

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Huang et al. (2024) studied this question.

synapsesocial.com/papers/68e72954b6db6435876a2cd8https://doi.org/10.1609/aaai.v38i21.30570
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities2024
  2. 2Using large language models in the analysis of urban sound environments2024
  3. 3A comprehensive overview of audio language models2025
  4. 4ChatGPT Alternative Solutions: Large Language Models Survey2024 · 11 citations
  5. 5MedPodGPT: A multilingual audio-augmented large language model for medical research and education2024 · 5 citations