Automated system transcribes and summarizes YouTube videos effectively, indicating its potential for various applications.
This study developed an automated system for transcribing and summarizing YouTube video content to help users efficiently access important information from long-duration videos. The system only requires a video link as input, then automatically extracts the audio, transcribes it using the Whisper Small model, and generates a summary using the LexRank algorithm, which selects key sentences based on graph centrality. Transcription quality is evaluated using the Word Error Rate (WER) metric, with an average score of 0.3703 or approximately 37%, indicating a fairly good level of accuracy. Meanwhile, the summarization evaluation using ROUGE metrics yielded average F1-Scores of 33% for ROUGE-1, 10% for ROUGE-2, and 19% for ROUGE-L, reflecting the relevance of the generated summaries to manual references. The average transcription processing time is around 0.17 seconds per word, while the summarization process takes less than 1 second. All results—including transcriptions, summaries, and evaluation metrics—are automatically saved in CSV format. This system demonstrates stable performance and holds strong potential for various video-based knowledge extraction applications, such as in education, journalism, research, and digital documentation.
No takes yet. Share an insight, caveat, or question.
Wicaksana et al. (2025) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: