Tools and systems for content-based access to multimedia and—image, video, audio, graphics, text, and any number of combinations—has increased in the last decade. We’ve seen a common theme of developing automatic analysis techniques for deriving metadata (data describing information in the content at both syntactic and semantic levels). Such metadata facilitates developing innovative tools and systems for multimedia information retrieval, summarization, delivery, and manipulation. Many interesting demonstrations of potential applications and services have emerged—finding images visually similar to a chosen picture (or sketch); summarizing videos with thumbnails of keyframes; finding video clips of a specific event, story, or person; and producing a two-minute skim of an hour-long program. (Audio–visual skims are condensed media clips that summarize information in the content.) There’s much excitement and buzz created by these fancy applications. But people are always asking, What will be engineering’s Holy Grail for content-based media analysis in practical applications? My response is that the answer is complex. It’s less about a specific algorithm or service than a rigorous methodology to formulate and evaluate content-based analysis research. Content chain To evaluate content-based research methodologies, we must first consider who the intended users are and whether alternative solutions exist. Hence, it’s important to consider each solution within the context of the content chain, the process starting with acquisition, followed by production, processing, and finally consumption. Figure 1 shows a diagram depicting the content chain and the relationships among different components. Each piece exists in part of the Figure 1. The end-toend content chain, content-based research principles, and potential applications.
No takes yet. Share an insight, caveat, or question.
Shih‐Fu Chang (2002) studied this question.
Synapse has enriched one closely related paper. Consider it for comparative context: