PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 9, 2015Proceedings of the National Academy of Sciences341 citationsOpen Access

Visual Turing test for computer vision systems

DGDonald GemanJohns Hopkins UniversitySGStuart GemanBrown UniversityNHNeil HallonquistJohns Hopkins University Applied Physics Laboratory

Key Points

  • To develop an operator-assisted visual Turing test that evaluates computer vision systems through sequential, story-based binary questions rather than standard object localization metrics.
  • Designed a query engine that generates stochastic sequences of binary questions based on statistical constraints learned from training data.
  • Implemented a human-assisted 'just-in-time truthing' process to filter ambiguous questions and provide ground-truth labels sequentially.
  • Structured query streams to follow visual storylines progressing from object identification to attribute evaluation and inter-object relationships.
  • Generated question sequences where conditional outcomes remain balanced and unpredictable (~50% positive/negative) based on prior context.
  • Established an evaluation benchmark isolated strictly to vision capabilities by utilizing deterministic binary parsing without requiring natural language processing modules.

Abstract

Today, computer vision systems are tested by their accuracy in detecting and localizing instances of objects. As an alternative, and motivated by the ability of humans to provide far richer descriptions and even tell a story about an image, we construct a "visual Turing test": an operator-assisted device that produces a stochastic sequence of binary questions from a given test image. The query engine proposes a question; the operator either provides the correct answer or rejects the question as ambiguous; the engine proposes the next question ("just-in-time truthing"). The test is then administered to the computer-vision system, one question at a time. After the system's answer is recorded, the system is provided the correct answer and the next question. Parsing is trivial and deterministic; the system being tested requires no natural language processing. The query engine employs statistical constraints, learned from a training set, to produce questions with essentially unpredictable answers-the answer to a question, given the history of questions and their correct answers, is nearly equally likely to be positive or negative. In this sense, the test is only about vision. The system is designed to produce streams of questions that follow natural story lines, from the instantiation of a unique object, through an exploration of its properties, and on to its relationships with other uniquely instantiated objects.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Geman et al. (2015) studied this question.

synapsesocial.com/papers/6a0ac95e334bc3615dac9edchttps://doi.org/10.1073/pnas.1422953112
Ask AI
Helpful
Bookmark
Share
View Full Paper