Language is increasingly being used to define rich visual recognition problems with supporting image collections sourced from the web. Structured prediction models are used in these tasks to take advantage of correlations between co-occurring labels and visual input but risk inadvertently encoding social biases found in web corpora.
No takes yet. Share an insight, caveat, or question.
Zhao et al. (2017) studied this question.
Synapse has enriched 4 closely related papers on similar clinical questions. Consider them for comparative context: