This paper proposes a novel framework for generating lingual descriptions of indoor scenes. Whereas substantial efforts have been made to tackle this problem, previous ap-proaches focusing primarily on generating a single sentence for each image, which is not sufficient for describing complex scenes. We attempt to go beyond this, by generating co-herent descriptions with multiple sentences. Our approach is distinguished from conven-tional ones in several aspects: (1) a 3D visual parsing system that jointly infers objects, attributes, and relations; (2) a generative grammar learned automatically from training text; and (3) a text generation algorithm that takes into account coherence among sen-tences. Experiments on the NYU-v2 dataset show that our framework is able to generate natural multi-sentence descriptions, outperforming those produced by a baseline. 1
No takes yet. Share an insight, caveat, or question.
Lin et al. (2015) studied this question.
Synapse has enriched 3 closely related papers on similar clinical questions. Consider them for comparative context: