Key points are not available for this paper at this time.
Photo editing can be a challenging task, and it becomes even more difficult on the small, portable screens of mobile devices that are now frequently used to capture and edit images. To address this problem we present PixelTone, a multimodal photo editing interface that combines speech and direct manipulation. In this video, we demonstrate how our system uses natural language for expressing users' desired changes to an image. We also demonstrate how we combine natural language and touch gestures for creating named references and sketching to localize image operations to specific regions.
Linder et al. (Sat,) studied this question.