GroundCap: A visually grounded image captioning dataset with object and action identification | Synapse