What is the clinical evidence from this study?

Study design: Other. Population: Robot-assisted surgery. Intervention: Fusion-KVE model vs. Single-input state estimation models. Primary outcome: Frame-wise state estimation accuracy on RIOUS dataset.

February 7, 2020Open Access

Temporal Segmentation of Surgical Sub-tasks through Deep Learning with Multiple Data Sources

Key Result

The Fusion-KVE model, incorporating kinematics, vision, and system events, achieved a frame-wise state estimation accuracy of 89.4% on the RIOUS dataset, improving upon single-input models.

Structured PICO

Does the Fusion-KVE model improve frame-wise state estimation accuracy in robot-assisted surgical tasks compared to existing models?

Population

JHU-ISI Gesture and Skill Assessment Working Set (JIGSAWS) and robotic intra-operative ultrasound (RIOUS) imaging datasets created using the da Vinci Xi surgical system

Intervention

Fusion-KVE (a unified surgical state estimation model incorporating Kinematics, Vision, and system Events)

Comparator

Other state-of-the-art surgical state estimation models

Outcome

Frame-wise state estimation accuracysurrogate

The proposed Fusion-KVE deep learning model improves the accuracy of temporal segmentation of surgical sub-tasks in robot-assisted surgeries by integrating multiple data sources.

Main Result

Absolute Event Rate: 89.4% vs 78.4%

Limitations

Running multiple state estimation models at the same time requires higher computing power
Real-time state estimation limits the amount of data available to the model (causal setting)

Abstract

Many tasks in robot-assisted surgeries (RAS) can be represented by finite-state machines (FSMs), where each state represents either an action (such as picking up a needle) or an observation (such as bleeding). A crucial step towards the automation of such surgical tasks is the temporal perception of the current surgical scene, which requires a real-time estimation of the states in the FSMs. The objective of this work is to estimate the current state of the surgical task based on the actions performed or events occurred as the task progresses. We propose Fusion-KVE, a unified surgical state estimation model that incorporates multiple data sources including the Kinematics, Vision, and system Events. Additionally, we examine the strengths and weaknesses of different state estimation models in segmenting states with different representative features or levels of granularity. We evaluate our model on the JHU-ISI Gesture and Skill Assessment Working Set (JIGSAWS), as well as a more complex dataset involving robotic intra-operative ultrasound (RIOUS) imaging, created using the da Vinci Xi surgical system. Our model achieves a superior frame-wise state estimation accuracy up to 89.4%, which improves the state-of-the-art surgical state estimation models in both JIGSAWS suturing dataset and our RIOUS dataset.

Mark Helpful

Bookmark

Relay

View Full Paper