PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 1, 20261 citationsOpen Access

UPxSocio at NTCIR-18 MedNLP-CHAT Task: Similarity-Based Few-Shot Example Selection for Prompt-Based Detection

MSMichael Van SupranesMBMartin Augustine BorlonganJLJoseph Ryan Lansangan

Key Points

  • This research aims to enhance the detection of medical, ethical, and legal risks in chatbot-generated responses using a few-shot learning approach.
  • Developed a two-step prompt-based classification framework with the Gemini-1.5-flash model.
  • Generated support statements to guide reasoning for few-shot prompt classification.
  • Evaluated methods on English versions of the Japanese and German subtasks, varying example selection and label distribution.
  • Conducted ablation studies on 24 prompt variants to understand performance influences.
  • Achieved strong performance in detecting medical risks, particularly in the German subtask.
  • Ethical and legal risk detection proved to be more challenging.
  • Logistic regression and CHAID analyses highlighted interactions between language, example similarity, and selection methods.
  • Higher similarity improved classification of risk-present cases but reduced accuracy for risk-absent cases.
  • Identified that the $k$-nearest method was more effective under high similarity, while $k$-spread offered balanced results.

Abstract

This paper presents our submission to the MedNLP-CHAT Task at NTCIR-18, which focuses on detecting medical, ethical, and legal risks in chatbot-generated responses. We propose a two-step prompt-based classification framework using the Gemini-1. 5-flash model. The method first generates support statements to guide reasoning, which are then integrated into a few-shot prompt for final classification. We evaluated our approach on the English versions of the Japanese and German subtasks, submitting two systems per subtask that varied in example selection strategy and label distribution. Our systems achieved strong performance in detecting medical risks—particularly in the German subtask—while ethical and legal risks were more challenging. To better understand the design factors influencing performance, we conducted ablation studies across 24 prompt variants. Logistic regression and CHAID analyses revealed that accuracy depends on complex interactions between subtask language, example similarity, actual label, and selection method. Higher similarity improves classification of risk-present cases but harms performance on risk-absent cases, indicating a trade-off between recall and false positives. The k-nearest method was more effective under high similarity, while k-spread offered balanced results across classes. Although the two-step prompting strategy did not show a statistically significant advantage overall, the best-performing configuration used five support statements, with diminishing gains beyond that. Our findings suggest that optimized prompt design, particularly with controlled support and example selection, can improve risk detection without requiring large-scale training or high computational resources.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Supranes et al. (2025) studied this question.

synapsesocial.com/papers/69cd79915652765b073a66e5https://doi.org/10.20736/0002002055
Ask AI
Helpful
Bookmark
Share
View Full Paper