The SCU-1 team participated in the "Detection of Argument Temporal References in Earnings Conference Calls" subtask of the NTCIR-18 FinArg-2 task. This study reports our approach to solving the problem and discusses the official results. We analyze the impact of step-by-step reasoning, model collaboration, and prompt design on the classification performance of large language models (LLMs). Through a series of experiments, we found that providing detailed explanations and incorporating previous model predictions significantly improved classification accuracy. Additionally, we compared different LLM discussion mechanisms and prompt design strategies, revealing that allowing models to reference each other and reason based on prior outputs effectively enhances decision-making quality. Run 3, which included complete reasoning steps and prior model outputs, achieved the best performance, highlighting the advantages of cross-model reference and optimized prompt design. These findings offer new directions for improving LLM-based classification tasks.
Ho et al. (Fri,) studied this question.