PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 13, 2026Journal of the Korea Society of Computer and Information0 citations

A Comparative Study on the Performance of GPT-4o-mini, Claude 4 Sonnet, and Gemini 2.5 Flash Models Using the Prompt Runner Framework

View Full Paper
MLMisun Lee

Key Points

  • This study aims to compare the performance of various AI models based on prompt types using the Prompt Runner framework.
  • Used a dataset of 90 prompts across 9 types to evaluate different performance metrics.
  • Measured accuracy using SBERT embedding-based cosine similarity.
  • Evaluated consistency and logic with specific weightings based on accuracy.
  • Assessed creativity through a weighted sum of novelty, diversity, and fluency.
  • Claude 4 Sonnet excelled in logic (0.58) and creativity (0.44).
  • GPT-4o-mini showed fast response times.
  • Gemini 2.5 Flash performed well in accuracy (0.66) and consistency (0.62).
  • Claude 4 Sonnet demonstrated a stable performance balance between overall effectiveness and response time.

Abstract

본 연구는 Prompt Runner 프레임워크를 기반으로 GPT-4o-mini, Claude 4 Sonnet, Gemini 2.5 Flash 모델을 대상으로 프롬프트 유형별 성능을 비교하였다. 총 9개 유형, 90문항으로 구성된 프롬프트 데이터셋을 활용하여 정확도(Accuracy), 일관성(Consistency), 논리성(Logic), 창의성(Creativity), 응답 시간(Response Time)을 평가하였다. 정확도는 SBERT 임베딩 기반 코사인 유사도로 산출하였으며, 일관성과 논리성은 각각 정확도에 0.95, 0.9의 가중치를 적용하였다. 창의성은 새로움, 다양성, 유창성의 가중합(0.5N+0.3D+0.2F)으로 산출하였다. 분석 결과 Claude 4 Sonnet은 논리성(0.58)과 창의성(0.44)에서 우수한 성능을 보였으며, GPT-4o-mini는 빠른 응답시간을 나타냈고, Gemini 2.5 Flash는 정확도(0.66)와 일관성(0.62)에서 높은 성능을 보였다. 특히, Claude 4 Sonnet은 전반적인 성능과 응답 시간 간의 균형 측면에서 가장 안정적이고 일관된 성능을 보여, 효율성과 품질을 동시에 확보한 모델로 평가되었다. 본 연구를 통해 LLM 성능평가에서 각 AI 모델의 API 기반 정량적 성능지표를 비교 분석함으로써, 프롬프트 유형별 모델 특성과 성능 차이를 체계적으로 규명하였다.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Misun Lee (2025) studied this question.

synapsesocial.com/papers/69b3aad702a1e69014ccb8e1https://doi.org/10.9708/jksci.2026.31.02.043
Ask AI
Helpful
Bookmark
Share
View Full Paper