PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
September 28, 20250 citationsOpen Access

Is GPT-OSS Good? A Comprehensive Evaluation of OpenAI's Latest Open Source Models

View Full Paper
ZBZhuming BiKCKeyu ChenCTChih-Kuo Tseng

Key Points

  • gpt-oss-20B consistently outperformed gpt-oss-120B in benchmarks like HumanEval and MMLU, indicating its superior efficiency.
  • The evaluation included comparisons against large language models with parameters ranging from 14.7B to 235B, showcasing various architectures.
  • Analyses utilized statistical validation such as McNemars test, confirming the reliability of performance findings in a standardized setting.
  • These results imply that scaling in sparse architectures may not equate to proportional gains, emphasizing the need for optimized model strategies.

Abstract

In August 2025, OpenAI released GPT-OSS models, its first open weight large language models since GPT-2 in 2019, comprising two mixture of experts architectures with 120B and 20B parameters. We evaluated both variants against six contemporary open source large language models ranging from 14.7B to 235B parameters, representing both dense and sparse designs, across ten benchmarks covering general knowledge, mathematical reasoning, code generation, multilingual understanding, and conversational ability. All models were tested in unquantised form under standardised inference settings, with statistical validation using McNemars test and effect size analysis. Results show that gpt-oss-20B consistently outperforms gpt-oss-120B on several benchmarks, such as HumanEval and MMLU, despite requiring substantially less memory and energy per response. Both models demonstrate mid-tier overall performance within the current open source landscape, with relative strength in code generation and notable weaknesses in multilingual tasks. These findings provide empirical evidence that scaling in sparse architectures may not yield proportional performance gains, underscoring the need for further investigation into optimisation strategies and informing more efficient model selection for future open source deployments.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Bi et al. (2025) studied this question.

synapsesocial.com/papers/68d913a34ddcf71ba560b8b7https://doi.org/10.48550/arxiv.2508.12461
Ask AI
Helpful
Bookmark
Share
View Full Paper