PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 29, 20260 citationsOpen Access

The Scaffold Effect: Agent Harnesses Introduce Larger Cost Variance Than Model Upgrades in Terminal Benchmark Evaluation

View Full Paper
NVNaman Vats

Abstract

This preprint presents a controlled empirical study of how agent harness scaffolding affects terminal benchmark evaluation independently of model choice. We evaluate two frontier models, Qwen 3.6 Plus and MiniMax M2.5, across three agent harnesses: Goose, OpenCode, and OpenHands-SDK on 50 tasks from Terminal-Bench Pro. The study finds that harness choice produces up to a 40× difference in tokens per solved task, while pass-rate differences remain within 2–10 percentage points. The results suggest that agent benchmark reporting should treat harness-model pairs, rather than models alone, as the unit of comparison. The accompanying GitHub repository contains the selected task list, harness configurations, raw trial logs, aggregate snapshots, schemas, and analysis scripts. Code and data: https://github.com/namanvats/scaffold-effects

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Naman Vats (2026) studied this question.

synapsesocial.com/papers/69f1545d879cb923c49448behttps://doi.org/10.5281/zenodo.19819491
Ask AI
Helpful
Bookmark
Share
View Full Paper