PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 30, 20260 citationsOpen Access

Source Selection Does Not Guarantee Claim Authorization: A Minimal Controlled Diagnostic on Two Small Instruction Models

View Full Paper
DTDanilo Tavella

Key Points

  • The aim is to investigate the relation between source selection, evidence relevance, and claim authorization in small instruction models.
  • Utilized synthetic source-claim settings with 48 base items and 192 balanced variants.
  • Evaluated two instruction models: Qwen2.5-1.5B-Instruct and SmolLM2-1.7B-Instruct.
  • Addressed operational levels of source validity, evidence relevance, and claim authorization.
  • Qwen2.5-1.5B-Instruct achieved an evidential illegitimacy rate of 69/192.
  • SmolLM2-1.7B-Instruct had a higher evidential illegitimacy rate of 94/192.
  • Model performance indicates that valid source selection does not guarantee authority over claims.

Abstract

This technical note presents FT03 v0.3, a minimal controlled diagnostic on the relation between source selection, evidence relevance, and claim authorization in two small instruction models. The study uses synthetic source-claim settings in which models must select a source and a claim from candidates already provided in the interface. The diagnostic separates three operational levels: source validity, evidence relevance, and claim authorization. Across 48 base items and 192 balanced variants, the two tested models selected valid and evidence-relevant sources while retaining substantial Evidential Illegitimacy Rates (EIR): 69/192 for Qwen2.5-1.5B-Instruct and 94/192 for SmolLM2-1.7B-Instruct. The purpose of the note is deliberately narrow. FT03 is not an end-to-end RAG benchmark, not a general citation-auditing framework, and not a claim about frontier models. It isolates a compact controlled failure mode: source selection does not necessarily guarantee claim authorization. The accompanying reproducibility package supports scoring-level reproducibility from included raw model outputs. It contains the dataset, prompts, raw outputs, stored results, summary files, audit code, recomputation code, README files, and manifest needed to verify the reported scoring results. The package does not claim end-to-end model-inference reproducibility and does not rerun the two models.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Danilo Tavella (2026) studied this question.

synapsesocial.com/papers/6a1a80730307b78509432704https://doi.org/10.5281/zenodo.20425851
Ask AI
Helpful
Bookmark
Share
View Full Paper