PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
March 14, 2026International Journal of Information Management Data Insights0 citationsOpen Access

Benchmarking Large Language Models on Arabic parsing

View Full Paper
KAKamal Al-SabahiABAmer Fael Mohammed BalhafMSMohamed Moheb Zagloul Al Shamy

Key Points

  • This research evaluates how well Large Language Models parse Arabic sentences, focusing on unique linguistic challenges.
  • Evaluated leading LLMs on a human-annotated dataset of 1100 Arabic sentences.
  • Used an evaluation framework to assess syntactic and morphological features.
  • Conducted cross-model comparisons under matched multi-shot prompting.
  • Claude-3.5-Sonnet achieved the highest F1 score of 0.84, followed closely by GPT-4o at 0.83.
  • Multi-shot prompting improved accuracy by up to 18% in complex categories.
  • Open-source models displayed substantial performance gaps compared to proprietary models.

Abstract

Parsing Arabic sentences, specifically i c rāb , poses unique challenges due to the language’s intricate morphology, diverse syntactic structures, and rich contextual nuances. This study evaluates the performance of leading general-purpose Large Language Models (LLMs) in Arabic i c rāb parsing using a novel human-annotated dataset, systematically covering various grammatical phenomena. A tailored evaluation framework assesses performance across detailed syntactic and morphological features. Under matched multi-shot prompting (basis for cross-model comparisons), Claude-3.5-Sonnet achieved the highest overall F1 score (0.84), followed by GPT-4o (0.83) and Gemini-1.5-Pro (0.77). Conversely, less advanced models such as Claude2.1 and GPT-3.5-turbo struggled with complex constructions, highlighting persistent linguistic limitations. Multi-shot prompting substantially improved accuracy across proprietary models, yielding improvements of up to 18% in complex categories and underscoring the value of in-context learning. Additionally, evaluations of open-source models (DeepSeek-chat-v3-0324 and LLaMA-4-scout) established baseline performance levels confirming substantial gaps compared to proprietary models. The findings reveal ongoing challenges like diacritic sensitivity and semantic ambiguity while establishing a robust benchmark for Arabic grammatical parsing in general-purpose LLMs. All resources (dataset, codebase, and evaluation outputs) are available at https://github.com/alsabahi2030/Arabic-LLM-Parsing . • Human-annotated dataset of 1100 Arabic sentences across 11 categories and 34 subtypes • Tailored evaluation framework assessing Arabic syntactic and morphological parsing • Comparative analysis of Claude, GPT, Gemini, DeepSeek, and LLaMA-4 on Arabic • Multi-shot prompting boosts parsing accuracy considerably in complex categories

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Al-Sabahi et al. (2026) studied this question.

synapsesocial.com/papers/69b4b9db18185d8a39801ecbhttps://doi.org/10.1016/j.jjimei.2026.100404
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1ArabianGPT: Native Arabic GPT-based Large Language Model2024 · 3 citations
  2. 2ArabianGPT: Native Arabic GPT-based Large Language Model2024 · 20 citations
  3. 3Towards automated evaluation of Arabic research: a study on the efficacy of large language models in analyzing the quality of Arabic academic papers2026
  4. 4Fine-Tuning Arabic Large Language Models for improved multi-turn dialogue: A blueprint for synthetic data generation and benchmarking2026
  5. 5A bilingual benchmark for evaluating large language models2024 · 11 citations