PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
May 29, 20260 citationsOpen Access

LLMP6: A six-layer composable architecture for large language model applications

View Full Paper
YLYangxiu Liu

Key Points

  • This research aims to introduce LLMP6, a six-layer architecture for improving large language model applications.
  • Developed a six-layer structured architecture for LLM applications.
  • Utilized a foundational large model in the Core layer and lightweight specialized models in other layers.
  • Evaluated performance through experimental results focusing on response times, throughput, and latency.
  • DeepSeek's response time decreased by 44.7%, while Kimi's decreased by 17.2%.
  • Throughput increased by 86.2% for DeepSeek and 21.4% for Kimi.
  • DeepSeek's F1 score improved from 77.6% to 79.8%, and Kimi's from 68.5% to 69.8%.

Abstract

This paper proposes LLMP6, a six-layer composable architecture designed for large language model applications. The architecture decomposes the LLM application workflow into six standardized layers—Assign, Core, Tool, Filter, Check, and Show—each supporting independent model selection. The Core layer utilizes a foundational large model, while the remaining five layers are implemented with lightweight specialized small models. Key innovations include heterogeneous model collaboration, parallel/serial execution modes, automatic optimization mechanisms, and core iteration processes. Experimental results demonstrate that under the LLMP6-Full configuration, the response times of DeepSeek and Kimi were reduced by 44.7% and 17.2%, respectively; throughput increased by 86.2% and 21.4%; P99 latency decreased by 77.4% and 51.0%; latency jitter decreased by 74.7% and 38.0%; the coefficient of variation decreased by 64.3% and 38.3%; DeepSeek's F1 score improved from 77.6% to 79.8%, while Kimi's F1 score rose from 68.5% to 69.8%. Changes in accuracy and quality scores remained below 1.5%, with output diversity and readability maintained consistently, achieving an optimal balance between efficiency and quality.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Yangxiu Liu (2026) studied this question.

synapsesocial.com/papers/6a192ee7fab5b468c441827fhttps://doi.org/10.5281/zenodo.20408748
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Optimizing Large Language Models in Distributed Environments: A Holistic Approach to Efficiency, Ethics, and Governance2025
  2. 2Mapping the LLM Landscape: A Cross-Family Survey of Architectures, Alignment Methods, and Benchmark Performance2026 · 6 citations
  3. 3LLM4Perf: Large Language Models Are Effective Samplers for Multi-Objective Performance Modeling2025
  4. 4Large Language Models: From Internal Architecture and Distributed Training Optimisation to Adaptation Strategies2026 · 1 citations
  5. 5SCALABLE REAL-TIME FEATURE ENGINEERING PIPELINES FOR LARGE LANGUAGE MODEL TRAINING: A DISTRIBUTED SYSTEMS APPROACH2025