Argania spinosa (L. ) Skeels, commonly known as the argan tree, is a species of significant ecological and economic importance endemic to southwestern Morocco. This drought-tolerant tree occupies approximately 900, 000 hectares and has been recognized by UNESCO as a biosphere reserve since 1998 1. The species is particularly valued for its oil-rich seeds, which produce argan oil used in culinary, cosmetic, and medicinal applications due to its high content of unsaturated fatty acids and antioxidants 1. Despite the economic importance of A. spinosa, comprehensive genomic resources have been limited. Previous genomic studies have focused primarily on organellar genomes, with the chloroplast genome (158, 848 bp) and mitochondrial genome (707, 441 bp) being sequenced and characterized 2, 3, 4. Initial nuclear genome assemblies have provided basic structural information 5, with a recent annotation published by Rupp et al. 6 describing 62, 590 predicted genes; however, that work, and earlier draft assemblies lacked comprehensive characterization of non-coding RNAs, detailed repetitive element landscape, and systematic pathway-level functional annotation. In this Data Descriptor we provide an updated scaffold-level assembly and a comprehensive structural and functional annotation for the “Argan Amghar” nuclear genome. The same individual was originally sequenced by our group using Illumina HiSeq X Ten technology, and the raw reads were deposited in the NCBI Sequence Read Archive under BioProject PRJNA294096 5. In the present work we re-use these reads to generate a curated scaffold-level reference, and an integrated annotation tailored to downstream comparative and functional genomics. Our assembly spans 690 Mbp in 186, 325 scaffolds with an N50 of 25 Mbp and an L50 of 11 large macro-scaffolds. Genome size estimates from k-mer analysis (698 Mbp), together with previous flow-cytometry estimates, are consistent with this assembly size. The final gene set contains 51, 078 protein-coding genes, representing a balanced gene repertoire that lies between earlier, more expansive annotations (~62, 590 genes) and the more conservative predictions for closely related Sapotaceae and Ericales genomes. We additionally annotate 2, 081 non-coding RNA genes and extensive repetitive elements covering 53. 0% of the assembly. Functional annotation combines orthology-based predictions from eggNOG-mapper v2 7, protein-domain and motif assignments from InterProScan 8, 9, and BLASTp 10 searches against UniProtKB/Swiss-Prot 11, together with pathway mapping using the Plant Metabolic Network (PMN) 12. These yields curated functional descriptions and domain architectures for 32, 785 genes and UniProt-validated annotations for 25, 484 proteins. Of these, 1, 906 are classified as putative transcription factors using the Transcription Factor Prediction pipeline from PlantTFDB v5. 0 13, and 7, 044 are linked to metabolic pathways, including those involved in lipid and tocopherol metabolism. Our dataset provides a deeply curated and well-documented annotation of the widely used scaffold-level assembly GCA₀03260245. 2. Researchers who have already adopted this assembly (Such as Rupp et al. 6) can directly reuse our GFF3, and proteome files without re-mapping to a different reference. The primary applications for this dataset include Comparative genomics studies within Ericales and Sapotaceae, identification of genes underlying drought tolerance and oil biosynthesis, development of molecular markers for breeding programs, conservation genetics applications, and biotechnological research aimed at understanding and optimizing argan oil production. The dataset is particularly valuable for researchers studying lipid metabolism, tocopherol biosynthesis, stress tolerance mechanisms, and evolutionary genomics of economically important trees.
Azami et al. (Tue,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: