Dear Editor Collective behavior refers to the phenomenon where massive individuals follow simple rules to generate complex swarm intelligence (Theraulaz et al., 2003; Sumpter, 2006; Ballerini et al., 2008). Research on such behaviors enhances our understanding of the behavioral decision-making mechanisms and evolutionary patterns of social animals (Petit Wang et al., 2020; Ioannou Prakash et al., 2014; Virágh et al., 2014). The honeybee (Apis mellifera), a well-known eusocial insect vital for pollination, represents a model species for the study of collective behavior (von Frisch, 1967; Hasenjager et al., 2022). However, investigations of honeybee collective behaviors are hindered by the difficulties in tracking and analyzing individual behaviors—due to large colony sizes, small body sizes, and frequent obstructions in dark, crowded hives (Bozek et al., 2021). Thus, clear, accurate and non-invasive recording of honeybee in-hive collective behavior remains challenging (Ulrich et al., 2024). Since Karl von Frisch used five colors to manually observe in-hive honeybee behavior in 1960s (von Frisch, 1967), researchers have continuously improved marking and observation methods (Camazine, 1991; Seeley et al., 1991; Seeley et al., 2000; Knauer et al., 2005). With advances in image processing and computer vision, research on honeybee collective behavior has entered a stage of automated identification and tracking. Knauer Bozek et al., 2021) avoid the time-consuming tagging process but face re-identification issues when bees exit and re-enter the field of view, leading to trajectory misassignment (Bozek et al., 2021). Third, the commonly used barcode or QR tags (Crall et al., 2015; Wario et al., 2015; Blut et al., 2017; Boenisch et al., 2018; Gernat et al., 2023; Wu et al., 2025) rely on multiple bits. An individual's identifier is determined by analyzing information across all such bits. However, occlusion of some of the bits (common in hives) causes misidentification. Last but not least, barcode tags cannot be manually checked or verified, making them inconvenient for research that requires retrieving specific individuals for downstream experiments. To address these challenges, we developed a STMIIT (Symbol Tag for Massive Insect Identification and Tracking) system, whose structure and workflow are detailed below. To expand marking capacity and facilitate both machine recognition and manual verification, we designed a novel set of 3-mm-diameter circular symbol tags (Fig. 1A, B) combining letters, numbers and shapes, capable of marking 10 000 insect individuals simultaneously (File S1). We also modified a single-comb observation hive (Fig. 1C, D) for long-term, precise and non-invasive video recording of in-hive behaviors under infrared light conditions (details in File S2). For accurate auto-identification and continuous tracking of massive bee individuals, we established a PyTorch-based software environment using YOLOv11 as the core algorithm. A three staged detection-and-recognition model was implemented to detect the circular symbol tag (CST) region, recognize the triangular orientation marker (TOM), and identify the characters. Because random honeybee movement results in inconsistent character orientations, arbitrarily oriented tags were rotated to upright position before character identification. The schematic process is in Fig. 2 and detailed algorithm in File S2. A single-comb colony with ∼3000 tagged workers and one egg-laying queen was established, and in-hive behavior video data were collected (screenshot see Fig. 3A). To train the STMIIT models: (1) Frames were extracted from the videos to construct Dataset-A for training the circular symbol tag (CST) detection model (Fig. 3B). In this study, 108 000 frames in total were extracted, 400 representative images with 1071 tagged bees were selected, and 56 933 cropped tag images were generated by the CST detection model. (2) Dataset-B was built by selecting the cropped tag images to train the triangular orientation marker (TOM) recognition model. In this study, we chose 3880 tag images including intact, incomplete, blurred, or deformed TOMs (Fig. 3C) commonly occurred in real-world conditions. (3) Dataset-C was constructed to train the character identification model. In the present study, 7500 tag images were selected, covering all 50 numbers, letters, and shapes (Fig. 3D) in various possible states (details in Files S2 and S3 and for output, see Video S1). To record and analyze the behavioral trajectories of tagged honeybees on the comb, we established a two-dimensional coordinate system based on the video resolution and field-of-view physical dimensions for precise cell localization. Once the characters on a tag have been identified, the system assigns a unique ID to this tag. The ID is then stored in a MySQL database together with the tag's characters, spatial coordinates, timestamp, rotation angle, and other related information (storage format shown in Table S1). Subsequently, the trajectory of the tagged honeybee within a given time period can be obtained by retrieving its spatial coordinates from the database and linking them sequentially according to their timestamps (see in Fig. 4). This ensures the robustness of long-term tracking for individual trajectories. Further details regarding trajectory continuity are provided in Files S2 and S3. The CST detection model trained on 1071 honeybees achieved a mean average precision at an intersection over union threshold of 0.5 (mAP@50) of 98.8%, with precision and recall of 96.6% and 97.7%, respectively; the TOM recognition model achieved a mAP@50 of 99.5%, and both precision and recall were 99.7%; the character identification model achieved 97.8% mAP@50, with both precision and recall reaching 96.8% (Table S2). To assess real-world performance of the STMIIT system, 10 honeybee in-hive activity videos were processed (300 frames for each, 3000 frames in total) for manual verification, yielding an overall recognition accuracy of 93.6% (Table S3). The computer hardware configuration used in this study is detailed in File S2, with the processing speed of STMIIT provided in Table S4. The STMIIT system is well-suited for two-dimensional planar research scenarios. In practical implementation, key considerations include tag-to-field-of-view ratio and tag-background color contrast. In our experiment, 3-mm-diameter tags paired with a 3840 × 2160-pixel camera and 260 × 150-mm field of view (fixed distance/focal length) yielded 44.3 × 43.2-pixel tags and 73.8 × 72.0-pixel cells—sufficient for manual identification. Lens distortion correction ensured consistent pixel dimensions across the field of view, supporting precise tag localization and tracking. When applying to other insects or experimental settings, clear imaging and precise trajectory tracking can be achieved by adjusting the tag size, the imaging resolution, along with the corresponding software parameters. While our data were collected under 940-nm infrared illumination, the system is also suitable for visible-light conditions, with expected better performance in recognition and tracking. In collective behavior studies of social insects, crowded environments such as bee-bee or bee-cell friction, which was almost an unavoidable cause to limit the experimental durability. In our study, 5.6% tags were observed detached around 10 days, which may compromise the integrity and accuracy of a corresponding portion of the experimental data. To improve tag retention, we are optimizing adhesive type, printing paper, tagging technique, and pressure parameters. We successfully developed an automated identification and tracking system for massive insects based on the YOLOv11 model via taking honeybees as an example. Its 93.6% overall recognition accuracy in practical applications is comparable to similar techniques (Boenisch et al., 2018; Knauer et al., 2022; Wu et al., 2025), with stable and accurate trajectory tracking in complex hive environments. In terms of encoding capacity, as shown in Table 1, STMIIT's 10 000 unique identifiers outperform 11-bit binary square tags (2048 theoretical capacity; Gernat et al., 2018), 12-bit binary circular tags (4096 theoretical capacity; Wario et al., 2015) and 25-bit binary square tags (7515 viable tags; Crall et al., 2015). Concerning the issue of tag recognition accuracy under occlusion, we face a challenge that no comparable data is available for quantitative comparison and analysis with existing tags. Nevertheless, from a theoretical standpoint, the recognition accuracy of barcode tags depends entirely on the correct identification of the bits that constitute the tags. Since existing barcode tags typically consist of a number of small bits, STMIIT tags demonstrate better, or at least comparable, identifiability and recognition accuracy compared to barcode tags, under identical occlusion angles and occlusion levels. In practice, we have indeed observed several cases where moderate occlusion does not compromise the identification accuracy of STMIIT tags (as illustrated in Fig. S1). Under equivalent conditions, however, the difficulty faced by barcode-type tags is evident. Another significant advantage over barcode/QR tags is the combination of numbers, letters, and shapes, enabling manual recognition of specific treatment groups (e.g., age, genetic background, gene modification). This facilitates downstream analysis and sample retrieval, broadening the system's applicability for in-depth research on collective behaviors of social insects. We wish to express our sincere gratitude to Haiou Kuang and Danyin Zhou in Department of Apiculture in Yunnan Agricultural University, Zachary Huang in Michigan State University and Deqing Zong in Sericultural and Apicultural Research Institute in Yunnan Academy of Agricultural Sciences for their constructive suggestions and insightful comments. This work was supported by the National Key Research and Development Program of China (2022YFD1601907) and the China Agriculture Research System of MOF and MARA (CARS-44-KXJ13). The authors declare they have no competing interests. Table S1 Data format in MySQL database. Table S2 Performance metrics of models at each stage on the independent test sets. Table S3 Results of characters identification. Table S4 Processing time per frame for a 5 min video segment. Fig. S1 Screenshots of moderately occluded but accurately identifiable tags. Please note: The publisher is not responsible for the content or functionality of any supporting information supplied by the authors. Any queries (other than missing content) should be directed to the corresponding author for the article.
Wang et al. (Sun,) studied this question.