On predefined and synthetic datasets with controlled noise, we test Borda, Copeland, Footrule, Kemeny-Young, Median Rank, PageRank, Plackett-Luce, Reciprocal Rank Fusion, and Schulze, varying numbers of alternatives and rankings to assess scalability. Positional and simple pairwise rules tend to agree and reward consistently strong options; distance-based and probabilistic models can shift winners toward items closest to the average order or with higher inferred worth. Under low disagreement, most methods yield similar consensus, permitting flexible choice; under high disagreement, rankings diverge and runtimes spread widely. Borda, Median, RRF, and Schulze scale well; Kemeny-Young and Plackett-Luce become costly. We provide a comparative, mechanism-aware view of aggregation choices, linking robustness and computational feasibility to disagreement regimes to guide practical MCDA in real decision settings. Our code is available at https://github.com/Valdecy/pyRankMCDA .
Pereira et al. (Thu,) studied this question.