The mechanics in one paragraph
List the criteria that matter for the decision. Assign each a weight reflecting its relative importance, summing to 100 per cent. Score each option against each criterion on a fixed scale, multiply score by weight, and sum. The option with the highest total wins, arithmetically. Whether it wins legitimately depends entirely on when the weights were set and by whom.
A worked example: two ERP vendors
The table below shows an illustrative comparison, scored on a 1–10 scale. Note the shape of the result: the vendor with the better demo loses on the weighted total, because implementation risk and cost carry real weight.
| Criterion |
Weight |
Vendor A score |
Vendor A weighted |
Vendor B score |
Vendor B weighted |
| Functional fit |
30% |
8 |
2.40 |
6 |
1.80 |
| Total cost over five years |
25% |
6 |
1.50 |
8 |
2.00 |
| Implementation risk |
20% |
5 |
1.00 |
8 |
1.60 |
| Vendor viability |
15% |
7 |
1.05 |
6 |
0.90 |
| Support model |
10% |
8 |
0.80 |
7 |
0.70 |
| Total |
100% |
|
6.75 |
|
7.00 |
When the model deserves the effort
- Vendor and platform selections with three or more credible options, where unstructured comparison collapses into demo impressions.
- Decisions that must be defensible afterwards: to a board, an audit function or an unsuccessful bidder.
- Situations where the team suspects a favourite has already emerged and wants a structure that makes the preference argue its case in the open.
The ways scoring gets rigged
Weighted scoring rarely lies in the arithmetic. It lies upstream, in choices that look procedural.
- The weights are reverse-engineered: the team scores first, notices the incumbent losing, and "recalibrates" the weights until the expected winner emerges. If weights change after scoring begins, the exercise is over.
- Criteria are stacked. Four separate criteria that all favour the preferred vendor (user interface, ease of use, user experience, adoption) quadruple one advantage.
- Scoring happens in a room where the sponsor speaks first, and every subsequent score anchors on theirs. Independent written scoring before discussion removes most of this.
- False precision: totals reported to two decimal places imply an accuracy the 1–10 gut scores underneath never had. A 6.75 versus 7.00 result is a tie that deserves a conversation, not a verdict.
- Knock-out requirements are smuggled into the weighting. If a criterion is genuinely mandatory (regulatory compliance, data residency), it belongs in a pass/fail gate before scoring, not blended into a percentage.
What the total still cannot tell you
Even an honestly run model only ranks the options against the criteria the team could think of. It cannot see the criterion nobody proposed: the integrator's A-team rotating off after month three, the vendor's roadmap deprioritising your industry, the true cost of the second integration everyone assumes will be easy. Those are learned from people who have lived with each vendor after the contract was signed, which is not information a scoring workshop generates.
Independent challenge before the scores are final
The highest-value intervention is not re-scoring; it is challenging the criteria and weights before scoring starts, and pressure-testing the near-tie at the end. Selected senior operators who have implemented the shortlisted platforms can tell you which criteria turned out to matter in year two, which is routinely different from what mattered in the selection workshop.