Field note 25

Combine percentages from counts, not from averages

Six unequal-width red and charcoal print impressions form a staggered stack on an off-white paper surface.
Combine counts, then calculate percentages

A report shows one branch at 90% completion and another at 40%. Calling the combined result 65% answers a question about the average branch. It does not necessarily answer how many eligible people completed the work.

For an operations analyst preparing a pooled completion percentage, keep the completed and eligible counts. Check that the groups describe the same measure and do not overlap. Then divide the combined completed count by the combined eligible count. If a group lacks the counts or uses a different definition, keep it outside that calculation and say so.

The distinction is mathematical, not an AI judgement. The following MAJLS checklist is a proposed reporting control, not evidence of a deployed integration or a client result.

Know which question the percentage answers

CDC's archived teaching material defines a proportion as a part compared with the whole, with the numerator included in the denominator; it can be expressed as a percentage.[2] We use that elementary definition here for learning completions, not for epidemiological advice.

NIST's 1997 Dataplot reference gives the weighted mean as the sum of each value multiplied by its weight, divided by the sum of the weights.[1] For disjoint groups measuring the same completion proportion, weighting each unrounded group proportion by its eligible count gives the pooled proportion. This follows because eligible count multiplied by completed/eligible recovers the completed count.

An equal-weight mean of branch percentages is still a valid answer to a different question: "What is the average branch percentage?" Label it that way if it is the intended metric. Do not silently substitute it for "What percentage of eligible people completed?"

This note covers simple observed proportions. Survey weights, repeated observations, exposure-time rates and uncertainty estimates require their own methods. Do not generalise this worksheet to them.

MAJLS pooled-proportion checklist

Before combining rows, the report owner should confirm:

Keep the source snapshot identifier, row locator, counts, definition, inclusion decision and human owner. Calculate from counts rather than rounded displayed percentages. Retain the exact fraction until formatting the final output.

An AI preparation role could draft this register from an authorised extract; deterministic checks should perform the count validation and arithmetic. This is a proposed division of work, not a claim that MAJLS has executed it.

Filled example: three branches ready, two unresolved

Every branch, person, source and number below is illustrative. Training owner Dana has defined one completion measure: eligible staff who finished Module M1 by the end of the October 1–7 reporting period. Snapshot TR-11 assigns each person to one branch. Its underlying roster is stipulated disjoint for this example; totals alone would not prove that in a real report.

Row / locator Completed Eligible Definition check Preparation decision
A / TR-11:A 9 10 M1 finished, agreed period Include; 90%
B / TR-11:B 36 90 M1 finished, agreed period Include; 40%
C / TR-11:C 0 0 Same rule, nobody eligible Retain counts; percentage undefined
D / TR-11:D 4 Unknown Denominator missing Hold; Dana requests roster
E / TR-11:E 8 10 Counts M1 opened, not finished Hold; Dana requests completion count

For A and B, the equal-weight mean is (90% + 40%)/2 = 65%. The pooled calculation is (9 + 36)/(10 + 90) = 45/100 = 45%. The equal-weight result is 20 percentage points higher. Branch B has nine times the eligible count, so it must carry nine times the weight in the person-level calculation.

C adds zero to both counts and cannot change that pooled result. Its own percentage is undefined, not 0%. If every included branch had no eligible people, there would be no pooled percentage to report.

The internal draft should read: "In branches A–C, 45 of 100 eligible staff completed M1 in October 1–7 (45%). Branch D lacks an eligible count; branch E supplied openings rather than completions. Those branches are unresolved and excluded."

It should not read "45% of all staff completed training." Dana has neither resolved D and E nor established that eligible staff represent all staff. Do not publish an organisation-wide total while those holds remain.

Test the preparation before using it

For a bounded read-only pilot, record today's baseline: time to prepare the register, arithmetic corrections and wrongly included rows. The baseline is a measurement to collect, not a saving assumed here.

Use the five rows above plus a duplicate-person fixture, a completed-greater-than-eligible row and an all-zero cohort. The proposed targets are exact 45/100 reproduction, every deliberate defect held, no invented counts and no changed source rows. Stop on a missing roster, overlap, definition conflict or failed arithmetic check; do not let a model's confidence clear the hold.

Dana, the named human report owner, must approve consequential changes to cohort definitions and release of the management report. The preparation role cannot approve those actions. A second analyst should reproduce the included counts before Dana accepts the pilot; exit requires every test to match its expected result and both unresolved branches to have explicit owners.

The next useful step is to replace one bare percentage column with completed and eligible counts. That makes the calculation reproducible and leaves the reporting decision with the person responsible for it.

Sources

  1. Weighted mean | NIST Dataplot Reference Manual (1997)
  2. Principles of Epidemiology, Lesson 3: Frequency measures | CDC (archived)

Discuss the process with MAJLS ↗