Flow Cytometry Sort Purity and Yield: Calculation and Post-Sort Check

flow cytometry sort purity yield calculation post-sort checkAugust 7, 2026

Four hours on the sorter, 1.55 million cells in the collection tube, and the post-sort check reads 96.2% pure. Your PI asks what the yield was. There are two numbers people call yield in flow cytometry, they differ by a factor that depends on your sort purity, and only one of them tells you whether to change anything next time. This walks the sort purity and yield calculation all the way through with real numbers, then covers what the post-sort check can and cannot resolve.

Purity and yield are two different calculations

Purity is a property of the tube you ended up with: of the cells sitting in the collection tube, what fraction are the ones you wanted. You measure it, you do not predict it. The measurement is a re-analysis — run an aliquot of the sorted cells back through a cytometer and apply the same gate you sorted on.

Yield (used interchangeably with recovery) is a property of the run: of the target cells you put into the sorter, what fraction came out the other end alive and in the tube. It is a ratio of two counts, and purity is one of the terms in it, because a raw tube count includes the contaminants.

Conflating them is the common error. A 99% pure tube holding 400,000 cells out of 2 million input targets is a 20% yield and a failed experiment. A 92% pure tube holding 1.8 million is usually the better outcome.

The formulas

Purity, from the post-sort re-analysis:

$P = \frac{n_{\text{target}}}{n_{\text{total}}}$

where \(n_{\text{target}}\) is the events falling inside the sort gate on re-analysis and \(n_{\text{total}}\) is all events acquired in that re-analysis.

Yield, as the University of Virginia flow core states it, is the number of cells in the collection tube times the percent purity, divided by the number of cells you started with times the percent of the target population:

$Y = \frac{N_{\text{collected}} \times P}{N_{\text{input}} \times f}$
Key Formula $Y = \frac{N_{\text{collected}} \times P}{N_{\text{input}} \times f}$ Numerator: target cells actually recovered. Denominator: target cells you loaded. \(f\) is the frequency of the target population in the pre-sort sample, measured on the same gate you sorted on.

The structure is worth reading once slowly. Both the numerator and the denominator convert a total-cell count into a target-cell count by multiplying by a fraction. Purity does that job on the output side, \(f\) does it on the input side. Skip either multiplication and the answer is wrong in a direction that flatters the sort.

Worked example

Worked Example Input: 2.0 × 107 PBMC. Pre-sort analysis puts the target population at 10.0% of live singlets, so \(f\) = 0.100. After the sort, a hemocytometer count of the collection tube gives 1.55 × 106 cells. The post-sort re-analysis reads 96.2% inside the sort gate.

Target cells loaded:

$N_{\text{input}} \times f = 2.0 \times 10^7 \times 0.100 = 2.0 \times 10^6$

Target cells recovered:

$N_{\text{collected}} \times P = 1.55 \times 10^6 \times 0.962 = 1.491 \times 10^6$

Yield:

$Y = \frac{1.491 \times 10^6}{2.0 \times 10^6} = 0.746 = 74.6\%$

That number lands just under the 75–90% band a well-run core typically reports for a routine sort, which makes it a normal result rather than a problem to chase. Note what happens if you skip the purity term: 1.55 × 106 over 2.0 × 106 reads 77.5%, and you would have reported a yield nearly three points higher than the sort delivered.

Where the missing quarter went

Roughly 510,000 target cells did not make it to the tube. They are lost in three places, and the three have different fixes.

  • Electronic aborts. Two cells arrive in the laser interrogation point too close together to be resolved as separate events. The electronics discard both, so each abort costs more than one cell. Fix by dropping the event rate — the coincidence arithmetic behind this is worked through in UNC’s Practical Issues in High-Speed Cell Sorting.
  • Sort conflicts. A target and a non-target fall inside the same or overlapping sort envelope. In purity mode the sorter discards the droplet rather than risk the contaminant, so the target goes to waste. Fix by dropping the event rate or changing sort mode.
  • Handling loss. Cell death before and after the sort, and adherence to the walls of the collection tube. Fix with buffer composition, tube coating, and getting the cells spun down promptly.

You can tell the first two apart from the third with one comparison. The sorter reports how many events it deflected. If that count is close to your tube count, the losses happened upstream in the fluidics and the sorter already knows about them. If the sorter says it deflected 1.9 × 106 and your tube holds 1.55 × 106, about 18% disappeared after deflection, which is a handling problem and no amount of event-rate tuning will touch it.

Sort mode sets the floor on how much the first two cost. Purity mode rejects target events sitting within roughly half a droplet of a non-target. Single-cell mode is stricter still, roughly halving the predicted efficiency relative to purity mode. That is the price of accurate counts when you are depositing into wells. Yield or recovery mode deflects the droplet regardless of what else is in it and approaches 100% efficiency, paying for it in purity.

The post-sort purity check, and how many events it needs

The common instruction is to re-run about 5,000 sorted events for a purity check. That number is not arbitrary, and it is worth knowing what it buys you, because a purity measured on 500 events cannot support the conclusions people draw from it.

A purity check is a proportion estimated from a sample, so the 95% confidence interval half-width is:

$\pm 1.96 \sqrt{\frac{P(1-P)}{n}}$

At \(P\) = 0.96 and \(n\) = 5,000, that is \(1.96 \sqrt{0.96 \times 0.04 / 5000}\) = 1.96 × 0.00277 = 0.0054, so the interval is 96.0% ± 0.54%. At \(n\) = 500 the same 96% carries ± 1.72%, spanning 94.3% to 97.7%.

The practical consequence: at 500 events you cannot distinguish a 96% sort from a 98% sort, and at 5,000 events you can. If your downstream assay has a contamination threshold — a functional readout that breaks above 5% non-target, say — acquire enough events that the interval sits entirely on one side of it. The same counting-statistics reasoning governs how many events a rare population needs before its frequency means anything.

Where these numbers break

Four failure modes turn an arithmetically correct answer into a wrong one.

The re-analysis gate drifts from the sort gate. Purity is only meaningful against the gate you actually sorted on. Re-drawing it by eye on the post-sort sample produces a number that measures your gate-drawing, not the sort.

Cells die between the sort and the check. Dead cells lose forward scatter and drop out of the scatter gate, which quietly inflates or deflates purity depending on which population is dying. Run the check promptly and put a viability dye such as DAPI or PI in the re-analysis so you can separate a real purity number from a death artifact.

Doublets are counted once and delivered twice. A target bound to a non-target registers as one event on the sorter and lands in the tube as two cells. Your hemocytometer count sees both. The re-analysis catches it only if the FSC-A versus FSC-H singlet gate is applied to the re-analysis too, not just to the sort.

\(f\) is measured on the wrong denominator. If the pre-sort frequency was taken as a percentage of all events and the sort gate hung off live singlets, the two are not the same fraction and the yield is wrong by their ratio. Report \(f\) with its parent gate named.

Planning the next sort backwards

The same formula, rearranged, sizes the input:

$N_{\text{input}} = \frac{N_{\text{needed}}}{f \times Y_{\text{expected}}}$

Cores commonly tell users to plan on 50% yield rather than the 75–90% they usually achieve, which builds in the margin for a sort that runs badly. For 1 × 106 target cells at a 10% frequency:

$N_{\text{input}} = \frac{1 \times 10^6}{0.100 \times 0.50} = 2.0 \times 10^7$

Two × 107 cells is exactly what a core will quote you for that request, and now you can see which assumption you are allowed to relax. If the population is rarer, the arithmetic gets unforgiving fast: at \(f\) = 0.001, the same 106 target cells needs 2 × 109 input, which is usually the point where the answer is enrichment before the sort rather than a longer sort.

Purity has a ceiling that the sort mode cannot lift, and that ceiling is set at panel-design time. If the marker you are sorting on sits in a channel that a bright neighbour spreads into, the two populations never separate cleanly enough for the gate to be tight, and every sort mode inherits that. The fluorophore spectrum viewer shows how much a neighbouring dye emits into your sort channel on a given instrument configuration, which is the check worth running before the cells are already on ice.

Try Cytomaton

AI-assisted flow cytometry analysis that learns your gating style. Free during beta.

Join the beta