What are the measures of dispersion? Discover how to calculate and interpret data variability in biostatics using simple formula and examples
Measures of Dispersion
- Dispersion refers to the extent to which a dataset is spread out, measuring the variability of data points.
- High dispersion indicates that the data points are widely spread.
- Low dispersion shows that the data points cluster closely around a central value.
Key Measures of Dispersion:
-
Range
- The difference between the maximum and minimum values.
- Simple to calculate but sensitive to outliers.
- Formula:
-
- Range=Maximum value−Minimum value
- 𝑹𝒂𝒏𝒈𝒆=𝑴𝒂𝒙𝒊𝒎𝒖𝒎 𝒗𝒂𝒍𝒖𝒆−𝑴𝒊𝒏𝒊𝒎𝒖𝒎 𝒗𝒂𝒍𝒖𝒆
-
-
Interquartile Range (IQR)
- The difference between the 75th percentile (Q3) and the 25th percentile (Q1).
- Represents the spread of the middle 50% of data.
- Less sensitive to outliers.
-
Variance
- The average of the squared differences from the mean.
- Gives an idea of how much data points vary from the mean, but in squared units.
- Less intuitive due to units being squared.
-
Standard Deviation
- The square root of the variance.
- Returns dispersion to the original units of measurement.
- Widely used and more interpretable than variance.
-
Mean Absolute Deviation (MAD)
- The average of the absolute differences between each data point and the mean.
- A direct and intuitive measure of spread.
-
Coefficient of Variation (CV)
- A standardized measure of dispersion.
- Formula:
- CV=Standard DeviationMean
- Useful for comparing datasets with different units or scales.
Advertisements
Example: Standard Deviation Calculation
- A pharmaceutical company produces tablets intended to contain 100 mg of an active pharmaceutical ingredient (API).
- A sample shows the following API concentrations (in mg):
Data:
- 98, 102, 99, 101, 100, 98, 103, 97, 102, 100
Advertisements
Steps:
- Mean concentration = 100 mg
- Differences from the mean:
- −2,2,−1,1,0,−2,3,−3,2,0
- Squared differences:
- 4,4,1,1,0,4,9,9,4
- Variance:
Variance:
(4+4+1+1+0+4+9+9+4+0)10=3610=3.6 mg24+4+1+1+0+4+9+9+4+010=3610=3.6 mg2
Interpretation:
- A standard deviation of 1.9 mg shows the variability in API concentration.
- A lower standard deviation indicates better consistency, crucial for drug safety and efficacy.
Range
- The range is the simplest measure of dispersion.
- It is calculated as the difference between the maximum and minimum values.
- Provides a quick sense of spread but is sensitive to outliers.
Advertisements
-
Range in a Discrete Series
Discrete data consist of separate, distinct values.
Example:
- Number of prescriptions filled per day:
Advertisements
| Day | Prescriptions Filled |
| Monday | 45 |
| Tuesday | 50 |
| Wednesday | 40 |
| Thursday | 55 |
| Friday | 48 |
| Saturday | 60 |
| Sunday | 43 |
- Calculation:
- Maximum value = 60 (Saturday)
- Minimum value = 40 (Wednesday)
- Range = 60 − 40 = 20
- Interpretation:
There’s a difference of 20 prescriptions between the busiest and slowest day.
-
Range in a Continuous Series
Advertisements
Continuous data can take any value within a range and are often grouped into intervals.
Example:
- Reduction in blood pressure (mmHg):
| Blood Pressure Reduction (mmHg) | Number of Patients |
| 10–19 | 5 |
| 20–29 | 12 |
| 30–39 | 8 |
| 40–49 | 5 |
| 50–59 | 2 |
- Calculation:
- Minimum interval = 10–19 → use 10 as minimum boundary
- Maximum interval = 50–59 → use 59 as maximum boundary
- Range = 59 − 10 = 49 mmHg
- Interpretation:
Advertisements
The approximate spread in blood pressure reduction is 49 mmHg.
