Interquartile range - Misplaced Pages

(Redirected from Inter-quartile range) Measure of statistical dispersion "IQR" redirects here. For other uses, see IQR (disambiguation).

In descriptive statistics, the interquartile range (IQR) is a measure of statistical dispersion, which is the spread of the data. The IQR may also be called the midspread, middle 50%, fourth spread, or H‑spread. It is defined as the difference between the 75th and 25th percentiles of the data. To calculate the IQR, the data set is divided into quartiles, or four rank-ordered even parts via linear interpolation. These quartiles are denoted by Q₁ (also called the lower quartile), Q₂ (the median), and Q₃ (also called the upper quartile). The lower quartile corresponds with the 25th percentile and the upper quartile corresponds with the 75th percentile, so IQR = Q₃ − Q₁_.

The IQR is an example of a trimmed estimator, defined as the 25% trimmed range, which enhances the accuracy of dataset statistics by dropping lower contribution, outlying points. It is also used as a robust measure of scale It can be clearly visualized by the box on a box plot.

Use

Unlike total range, the interquartile range has a breakdown point of 25% and is thus often preferred to the total range.

The IQR is used to build box plots, simple graphical representations of a probability distribution.

The IQR is used in businesses as a marker for their income rates.

For a symmetric distribution (where the median equals the midhinge, the average of the first and third quartiles), half the IQR equals the median absolute deviation (MAD).

The median is the corresponding measure of central tendency.

The IQR can be used to identify outliers (see below). The IQR also may indicate the skewness of the dataset.

The quartile deviation or semi-interquartile range is defined as half the IQR.

Algorithm

The IQR of a set of values is calculated as the difference between the upper and lower quartiles, Q₃ and Q₁. Each quartile is a median calculated as follows.

Given an even 2n or odd 2n+1 number of values

first quartile Q₁ = median of the n smallest values

third quartile Q₃ = median of the n largest values

The second quartile Q₂ is the same as the ordinary median.

Examples

Data set in a table

The following table has 13 rows, and follows the rules for the odd number of entries.

i	x	Median	Quartile
1	7	Q₂=87 (median of whole table)	Q₁=31 (median of lower half, from row 1 to 6)
2	7
3	31
4	31
5	47
6	75
7	87
8	115		Q₃=119 (median of upper half, from row 8 to 13)
9	116
10	119
11	119
12	155
13	177

For the data in this table the interquartile range is IQR = Q₃ − Q₁ = 119 - 31 = 88.

Data set in a plain-text box plot

                             +−−−−−+−+
               * |−−−−−−−−−−−|     | |−−−−−−−−−−−|
                             +−−−−−+−+
 +−−−+−−−+−−−+−−−+−−−+−−−+−−−+−−−+−−−+−−−+−−−+−−−+   Number line
 0   1   2   3   4   5   6   7   8   9   10  11  12

For the data set in this box plot:

Lower (first) quartile Q₁ = 7
Median (second quartile) Q₂ = 8.5
Upper (third) quartile Q₃ = 9
Interquartile range, IQR = Q₃ - Q₁ = 2
Lower 1.5*IQR whisker = Q₁ - 1.5 * IQR = 7 - 3 = 4. (If there is no data point at 4, then the lowest point greater than 4.)
Upper 1.5*IQR whisker = Q₃ + 1.5 * IQR = 9 + 3 = 12. (If there is no data point at 12, then the highest point less than 12.)
Pattern of latter two bullet points: If there are no data points at the true quartiles, use data points slightly "inland" (closer to the median) from the actual quartiles.

This means the 1.5*IQR whiskers can be uneven in lengths. The median, minimum, maximum, and the first and third quartile constitute the Five-number summary.

Distributions

The interquartile range of a continuous distribution can be calculated by integrating the probability density function (which yields the cumulative distribution function—any other means of calculating the CDF will also work). The lower quartile, Q₁, is a number such that integral of the PDF from -∞ to Q₁ equals 0.25, while the upper quartile, Q₃, is such a number that the integral from -∞ to Q₃ equals 0.75; in terms of the CDF, the quartiles can be defined as follows:

Q_{1}={\text{CDF}}^{-1}(0.25),

Q_{3}={\text{CDF}}^{-1}(0.75),

where CDF is the quantile function.

The interquartile range and median of some common distributions are shown below

Distribution	Median	IQR
Normal	μ	2 Φ(0.75)σ ≈ 1.349σ ≈ (27/20)σ
Laplace	μ	2b ln(2) ≈ 1.386b
Cauchy	μ	2γ

Interquartile range test for normality of distribution

The IQR, mean, and standard deviation of a population P can be used in a simple test of whether or not P is normally distributed, or Gaussian. If P is normally distributed, then the standard score of the first quartile, z₁, is −0.67, and the standard score of the third quartile, z₃, is +0.67. Given mean = ${\bar {P}}$ and standard deviation = σ for P, if P is normally distributed, the first quartile

Q_{1}=(\sigma \,z_{1})+{\bar {P}}

and the third quartile

Q_{3}=(\sigma \,z_{3})+{\bar {P}}

If the actual values of the first or third quartiles differ substantially from the calculated values, P is not normally distributed. However, a normal distribution can be trivially perturbed to maintain its Q1 and Q2 std. scores at 0.67 and −0.67 and not be normally distributed (so the above test would produce a false positive). A better test of normality, such as Q–Q plot would be indicated here.

Outliers

The interquartile range is often used to find outliers in data. Outliers here are defined as observations that fall below Q1 − 1.5 IQR or above Q3 + 1.5 IQR. In a boxplot, the highest and lowest occurring value within this limit are indicated by whiskers of the box (frequently with an additional bar at the end of the whisker) and any outliers as individual points.

References

^ Dekking, Frederik Michel; Kraaikamp, Cornelis; Lopuhaä, Hen Paul; Meester, Ludolf Erwin (2005). A Modern Introduction to Probability and Statistics. Springer Texts in Statistics. London: Springer London. doi:10.1007/1-84628-168-7. ISBN 978-1-85233-896-1.
Upton, Graham; Cook, Ian (1996). Understanding Statistics. Oxford University Press. p. 55. ISBN 0-19-914391-9.
Zwillinger, D., Kokoska, S. (2000) CRC Standard Probability and Statistics Tables and Formulae, CRC Press. ISBN 1-58488-059-7 page 18.
Ross, Sheldon (2010). Introductory Statistics. Burlington, MA: Elsevier. pp. 103–104. ISBN 978-0-12-374388-6.
^ Kaltenbach, Hans-Michael (2012). A concise guide to statistics. Heidelberg: Springer. ISBN 978-3-642-23502-3. OCLC 763157853.
Rousseeuw, Peter J.; Croux, Christophe (1992). Y. Dodge (ed.). "Explicit Scale Estimators with High Breakdown Point" (PDF). L1-Statistical Analysis and Related Methods. Amsterdam: North-Holland. pp. 77–92.
Yule, G. Udny (1911). An Introduction to the Theory of Statistics. Charles Griffin and Company. pp. 147–148.
^ Bertil., Westergren (1988). Beta mathematics handbook : concepts, theorems, methods, algorithms, formulas, graphs, tables. Studentlitteratur. p. 348. ISBN 9144250517. OCLC 18454776.
Dekking, Kraaikamp, Lopuhaä & Meester, pp. 235–237

External links

Media related to Interquartile range at Wikimedia Commons

Statistics

Descriptive statistics

Continuous data

Center	Mean Arithmetic Arithmetic-Geometric Contraharmonic Cubic Generalized/power Geometric Harmonic Heronian Heinz Lehmer Median Mode
Dispersion	Average absolute deviation Coefficient of variation Interquartile range Percentile Range Standard deviation Variance
Shape	Central limit theorem Moments Kurtosis L-moments Skewness

Count data

Index of dispersion

Summary tables

Dependence

Graphics

Data collection

Study design	Effect size Missing data Optimal design Population Replication Sample size determination Statistic Statistical power
Survey methodology	Sampling Cluster Stratified Opinion poll Questionnaire Standard error
Controlled experiments	Blocking Factorial experiment Interaction Random assignment Randomized controlled trial Randomized experiment Scientific control
Adaptive designs	Adaptive clinical trial Stochastic approximation Up-and-down designs
Observational studies	Cohort study Cross-sectional study Natural experiment Quasi-experiment

Statistical inference

Statistical theory

Frequentist inference

Point estimation	Estimating equations Maximum likelihood Method of moments M-estimator Minimum distance Unbiased estimators Mean-unbiased minimum-variance Rao–Blackwellization Lehmann–Scheffé theorem Median unbiased Plug-in
Interval estimation	Confidence interval Pivot Likelihood interval Prediction interval Tolerance interval Resampling Bootstrap Jackknife
Testing hypotheses	1- & 2-tails Power Uniformly most powerful test Permutation test Randomization test Multiple comparisons
Parametric tests	Likelihood-ratio Score/Lagrange multiplier Wald

Specific tests

Z-test (normal) Student's t-test F-test
Goodness of fit	Chi-squared G-test Kolmogorov–Smirnov Anderson–Darling Lilliefors Jarque–Bera Normality (Shapiro–Wilk) Likelihood-ratio test Model selection Cross validation AIC BIC
Rank statistics	Sign Sample median Signed rank (Wilcoxon) Hodges–Lehmann estimator Rank sum (Mann–Whitney) Nonparametric anova 1-way (Kruskal–Wallis) 2-way (Friedman) Ordered alternative (Jonckheere–Terpstra) Van der Waerden test

Bayesian inference

Correlation	Pearson product-moment Partial correlation Confounding variable Coefficient of determination
Regression analysis	Errors and residuals Regression validation Mixed effects models Simultaneous equations models Multivariate adaptive regression splines (MARS)
Linear regression	Simple linear regression Ordinary least squares General linear model Bayesian regression
Non-standard predictors	Nonlinear regression Nonparametric Semiparametric Isotonic Robust Homoscedasticity and Heteroscedasticity
Generalized linear model	Exponential families Logistic (Bernoulli) / Binomial / Poisson regressions
Partition of variance	Analysis of variance (ANOVA, anova) Analysis of covariance Multivariate ANOVA Degrees of freedom

Categorical / Multivariate / Time-series / Survival analysis

Categorical

Multivariate

Time-series

General	Decomposition Trend Stationarity Seasonal adjustment Exponential smoothing Cointegration Structural break Granger causality
Specific tests	Dickey–Fuller Johansen Q-statistic (Ljung–Box) Durbin–Watson Breusch–Godfrey
Time domain	Autocorrelation (ACF) partial (PACF) Cross-correlation (XCF) ARMA model ARIMA model (Box–Jenkins) Autoregressive conditional heteroskedasticity (ARCH) Vector autoregression (VAR)
Frequency domain	Spectral density estimation Fourier analysis Least-squares spectral analysis Wavelet Whittle likelihood

Survival

Survival function	Kaplan–Meier estimator (product limit) Proportional hazards models Accelerated failure time (AFT) model First hitting time
Hazard function	Nelson–Aalen estimator
Test	Log-rank test

Applications

Biostatistics	Bioinformatics Clinical trials / studies Epidemiology Medical statistics
Engineering statistics	Chemometrics Methods engineering Probabilistic design Process / quality control Reliability System identification
Social statistics	Actuarial science Census Crime statistics Demography Econometrics Jurimetrics National accounts Official statistics Population statistics Psychometrics
Spatial statistics	Cartography Environmental statistics Geographic information system Geostatistics Kriging

Category:

Scale statistics

Use