The chi-square ($\chi^2$) test is a family of statistical tests for categorical count data. These tests compare observed frequencies with the frequencies expected under a null hypothesis.
A large difference between observed and expected counts provides evidence that the null hypothesis may not adequately explain the data.
Types of Chi-Square Tests:
Although these tests address different questions, they use the same basic idea:
\chi^2
=
\sum
\frac{(\text{Observed}-\text{Expected})^2}
{\text{Expected}}.
Categorical data consist of observations classified into groups or categories rather than measured on a numerical scale.
Categorical variables may be:
For a single categorical variable, each observation belongs to one category. When several categorical variables are recorded, each observation belongs to one category for each variable.
For example, consider Titanic passengers classified by survival status and ticket class:
Each passenger has one survival status and one ticket-class category.
A contingency table, also called a cross-tabulation or crosstab, summarizes counts for combinations of categorical variables.
For example, Titanic survival data can be represented in a $2\times4$ contingency table:
| First Class | Second Class | Third Class | Crew | Total | |
| Survived | $a$ | $b$ | $c$ | $d$ | $S$ |
| Died | $e$ | $f$ | $g$ | $h$ | $D$ |
| Total | 325 | 285 | 706 | 885 | 2,201 |
Here, $a$ through $h$ represent the observed counts in each cell.
The row and column totals are called marginal totals. They are used to calculate expected counts in tests of homogeneity and independence.
A chi-square goodness-of-fit test compares observed counts for one categorical variable with the counts expected under a specified distribution.
The observed and expected counts do not need to match exactly under $H_0$; some difference is expected because of sampling variation.
Suppose we want to test whether the color distribution of M&Ms is consistent with a distribution reported for 2008.
2008 Expected Color Distribution:
| Color | Percentage (%) |
| Blue | 24 |
| Orange | 20 |
| Green | 16 |
| Yellow | 14 |
| Red | 13 |
| Brown | 13 |
Observed Counts: From a sample of 410 M&Ms:
| Color | Count |
| Blue | 105 |
| Orange | 91 |
| Green | 70 |
| Yellow | 50 |
| Red | 45 |
| Brown | 49 |
The hypotheses are:
H_0:
(p_{\text{blue}},p_{\text{orange}},p_{\text{green}},
p_{\text{yellow}},p_{\text{red}},p_{\text{brown}})
=
(0.24,0.20,0.16,0.14,0.13,0.13)
versus the alternative that at least one population proportion differs.
For each category, the expected count is:
E_i=Np_i
where:
For blue M&Ms:
E_{\text{blue}}
=
410(0.24)
=
98.4.
The complete expected counts are:
| Color | Observed $O_i$ | Expected $E_i$ |
| Blue | 105 | 98.4 |
| Orange | 91 | 82.0 |
| Green | 70 | 65.6 |
| Yellow | 50 | 57.4 |
| Red | 45 | 53.3 |
| Brown | 49 | 53.3 |
The expected counts also sum to 410.
The goodness-of-fit statistic is:
\chi^2
=
\sum_{i=1}^{k}
\frac{(O_i-E_i)^2}{E_i}
where:
For these data:
\chi^2\approx4.32.
Each term measures the difference between an observed count and its expected value. Larger differences contribute more to the total statistic.
When all expected proportions are specified in advance:
df=k-1.
For 6 colors:
df=6-1=5.
Suppose the significance level is:
\alpha=0.05.
Using the chi-square distribution with $df=5$, the critical value is approximately:
\chi^2_{0.95,5}\approx11.07.
Reject $H_0$ if:
\chi^2_{\text{observed}}>11.07.
Equivalently, using a p-value, reject $H_0$ when:
p<\alpha.
For this example:
Because the p-value is greater than 0.05, the sample does not provide sufficient evidence that the color distribution differs from the specified 2008 proportions.
Failing to reject $H_0$ does not prove that the distribution is unchanged. It means that the observed differences are not large enough to provide evidence against the specified distribution at the chosen significance level.
Analysis Results:
Based on this sample, there is insufficient evidence to conclude that the M&M color distribution differs from the specified 2008 distribution.
A chi-square test of homogeneity compares the distribution of a categorical variable across several populations or groups.
Suppose we want to test whether survival rates are the same across Titanic ticket classes.
Data Summary:
| Survived | Died | Total | |
| First Class | 203 | 122 | 325 |
| Second Class | 118 | 167 | 285 |
| Third Class | 178 | 528 | 706 |
| Crew | 212 | 673 | 885 |
| Total | 711 | 1,490 | 2,201 |
The null hypothesis is that the survival distribution is the same across all four groups.
If the survival distribution were the same across groups, the expected count in cell $(i,j)$ would be:
E_{ij}
=
\frac{
(\text{Row Total}_i)
(\text{Column Total}_j)
}{
\text{Grand Total}
}.
For First Class survivors:
E_{11}
=
\frac{325(711)}{2201}
\approx104.99.
The observed number of First Class survivors is 203, which is much larger than this expected count.
The statistic is:
\chi^2
=
\sum_{i=1}^{r}
\sum_{j=1}^{c}
\frac{(O_{ij}-E_{ij})^2}{E_{ij}}.
Here:
Summing across all eight cells gives:
\chi^2\approx190.40.
For a contingency table:
df=(r-1)(c-1).
Therefore:
df
=
(4-1)(2-1)
=
3.
At:
\alpha=0.05,
the critical value for $df=3$ is approximately:
\chi^2_{0.95,3}=7.81.
Because:
190.40>7.81,
we reject the null hypothesis.
The p-value is also extremely small:
p\approx5.0\times10^{-41}.
The data provide very strong evidence that survival rates were not the same across the four Titanic groups.
In other words, survival and ticket-class group are associated in these data. The chi-square test does not explain why the groups differ.
Analysis Results:
The data provide strong evidence that survival rates differed among First Class, Second Class, Third Class, and Crew.
A chi-square test of independence examines whether two categorical variables are associated within a population.
If two variables are independent, knowing the value of one does not change the distribution of the other.
Suppose we survey individuals to investigate whether gender is associated with voting preference.
Data Summary:
| Liberal | Conservative | Total | |
| Male | 40 | 60 | 100 |
| Female | 70 | 30 | 100 |
| Total | 110 | 90 | 200 |
If the variables were independent, both gender groups would have the same voting-preference distribution apart from sampling variation.
Under independence:
E_{ij}
=
\frac{
(\text{Row Total}_i)
(\text{Column Total}_j)
}{
\text{Grand Total}
}.
For Male/Liberal:
E_{11}
=
\frac{100(110)}{200}
=
55.
The expected counts are therefore:
| Liberal | Conservative | |
| Male | 55 | 45 |
| Female | 55 | 45 |
The Pearson chi-square statistic is:
\chi^2
=
\sum_{i=1}^{2}
\sum_{j=1}^{2}
\frac{(O_{ij}-E_{ij})^2}{E_{ij}}.
For these data:
\chi^2\approx18.18.
df
=
(2-1)(2-1)
=
1.
For a $2\times2$ table, Yates' continuity correction may be applied:
\chi^2_{\text{Yates}}
=
\sum
\frac{(|O_{ij}-E_{ij}|-0.5)^2}{E_{ij}}.
For these data:
\chi^2_{\text{Yates}}
\approx16.99.
with:
p\approx3.76\times10^{-5}.
Without the correction, the Pearson test gives:
\chi^2\approx18.18,
\qquad
p\approx2.01\times10^{-5}.
Both lead to the same conclusion.
At:
\alpha=0.05,
the chi-square critical value with one degree of freedom is approximately:
3.84.
Using either statistic:
\chi^2>3.84,
so we reject $H_0$.
The data provide strong evidence of an association between gender and voting preference in the sampled population.
This indicates association, not necessarily causation.
Analysis Results:
Using Yates' continuity correction:
There is strong evidence of an association between gender and voting preference in the observed data.
The chi-square tests of homogeneity and independence use the same test statistic and expected-count formula. Their main difference lies in the research question and how the data are collected.
The population structure differs:
The research question also differs:
The calculations are otherwise essentially the same once the contingency table has been constructed.
For chi-square tests to be valid, several conditions should be checked:
The data are counts. The test is applied to observed frequencies in categories.
Categories are mutually exclusive. Each observation contributes to one relevant category or table cell.
Observations are independent. One observation should not determine or duplicate another.
The sampling design should support the intended inference.
Expected counts should not be too small. A common rule of thumb is that expected counts should generally be at least 5.
These conditions help ensure that the chi-square distribution provides a reasonable approximation to the sampling distribution of the test statistic.
\chi^2
=
\sum\frac{(O-E)^2}{E}.
A chi-square test tells us whether the observed differences are larger than would reasonably be expected under the null hypothesis. It does not, by itself, explain why an association exists or establish causation.