ants %>%ggplot(aes(y = Injury, fill = Outcome)) +geom_bar(position =position_fill(reverse =TRUE)) +#reverse equals true puts carried back on the leftscale_y_discrete(limits =c("Heavily injured", "Lightly injured")) +#this line puts light injury as the top row to match our contingency table formattinglabs(x ="Proportion", y ="Injury severity",title ="Carrying outcome by injury severity") +scale_fill_manual(values =c("Carried back"="#7FBB9E", # Neptune Green"Carried away"="#60373D", # Red Mahogany"Not carried"="#133955"# Poseidon ))
Are injury severity and carrying outcome related?
If there were no association, we would expect approximately the same proportions of the three outcomes in both injury groups.
Are the differences in this sample evidence of an association, or could they reflect random variation?
Typical questions with \(r\times c\) contingency tables
Is there an association between the row variable and the column variable?
Here, \(r=2\) injury groups and \(c=3\) carrying outcomes.
How strong is the association?
\(H_0:\) injury severity and carrying outcome are independent.
\(H_A:\) injury severity and carrying outcome are associated.
Tests for association
Fisher’s exact test
Fisher’s exact test
Fisher’s exact test assesses whether two categorical variables are associated.
It does not require large expected cell counts, so it is useful when a contingency table includes small counts, as is the case here.
The test can be computationally expensive for large tables.
How Fisher’s exact test works
We condition on the observed margins, the row and column totals:
50 lightly injured ants and 30 heavily injured ants;
50 carried back, 5 carried away, and 25 not carried.
Consider all possible tables with those same margins. Under \(H_0\), each table has a probability.
The two-sided p-value sums the probabilities of tables whose null probabilities are less than or equal to that of our observed table.
Tables with the same margins
Observed table
Injury
Carried back
Carried away
Not carried
Total
Lightly injured
45
0
5
50
Heavily injured
5
5
20
30
Total
50
5
25
80
A more extreme table with the same margins
Injury
Carried back
Carried away
Not carried
Total
Lightly injured
50
0
0
50
Heavily injured
0
5
25
30
Total
50
5
25
80
Fisher’s exact test for the ant data
ant_tab #view the table
#>
#> Carried back Carried away Not carried
#> Lightly injured 45 0 5
#> Heavily injured 5 5 20
ant_fisher <-fisher.test(ant_tab)ant_fisher
#>
#> Fisher's Exact Test for Count Data
#>
#> data: ant_tab
#> p-value = 2.576e-11
#> alternative hypothesis: two.sided
The p-value is approximately \(2.58\times10^{-11}\).
There is strong evidence of an association between injury severity and carrying outcome.
\(\chi^2\) (chi-squared) test
\(\chi^2\) test
The chi-squared test compares observed cell counts with the counts expected if \(H_0\) were true.
Its p-value uses an approximation that depends on the expected counts. For this course, use the guideline that every expected cell count should be at least 5.
We will calculate the expected counts and test statistic for the ant table, then check that guideline.
Marginal probabilities
Injury
Carried back
Carried away
Not carried
Total
Lightly injured
45
0
5
50
Heavily injured
5
5
20
30
Total
50
5
25
80
Injury severity:
\(P(\text{lightly injured})=50/80\)
\(P(\text{heavily injured})=30/80\)
Carrying outcome:
\(P(\text{back})=50/80\)
\(P(\text{away})=5/80\)
\(P(\text{not carried})=25/80\).
Expected counts under independence
Suppose \(H_0\) is true: injury severity and carrying outcome are independent.
What is the probability that an ant is lightly injured and carried back? Recall that under independence, \(P(A \cap B)=P(A)P(B)\).
The expected counts for carried away are 3.125 and 1.875, both below 5.
Use Fisher’s exact test for the main conclusion. The chi-squared calculations illustrate the method, but the large-sample approximation may be unreliable here.
The \(\chi^2\) distribution
Under \(H_0\), the test statistic is approximately chi-squared when expected counts are sufficiently large.
The degrees of freedom are \((r-1)(c-1)=(2-1)(3-1)=2\).
The distribution is asymmetric and takes nonnegative values.
Larger statistics provide more evidence against \(H_0\), so the p-value is an area in the right tail.
Calculating the \(\chi^2\) statistic in R
observed_chisq_statistic <- ants %>%specify(Injury ~ Outcome) %>%calculate(stat ="Chisq")observed_chisq_statistic
#>
#> Fisher's Exact Test for Count Data
#>
#> data: ant_binary_tab
#> p-value = 3.476e-11
#> alternative hypothesis: true odds ratio is not equal to 1
#> 95 percent confidence interval:
#> 10.29032 212.82087
#> sample estimates:
#> odds ratio
#> 41.32031
The estimated OR (conditional MLE) is 41.3, with a 95% CI of (10.3, 212.8).
The CI excludes 1, providing evidence that the odds of being carried back differ by injury severity.
fisher.test() reports a conditional maximum likelihood estimate, which differs slightly from the sample cross-product estimate of 45.
Practice: pneumococcal serotypes and survival
Practice: the serotype data
A Scottish study examined the association between pneumococcal serotype and survival following infection.
Serotype
Survived
Died
Total
Serotype 10
37
7
44
Serotype 15
60
12
72
Serotype 20
97
9
106
Serotype 31
24
10
34
Total
218
38
256
Identify the two categorical variables and the dimensions of this table.
State the null and alternative hypotheses for a test of association.
Practice: expected counts
Serotype
Survived
Died
Total
Serotype 10
37
7
44
Serotype 15
60
12
72
Serotype 20
97
9
106
Serotype 31
24
10
34
Total
218
38
256
Calculate the expected number of deaths for serotype 31 under \(H_0\).
Calculate the other expected counts. Is the chi-squared approximation appropriate under our course guideline?