This review summarizes the salient ideas from Chapters 1–2. Use the conceptual questions to connect calculations with interpretation and to practice applying these ideas in unfamiliar settings. The topic list describes material covered; it does not imply that every topic will appear on the 50-minute exam.
Independence and the expected counts under independence.
Marginal distributions and their role in calculating expected counts:
\[ E_{ij}=\frac{n_{i+}n_{+j}}{n}, \]
where \(n_{i+}\) and \(n_{+j}\) are the row and column totals and \(n\) is the total sample size.
Pearson’s statistic:
\[ \chi^2=\sum\frac{(O_{ij}-E_{ij})^2}{E_{ij}}. \]
The likelihood-ratio statistic and its relationship to likelihood-ratio inference:
\[ G^2=2\sum_{i,j}O_{ij}\log\left(\frac{O_{ij}}{E_{ij}}\right), \]
Degrees of freedom for testing independence in an \(r\times c\) table: \((r-1)(c-1)\)
Cell contributions and residuals: identifying where observed data depart from the independence model.
Standardized residuals and their interpretation (but you will not need to compute them or remember the formula)
Why \(\chi^2\) and \(G^2\) generally differ numerically but often lead to similar inferential conclusions.
Conditions for chi-square approximations and the role of Fisher’s exact test when those approximations are questionable.
Fisher’s exact test
What a \(p\)-value does and does not tell us about association, particularly in the context of \(\chi^2\).
Three-way contingency tables
Collapsing over a variable to obtain a marginal table.
Conditioning or stratifying on a variable to obtain conditional tables.
Marginal versus conditional odds ratios:
\[ \theta_{XY} \qquad\text{versus}\qquad \theta_{XY\mid Z=k}. \]
Marginal versus conditional independence
Comparing associations across levels of \(Z\)
Homogeneous versus heterogeneous association
Why a marginal association can differ from conditional associations, even when the conditional associations are homogeneous.
Simpson’s paradox as an extreme example in which marginal and conditional associations have opposite directions.
Reasoning in context: identifying which table and comparison answer the question, and explaining what is lost when only the marginal association is reported.
These eight questions can be applied to many different problems. Practice explaining your reasoning in context, using calculations or statistical output as evidence when appropriate.
What quantity are you estimating or testing?
Given a research question and categorical data, identify the parameter or hypothesis that actually answers the question. Is the relevant object a probability, risk difference, risk ratio, odds ratio, conditional odds ratio, or hypothesis of independence? What population or comparison does it describe?
What is the difference between the magnitude of an association and the evidence for an association?
Explain why an odds ratio and a chi-square test statistic answer different questions. Could a weak association provide strong statistical evidence? Could a strong estimated association provide weak statistical evidence? Explain the role of sample size.
What does independence predict the data should look like?
Starting from the marginal distributions of a contingency table, explain what counts you would expect under independence. What does a large discrepancy between \(O\) and \(E\) tell you? What doesn’t it tell you? Consider the size of the discrepancy relative to expected sampling variability e.g., is \(O - E = 10\) considered large?
Where is the evidence against independence coming from?
Suppose a contingency table has a statistically significant \(X^2\) test. How would you determine which cells or comparisons contribute to the departure from independence? What information do residuals provide that the overall test statistic does not?
How are the inferential procedures we’ve studied related?
Explain the conceptual relationship among Wald, score, and likelihood-ratio inference, including the connection between score and Pearson \(\chi^2\) tests as well as LRT and \(G^2\). Why might different procedures produce somewhat different numerical answers but broadly similar conclusions when the large-sample approximations are adequate? When might those differences become important?
What changes when you change the comparison?
If you reverse the exposure groups, redefine which outcome is the event, or otherwise change the reference category, determine what happens to RD, RR, and OR and whether the substantive evidence for association has changed. Distinguish reversing the group comparison from changing the event; these do not affect every measure in the same way.
What information is lost when we collapse a table?
Given \(X\), \(Y\), and a third variable \(Z\), distinguish the marginal association between \(X\) and \(Y\) from their association conditional on \(Z\). What question does each answer? Why can the marginal and conditional associations differ? Does having similar conditional odds ratios across strata guarantee agreement with the marginal odds ratio?
What conclusion is actually justified by the analysis?
Given statistical output or competing analyses of the same data, state what can be concluded in context. Distinguish evidence against independence from a failure to reject independence, effect magnitude from evidence, and marginal from conditional claims. Explain why a nonsignificant result does not establish independence and why an observed association alone does not establish causation.