Statistical framework (parameter vs statistic)
Definitions (bold words from slides)
Quantitative vs Categorical variables
What is a distribution?
Data visualizations
Numerical summaries
Tables and Odds
Z-scores
Regression
For the exam, be prepared to use the following questions to reason about unfamiliar examples. You may receive graphs, tables, numerical summaries, or regression output. You will not need to read or write R code. Support your conclusions with relevant evidence and interpret quantities in context.
How do the types of observations, variable types, and
question of interest determine an appropriate statistical
summary?
Be able to choose and justify a graph, table, or numerical summary.
Explain what its entries, points, or groups represent.
What can we learn about a distribution from its center,
spread, and shape—and what might a summary hide?
Consider skewness, outliers, multiple groups, percentiles, and the
choice between mean/standard deviation and median/IQR. Consider the Datasaurus
Dozen and what that says about the need for both plots and numerical
summaries
When a question asks for a proportion or probability,
which observations belong in the denominator?
Distinguish marginal, joint, and conditional probabilities. Also be able
to distinguish questions like, “What is the probability of being a large
and private school” vs “What is the probability of being a large school
given that is is private”
What would it mean for two variables to be associated,
and what evidence would show that association?
Consider situations with two categorical, two quantitative, and a
quantitative and a categorical. Explain what independence would look
like and what a graph or numerical measure can establish. This gets at
the question: “If I know X, does that tell me anything else about
Y?”
How do probability, odds, and an odds ratio express
different information?
Interpret an odds ratio with its event and comparison groups clearly
specified. Explain what changes when the groups or event are reversed.
Be able report the direction of comparison so that the odds ratio is
reported as \(\theta > 1\)
How does standardization help us compare observations
from different distributions?
Interpret a z-score using its sign, magnitude, and reference group.
Explain why we might want to discuss observations on a standardized
scale rather than the measured one.
What does a regression model predict, and how should its
coefficients be interpreted?
Distinguish quantitative and categorical predictors. Explain the
reference category, the meaning of “holding other predictors fixed,” and
what changes when the reference category changes.
How can we judge whether a regression model provides
useful predictions?
Consider the plot, unexplained variation, \(R^2\), regression toward the mean, and the
range of observed predictor values. Distinguish predicted averages from
individual outcomes.
What conclusions do the data support, and what remains
uncertain?
Distinguish a sample statistic from a population parameter. Explain why
repeated samples can differ and what the Law of Large Numbers does and
does not promise.
Formulas for correlation or standard deviation
R programming