Topics List

Statistical framework (parameter vs statistic)

Definitions (bold words from slides)

Quantitative vs Categorical variables

What is a distribution?

Data visualizations

Numerical summaries

Tables and Odds

Z-scores

Regression

Questions List

For the exam, be prepared to use the following questions to reason about unfamiliar examples. You may receive graphs, tables, numerical summaries, or regression output. You will not need to read or write R code. Support your conclusions with relevant evidence and interpret quantities in context.

  1. How do the types of observations, variable types, and question of interest determine an appropriate statistical summary?
    Be able to choose and justify a graph, table, or numerical summary. Explain what its entries, points, or groups represent.

  2. What can we learn about a distribution from its center, spread, and shape—and what might a summary hide?
    Consider skewness, outliers, multiple groups, percentiles, and the choice between mean/standard deviation and median/IQR. Consider the Datasaurus Dozen and what that says about the need for both plots and numerical summaries

  3. When a question asks for a proportion or probability, which observations belong in the denominator?
    Distinguish marginal, joint, and conditional probabilities. Also be able to distinguish questions like, “What is the probability of being a large and private school” vs “What is the probability of being a large school given that is is private”

  4. What would it mean for two variables to be associated, and what evidence would show that association?
    Consider situations with two categorical, two quantitative, and a quantitative and a categorical. Explain what independence would look like and what a graph or numerical measure can establish. This gets at the question: “If I know X, does that tell me anything else about Y?”

  5. How do probability, odds, and an odds ratio express different information?
    Interpret an odds ratio with its event and comparison groups clearly specified. Explain what changes when the groups or event are reversed. Be able report the direction of comparison so that the odds ratio is reported as \(\theta > 1\)

  6. How does standardization help us compare observations from different distributions?
    Interpret a z-score using its sign, magnitude, and reference group. Explain why we might want to discuss observations on a standardized scale rather than the measured one.

  7. What does a regression model predict, and how should its coefficients be interpreted?
    Distinguish quantitative and categorical predictors. Explain the reference category, the meaning of “holding other predictors fixed,” and what changes when the reference category changes.

  8. How can we judge whether a regression model provides useful predictions?
    Consider the plot, unexplained variation, \(R^2\), regression toward the mean, and the range of observed predictor values. Distinguish predicted averages from individual outcomes.

  9. What conclusions do the data support, and what remains uncertain?
    Distinguish a sample statistic from a population parameter. Explain why repeated samples can differ and what the Law of Large Numbers does and does not promise.

Do not need to know

Formulas for correlation or standard deviation

R programming