Methods in Political Science

Multiple Regression Analysis

Sirus Dehdari

Department of Political Science, Stockholm University

Last week:

  • Is there a relationship between two variables?
  • Example: is there a relationship between indoor temperature and a household’s monthly income?
  • How well does the model fit our data?
  • Statistical inference: are we confident that a relationship actually exists in the population?
  • Confidence interval for the estimated slope coefficient, \(b\)

Third week’s learning objectives

  • Understand what multiple regression analysis is and why it’s used
  • How to interpret the slope coefficient in a multiple regression model
  • Causal inference
  • “Economic” significance
  • Chapter 5.2 in Österman and Folke

Remember: bivariate regression model

  • Last week we talked about bivariate regression models
  • We used one independent variable to explain the variation in the dependent variable
  • Can’t we use more than one independent variable to explain the variation in the dependent variable?
  • In other words: can’t the variation in the dependent variable depend on more than one independent variable?
  • A regression model with more than one independent variable is called a multiple regression model

Including more independent variables

  • Remember the equation for a bivariate regression model:

\[Y_i = a + b_1X_{1i} + e_i\]

  • Note the index 1 on \(b\) and \(X\): they are indicators for the first variable
  • We can include more independent variables

\[Y_i = a + b_1X_{1i} + b_2X_{2i} + e_i\]

  • Now we’ve included a second independent variable, \(X_2\), in the model

Multiple regression model

  • If there is a relationship between \(X_2\) and \(Y\) (i.e. \(b_2\neq 0\)), the multiple regression model will explain even more of the variation in \(Y\)
  • There may be even more variables that help us explain the variation in \(Y\)

\[Y_i = a + b_1X_{1i} + b_2X_{2i} + b_3X_{3i} + \ldots + b_kX_{ki} + \ldots + e_i\]

  • Here we’ve included a total of \(k\) different independent variables
  • But why include more than one variable at all?

Why multiple regression models?

  • The two most common reasons we add more independent variables:
    • Explain more of the variation in the dependent variable \(Y\)
    • Rule out that the relationship between \(Y\) and \(X_1\) is affected by some other variable
  • In research where we’re interested in the outcome (the dependent variable), we often want to explain as much of its variation as possible
  • Example: What determines voters’ propensity to vote?
  • Variables that can explain turnout: income, education level, citizenship, marital status, age, etc.

Explaining the variation in the outcome

  • By explaining the variation in the outcome, we believe we understand this specific phenomenon
  • If there’s still unexplained variation in, say, turnout, we need to develop new or better theories
  • Example: many municipalities want to increase early voting
  • A regression model that explains a large share of the variation in the propensity to vote early can help municipalities promote early voting

Can we explain too much?

  • It’s usually valuable to explain as much as possible of the variation in the dependent variable
  • However, variables shouldn’t be included at random: there should be a reason for including a variable
  • We won’t go into the technical details of why we should avoid adding too many independent variables
  • A rule of thumb: only add variables you believe are related to the dependent variable

Causal relationships

  • So far we’ve only considered the relationship between \(Y\) and \(X\) as nothing more than a correlation
  • But we often want to understand whether there’s a causal relationship between these variables
  • Is it \(X\) that causes the change in \(Y\)?
  • This is the main reason we estimate multiple regression models
  • \(X\) can cause \(Y\), but there may be other variables that affect both \(X\) and \(Y\)

A third factor: confounder

  • \(X\) causes \(Y\), but the confounder (also called a confounding factor) \(c\) causes both \(X\) and \(Y\)

Example: political participation

  • If we don’t include \(c\) in our model, we may over- or underestimate the (causal) relationship between \(X\) and \(Y\)
  • Example: suppose we believe regional out-migration has a positive effect on support for populist parties
  • Theory: in regions where a larger share of the population moves away, support for populist parties is higher
  • But one reason people move is high unemployment, which itself affects support for populist parties
  • If we don’t control for unemployment, we will overestimate the effect of out-migration on populist parties’ vote share

Room temperatures again

  • Another example: think about the scenario where we want to explain households’ indoor temperature
  • Last week we proposed a theory with the following prediction: households with high income are more likely to have higher room temperatures
  • What’s a third factor that potentially affects both income and indoor temperature?
  • Maybe education level?

Upward and downward biases

  • Assume years of education has a positive effect on income but a negative effect on temperature
  • Not controlling for education level gives us a downward bias on the effect income has on households’ indoor temperature

Example of downward bias

  • \(c\) has a positive effect on \(X\) and a negative effect on \(Y\), while \(X\) has a positive effect on \(Y\)

Other examples

  • The combination of the relationship between \(c\) and \(X\), and \(c\) and \(Y\), is what determines the direction of the bias
  • Upward bias: without controlling for \(c\), the estimated effect of \(X\) on \(Y\) will be more positive than the true effect
  • Downward bias: without controlling for \(c\), the estimated effect of \(X\) on \(Y\) will be more negative than the true effect

Interpreting the slope coefficients

  • In the bivariate case: we interpret \(b\) as the expected increase in the dependent variable when the independent variable increases by one unit
  • In the multiple regression case, we add “holding everything else constant”
  • Consider the second column in the table on the previous slide: an increase of 1000 SEK in monthly income is expected to increase room temperature by 0.187 degrees Celsius, holding everything else constant
  • “Holding everything else constant” here means the effect income has on room temperature while holding all other variables constant

Interpreting the size of the coefficient

  • How do we know if the effect we’re estimating is large or small?
  • Last week we distinguished between good fit (high \(R^2\) and/or low Root-MSE) and a strong relationship (large positive/negative \(b\))
  • However, large positive or negative estimates of \(b\) don’t necessarily tell us whether the relationship is “meaningful”
  • Example: suppose we estimate the relationship between local unemployment and support for populist parties
  • In our regression model we estimate \(b =\) 0.2: seems like a small effect, right?

Large/small change in \(Y\) and/or \(X\)

  • Suppose local unemployment varies from 2 to 20%
  • At the same time, the local vote share for the populist party varies between 3 and 6%
  • A one-unit increase in unemployment is therefore a small change relative to the range of the independent variable
  • Its effect, 0.2, is fairly large compared to the total vote share for populist parties
Change in X
Small Large
Change in Y Small Moderate Weak
Large Strong Moderate

Range of \(X\) and \(Y\)

  • So what counts as a large change in \(X\) and \(Y\)?
  • One way is to look at the spread of these variables
  • Consider once again the room-temperature example:
(1) (2)
Intercept 18.363 17.039
(0.804) (1.063)
Income 0.015 0.187
(0.065) (0.098)
Years of education -0.208
(0.111)
Observations 100 100
Adj. \(R^2\) 0.00 0.02
Root MSE 1.99 1.96

How much is one unit?

  • Suppose we have the following table of descriptive statistics:
Variable Mean s.d. Min Max
Temperature 18.71 2.01 14.30 22.80
Income 22.36 5.82 9.78 36.98
Years of education 12.06 5.13 0.00 26
  • The estimated slope coefficient for income is 0.187, and it’s statistically significant at the 90% confidence level (why?)
  • The effect of income on temperature is 0.187, which is small relative to the range of the dependent variable, which is 22.80 \(-\) 14.3 \(=\) 8.5
  • The range of income is 36.98 \(-\) 9.78 \(=\) 27.2, so a one (1) unit increase in income is fairly small compared to its range
  • A small change in \(X\) gives a small change in \(Y\): Moderate effect

Model fit with multiple variables

  • Using \(R^2\) to assess fit is problematic in a multiple regression model
  • We can always increase, or at least not decrease, \(R^2\) by including more variables in our model
  • We prefer as small a model as possible that explains as much as possible
  • Adjusted \(R\)-squared “penalizes” the inclusion of independent variables
  • The interpretation of Root-MSE remains the same

Bad controls

  • Earlier we said we must control for confounders
  • Consider this example instead:
  • We now have a mediator: \(m\)

Bad controls: an example

  • \(X\) has a direct effect on \(Y\) but also an indirect one through its effect on \(m\)
  • If we include \(m\) in the model, we “absorb” part of the (total) effect \(X\) has on \(Y\)
  • Example: suppose there are two separate labor market sectors: a high-skill and a low-skill sector
  • The high-skill sector pays a much higher wage than the low-skill sector
  • Suppose it’s nearly impossible to get hired in the high-skill sector without an advanced degree, and that there’s practically no wage variation within each sector

Friends don’t let friends include bad controls

  • Suppose we want to examine the effect of education level on wages
  • If we control for sector, we will estimate no or a very small effect of education level on wages
  • Remember the interpretation in a multiple regression: the effect of education on wages while holding sector constant
  • Within each sector, almost all workers have the same (or no) degree and earn the same wage
  • Sector is, in this case, a mediator

Controlling: what to do and not do

  • Think of the causal link as a stream
  • If the confounder is upstream, i.e. it was determined before the variable you’re interested in: include it
  • If the confounder is downstream, i.e. it was determined after the variable you’re interested in: don’t include it
  • What do you do if the variable you’re interested in and the confounder affect each other?
  • Adding controls to identify the causal link doesn’t always work: we may need to look for random variation / run experiments instead

Example exam question

  • Understanding regression analysis lets you interpret results in regression tables
  • Let’s go back to the table we saw earlier

Interpreting the results: no controls

Interpreting the results: with control variables

  • Now let’s look at the second column