Download Lecture 7

Document related concepts
no text concepts found
Transcript
Lecture 7:
Multiple
Regression
(Chapter 6)
Copyright © 2006 Pearson Addison-Wesley. All rights reserved.
Agenda for Today
• Review Regression with a Single
Variable
– From Chapters 3 and 4
• Multiple Regression
– From Chapter 6.1–6.3
– Note: we will defer coverage of the material
on polynomials from Chapter 6.1 and
dummy variables from Chapter 6.3
Copyright © 2006 Pearson Addison-Wesley. All rights reserved.
7-2
Review
• As a starting place, we need to write down all
our assumptions about the way the
underlying process works, and about how that
process led to our data.
• These assumptions are called the “Data
Generating Process.”
• Then we can derive estimators that have
good properties for the Data Generating
Process we have assumed.
Copyright © 2006 Pearson Addison-Wesley. All rights reserved.
7-3
Review: The Gauss–Markov DGP
• From Chapter 3: no intercept
• Y = bX + e
• E(ei ) = 0
• Var(ei ) = s 2
• Cov(ei ,ej ) = 0, for i ≠ j
• X ’s fixed across samples (so we can treat
them like constants).
• We want to estimate b
Copyright © 2006 Pearson Addison-Wesley. All rights reserved.
7-4
Review
• We will focus on linear estimators.
• Linear estimator: a weighted sum of the Y ’s.
ˆ
b
Copyright © 2006 Pearson Addison-Wesley. All rights reserved.
wY
 ii
7-5
Review (cont.)
Yi  b X i  e i
E(e i )  0
Var(e i )  s
2
Cov(e i ,e j )  0, for i  j
X's fixed across samples
(so we can treat it as a constant).
Copyright © 2006 Pearson Addison-Wesley. All rights reserved.
7-6
Review (cont.)
n
bˆ   wY
i i
i 1
n
E ( bˆ )  b  wi X i
i 1
n
A linear estimator is unbiased if
w X
i 1
Copyright © 2006 Pearson Addison-Wesley. All rights reserved.
i
i
 1.
7-7
Review (cont.)
Yi  b X i  e i E(e i )  0
Var(e i )  s 2 Cov(e i ,e j )  0, for i  j
X's fixed across samples
n
A linear estimator is unbiased if  wi X i  1.
i1
Many linear estimators will be unbiased.
The Best Linear Unbiased Estimator (BLUE) is
the unbiased linear estimator with the
smallest cross-sample variance.
Copyright © 2006 Pearson Addison-Wesley. All rights reserved.
7-8
Review: BLUE Estimators
• Ordinary Least Squares (OLS) is BLUE
for our Gauss–Markov DGP.
• This result is called the “Gauss–
Markov Theorem.”
Copyright © 2006 Pearson Addison-Wesley. All rights reserved.
7-9
Review: BLUE Estimators (cont.)
• OLS is a very good strategy for the
Gauss–Markov DGP.
• OLS is unbiased: our guesses are right
on average.
• OLS is efficient: it has the smallest
possible variance (for unbiased linear
estimators).
Copyright © 2006 Pearson Addison-Wesley. All rights reserved.
7-10
Review: BLUE Estimators (cont.)
• Our guesses will tend to be close to
right (or at least as close to right as we
can get).
• WARNING: the minimum variance could
still be pretty large!
Copyright © 2006 Pearson Addison-Wesley. All rights reserved.
7-11
Gauss–Markov with an Intercept
(Chapter 4)
Yi  b0  b1 X i  e i (i  1...n)
E(e i )  0
Var(e i )  s
2
Cov(e i ,e j )  0, i  j
X's fixed across samples.
All we have done is add a b0 .
Copyright © 2006 Pearson Addison-Wesley. All rights reserved.
7-12
BLUE Estimator of b1
( X i  X )(Yi  Y )
ˆ
b1   n
2
(X j  X )
j 1
• OLS for the DGP with an intercept.
• OLS is the Best (minimum variance)
Linear Unbiased Estimator for the
Gauss–Markov DGP with an intercept.
Copyright © 2006 Pearson Addison-Wesley. All rights reserved.
7-13
BLUE Estimator of b0
• The easiest way to estimate the intercept:
bˆ0  Y  bˆ1 X
• Notice that the fitted regression line always
goes through the point.
( X ,Y )
• Our fitted regression line passes through “the
middle of the data.”
Copyright © 2006 Pearson Addison-Wesley. All rights reserved.
7-14
Review
• We can use residuals to calculate s2, the
estimated variance of the error terms from
our DGP.
• We can use s2 to calculate the estimated
standard error of our estimator.
• The e.s.e. gives us a sense of how precise
our estimate is. We can use Confidence
Intervals to gauge precision.
Copyright © 2006 Pearson Addison-Wesley. All rights reserved.
7-15
Multiple Regression (Chapter 6)
• We know how to regress Y on a
constant and a single X variable
Y  b0  b1 ·X  e
• b1 is the change in Y from a 1-unit
change in X
Copyright © 2006 Pearson Addison-Wesley. All rights reserved.
7-16
Multiple Regression (cont.)
• Usually we will want to include more
than one independent variable.
• How can we extend our procedures to
permit multiple X variables?
Copyright © 2006 Pearson Addison-Wesley. All rights reserved.
7-17
Example: Growth
• Barro on growth:
– Examine cross-country growth from
1960–1985
– Consider numerous determinants
– (see On-line Extension 1: A Medley of Hits)
Copyright © 2006 Pearson Addison-Wesley. All rights reserved.
7-18
Example: Growth (cont.)
• Wealthy nations enjoy a real output per
worker thirty times greater than that in
poor nations.
• Will poorer countries eventually “catch up,”
or is the difference in productivity widening
over time?
• A theory of “catch up” suggests that lower
GDP per capita in 1960 predicts higher
growth rates between 1960 and 1985.
Copyright © 2006 Pearson Addison-Wesley. All rights reserved.
7-19
Example: Growth (cont.)
• Barro (1991) showed that in a univariate
relationship, lower GDP per capita in
1960 predicts lower growth rates
between 1960 and 1985.
• Looking just at GDP per capita in 1960
and growth rates, we see evidence of
divergence, NOT convergence.
Copyright © 2006 Pearson Addison-Wesley. All rights reserved.
7-20
Hit Figure Ext.1.1 Per Capita GDP Growth
Rates (1960–1985) vs. 1960 Per Capita GDP.
Copyright © 2006 Pearson Addison-Wesley. All rights reserved.
7-21
Example: Growth (cont.)
• Perhaps there is some tendency towards
“catch up” that is overshadowed by
other factors:
–
–
–
–
Human capital
Government consumption
Political instability
Price distortions
• Can we include these variables in
the regression?
Copyright © 2006 Pearson Addison-Wesley. All rights reserved.
7-22
Gauss–Markov DGP with Multiple X ’s
Y  b0  b1 X 1i  b2 X 2i  bk X ki  e i
E(e i )  0
Var(e i )  s
2
Cov(e i , e j )  0, for i  j
X 1  X k fixed across samples (so we can
treat them like constants).
Copyright © 2006 Pearson Addison-Wesley. All rights reserved.
7-23
Expectation of a Linear Estimator
• What is the expectation of a
linear estimator?
Copyright © 2006 Pearson Addison-Wesley. All rights reserved.
7-24
Expectation of Linear Estimators (cont.)
Yi  b 0  b1 X 1i  b 2 X 2i  ..  b k X ki  e i
E (e i )  0, Var (e i )  s 2 , Cov(e i , e j )  0 for i  j
X 's fixed across samples (so we can treat them as constants).
n
n
i 1
i 1
E ( bˆ )  E ( wY
i i )   wi E (Yi )
n
  wi E ( b 0  b1 X 1i  b 2 X 2i  ..  b k X ki  e i )
i 1
n
  wi [ E ( b 0 )  E ( b1 X 1i )  ..  E ( b k X ki )  E (e i )]
i 1
 b o wi  b1wi X 1i  ..  b k wi X ki
Copyright © 2006 Pearson Addison-Wesley. All rights reserved.
7-25
Expectation of Linear Estimators (cont.)
Yi  b 0  b1 X 1i  b 2 X 2i  ..  b k X ki  e i
E (e i )  0, Var (e i )  s 2 , Cov(e i , e j )  0 for i  j
X 's fixed across samples (so we can treat them as constants).
n
bˆ   wY
i i
i 1
E ( bˆ )  b o wi  b1wi X 1i  ..  b k wi X ki
What are the conditions for an unbiased estimator
of b 0? Of b1?
Copyright © 2006 Pearson Addison-Wesley. All rights reserved.
7-26
Expectation of Linear Estimators (cont.)
E ( bˆ )  b o wi  b1wi X 1i  ..  b k wi X ki
What are the conditions for an unbiased estimator
of b 0?
wi  1
wi X 1i  0
wi X 2i  0,..., wi X ki  0
Copyright © 2006 Pearson Addison-Wesley. All rights reserved.
7-27
Expectation of Linear Estimators (cont.)
E ( bˆ )  b o wi  b1wi X 1i  ..  b k wi X ki
What are the conditions for an unbiased estimator
of b1 ?
wi  0
wi X 1i  1
wi X 2i  0,..., wi X ki  0
When we have k X 's, plus a constant,
we need k  1 unbiasedness conditions.
Copyright © 2006 Pearson Addison-Wesley. All rights reserved.
7-28
BLUE Estimators
• Ordinary Least Squares is still BLUE
• The OLS formula for multiple X ’s
requires matrix algebra, but is very
similar to the formula for a single X
Copyright © 2006 Pearson Addison-Wesley. All rights reserved.
7-29
BLUE Estimators (cont.)
• Intuitions from the single variable
formulas tend to generalize to
multiple variables.
• We’ll trust the computer to get the
formulas right.
• Let’s focus on interpretation.
Copyright © 2006 Pearson Addison-Wesley. All rights reserved.
7-30
Single Variable Regression
Copyright © 2006 Pearson Addison-Wesley. All rights reserved.
7-31
Multiple Regression
Y  b0  b1 X1  e
• b1 is the change in Y from a 1-unit
change in X1
Copyright © 2006 Pearson Addison-Wesley. All rights reserved.
7-32
Multiple Regression (cont.)
Y  b0  b1 X1  b2 X2  bk X k  e
• How can we interpret b1 now?
• b1 is the change in Y from a 1-unit
change in X1 , holding X2…Xk FIXED
Copyright © 2006 Pearson Addison-Wesley. All rights reserved.
7-33
Multiple Regression (cont.)
Copyright © 2006 Pearson Addison-Wesley. All rights reserved.
7-34
Multiple Regression (cont.)
• How do we implement multiple
regression with our software?
Copyright © 2006 Pearson Addison-Wesley. All rights reserved.
7-35
Example: Growth
• Regress GDP growth from 1960–1985 on
– GDP per capita in 1960 (GDP60)
– Primary school enrollment in 1960 (PRIM60)
– Secondary school enrollment in 1960 (SEC60)
– Government spending as a share of GDP (G/Y)
– Number of coups per year (REV)
– Number of assassinations per year (ASSASSIN)
– Measure of Investment Price Distortions
(PPI60DEV)
Copyright © 2006 Pearson Addison-Wesley. All rights reserved.
7-36
Hit Table Ext.1.1 A Multiple Regression
Model of per Capita GDP Growth.
Copyright © 2006 Pearson Addison-Wesley. All rights reserved.
7-37
Example: Growth (cont.)
GDP Growth  3.02 – 0.008·GDP60
 0.025·PRIM 60  0.031·SEC 60
- 0.119·G / Y –1.950·REV
- 3.330·ASSASSIN – 0.014·PPI 60 DEV
• A 1-unit increase in GDP in 1960 predicts a
0.008 unit decrease in GDP growth, holding
fixed the level of PRIM60, SEC60, G/Y, REV,
ASSASSIN, and PPI60DEV.
Copyright © 2006 Pearson Addison-Wesley. All rights reserved.
7-38
Example: Growth (cont.)
• Before we controlled for other variables,
we found a POSITIVE relationship
between growth and GDP per capita
in 1960.
• After controlling for measures of human
capital and political stability, the
relationship is negative, in accordance
with “catch up” theory.
Copyright © 2006 Pearson Addison-Wesley. All rights reserved.
7-39
Example: Growth (cont.)
• Countries with high values of GDP per capita
in 1960 ALSO had high values of schooling
and a low number of coups/assassinations.
• Part of the relationship between growth and
GDP per capita is actually reflecting the
influence of schooling and political stability.
• Holding those other variables constant lets us
isolate the effect of just GDP per capita.
Copyright © 2006 Pearson Addison-Wesley. All rights reserved.
7-40
Example: Growth
• The Growth of GDP from 1960–1985
was higher:
1. The lower starting GDP, and
2. The higher the initial level of human capital.
• Poor countries tended to “catch up” to richer
countries as long as the poor country began
with a comparable level of human capital, but
not otherwise.
Copyright © 2006 Pearson Addison-Wesley. All rights reserved.
7-41
Example: Growth (cont.)
• Bigger government consumption is
correlated with lower growth; bigger
government investment is only weakly
correlated with growth.
• Politically unstable countries tended to have
weaker growth.
• Price distortions are negatively related
to growth.
Copyright © 2006 Pearson Addison-Wesley. All rights reserved.
7-42
Example: Growth (cont.)
• The analysis leaves largely unexplained
the very slow growth of Sub-Saharan
African countries and Latin American
countries.
Copyright © 2006 Pearson Addison-Wesley. All rights reserved.
7-43
Example: Earnings Equations
(Chapter 6.3)
• We want to know the benefit in increased
earnings for a black woman from getting one
more year of education.
• We are concerned that years of education
and years of work experience may be related.
• We want to control for years of work
experience. What is the benefit of an extra
year of education, holding constant the years
of work experience?
Copyright © 2006 Pearson Addison-Wesley. All rights reserved.
7-44
Example: Earnings Equations (cont.)
ln(earnings) 
b0  b1 ·Education  b2 ·Experience  e
• Using data for black women from the National
Longitudinal Survey of Youth, we estimate
ln(earnings) 
6.77  0.15·Education  0.05·Experience
(0.182)
(0.012)
(0.010)
where we place standard errors for each
coefficient in parentheses underneath.
Copyright © 2006 Pearson Addison-Wesley. All rights reserved.
7-45
Example: Earnings Equations (cont.)
ln(earnings) 
6.77  0.15·Education  0.05·Experience
(0.182)
(0.012)
(0.010)
• We predict that, holding years of experience
fixed, a 1-year increase in education leads to
a 15% increase in earnings.
• Remember, when we log Y but not X, we
multiply b by 100 to translate into %.
Copyright © 2006 Pearson Addison-Wesley. All rights reserved.
7-46
TABLE 6.1 OLS Estimation of Black
Women’s Wage Equation
Copyright © 2006 Pearson Addison-Wesley. All rights reserved.
7-47
Example: Earnings Equations
• Our sample size of 1,855 is large
enough to approximate the t distribution
with the standard normal distribution.
We can use 1.96 as our multiplier when
computing a 95% Confidence Interval.
• The 95% C.I. for the coefficient on
education is…
Copyright © 2006 Pearson Addison-Wesley. All rights reserved.
7-48
Example: Earnings Equations (cont.)
95%
C.I .
 bˆeducation  1.96 e.s.e.( bˆeducation )
 0.150  1.96 0.012
 (0.126, 0.174)
Copyright © 2006 Pearson Addison-Wesley. All rights reserved.
7-49
Example: Earnings Equations (cont.)
• We predict that, holding years of experience
fixed, a 1-year increase in education leads to
a 15% increase in earnings.
• The effect could plausibly range from a 13%
increase in earnings to a 17% increase.
• For most purposes, this estimate is probably
reasonably precise.
Copyright © 2006 Pearson Addison-Wesley. All rights reserved.
7-50
Example: Earnings Equations (cont.)
• We predict that, holding years of experience
fixed, a 1-year increase in education leads to
a 15% increase in earnings.
• That is, a 1-year increase in education leads
to a 15% increase in earnings for a black
woman with 0 years of experience, the
same as for a black woman with 15 years
of experience.
Copyright © 2006 Pearson Addison-Wesley. All rights reserved.
7-51
Example: Earnings Equations (cont.)
• Do we really think that the benefit of more
education is the same, regardless of how
much experience a worker has?
• Do we really think that the benefit of 10th
grade is the same as the benefit of finishing
your senior year of college?
• We looked only at black women. Can we
directly compare the benefits of education for
workers of different races or genders?
Copyright © 2006 Pearson Addison-Wesley. All rights reserved.
7-52
Example: Earnings Equations (cont.)
• Later on, we will learn how to make our
DGP more flexible by adding various
other nonlinearities. We have already
learned one trick for introducing
nonlinearities: taking logs.
Copyright © 2006 Pearson Addison-Wesley. All rights reserved.
7-53
Example: Earnings Equations (cont.)
• Given the specification we used, are we
convinced that black women’s earnings
increase when they get more
education? Are we convinced that
earnings increase by at least 10%? By
at least 14%?
Copyright © 2006 Pearson Addison-Wesley. All rights reserved.
7-54
Review
• Ordinary Least Squares is BLUE for the
Gauss–Markov DGP, for both univariate
and multivariate analyses.
Copyright © 2006 Pearson Addison-Wesley. All rights reserved.
7-55
Gauss–Markov DGP with Multiple X ’s
Y  b0  b1 X 1i  b2 X 2i  bk X ki  e i
E(e i )  0
Var(e i )  s
2
Cov(e i , e j )  0, for i  j
X 1  X k fixed across samples (so we can
treat them like constants).
Copyright © 2006 Pearson Addison-Wesley. All rights reserved.
7-56
Review
E ( bˆ )  b o wi  b1wi X 1i  ..  b k wi X ki
What are the conditions for an unbiased estimator
of b 0?
wi  1
wi X 1i  0
wi X 2i  0,..., wi X ki  0
Copyright © 2006 Pearson Addison-Wesley. All rights reserved.
7-57
Review (cont.)
E ( bˆ )  b o wi  b1wi X 1i  ..  b k wi X ki
What are the conditions for an unbiased estimator
of b1 ?
wi  0
wi X 1i  1
wi X 2i  0,..., wi X ki  0
When we have k X 's, plus a constant,
we need k  1 unbiasedness conditions.
Copyright © 2006 Pearson Addison-Wesley. All rights reserved.
7-58
Review: Multiple Regression
Y  b0  b1 X1  e
• b1 is the change in Y from a 1-unit
change in X1
Copyright © 2006 Pearson Addison-Wesley. All rights reserved.
7-59
Review: Multiple Regression (cont.)
Y  b0  b1 X1  b2 X2  bk X k  e
• How can we interpret b1 now?
• b1 is the change in Y from a 1-unit
change in X1, holding X2…Xk FIXED
Copyright © 2006 Pearson Addison-Wesley. All rights reserved.
7-60
Multiple Regression
Copyright © 2006 Pearson Addison-Wesley. All rights reserved.
7-61