Download CN 13.1A Two-Sample t Test and Interval for Means

Survey
yes no Was this document useful for you?
   Thank you for your participation!

* Your assessment is very important for improving the work of artificial intelligence, which forms the content of this project

Document related concepts
no text concepts found
Transcript
AP Statistics
Notes
Name: ____________
Date: _____________
Lesson 13.1A: Two-Sample t Test of Hypothesis and
Confidence Interval for the Difference
between Population Means
Learning Targets:
A: Determine whether a problem requires inference about comparing means or proportions.
B: Recognize from the design of a study whether one-sample t, paired t, or two-sample t
procedures are needed.
C: Calculate and interpret a confidence interval for the difference between two means.
D: Test the hypothesis that two populations have equal means against either a one-sided or
two-sided alternative.
E: Recognize when the two-sample t procedures are appropriate in practice.
I.
Two-Sample t Test of Hypothesis for the Difference between
Two Population Means, 1 and  2
In a study of the effect of college student employment on academic performance, the
following summary statistics for GPA were reported for a sample of students who
worked and for a sample of students who did not work (University of Central Florida
Undergraduate Research Journal, Spring 2005):
Sample
Size
Mean
GPA
Standard
Deviation
Students Who
Are Employed
184
3.12
0.485
Students Who
Are Not Employed
114
3.23
0.524
The samples were selected at random from working and nonworking students at the
University of Central Florida. Does this information support the hypothesis that for
students at this university, those who are not employed have a higher mean GPA than
those who are employed? Be PPCCI !!
Parameters: Population 1 with mean 1 and standard deviation  1
Population 2 with mean  2 and standard deviation  2 .
Hypotheses:
Procedure:
One-Sample / Two-Sample
Confidence Interval / Test of Hypothesis
Mean / Proportion
Sample:
Choose a SRS separately from each population, or from each treatment
group, if conducting a randomized comparative experiment.
Sample 1 has size n1 , mean x1 , and standard deviation s1 .
Sample 2 has size n 2 , mean x 2 , and standard deviation s2 .
Sampling Distribution of x1  x2 : Mean:  x1  x2 =  x1   x2
= 1   2
Variance:  2 x1  x2 =  2 x1   2 x2 =
Standard Deviation:  x1  x2 =
 12
n1
 12
n1


 22
n2
 22
n2
Conditions: S Two independent SRSs of sizes n1 and n 2 drawn from two distinct
populations. Independent samples means one sample has no
influence on the other.
I
Independent observations. Check if sampling without replacement.
Each population must be at least 10 times as large as its
corresponding sample  N1  10n1 and N 2  10n2 .
N The sampling distribution of x1  x2 is exactly normal if
both populations are normal.
The sampling distribution of x1  x2 is approximately normal if
both samples are large  n1  30 and n2  30 .
Check that the two distributions have similar shapes and that the data
have no strong outliers.
Calculations:
Test Statistic:
z
( x1  x2 )  ( 1   2 )

2
1
n1
OR
t


2
2
n2
( x1  x2 )  ( 1   2 )
2
1
(  12 and  22 known)
2
2
s
s

n1 n2
(  12 and  22 unknown)
With degrees of freedom equal to the smaller of
n1  1 or n2  1 . (Conservative approach.)
P-Value: P-value = P(t
)
(Sketch required!!)
Interpretation: (Interpret results in the context of the problem.)
II.
Two-Sample t Confidence Interval
Estimate with 95% confidence the difference between the mean GPA for students who
are not employed and the mean GPA for students who are employed (at the University
of Central Florida).
Estimate  t  Standard Error of the Estimate
( x1  x2 )  t *
s12 s22

n1 n2
Interpret this confidence interval in the context of the problem.
III.
Options for Determining Degrees of Freedom
In our work above, we determined the degrees of freedom for a two-sample t procedure
by considering the smaller of n 1 - 1 and n 2 - 1. This is a conservative approach,
meaning that it can give us a higher P-value and a wider confidence interval than are
actually true.
The TI-83/84/89 and Minitab use a very accurate approximation to the t-distribution with
degrees of freedom determined from the following formula:
df =
 s12 s 22 
  
 n1 n2 
2
2
1  s12 
1  s 22 
  
 
n1  1  n1 
n2  1  n2 
2
Let’s use this formula to compute degrees of freedom for the mean GPA and
student employment status from the last lesson.
IV.
Warning Label for using the Two-Sample t -Procedures

The two-sample t-procedures are more robust than the one-sample
t-procedures, particularly when the distributions are not symmetric.

When the samples are the same size and the two populations being compared
have distributions with similar shapes, the two-sample t-procedures are very
accurate, even when the samples are small.

When the two population distributions have different shapes, then larger samples
are needed.
Conditions:

The assumption that the data are SRS’s from two populations is more important
than the assumption that the population distributions are normal.

If the sum of the sample sizes is very small (n 1 + n 2 < 15), use the t procedures
only if the data are close to normal. If the data are clearly not normal or if outliers
are present, do not use t.

If the sum of the sample sizes is not very small (15 < n 1 + n 2 < 30), use the t
procedures except in the presence of outliers or strong skewness.

If the sum of the sample sizes is large (n 1 + n 2 > 30), the t procedures can be
used even in the case of strongly skewed distributions. Outliers should always
be examined!
Related documents