Statsvetenskapliga institutionen
Sirus Dehdari — sirus.dehdari@statsvet.su.se
Try to solve each problem on your own before checking the solution.
A nutritionist working with a rural elementary school is evaluating whether the students are, on average, meeting national height benchmarks. According to national data, the average height for children in this age group is 150 cm. To assess the situation, the nutritionist measures the height of a random sample of 100 students at the school. The sample yields an average height of 147.5 cm, with a sample standard deviation of 12 cm.
Based on this information, can we conclude at the 95% confidence level that the average height at the school differs from the national average?
Solution
To determine whether \(\bar x\), the estimate of the population mean, is statistically significantly different from 150, we construct a 95% confidence interval.
We use the following formula:
\[\bar x \pm t \times \hat\sigma_{\bar x} = \bar x \pm t \times \frac{s}{\sqrt{n}}\]
\[\hat\sigma_{\bar x} = \frac{s}{\sqrt{n}} = \frac{12}{\sqrt{100}} = 1.2\]
Using \(t = 1.96\), the confidence interval is:
\[147.5 \pm 1.96 \times 1.2 = 147.5 \pm 2.352 \;\Rightarrow\; CI = [145.15,\ 149.85]\]
Since the 95% confidence interval does not cover 150, we conclude that the average height at the school is statistically different from the national average.
City planners are evaluating traffic congestion on Maple Avenue, a major road leading into the city center. They believe that the average number of cars passing through during the 7–9 AM rush hour has increased compared to the usual volume of 800 vehicles. To investigate this, they count the number of cars passing through the intersection over 64 randomly selected weekdays. The sample yields an average of 826 cars per morning, with a sample standard deviation of 80 cars.
Using a 95% confidence level, can the planners conclude that the average number of cars on Maple Avenue during rush hour is different from 800?
Solution
To determine whether \(\bar x\), the estimate of the population mean, is statistically significantly different from 800, we construct a 95% confidence interval.
\[\hat\sigma_{\bar x} = \frac{s}{\sqrt{n}} = \frac{80}{\sqrt{64}} = 10\]
Using \(t = 1.96\), the confidence interval is:
\[826 \pm 1.96 \times 10 = 826 \pm 19.6 \;\Rightarrow\; CI = [806.4,\ 845.6]\]
Since the 95% confidence interval does not cover 800, we conclude that the average number of cars during rush hour is statistically different from 800.
A public health researcher is studying the daily caffeine intake of university students. According to the Public Health Agency, the recommended maximum caffeine intake for adults is 400 mg per day.
To assess whether students exceed this recommended amount, the researcher surveys 49 randomly selected students and finds that their average daily caffeine consumption is 418 mg, with a sample standard deviation of 85 mg.
Based on this sample, can we conclude that the average caffeine intake among students differs from the recommended maximum of 400 mg?
Solution
First, calculate the estimated standard error:
\[\hat\sigma_{\bar x} = \frac{s}{\sqrt{n}} = \frac{85}{\sqrt{49}} = \frac{85}{7} = 12.14\]
Using \(t = 1.96\), the 95% confidence interval is:
\[418 \pm 1.96 \times 12.14 = 418 \pm 23.79 \;\Rightarrow\; CI_{95\%} = [394.21,\ 441.79]\]
Since the 95% confidence interval covers 400, we cannot conclude that the average caffeine intake is statistically different from the recommended amount at the 95% confidence level.
We now test using a 90% confidence level. Using \(t = 1.645\):
\[418 \pm 1.645 \times 12.14 = 418 \pm 19.97 \;\Rightarrow\; CI_{90\%} = [398.03,\ 437.97]\]
Since the 90% confidence interval covers 400, we also cannot conclude that the difference is statistically significant at the 90% confidence level.
A transportation researcher wants to evaluate whether students at a suburban university have longer daily commutes than the commonly accepted average of 35 minutes. To investigate, the researcher surveys a random sample of 36 students and finds that their average commute time is 38.2 minutes, with a sample standard deviation of 9 minutes.
Can we conclude that the average student commute time differs from 35 minutes?
Solution
First, calculate the estimated standard error:
\[\hat\sigma_{\bar x} = \frac{9}{\sqrt{36}} = \frac{9}{6} = 1.5\]
Using \(t = 1.96\), the 95% confidence interval is:
\[38.2 \pm 1.96 \times 1.5 = 38.2 \pm 2.94 \;\Rightarrow\; CI_{95\%} = [35.26,\ 41.14]\]
Since the 95% confidence interval does not cover 35, we can conclude that the average commute time differs from 35 minutes at the 95% confidence level.
An education researcher is investigating whether the average number of hours high school students spend reading per week has declined from the previously reported national average of 12 hours. The researcher surveys 49 students and finds an average reading time of 10 hours per week, with a sample standard deviation of 4 hours.
Can we conclude that the average reading time is statistically significantly different from 12 hours?
Solution
First, calculate the estimated standard error:
\[\hat\sigma_{\bar x} = \frac{4}{\sqrt{49}} = \frac{4}{7} \approx 0.57\]
Using \(t = 1.96\), the 95% confidence interval is:
\[10 \pm 1.96 \times 0.57 = 10 \pm 1.12 \;\Rightarrow\; CI_{95\%} = [8.88,\ 11.12]\]
Since the 95% confidence interval does not cover 12, we conclude the average reading time is statistically significantly different from 12 at the 95% confidence level.
Next, we test at the 99% confidence level using \(t = 2.58\):
\[10 \pm 2.58 \times 0.57 = 10 \pm 1.47 \;\Rightarrow\; CI_{99\%} = [8.53,\ 11.47]\]
Since the 99% confidence interval does not cover 12, the difference is statistically significant even at the 99% confidence level.
Conclusion: The average reading time has significantly declined from the national average of 12 hours per week.
A researcher conducts a survey of 400 randomly selected voters to estimate support for the Social Democrats ahead of an upcoming election. Out of these respondents, 120 say they would vote for the Social Democrats if the election were held today.
Is there statistically significant evidence, at the 95% confidence level, that the true support for the Social Democrats in the population is different from 30%?
Solution
We use the following formula for a confidence interval for a proportion:
\[\hat p \pm z \times \sqrt{\frac{\hat p (1-\hat p)}{n}}\]
where \(\hat p\) is the sample proportion, \(n\) is the sample size, and \(z = 1.96\) for a 95% confidence level.
Calculate the sample proportion:
\[\hat p = \frac{120}{400} = 0.30\]
Calculate the standard error:
\[SE = \sqrt{\frac{0.30 \times (1-0.30)}{400}} = \sqrt{\frac{0.21}{400}} = \sqrt{0.000525} \approx 0.0229\]
Construct the 95% confidence interval:
\[0.30 \pm 1.96 \times 0.0229 = 0.30 \pm 0.0449 \;\Rightarrow\; CI_{95\%} = [0.2551,\ 0.3449]\]
Since the 95% confidence interval covers 30% (0.30), we cannot conclude that the true support differs from 30%.
Conclusion: There is no statistically significant evidence at the 95% confidence level that support for the Social Democrats is different from 30%.
A researcher wants to study the relationship between the number of hours students spend studying per week and their final exam scores (measured as percentages). The researcher estimates a simple linear regression on data from 50 students. The results are summarized below:
| Intercept | 55.0 |
| (5.0) | |
| Hours studied | 3.2 |
| (0.8) | |
| Observations | 50 |
| R² | 0.30 |
(Standard errors in parentheses)
Answer the following questions:
Solution
a) We interpret the slope coefficient as: a one hour increase in the hours studied is expected to increase exam scores by 3.2 percentage points.
b) To check whether the relationship is statistically significant, we construct a 95% confidence interval for the slope coefficient using:
\[\hat b \pm t \times \hat\sigma_{\hat b} = 3.2 \pm 1.96 \times 0.8 = 3.2 \pm 1.57 \;\Rightarrow\; CI_{95\%} = [1.63,\ 4.77]\]
Since the confidence interval does not cover 0, the slope is statistically significant at the 95% confidence level.
We further check the 99% confidence interval with \(t = 2.58\):
\[3.2 \pm 2.58 \times 0.8 = 3.2 \pm 2.06 \;\Rightarrow\; CI_{99\%} = [1.14,\ 5.26]\]
0 is also not covered by the 99% interval, so the slope is significant at the 99% confidence level as well.
c) Based on the answers in a) and b), we conclude that there is a statistically significant positive relationship between hours studied and exam scores.
An economist investigates the relationship between education and economic development. Specifically, she examines whether countries with higher average years of schooling tend to have higher levels of GDP per capita (measured in thousands of USD). She collects data from 60 countries and estimates a simple linear regression where the dependent variable is GDP per capita, and the independent variable is average years of schooling. The regression output is presented below:
| Intercept | 5.2 |
| (2.3) | |
| Years of schooling | 1.8 |
| (0.5) | |
| Observations | 60 |
| R² | 0.42 |
(Standard errors in parentheses)
Answer the following questions:
Solution
a) The slope coefficient of 1.8 means that each additional year of schooling in a country is expected to increase GDP per capita by $1,800.
b) We use the following formula to construct a confidence interval for the slope coefficient:
\[b \pm t \times \hat\sigma_b = 1.8 \pm t \times 0.5\]
Constructing the 95% confidence interval with \(t = 1.96\):
\[1.8 \pm 1.96 \times 0.5 = 1.8 \pm 0.98 \;\Rightarrow\; CI_{95\%} = [0.82,\ 2.78]\]
Since the confidence interval does not cover 0, the relationship is statistically significant at the 95% confidence level.
We also test at the 99% level, using \(t = 2.58\):
\[1.8 \pm 2.58 \times 0.5 = 1.8 \pm 1.29 \;\Rightarrow\; CI_{99\%} = [0.51,\ 3.09]\]
The confidence interval still does not cover 0, so the relationship is statistically significant at the 99% confidence level as well.
c) Based on the answers in a) and b), we conclude that there is a statistically significant positive relationship between average years of schooling and GDP per capita. Countries with more educated populations tend to have higher levels of economic development.
A political scientist wants to understand whether citizens’ trust in government is associated with voter turnout in national elections. She gathers data from 45 democratic countries. The dependent variable is voter turnout (in percent), and the independent variable is the average level of trust in government institutions (on a scale from 0 to 10, based on survey data).
She estimates a bivariate linear regression model, and the results are presented below:
| Intercept | 52.1 |
| (5.7) | |
| Trust in government | 1.3 |
| (0.9) | |
| Observations | 45 |
| R² | 0.18 |
(Standard errors in parentheses)
Answer the following questions:
Solution
a) The slope coefficient of 1.3 means that a one-unit increase in trust in government is expected to increase voter turnout by 1.3 percentage points.
b) We use the following formula to construct a confidence interval for the slope coefficient:
\[b \pm t \times \hat\sigma_b = 1.3 \pm t \times 0.9\]
Constructing the 95% confidence interval using \(t = 1.96\):
\[1.3 \pm 1.96 \times 0.9 = 1.3 \pm 1.76 \;\Rightarrow\; CI_{95\%} = [-0.46,\ 3.06]\]
Since the confidence interval covers 0, the relationship is not statistically significant at the 95% confidence level.
Now, construct the 90% confidence interval using \(t = 1.645\):
\[1.3 \pm 1.645 \times 0.9 = 1.3 \pm 1.48 \;\Rightarrow\; CI_{90\%} = [-0.18,\ 2.78]\]
The 90% confidence interval does not cover 0, so we conclude that the relationship is statistically significant at the 90% confidence level.
c) Based on the answers in a) and b), we conclude that there is a positive relationship between trust in government and voter turnout, and that this relationship is statistically significant at the 90% confidence level (but not at the 95% level).
A sociologist conducts a national survey asking 200 adults about their highest attained level of education. The results are summarized in the table below:
| Education level | Number of respondents |
|---|---|
| No formal education | 12 |
| Primary education | 58 |
| Secondary education | 54 |
| Tertiary education | 42 |
| Postgraduate education | 34 |
| Total | 200 |
Solution
a) Since the variable “highest attained education level” is ordinal, it does not make sense to compute the mean. The mean assumes that the distances between categories are meaningful and equally spaced, which we cannot assume for ordinal data.
However, both the median and the mode make sense. The median identifies the middle category when responses are ordered, and the mode identifies the most frequently occurring category.
b) To compute the median, we identify the middle observation in the ordered list of 200 respondents. Since the number is even, we look at the 100th and 101st observations.
Cumulative totals: No formal education: 12. Primary education: 12 + 58 = 70. Secondary education: 70 + 54 = 124. Tertiary education: 124 + 42 = 166. Postgraduate education: 166 + 34 = 200.
Both the 100th and 101st observations fall in the category “Secondary education”, since this category includes positions 71 through 124.
Therefore, the median is Secondary education.
To find the mode, we identify the category with the highest number of respondents. The most frequent response is “Primary education”, with 58 respondents. The mode is Primary education.
A researcher is studying how university students access political news. Each student in a sample of 210 was asked: “Which of the following sources do you primarily rely on for political news?” The responses are summarized in the table below:
| Primary news source | Number of respondents |
|---|---|
| Public service broadcasting (e.g., SVT, SR) | 68 |
| Newspapers (print or digital) | 47 |
| Social media platforms | 44 |
| Commercial TV channels | 19 |
| Online news aggregators (e.g., Omni) | 12 |
| Podcasts | 11 |
| Other | 9 |
| Total | 210 |
Solution
a) The variable “primary news source” is nominal, meaning its categories have no natural numerical or ordinal order. This makes the mean and median inappropriate for analysis. Only the mode is a suitable measure of central tendency for nominal variables, as it identifies the most frequently occurring category.
b) The category with the highest number of respondents is “Public service broadcasting,” with 68 students selecting it as their primary source. Thus, the mode is Public service broadcasting.