MM207 Milestone 4

MM207 Milestone 4

Name

Purdue University Globle

MM207 Statistics

Prof. Name

Date

MM207 Milestone 4: Scatterplots, Correlation, Regression, and Causation

MM207 Milestone 4 focuses on understanding how scatterplots, correlation, regression, and causation are used to examine relationships between quantitative variables. Scatterplots help visualize patterns, correlation measures the direction and strength of a linear association, regression is used to model and predict outcomes, and causation requires evidence that a change in one variable directly produces a change in another. Understanding the differences among these concepts is essential for interpreting statistical results accurately and avoiding misleading conclusions.

Understanding Explanatory and Response Variables

A statistical analysis involving two quantitative variables typically identifies one as the explanatory variable and the other as the response variable. The explanatory variable is the factor used to explain or predict changes in another variable, while the response variable is the outcome being measured.

In a scatterplot, the explanatory variable is generally placed on the horizontal x-axis, and the response variable is placed on the vertical y-axis. This arrangement makes it easier to examine whether changes in the explanatory variable are associated with changes in the response variable.

Example of Explanatory and Response Variables

Suppose researchers examine the relationship between voltage and motor rotation speed. Voltage would be the explanatory variable because it is used to explain or predict changes in motor speed. Motor rotation speed would be the response variable because it represents the outcome being measured.

How to Identify Variables in a Scatterplot

When interpreting a scatterplot, first identify the variables represented on each axis. Look at the label on the horizontal axis to determine the explanatory variable, then examine the vertical axis to identify the response variable.

A useful approach is to ask: Which variable is being used to explain or predict the other variable? The answer is generally the explanatory variable.

After identifying both variables, examine whether changes in the x-variable appear to correspond with changes in the y-variable. This provides the foundation for describing the relationship.

What Is a Scatterplot?

A scatterplot is a graph that displays pairs of quantitative observations as individual points. Each point represents the values of two variables for one observation.

Scatterplots are useful because they allow researchers to see patterns that may not be obvious from a table of numbers. They can reveal whether variables appear to have a positive, negative, linear, or nonlinear relationship and can also highlight unusual observations.

When analyzing a scatterplot, focus on four major characteristics: form, direction, strength, and unusual observations.

Form of a Scatterplot

The form describes the general shape of the relationship between the variables. Common forms include linear and curved relationships.

A linear relationship follows an approximately straight-line pattern. A nonlinear relationship follows a curve or another pattern that cannot be adequately represented by a straight line.

Direction of a Scatterplot

Direction describes how the response variable tends to change as the explanatory variable increases.

A positive association occurs when larger values of one variable tend to be associated with larger values of the other. A negative association occurs when larger values of one variable tend to be associated with smaller values of the other.

Strength of a Scatterplot

Strength describes how closely the observations follow the overall pattern. Points tightly grouped around a trend indicate a stronger association, while widely scattered points indicate a weaker association.

Outliers or unusual observations should also be identified because they can influence statistical calculations and regression results.

What Is Correlation?

Correlation measures the strength and direction of a linear relationship between two quantitative variables. The correlation coefficient is represented by r and ranges from -1 to +1.

Correlation coefficientInterpretation
+1Perfect positive linear correlation
0No linear correlation
-1Perfect negative linear correlation

A value of r close to +1 indicates a strong positive linear association. A value close to -1 indicates a strong negative linear association. A value near 0 indicates little or no linear association.

Correlation describes an association between variables; it does not by itself establish that one variable causes the other.

Positive Correlation

A positive correlation occurs when higher values of one variable tend to occur with higher values of the other variable.

Examples include study time and examination scores, years of work experience and salary, or advertising expenditure and sales revenue. These examples describe associations and should not automatically be interpreted as proof of a causal relationship.

Negative Correlation

A negative correlation occurs when higher values of one variable tend to occur with lower values of the other variable.

For example, as the number of miles driven increases, the amount of fuel remaining in a vehicle generally decreases. This creates a negative association between miles driven and remaining fuel.

How to Interpret a Correlation Coefficient

The sign of r indicates the direction of the linear relationship, while the absolute value of r indicates its strength.

For example, r = 0.90 represents a strong positive linear association, whereas r = -0.90 represents a strong negative linear association. Both have the same strength because their absolute values are equal, but their directions are opposite.

Correlation should always be interpreted in context rather than relying only on a numerical value.

What Is Regression Analysis?

Regression analysis is a statistical method used to describe the relationship between variables and predict the value of a response variable from one or more explanatory variables.

In simple linear regression, the regression equation is commonly written as:

ŷ = a + bx

Here, ŷ represents the predicted response value, a is the y-intercept, b is the slope, and x represents the explanatory variable.

Regression is commonly applied to areas such as healthcare, business, economics, education, engineering, finance, and marketing. For example, a regression model could be used to estimate sales based on advertising expenditure or predict an outcome based on a measurable risk factor.

Example of a Regression Prediction

Consider the regression equation:

ŷ = 0.375x + 1.33

Suppose x represents hours of sleep and ŷ represents predicted GPA. If a student sleeps 2.5 hours, the predicted GPA is:

ŷ = 0.375(2.5) + 1.33

ŷ ≈ 2.27

Based on this model, the predicted GPA is approximately 2.27. This is a model-based prediction, not a guarantee of the student’s actual GPA.

What Is a Best-Fit Line?

A best-fit line, also called a regression line, summarizes the overall linear trend in a scatterplot. Rather than connecting individual observations, the line represents the expected relationship between the explanatory and response variables.

A best-fit line can be used to make predictions. For example, a regression model could estimate sales based on advertising expenditure or predict athletic performance based on a measurable characteristic.

Understanding Slope in Regression

The slope describes the expected change in the response variable for a one-unit increase in the explanatory variable.

The slope between two points can be calculated as:

Slope = (y₂ − y₁) / (x₂ − x₁)

For example, suppose a hamster weighs 0.5 pounds and has a weekly feeding cost of $2, while a Labrador weighs 62.5 pounds and has a weekly feeding cost of $10.

The slope is:

(10 − 2) / (62.5 − 0.5) = 8 / 62 ≈ 0.13

This means the average feeding cost increases by approximately $0.13 for each additional pound within the context of these two observations.

Interpreting a Regression Slope

Consider the regression equation:

ŷ = 17 + 0.8x

If x represents a baby’s age in months and ŷ represents predicted weight in pounds, the slope of 0.8 means that the predicted weight increases by an average of 0.8 pounds for each additional month of age, according to the model.

The interpretation of a slope should always include the units of both variables.

Correlation Does Not Equal Causation

One of the most important principles in statistics is that correlation does not prove causation. Two variables can be strongly associated even when changes in one variable do not directly cause changes in the other.

An observed relationship may be explained by a third variable, known as a lurking variable, or by other factors that have not been controlled.

Ice Cream Sales and Drowning Example

Ice cream sales and drowning incidents may both increase during the summer. This does not mean that purchasing ice cream causes drowning.

A likely explanation is that warmer weather increases both ice cream consumption and participation in swimming and other outdoor activities. Temperature and seasonal behavior therefore provide alternative explanations for the observed association.

How Researchers Establish Causation

Establishing a cause-and-effect relationship requires stronger evidence than simply observing that two variables are associated.

Researchers may use controlled experiments and randomized assignment to reduce the influence of alternative explanations. A cause must also occur before its effect, and researchers need evidence that competing explanations have been adequately addressed.

Randomized controlled trials are particularly valuable for evaluating causal effects because random assignment helps balance other characteristics among treatment groups.

Observational studies can identify important associations, but they generally require additional evidence and careful study design before causal conclusions can be made.

Understanding Nonlinear Relationships

Correlation measures linear association. Consequently, a correlation coefficient may fail to describe an important relationship when the data follow a curved pattern.

For example, a dataset may have a strong U-shaped or inverted-U-shaped relationship while producing a correlation near zero. This is why examining a scatterplot before interpreting correlation is important.

Nonlinear relationships can occur in areas such as population growth, compound interest, projectile motion, and seasonal measurements.

What Is R²?

The coefficient of determination, represented by R², describes the proportion of variation in the response variable that is accounted for by a linear regression model.

For example, if:

R² = 0.9846

then approximately 98.46% of the variation in the response variable is explained by the fitted linear model, while approximately 1.54% is not explained by the model.

A higher R² can indicate that a model explains more variation, but there is no single R² value that automatically makes a model appropriate. Researchers should also consider the research question, data quality, residuals, outliers, and assumptions of the regression model.

Relationship Between R² and Correlation

For simple linear regression involving two quantitative variables, the coefficient of determination is related to the correlation coefficient by:

R² = r²

If:

R² = 0.9846

then:

r ≈ ±0.992

The sign of r must be determined from the direction of the relationship. If the regression slope is positive, r ≈ +0.992; if the slope is negative, r ≈ -0.992.

What Are Residuals?

A residual represents the difference between an observed value and the value predicted by a regression model.

The formula is:

Residual = Actual value − Predicted value

Residuals help researchers evaluate how well a regression model represents the observed data.

Residual Example

Suppose the regression equation is:

ŷ = -210 + 5.6x

If a person is 72 inches tall, the predicted weight is:

ŷ = -210 + 5.6(72) = 193.2 pounds

If the actual weight is 182 pounds:

Residual = 182 − 193.2 = -11.2

The negative residual indicates that the actual weight is 11.2 pounds below the predicted value.

Why Residuals Matter

Residuals are useful for assessing the performance and appropriateness of a regression model. Researchers can examine residual patterns to identify problems such as nonlinearity, unequal variability, unusual observations, or other violations of model assumptions.

A smaller residual means an individual prediction is closer to the observed value, although residual size alone does not determine whether an entire regression model is appropriate.

Understanding Outliers and Influential Points

An outlier is an observation that appears unusually far from the overall pattern of the data. Outliers may occur because of measurement problems, data-entry errors, unusual circumstances, or genuine variation.

An influential observation is a data point that has a substantial effect on the fitted regression model. Some influential observations have extreme values of the explanatory variable and therefore have high leverage.

Outliers and influential observations should be investigated rather than automatically deleted. Researchers should determine whether the observation is valid and consider how it affects the conclusions.

What Is Least-Squares Regression?

The least-squares regression line is the line that minimizes the sum of the squared residuals between observed and predicted response values.

For simple linear regression, the slope can be calculated as:

b = r(sᵧ / sₓ)

The intercept is calculated as:

a = ȳ − b x̄

The resulting regression equation is:

ŷ = a + bx

Least-squares regression is widely used in business analytics, healthcare research, economics, engineering, finance, and other fields where researchers need to describe relationships or make predictions.

What Is a Lurking Variable?

A lurking variable is an unmeasured or unaccounted-for variable that influences the variables being studied and can create or distort an observed association.

For example, suppose a study finds that communities with more vehicles also have higher levels of health insurance coverage. Vehicle ownership itself may not explain the relationship. Household income, employment, education, or other socioeconomic factors could influence both vehicle ownership and access to health insurance.

Recognizing potential lurking variables is important because they can lead to incorrect interpretations of correlation and regression results.

Correlation vs. Regression

Correlation and regression are related but serve different purposes. Correlation summarizes the strength and direction of a linear association between two quantitative variables. Regression, on the other hand, models the relationship and can be used to predict the response variable from an explanatory variable.

Correlation does not distinguish between explanatory and response variables in the same way regression does. Regression also provides an equation that can be used for prediction within an appropriate range of the observed data.

Practical Applications of Correlation and Regression

Correlation and regression are used across many disciplines to analyze quantitative data and support evidence-based decisions. Applications include predicting housing prices, forecasting business revenue, evaluating healthcare outcomes, examining educational performance, analyzing sports statistics, studying climate patterns, and assessing financial risk.

The usefulness of these techniques depends on selecting appropriate variables, checking the data, examining graphs, evaluating model assumptions, and interpreting results within their proper context.

Frequently Asked Questions About MM207 Milestone 4

What is the explanatory variable?

The explanatory variable is the variable used to explain or predict changes in the response variable. In a typical scatterplot, it is placed on the horizontal x-axis.

What is the response variable?

The response variable is the outcome being measured or predicted. It is commonly displayed on the vertical y-axis of a scatterplot.

What does a scatterplot show?

A scatterplot displays paired quantitative observations and helps identify the form, direction, strength, and unusual features of a relationship between two variables.

What does correlation measure?

Correlation measures the strength and direction of a linear relationship between two quantitative variables. Its value ranges from -1 to +1.

Does correlation prove causation?

No. Correlation demonstrates an association, but it does not establish that one variable causes the other. Causal conclusions require stronger evidence and an appropriate research design.

What is a regression line?

A regression line is a mathematical model that summarizes a linear relationship between an explanatory variable and a response variable. It can also be used to generate predictions.

What does the slope of a regression line mean?

The slope represents the expected change in the response variable associated with a one-unit increase in the explanatory variable.

What does R² tell you?

R² represents the proportion of variation in the response variable that is explained by the fitted regression model.

Why are residuals important in regression?

Residuals show the difference between observed and predicted values. Examining residuals helps researchers evaluate prediction errors and determine whether a linear model is appropriate.

What is an outlier?

An outlier is an observation that is unusually distant from the overall pattern of the data. Outliers can affect correlation and regression results and should be investigated carefully.

What is a lurking variable?

A lurking variable is an unmeasured factor that may influence the variables being studied and create a misleading or distorted association.

Key Takeaways

MM207 Milestone 4 emphasizes the importance of interpreting relationships between quantitative variables correctly. Scatterplots provide a visual starting point, while correlation quantifies the direction and strength of a linear association. Regression extends this analysis by modeling relationships and making predictions.

The most important concepts to remember are:

  • Scatterplots display relationships between two quantitative variables.

  • The explanatory variable is generally placed on the x-axis.

  • The response variable is generally placed on the y-axis.

  • Correlation ranges from -1 to +1 and measures linear association.

  • A positive correlation indicates that variables tend to increase together.

  • A negative correlation indicates that one variable tends to decrease as the other increases.

  • Regression equations can be used to model relationships and make predictions.

  • The slope represents the expected change in the response variable for a one-unit change in the explanatory variable.

  • R² measures the proportion of response-variable variation explained by a regression model.

  • Residuals represent the difference between observed and predicted values.

  • Outliers and influential observations can substantially affect statistical models.

  • Correlation alone does not establish causation.

  • Lurking variables can produce misleading associations.

  • Scatterplots should be examined before relying on a correlation coefficient or regression model.

References

Freedman, D., Pisani, R., & Purves, R. (2007). Statistics (4th ed.). W. W. Norton & Company. https://wwnorton.com/

Moore, D. S., McCabe, G. P., & Craig, B. A. (2021). Introduction to the practice of statistics (10th ed.). W. H. Freeman. https://www.macmillanlearning.com/

National Institute of Standards and Technology. (2023). NIST/SEMATECH e-Handbook of statistical methodshttps://www.itl.nist.gov/div898/handbook/

OpenStax. (2023). Introductory statistics 2e. Rice University. https://openstax.org/details/books/introductory-statistics-2e

MM207 Milestone 4

Penn State Eberly College of Science. (2024). STAT 200: Elementary statisticshttps://online.stat.psu.edu/stat200/

UCLA Institute for Digital Research and Education. (2024). Statistical computing seminars and regression resourceshttps://stats.oarc.ucla.edu/