📧 hipravat@gmail.com Support

Quantitative Research Methods

Course: RES814 — Doctoral Program in Computer Science (Cybersecurity & Information Assurance) Tool: IBM SPSS Statistics | Focus Area: AI-Driven Cybersecurity Research


Overview

This course develops doctoral-level expertise in quantitative research methodology — from foundational variable classification and SPSS data entry, through descriptive statistics, hypothesis testing, nonparametric and parametric analysis, correlation, regression, and ANOVA — all applied to real datasets and research scenarios relevant to cybersecurity and organizational studies.


Foundations of Quantitative Research

What is Quantitative Research?

Quantitative research uses numerical data and statistical analysis to identify patterns, test hypotheses, and draw conclusions generalizable to larger populations.

Dimension Quantitative Qualitative
Paradigm Positivist Constructivist
Reasoning Deductive Inductive
Data Numerical Textual / observational
Goal Measure, test, generalize Understand, explore, interpret
Analysis Statistical (regression, ANOVA) Thematic, narrative
Sample Large, random Small, purposive

When to Choose Quantitative Methods

  • The research question involves measurable relationships or differences between variables
  • A hypothesis needs to be statistically tested
  • Findings must be generalizable beyond the study sample
  • Variables can be operationalized into numerical form

Maintaining a Human Perspective in Quantitative Research

Numbers represent people. To preserve the human dimension:

  • Embed the study in real organizational context — each survey response reflects a lived experience
  • Use meaningful variable labels in SPSS (e.g., "Employee Satisfaction Score", not just "var001")
  • Combine with qualitative methods (mixed methods) when statistics reveal what but not why
  • Report results in ways meaningful to stakeholders, not just statisticians
  • Handle missing data and outliers ethically — misrepresentation has human consequences

Variables and Measurement Levels

Types of Variables

Type Role Description
Independent Variable (IV) Predictor The variable manipulated or used to predict outcomes
Dependent Variable (DV) Outcome The variable measured to assess the effect of the IV
Control Variable Covariate Held constant to reduce confounding

Measurement Levels in SPSS

Nominal Data - Categories with no inherent order or ranking - Numbers serve as labels only, not values - Examples: Gender (1=Male, 2=Female), Cybersecurity incident type (1=Phishing, 2=Malware, 3=DoS) - Appropriate tests: Chi-square, frequency tables

Ordinal Data - Categories with a meaningful order, but unequal intervals between ranks - Examples: Satisfaction scale (1=Very Unsatisfied … 5=Very Satisfied), Military rank, Likert scale responses - Appropriate tests: Mann-Whitney U, Kruskal-Wallis, Spearman correlation

Scale (Interval/Ratio) Data - Continuous numeric data with equal intervals between values - Examples: Annual income ($), Age (years), Hours online per week, ACT test scores - Appropriate tests: t-tests, ANOVA, Pearson correlation, regression

Rule: The measurement level of your variable determines which statistical test is valid. Using the wrong test produces misleading results.

Variable Scenarios

1 IV → 1 DV

Training completion (trained vs. untrained) → Employee productivity score Analysis: Independent samples t-test

2 IVs → 1 DV

Work shift (day/night) + Experience level (novice/expert) → Job satisfaction score Analysis: Two-way ANOVA — examines individual effects and interaction effects

1 IV → Multiple DVs

Leadership style → Job satisfaction + employee retention + productivity Analysis: MANOVA (Multivariate ANOVA)


SPSS: Data Entry and Variable Definition

Setting Up Variables in SPSS

In Variable View, each column must be defined before analysis:

Field Purpose Example
Name Unique variable identifier rincdol, sex, age
Type Numeric, string, or date Numeric
Label Descriptive title "Respondent's Annual Income"
Values Codes for categorical data 1=Male, 2=Female
Missing Codes for absent data 99=No Response
Measure Nominal / Ordinal / Scale Scale

Why Correct Variable Definition Matters

  • SPSS will misinterpret categorical variables as numeric values if labels are not assigned
  • The Measure setting determines which statistical tests SPSS makes available
  • Properly labeled output is reproducible, transparent, and publication-ready
  • Incorrect measurement level assumptions invalidate statistical conclusions

Descriptive Statistics

Key Descriptive Measures

Measure What It Tells You
Mean Average value — sensitive to outliers
Median Middle value — robust to skewed data
Standard Deviation Spread of data around the mean
Skewness Direction and degree of distribution asymmetry
Minimum / Maximum Range of observed values

Income by Gender — SPSS Case Study (GSS Dataset)

Using the Explore command in SPSS to compare income (rincdol) by gender (sex):

Findings: - Males: M = $34,603.75, SD = $24,350.70 — positively skewed; high earners inflate the mean - Females: M = $27,729.70, SD = $23,321.62 — similar positive skew; more concentrated in lower income brackets - Both distributions are right-skewed → median is the more appropriate measure of central tendency

Visualizations used: Histograms, boxplots, stem-and-leaf plots

Interpretation: The mean overstates typical income for both groups due to high-income outliers. Gender-based income disparity is apparent in raw descriptive statistics.


Hypothesis Testing

Structure of Quantitative Hypotheses

Every quantitative research question generates two hypotheses:

  • Null Hypothesis (H₀): No significant difference or relationship exists between variables
  • Alternative Hypothesis (H₁): A significant difference or relationship exists

Decision rule: If p < .05, reject H₀ in favor of H₁ (statistically significant result)

Example: Gender and Annual Income

Research Question: Do men and women differ significantly in their average annual income?

H₀: μ_men = μ_women (no difference in mean income) H₁: μ_men ≠ μ_women (difference in mean income exists)

Test: Independent-samples t-test (continuous DV, two independent groups)

Results: - t(905) = 3.87, p < .001 - Males: M = $34,603.75 | Females: M = $27,729.70 - Levene's test for equality of variances: F(1, 905) = 1.62, p = .204 (equal variances assumed)

Conclusion: Reject H₀. Men and women differ significantly in annual income in the GSS sample.

Key Elements of a Sound Hypothesis

  1. Determinable variables — IV and DV clearly defined and measurable
  2. Specified population or sample — Who is being studied?
  3. Testable relationship — Can the prediction be statistically evaluated?
  4. Directionality — One-tailed (predicted direction) or two-tailed (any difference)?
  5. Falsifiability — Can the hypothesis be disproven by data?

Parametric Tests

Paired Samples T-Test

Compares means of the same group measured twice (pre/post design).

When to use: Same participants measured before and after an intervention

ACT Preparation Program Example: - 1993 ACT mean: M = 15.986 - 1994 ACT mean: M = 15.861 - Result: t(63) = 2.303, p = .025 - Conclusion: Scores decreased slightly after the program — not statistically meaningful improvement (effect size d = 0.29, small) - Practical significance does not follow from statistical significance alone

Independent Samples T-Test

Compares means of two unrelated groups.

When to use: Two independent groups on a continuous DV

Cholesterol and Survival Example: - Compared cholesterol levels (CHOL58) between patients who survived 10 years vs. those who did not - Result: No statistically significant difference found - Lesson: Even clinical datasets can yield null results — report them accurately

One-Way ANOVA

Compares means across three or more independent groups.

When to use: One nominal IV (3+ groups) and one continuous DV

ACT Scores Across School Groups: - Tests whether mean ACT scores differ significantly by school or group assignment - F-statistic measures the ratio of between-group variance to within-group variance - Significant F → at least one group mean differs; follow-up with post-hoc tests (Tukey, Bonferroni)

Between-Group vs. Within-Group Variance

Variance Type Description What It Reflects
Between-group Differences across group means Effect of the IV
Within-group Variability within each group Random error / individual differences

A significant ANOVA result means between-group variance is large relative to within-group variance.


Nonparametric Tests

When to Use Nonparametric Tests

Use nonparametric methods when data violates parametric assumptions: - Data is nominal or ordinal (e.g., Likert scales) - Small sample sizes prevent normality assumptions - Distribution is heavily skewed - Variables cannot be manipulated for ethical or practical reasons

Nonparametric Test Selection Guide

Research Scenario Appropriate Test
Association between two categorical variables Chi-Square Test of Independence
Comparing two independent groups on ordinal data Mann-Whitney U Test
Comparing same group before and after (ordinal) Wilcoxon Signed-Rank Test
Comparing 3+ independent groups on ordinal data Kruskal-Wallis Test

Chi-Square Test — Education and Loan Default

Research Question: Is there a statistically significant association between education level and loan default?

H₀: No significant association exists between education level and loan default status H₁: A significant association exists

Results: χ²(4, N = 700) = 11.492, p = .022 - Conclusion: Reject H₀ — education level significantly associated with loan default - Effect size: Cramer's V = .128 (small but significant) - Higher education → lower default rates; inverse relationship confirmed by linear-by-linear association

Chi-Square Test — Age and Political Orientation

Results: χ²(4, N = 60) = 12.667, p = .013 - Young adults: predominantly liberal - Middle-aged: predominantly moderate - Older adults: predominantly conservative - Linear-by-linear association (8.016, p = .005): political conservatism increases with age - Lambda = .263: knowing age category reduces classification error by ~26%


Correlation Analysis

Pearson Correlation Coefficient (r)

Measures the strength and direction of the linear relationship between two continuous variables.

r Value Interpretation
.00 – .19 Very weak
.20 – .39 Weak
.40 – .59 Moderate
.60 – .79 Strong
.80 – 1.00 Very strong

Correlation ≠ causation. A significant r only confirms a linear relationship.

Father's vs. Mother's Education (GSS Dataset)

Variables: maeduc (mother's education) → paeduc (father's education)

Results: - r = .637, p < .001 — strong positive correlation - R² = 0.406 — mother's education explains 40.6% of variance in father's education - Consistent with educational homogamy theory — couples tend to share similar educational levels


Linear Regression

Bivariate Linear Regression

Tests whether one continuous IV predicts one continuous DV.

Regression Equation: Y = b₀ + b₁X

Where: - Y = predicted outcome (DV) - b₀ = intercept (value of Y when X = 0) - b₁ = slope (change in Y for each 1-unit increase in X)

Interpreting R² (Coefficient of Determination)

R² = proportion of variance in the DV explained by the IV

  • R² = .40 → 40% of DV variance explained by IV
  • Remaining 60% is explained by other unmeasured variables
  • Statistical significance (p < .05) ≠ practical significance (high R²)
  • A significant relationship with low R² means the IV is a weak predictor; multiple regression may be needed

Wife's Education Predicted by Husband's Education (GSS)

Regression Equation: Wife's Education = 6.407 + 0.506(Husband's Education)

  • R = .561, R² = .314 — husband's education explains 31.4% of variance in wife's education
  • F(1, 607) = 278.086, p < .001
  • Slope b₁ = 0.506: each additional year of husband's education predicts ~0.5 more years of wife's education
  • Predicted value: If husband has 14 years of education → Wife's predicted education = 13.48 years

Years of Service and Productivity

Regression Equation: Predicted Productivity = 60 + 2.5(Years of Service)

  • Each additional year of service → +2.5 productivity points
  • However, if R² is low (e.g., .12), only 12% of productivity variation is explained by service years
  • Practical implication: Other factors (training, motivation, leadership) should be added via multiple regression

Secondary Data Analysis with SPSS

Using the General Social Survey (GSS)

The GSS is a reliable, nationally representative secondary dataset widely used in social science and organizational research.

Key variables used in this course: - sex — Gender (nominal) - rincdol — Annual income in dollars (scale) - useweb / webhrs — Internet use frequency and hours (nominal / scale) - maeduc / paeduc — Mother's / Father's education in years (scale) - husbeduc / wifeduc — Spouse education levels (scale)

Selecting the Right Test for Secondary Data

Research Question Variables Test
Do males and females differ in internet use? Sex (nominal) + useweb (nominal) Chi-Square
Do males and females differ in hours online? Sex (nominal) + webhrs (scale) Independent T-Test
Does mother's education predict father's? maeduc (scale) + paeduc (scale) Bivariate Regression
Do income levels differ by gender? Sex (nominal) + rincdol (scale) Independent T-Test

Literature Reviews in Quantitative Research

Quantitative vs. Qualitative Literature Reviews

Dimension Quantitative Qualitative
Organization By variable, theory, statistical relationships By theme, historical development, context
Language Objective, impersonal, results-focused Exploratory, interpretive
Theory role Informs predictions and hypotheses before data Emerges during/after data collection
Conclusion Leads to testable hypotheses Leads to research purpose and open questions
Researcher role Neutral, minimizes personal interpretation Declares positionality; reflexive

Building a Quantitative Literature Review

  1. Establish what is already known about the phenomenon
  2. Identify gaps — variables not studied, relationships not tested, populations not examined
  3. Organize around key variables and theoretical constructs
  4. Progress from general to specific — broader field → your specific research model
  5. Conclude with testable hypotheses or research questions

Ethical Considerations in Quantitative Research

Informed Consent Requirements

Participants must be informed of: - The study's purpose and scope - Any risks or benefits of participation - How their data will be stored and used - Their right to withdraw at any time without penalty - Confidentiality and anonymization procedures

Research Design Transparency

A vague research announcement (e.g., "We want to find the best leadership style — please participate") raises serious ethical and methodological concerns: - Participants cannot give meaningful informed consent without knowing the design - Without operationalizing "leadership style," results cannot be replicated or interpreted - Ambiguous sampling methods undermine external validity

Data Integrity

  • Properly classify variables to avoid misleading statistical output
  • Do not present ordinal data as scale data — this invalidates parametric assumptions
  • Report effect sizes alongside p-values — statistical significance alone is insufficient
  • Missing data must be handled systematically (listwise deletion, multiple imputation) and reported

Summary and Key Takeaways

  1. Variable classification determines your test — nominal, ordinal, and scale data each require different statistical approaches

  2. SPSS is a tool, not a decision-maker — understanding the assumptions behind each test is the researcher's responsibility

  3. Statistical significance ≠ practical significance — always report effect sizes (d, r, R², Cramer's V) alongside p-values

  4. Descriptive statistics first — always explore your data distributions before running inferential tests

  5. Parametric tests require assumptions — normality, equal variances, and scale data; when assumptions fail, use nonparametric alternatives

  6. Regression quantifies prediction — R² tells you how much variance is explained; low R² with significant p signals a real but weak relationship

  7. Secondary data (GSS) enables doctoral-level analysis — large representative datasets allow hypothesis testing without primary data collection

  8. Ethics are foundational — informed consent, transparency, and data integrity apply to every quantitative study


References

  • Creswell, J. W., & Creswell, J. D. (2014). Research Design: Qualitative, Quantitative, and Mixed Methods Approaches. SAGE.
  • Field, A. (2024). Discovering Statistics Using IBM SPSS Statistics (6th ed.). SAGE.
  • Gravetter, F. J., Wallnau, L. B., et al. (2021). Statistics for the Behavioral Sciences. Cengage.
  • Green, S. B. (2016). Using SPSS for Windows and Macintosh. Pearson.
  • Pallant, J. (2020). SPSS Survival Manual (7th ed.). McGraw-Hill.
  • Northouse, P. G. (2025). Leadership: Theory and Practice (9th ed.). SAGE.
  • Huck, S. W. (2012). Reading Statistics and Research (6th ed.). Pearson.
  • Ravitch, S. M., & Carl, N. M. (2019). Qualitative Research: Bridging the Conceptual, Theoretical, and Methodological. SAGE.