Table of Contents (Index)
- 01 What is Biostatistics & Statistical Inference?
- 02 Probability Distributions & Hypothesis Testing
- 03 Regression Models & Confounding Adjustment
- 04 Genome-Wide Association Studies (GWAS) Quality Control
- 05 Polygenic Risk Scores (PRS) & Genetic Liability
- 06 Mendelian Randomization & Causal Inference
What is Biostatistics & Statistical Inference?
Biostatistics is the application of statistical principles to questions in medicine, public health, and biology. Statistical inference allows researchers to draw meaningful conclusions about a broad population based on limited sample observations, accounting for inherent biological variability and measurement noise.
Probability Distributions & Hypothesis Testing
Biological measurements often follow specific distributions, such as the normal distribution or binomial distribution. Hypothesis testing provides a formal framework to evaluate evidence against a null hypothesis ($H_0$), quantified using p-values and confidence intervals.
Regression Models & Confounding Adjustment
To model relationships between risk factors and health outcomes, biostatisticians use generalized linear models (GLMs). Multiple linear regression handles continuous outcomes, logistic regression models binary disease status, and covariate adjustment prevents false associations driven by confounders like age or sex.
Genome-Wide Association Studies (GWAS) Quality Control
High-throughput genetic studies require strict preprocessing QC, filtering out markers with high missingness or deviations from Hardy-Weinberg equilibrium. Principal Component Analysis (PCA) is applied to control for population stratification and ancestral divergence.
Polygenic Risk Scores (PRS) & Genetic Liability
Polygenic risk scores aggregate thousands of genetic variants to estimate inherited susceptibility to complex traits. Utilizing Bayesian shrinkage models like LDpred2, our pipelines account for local linkage disequilibrium structures to maximize predictive performance.
Mendelian Randomization & Causal Inference
Mendelian randomization uses genetic variants as instrumental variables to deduce causal links between risk exposures and clinical endpoints, utilizing sensitivity estimators like MR-Egger to test for horizontal pleiotropy.