BIOS 600 - Spring 2026
Today, we’ll “collect” some data from the class and record on the whiteboard!
How many hours of sleep did you get last night? If you feel comfortable sharing, please draw a dot above the number corresponding to the number of hours of sleep you got. (Round to the nearest hour)
Draw each dot about 1 inch above the previous dot. Make dots about the same size. We’ll come back to this later in class.
No reading for today.
Lab 1 held today at 3:30pm in MHRC 0003.
Fill out the Getting to Know You Survey (Canvas assignments) by Friday at 11:59pm.
TA office hours will be Fridays 3:00-4:00pm in 6110 Mary Ellen Jones Building, or on Zoom.
My office hours will be Tuesdays 9:45am-10:45am. This week and next week, in my office, 3104A. Afterwards, in McG 3106E (Hoteling Room 2)
Reach out to the TA if you need help installing R/RStudio!
A process that converts data into useful information, whereby practitioners
The population is the group we’d like to learn something about:
What is the prevalence of diabetes among U.S. adults, and has it changed over time?
Is there a relationship between tumor type and five-year mortality in breast cancer patients?
Does the average amount of caffeine vary by vendor in 12 oz cups of coffee at UNC coffee shops?
If we had data from every unit in the population, we could just calculate what we wanted and be done!
Unfortunately, we (usually) have to settle with a sample from the population.
Based on our collective histogram, get in groups of 4 and discuss the following:
What is our population?
What is our sample?
What does our sample show? Do we think it’s representative of our population? Why or why not?
What limitations might we have? (How could we improve?)
Probability sampling (e.g., simple random sampling, stratified, cluster, or multi-stage sampling)
All units have a known chance of being selected
More likely to be generalizable
Can be more expensive and time-consuming
Non-probability sampling (e.g. quota, convenience, or snowball sampling)
Some units unable to be selected, with no way of knowing size or effect of sampling errors
Less generalizable to population of interest
More convenient and less costly
Experimental studies (e.g. randomized control trials (RCTs))
Researchers directly control exposures or treatment groups.
Ability to make causal statements
Less real-world applicability and generalizability
Observational studies (e.g. surveys)
Researchers do not assign exposures or treatments
Real-world setting with lower burden on participants
Inability to prove causality
Study: The effect of a new diet on weight loss in adults
Objective: To determine whether a new low-carbohydrate diet leads to significant weight loss compared to a standard diet in adults.
Study Design:
In small groups, discuss the following:
Depending on the study design and sampling methods, different types of bias occur (again, we don’t have the full population, only a sample).
Examples include:
seroprevalence = level of pathogen in a population
From Conclusion: “The estimated prevalence of SARS-CoV-2 antibodies in Santa Clara County implies that COVID-19 was likely more widespread than indicated by the number of cases in late March, 2020”.
Study recruited participants through Facebook ads, may have attracted people more concerned about COVID-19 or those who thought they were already exposed \(\rightarrow\) overestimation of seroprevalence (level of pathogen in a population)
Symptomatic/exposed individuals may have signed up to get tested, especially when testing was limited.
Participants may have recruited others in their network
Without a truly random sample, findings are less generalizable and reproducible
Study had high false positive rate (2 FP out of 401 known negative samples, about 0.5%), but with a wide confidence interval! (FP Rate could be above 1.2%)
Study reported only 50 positives out of ~3300 participants, so such a high FP rate could account for a significant portion - if not all - of the positive results. (\(\frac{2}{401}\cdot 3300 = 17\)) Yikes!
It could even be more! (re-calibrate the assay on a different set of 401 known negative samples, you may get 0/401 FP or even 6/401).
Results would imply much faster spread than previous pandemics
Not impossible, but a reason to be a bit more skeptical of results.
Study does not make its raw data, code, and full methodology publicly available
Without access to these, other researchers cannot replicate the study or verify its findings, undermining the reproducibility of the research
Transparency essential for the scientific process (particularly in high-stakes studies!)
Reproducibility: being able to take original data and code to reproduce all numerical findings.
Replicability: being able to independently repeat an entire study without use of the original data (generally with the same methods)
Best practices from the American Statistical Association:
End-to-end scripting of research
From the moment you “read in” your raw data, document the code you used for data cleaning, visualization, and analysis
Use of version control and documentation (we will not use this in our class)
Publication of code along with data


Fraudulent data can lead to unreliable research findings, waste resources, and erode trust in scientific integrity!
📊 Turning Data Into Insight
Biostatistics helps us move beyond raw numbers to identify meaningful patterns in health and disease.
🧪 Evaluating Interventions
From vaccines to health policies, biostatistics allows us to rigorously test what really works.
🌍 Promoting Health Equity
Careful analysis helps uncover disparities in health outcomes and guide solutions for underserved communities.
🔮 Forecasting & Prevention
Models help us anticipate outbreaks, chronic disease trends, and the impact of prevention strategies.
🏥 Improving Patient Care
Biostatistics drives evidence-based medicine by connecting research to better clinical decisions.
Step 1: Explore Google Scholar
Step 2: Share & Discuss
Form groups of 3. Each person shares the paper they found. As a group, discuss:
What is the paper studying? How are biostatistics being applied? Why is this meaningful or impactful for public health?
Step 3: Reflect & Submit
Take a screenshot of the abstract/front page of your chosen paper
Write a short reflection (3–4 sentences):
What do you find interesting/important about the paper you chose?
How does it show the value of biostatistics?
Submit the screenshot + reflection on Canvas
Due this Friday at 11:59pm
What is (bio)statistical inference?
Identifying the population of interest
Sampling types and biases
Reproducibility crisis: falsifying data
Biostatistics for public good!