Your first step into how data scientists actually think about numbers.
Welcome to Day 1. If you're looking to build real data science and analytics skills, statistics is where that has to start. Not because it's a prerequisite to check off, but because almost nothing else in this field works without it.
I'll be your guide through the fundamentals here, and the series will get more advanced as we go. Whether you're completely new to this or brushing up, the goal is the same: by the end of each post you should be able to actually use the idea, not just recognize the term. Let's get into it.
Statistics is the backbone of data science. Before you build a model, make a prediction, or draw a business conclusion, you need to understand your data: statistics is the toolkit for doing that.
Think of statistics as the language your data speaks. Without it, you're looking at a spreadsheet full of numbers. With it, you can find patterns, measure how uncertain you should be, and make decisions with some actual confidence behind them. Every major technique in data science, from linear regression to neural networks, sits on a statistical foundation.
A doctor doesn't examine every cell in your body to check your health. They look at a handful of key indicators (blood pressure, cholesterol, temperature) and use those to understand the full picture. Data scientists do the same thing with data: a few well-chosen summary statistics tell you far more than staring at every raw row ever will.
Statistics splits into two main branches, and almost every technique you'll learn in this series falls into one of them.
| Descriptive statistics | Inferential statistics | |
|---|---|---|
| Goal | Summarize and describe the data you already have | Draw conclusions about a population from a sample |
| Scope | The dataset in hand | A broader, unseen population |
| Output | Charts, averages, distributions | Predictions, hypothesis test results, confidence intervals |
| Example | "The average exam score in this class is 74." | "Students nationally likely score between 70–78." |
You'll almost always use them in that order: descriptive first, to understand what you're looking at, then inferential, to go beyond it. Skip the first step and the second one is built on sand.
Descriptive statistics is about making raw data understandable. It helps you collect, organize, summarize, and present data so its key features are visible at a glance. Instead of getting lost in thousands of rows, you see the shape of the thing.
Four groups of tools do most of the work here. Measures of central tendency (mean, median, and mode) tell you where the "center" of your data sits. Measures of dispersion (range, variance, standard deviation, interquartile range or IQR, and median absolute deviation) tell you how spread out the data is, and how much you should trust that center value. Frequency distributions tell you how often each value, or range of values, shows up. And graphs and tables (histograms, box plots, bar charts, scatter plots) turn all of the above into something you can actually see.
None of these tools mean much in isolation. The real skill is knowing which combination answers your actual question, which is exactly what the next example is about.
Say you collect the salaries (in USD/year) of 10 employees at a startup:
[40K, 45K, 48K, 50K, 52K, 55K, 60K, 65K, 70K, 200K]
The mean says the "typical" employee earns about $68.5K. The median says $53.5K. That's a $15K gap between two numbers that are both supposed to describe the same "average" employee, and it's entirely because of one person on that list.
The $200K salary is an outlier: a value far removed from the rest of the data. The mean uses every value in its calculation, so that one number drags it upward. The median only cares about the middle position in a sorted list, so it barely notices. This is why the median is usually a better "typical value" for skewed data like income, house prices, or app session times, where a few extreme values sit alongside a lot of ordinary ones.
The standard deviation makes the distortion even clearer. Above, it comes out to about $47K, which would normally tell you salaries are spread wildly, but nine of the ten people are within $30K of each other. Drop the outlier and recompute:
One data point took the standard deviation from $47K down to under $10K. That's the whole lesson in one number: a single outlier can dominate the mean and the standard deviation while leaving the median almost untouched. A good data scientist checks central tendency and dispersion together, and looks at the actual data before trusting either one.
Notice that the IQR ($15.25K) and the MAD ($7.5K) both stayed small without anyone removing the outlier. That's not a coincidence: IQR only looks at the middle 50% of the data (the gap between Q1 and Q3), so a single extreme value sitting far outside that range barely moves it. MAD works the same way, using the median as the center point and then taking the median of how far every value sits from it. Both are built to ignore exactly the kind of extreme value that inflates the mean and the standard deviation, which is why they're often better choices for skewed, real-world data.
Quartiles don't have a single universal definition. Pandas and numpy use linear interpolation by default, which is why the code above gives Q1 = 48.5K and Q3 = 63.75K. Split the same data into a lower and an upper half and take the median of each instead, a method plenty of introductory statistics courses teach, and you'd get Q1 = 48K and Q3 = 65K, for an IQR of 17K instead of 15.25K. Neither answer is wrong. They're just different conventions, so don't be alarmed if your textbook and your code disagree slightly on quartiles.
There are actually two versions of standard deviation. The population version divides by n (the count of values); the sample version divides by n−1. Pandas and most stats libraries default to the sample version, because a sample tends to slightly underestimate the true spread of the population it came from, and dividing by a smaller number corrects for that. On our 10 salaries, population std dev is $44.68K and sample std dev is $47.09K, close but not identical. This becomes important the moment you move from "describing this dataset" to "estimating something about a bigger population," which is exactly where inferential statistics picks up.
The fourth tool, frequency distributions, is just counting how often each value shows up. Say 20 users rated an app between 1 and 5 stars:
| Rating | Count | Share |
|---|---|---|
| 1 ★ | 0 | 0% |
| 2 ★ | 1 | 5% |
| 3 ★ | 4 | 20% |
| 4 ★ | 6 | 30% |
| 5 ★ | 9 | 45% |
Laid out as a table, this is a frequency distribution. Laid out as a bar chart, it's instantly readable: most people are happy, a handful are lukewarm, and almost nobody is unhappy. This is also where the mode comes from: the most frequent value in the data, 5 stars here. Mean and median describe a distribution's center; the mode describes its peak, and for categorical or heavily repeated data like ratings, it's often the most useful of the three.
This isn't just theory you check off before the "real" modeling starts. Descriptive statistics is where most of a data scientist's actual time goes.
In practice, this usually starts with one line of code. Running salaries.describe() on the series above hands you the mean, std dev, min, max, and quartiles in a single call, which is exactly why pandas' .describe() is often the very first thing a data scientist runs on a new dataset.
Reporting the mean alone, without checking the spread or plotting the data first, is one of the most common mistakes new data scientists make. If you'd only seen "average salary: $68.5K" without the raw numbers, you'd have walked away with a genuinely wrong impression of what this startup pays.
Descriptive statistics tells you about the data you have. Inferential statistics lets you go further and reason about data you don't have.
Population is the entire group you're actually interested in. Sample is the smaller, manageable subset of that group you actually collect data from. By carefully analyzing a sample, you make educated, mathematically grounded guesses about the much larger population, essential whenever collecting data from everyone is impossible or just impractical.
You can't survey every voter in a country before an election, and you can't test every light bulb a factory produces. Inferential statistics lets you draw reliable conclusions from a manageable sample instead.
Say an e-commerce company tests a new product recommendation algorithm. They show it to a randomly selected group of users for a month and measure purchase rates. Using inferential statistics, specifically a hypothesis test, they determine whether the improvement in that group is statistically significant: likely to hold across all their users, or just random chance.
A clinical trial works the same way. A new medication gets tested on 200 patients, and inferential statistics determines whether the observed benefit is strong and consistent enough to conclude the medication works for patients broadly, not just the 200 in the trial.
We'll build up hypothesis testing and confidence intervals properly over the next few posts, but it's worth seeing the actual computation once, even in miniature. Say you survey 1,000 people about monthly spend on a subscription app. The sample gives you a mean of $42 and a standard deviation of $15. How confident can you be about the true average across everyone, not just your 1,000 respondents?
That's it. That's a confidence interval. It says: if you ran this survey over and over, 95% of the intervals you'd compute this way would contain the true population average, not that there's a 95% chance the true value sits in this one specific range. That specific mix-up is common enough that we'll unpack it properly when confidence intervals get their own day.
One more note while we're here: with a sample of 1,000, the normal distribution used above and the more exact t-distribution give you practically the same critical value, 1.96 versus 1.9623, so reaching for the simpler normal approximation didn't cost any real accuracy. That gap only starts to matter with small samples, which is where the t-distribution earns its keep.
| Application | Technique |
|---|---|
| Testing if a new marketing strategy leads to higher sales | Hypothesis testing |
| Estimating a city's average income from a 1,000-person survey | Confidence intervals |
| Predicting housing prices from historical data | Regression analysis |
| Comparing test scores between two teaching methods | Two-sample t-test |
You can't fully understand inferential statistics without probability. It's the mathematical engine behind statistical inference, giving you a principled way to reason about uncertainty.
When someone says a result is "statistically significant at the 5% level," that's a probability statement. When you construct a 95% confidence interval, probability defines what that 95% actually means. When a spam filter flags an email, it's outputting a probability: if 2 out of every 100 emails it's seen labeled "spam" in the past looked like this one, it might assign this new email a 2% prior probability of being spam before it even reads the content.
Probability underpins statistical inference (how confident should you be in a conclusion), risk assessment (how likely is a bad outcome), predictive modeling (what's the most likely future value), and anomaly detection (how unusual is this data point, statistically speaking). It's a big enough topic that it gets its own dedicated post later in this series.
| Scenario | Descriptive | Inferential |
|---|---|---|
| Student exam scores | "The average score in Class A is 74, with a std dev of 8." | "Is Class A's performance significantly better than Class B's?" |
| Product sales | "Last month's sales peaked on weekends, with a median of $320/day." | "Will this weekend trend hold for next quarter?" |
| Healthcare | "40% of patients in this trial reported side effects." | "We estimate 35–45% of all patients on this drug experience side effects." |
| Customer survey | "In our survey of 500 customers, 65% rated us 4+ stars." | "We're 95% confident that between 64–72% of all customers would rate us 4+ stars." |
| App engagement | "Users in our sample opened the app 5.2 times a day on average." | "Does a new notification design actually increase daily opens, or is the change just noise?" |
Statistics isn't about memorizing formulas. It's about developing intuition for data. Descriptive statistics gets you to know your data deeply; inferential statistics lets you make confident, honest claims about things you haven't directly observed.
Start with descriptive statistics, always. Compute summary measures, plot distributions, check for outliers and missing values. Only once you actually understand what you're working with should you move on to inferential techniques and model building. Skipping this step is one of the most common mistakes new data scientists make. Now you've seen exactly why, with one $200K salary as proof.
A few questions to test whether this stuck. No pressure: the answers are one click away.
Because the mean uses every value in its calculation, so the single $200K outlier pulls it upward. The median only depends on the middle position in a sorted list, so it's far less sensitive to extreme values.
What's the spread, and what does the distribution actually look like? An average alone tells you almost nothing: the same "50" could mean tightly clustered values around 50, or a wildly skewed distribution being pulled toward one side, the same way the $68.5K average salary looked nothing like what most of the ten employees actually earned.
Inferential. It's using a sample (1,200 voters) to make a claim about a much larger population (the country), with an explicit range that accounts for uncertainty. A purely descriptive version would just report the poll's raw number: "58% of our 1,200 respondents support the policy."
We'll dive into Types of Data: the difference between numerical, categorical, ordinal, and nominal data, and why that distinction decides which statistical tools you're even allowed to use.