This guide combines everything from the ground up: what NumPy is and why it's so fast, how to create and manipulate arrays, all the mathematical and statistical operations, how to read and write real data files, and finally a full real-world case study analyzing ride-sharing data the way a working data scientist would. Throughout, you'll find visual diagrams so you can actually see what's happening to the numbers, no plotting code, just pictures made of text.
Part 1: Why NumPy Exists
Before any syntax, it helps to understand the problem NumPy solves. Python lists are flexible, but slow for math. NumPy arrays are the opposite: a little stricter, but blazingly fast.
The Core Advantages
NumPy gives us four big wins: memory efficiency, speed, a huge library of ready-made functions, and ease of use. The speed comes from the fact that NumPy is written in the C language under the hood.
To use it, we import it with its near-universal nickname np:
import numpy as np
The "Square Every Number" Problem
Suppose you have a list of numbers and want to square each one. With a plain Python list, this doesn't work the way you'd hope:
list1 = [10, 20, 30, 45, 65, 87, 99]
list1 ** 2 # ERROR — you can't square a whole list at once
You'd have to write a loop:
twice = []
for val in list1:
twice.append(val ** 2)
But with NumPy, you just... square the array. This is called element-wise operation, and it's the single biggest day-to-day convenience NumPy offers:
arr1 = np.array(list1) # convert the list into a NumPy array
print(arr1 ** 2) # squares every element at once — clean and simple
Python list (manual loop) NumPy array (element-wise)
[10, 20, 30, 45] [10, 20, 30, 45]
│ │ │ │ │ │ │ │
loop loop loop loop (all at once)
↓ ↓ ↓ ↓ ↓ ↓ ↓ ↓
[100,400,900,2025] [100,400,900,2025]
The Speed Advantage
The real story is speed. Using the %timeit tool to measure performance on a million numbers tells the tale:
l = range(1000000)
%timeit [i**2 for i in l] # Python list: ~25.3 milliseconds per loop
l_arr = np.array(range(1000000))
%timeit l_arr**2 # NumPy array: ~485 microseconds per loop
That's roughly 50x faster. Why such a dramatic difference?
- NumPy arrays are densely packed in memory because every element is the same type.
- NumPy uses SIMD instructions (Single Instruction, Multiple Data). One CPU instruction that operates on several numbers at once, like running four additions in the time it takes Python to run one.
- The core operations are implemented in C, which runs far faster than interpreted Python.
Part 2: How NumPy Works Under the Hood
This is the most important concept for understanding NumPy rather than just using it. The key question: why is a NumPy array so much faster than a Python list?
The Analogy: A Scattered Library vs. a Packed Bookshelf
In a Python list, you can mix types: integers, strings, floats, all together. But there's a hidden cost. The actual values live scattered all over your computer's memory (RAM). The list itself only holds references (think of them as addresses, R1, R2, R3...) pointing to where each value really lives.
So reading one number is a two-step trip: the computer looks at the list to find the address, then travels to that address in RAM to fetch the actual value.
A C array (which is what NumPy actually is) stores all the values of the same type packed together in one continuous block of memory. Reading a value is a single step. You go straight to it by its position.
PYTHON LIST (references scattered in memory)
list ──► [ R1 ][ R2 ][ R3 ][ R4 ]
│ │ │ │
▼ ▼ ▼ ▼
(10) (20) (30) (40) ← values scattered everywhere
two steps to read each value
NUMPY ARRAY (values packed together — a C array)
array ──► [ 10 ][ 20 ][ 30 ][ 40 ] ← all values side by side
one step to read each value
The trade-off: because everything is packed together by type, a NumPy array can only hold one data type (all integers, or all floats). It loses the flexibility of a Python list, but gains enormous speed. In short: NumPy arrays are not really Python lists. They are C arrays wearing a Python costume.
Part 3: Creating Arrays
From a List, and the Three "Dimensions"
The simplest way to make an array is from a list. And lists-within-lists make higher dimensions:
np.array([10, 20, 30]) # 1D array
np.array([[10, 20], [30, 40], [60, 70]]) # 2D array
NumPy has special names for arrays of different dimensions, borrowed from math:
1D = VECTOR 2D = MATRIX 3D+ = TENSOR
(a single row) (rows & columns) (a stack of matrices)
[10 20 30] [[10 20] ┌─────────┐
[30 40] ┌─────────┐│
[60 70]] │ [[..]] ││
│ [[..]] │┘
└─────────┘
Building Sequences with arange()
np.arange() is NumPy's version of Python's range() , which builds a sequence from a start to an end, stepping by some amount. The format is arange(start, end, step), and importantly, the end value is not included:
np.arange(1, 5) # [1 2 3 4] — stops before 5
np.arange(1, 5, 2) # [1 3] — step of 2
np.arange(1, 5, 0.5) # [1. 1.5 2. ... 4.5] — float steps allowed!
Unlike Python's range(), np.arange() happily accepts a floating-point step size, which is genuinely useful.
Ready-Made Array Builders
You'll often want an array pre-filled with a particular pattern. NumPy has a builder for each common case:
| Function | What it builds | Mental picture |
|---|---|---|
np.ones((2,3)) |
All ones | a grid full of 1s |
np.zeros((2,3)) |
All zeros | a blank grid |
np.full((2,3), 7) |
All the same value | a grid full of 7s |
np.identity(3) |
Identity matrix | 1s on the diagonal, 0s elsewhere |
np.eye(3) |
Like identity (more flexible) | 1s on a diagonal |
np.diag([1,2,3]) |
Values on the diagonal | your numbers down the diagonal |
np.random.random((2,2)) |
Random floats 0–1 | a grid of random decimals |
np.random.randint(10,51,size=15) |
Random integers in a range | random whole numbers |
The identity matrix is worth picturing, since it shows up constantly in math:
np.identity(3)
┌ ┐
│ 1 0 0 │
│ 0 1 0 │ ← 1s run down the diagonal, 0s everywhere else
│ 0 0 1 │
└ ┘
Part 4: Indexing and Slicing
Once you have an array, you need to grab pieces of it. Let's use a random array of 15 numbers:
mat1 = np.random.randint(10, 51, size=15)
Basic Access
mat1[5] # the element at position 5
mat1[16] # ERROR — index out of bounds (there's no position 16)
mat1[3:10] # positions 3 through 9
mat1[2:11:2] # positions 2 to 10, stepping by 2
mat1[::-1] # the whole array, reversed
mat1[[2,5,3,7]] # fancy indexing — grab specific positions in any order
Negative Indexing
NumPy lets you count from the end using negative numbers. -1 is the last element:
mat1[-1] # last element
mat1[-5:-1] # from 5th-from-end up to (not including) last
mat1[-1:-5:-1] # last four elements, in reverse
Positive index: 0 1 2 3 4
Array: [ 42 17 33 25 48 ]
Negative index: -5 -4 -3 -2 -1
▲ ▲
first mat1[-1] = 48
Part 5: Masking (Boolean / Fancy Indexing)
This is one of NumPy's superpowers and a concept you'll use every single day in data analysis: filtering an array by a condition.
When you write a comparison against an array, NumPy returns an array of True/False values. One per element, a boolean mask:
mat1 > 25 # [True, False, True, True, ...] (a mask)
Think of the mask as a stencil laid over your data. Where the stencil says True, the value shows through; where it says False, it's filtered out:
data: [ 42 17 33 25 48 12 ]
mask: [ T F T F T F ] ← (data > 25)
───────────────────────────
result: [ 42 33 48 ] ← only the True values kept
mat1[mat1 > 25] # keeps only values greater than 25
mat1[(mat1 > 25) & (mat1 < 45)] # combine conditions with &
When combining conditions, use & (and), | (or), and wrap each condition in its own parentheses. This is a strict rule in NumPy.
np.where — Finding Positions
Where masking returns the values, np.where returns the index positions where a condition is true:
np.where(mat1 > 25) # the positions of values > 25
np.where((mat1 > 25) & (mat1 < 45)) # positions matching both conditions
A second, very powerful form of np.where acts like an if/else for arrays. We'll use that in the case study at the end.
Part 6: Universal Functions (Math on Whole Arrays)
"Universal functions" (ufuncs) are functions that operate on every element of an array at once. They come in two flavors: those that take two arrays, and those that transform a single array.
Two-Array Math
Let's say we have two 3×3 arrays, arr3 and arr4:
np.add(arr3, arr4) # same as arr3 + arr4
np.add(arr3, 10) # adds 10 to every element
np.subtract(arr3, arr4) # same as arr3 - arr4
np.multiply(arr3, arr4) # ELEMENT-WISE multiply (not matrix multiply!)
np.divide(arr3, arr4) # same as arr3 / arr4
There's an important distinction here. np.multiply multiplies matching positions. Matrix multiplication (the real linear-algebra operation) is a different function, np.dot:
ELEMENT-WISE (np.multiply) MATRIX MULTIPLY (np.dot)
[a b] [e f] [a·e b·f] [a b] [e f] [ae+bg af+bh]
[c d] × [g h] = [c·g d·h] [c d] · [g h] = [ce+dg cf+dh]
each cell × matching cell rows combined with columns
np.dot(arr3, arr4) # matrix multiplication, same as arr3 @ arr4
Single-Array Transformations
These take one array (or number) and transform every value:
| Function | What it does | Example result |
|---|---|---|
np.log(10) |
natural log (base e) | 2.302 |
np.log2(10) |
log base 2 | 3.321 |
np.log10(10) |
log base 10 | 1.0 |
np.exp(5) |
e raised to the power | 148.4 |
np.sqrt(800) |
square root | 28.28 |
np.square(80) |
the square | 6400 |
np.abs(-10) |
absolute value | 10 |
np.ceil(4.56) |
round up | 5.0 |
np.floor(3.81) |
round down | 3.0 |
np.round(12.6781, 2) |
round to N decimals | 12.68 |
The ceiling/floor pair is easy to remember with a picture:
ceil → 5.0 (always rounds UP to next integer)
▲
4.56
▼
floor → 4.0 (always rounds DOWN)
Percentiles
A percentile tells you the value below which a given percentage of data falls. The 25th percentile is the value that 25% of the data sits below:
np.percentile([150,120,30,45,60,78,81,93,25,33,110], 25) # 25th percentile
np.percentile([...], 50) # 50th percentile = the median
np.percentile([...], 100) # 100th percentile = the maximum
Comparison and Validity Checks
These ufuncs compare arrays element-by-element or check for special values:
a = np.array([1, 2, 3])
b = np.array([1, 3, 2])
np.equal(a, b) # [True, False, False]
np.not_equal(a, b) # [False, True, True]
np.greater(a, b) # [False, False, True]
np.less(a, b) # [False, True, False]
NumPy has special values for "infinity" (np.inf) and "not a number / missing" (np.nan). You check for them with:
arr = np.array([1, 2, np.inf, 4, np.nan, -np.inf])
np.isfinite(arr) # True only for normal, finite numbers
np.isinf(arr) # True where the value is infinite
np.isnan(arr) # True where the value is missing (NaN)
Finally, np.any and np.all collapse a whole array into a single yes/no:
np.any(arr)→ True if at least one element is truthy (non-zero).np.all(arr)→ True only if every element is truthy.
arr = [1, 2, 0, 4]
np.any(arr) → True (at least one non-zero)
np.all(arr) → False (one value is 0, so not ALL are truthy)
Part 7: Aggregation: and the All-Important axis
Aggregation functions summarize an array down to fewer numbers: np.sum, np.min, np.max, np.mean. On their own, they collapse the entire array to a single value:
np.sum(arr3) # one number: the total of everything
np.mean(arr3) # one number: the average of everything
But the real power appears with the axis parameter, which controls which direction you collapse. This is the concept beginners struggle with most, so here's the picture:
axis=0 → collapse DOWN the columns (one result per column)
axis=1 → collapse ACROSS the rows (one result per row)
col0 col1 col2
row0 [ 1 2 3 ] ──► axis=1 sums across → 6
row1 [ 4 5 6 ] ──► axis=1 sums across → 15
│ │ │
▼ ▼ ▼
axis=0 5 7 9 (sums down each column)
np.sum(arr3, axis=0) # column totals → [5, 7, 9]
np.sum(arr3, axis=1) # row totals → [6, 15]
np.max(arr3, axis=0) # max of each column
np.mean(arr3, axis=1) # average of each row
A simple way to remember it: axis=0 squashes the rows together (vertical), axis=1 squashes the columns together (horizontal).
Part 8: Sorting
1D Arrays — Copy vs. In-Place
There's a subtle but important difference between two ways of sorting:
np.sort(arr_unsorted) # returns a SORTED COPY; original is untouched
arr_unsorted.sort() # sorts IN-PLACE; modifies the original, returns None
np.sort(a) a.sort()
┌─────────────┐ ┌─────────────┐
│ a unchanged │ │ a is now │
│ copy → new │ │ sorted │
└─────────────┘ └─────────────┘
"give me a sorted copy" "sort yourself"
argsort — Sorting by Position
np.argsort doesn't return the sorted values; it returns the indices that would sort the array. This is incredibly useful when you want to reorder one array based on the order of another (we'll use exactly this in the case study):
np.argsort(arr_unsorted_2) # e.g. [3, 0, 1, 2] — "put position 3 first, then 0..."
2D Arrays
For 2D arrays, sorting respects the axis direction:
arr_2d = np.array([[3, 10, 2],
[1, 5, 7],
[2, 7, 5]])
np.sort(arr_2d, axis=0) # sorts within each column
np.sort(arr_2d, axis=1) # sorts within each row
Part 9: Combining and Splitting Arrays
Stacking Arrays Together
There are several ways to glue arrays together. The two foundational ones are concatenation along an axis:
np.concatenate((arr3, arr4), axis=0) # stack top-and-bottom (vertically)
np.concatenate((arr3, arr4), axis=1) # stack side-by-side (horizontally)
NumPy provides friendlier named shortcuts that do the same thing:
vstack / concatenate axis=0 hstack / concatenate axis=1
(stack vertically) (stack horizontally)
[A A A] [A A A] [B B B]
[A A A] +B → [A A A] [B B B]
[B B B] ← B added below ▲ A and B side by side
[B B B]
| Function | What it does |
|---|---|
np.vstack((a, b)) |
stack vertically (same as concatenate axis=0) |
np.hstack((a, b)) |
stack horizontally (same as concatenate axis=1) |
np.column_stack((a, b)) |
stack 1D arrays as columns |
Splitting Arrays Apart
np.split breaks one array into several pieces. The second argument tells it how:
- An integer
nsplits the array into n equal parts. - A list of indices splits at those exact positions.
arr_big = np.random.randint(10, 51, size=(12, 10))
np.split(arr_big, 2, axis=1) # split into 2 along columns
np.split(arr_big, 4, axis=0) # split into 4 along rows
np.hsplit(arr_big, 2) # horizontal split shortcut
np.vsplit(arr_big, 4) # vertical split shortcut
np.hsplit(arr, 2) np.vsplit(arr, 2)
┌────┬────┐ ┌─────────┐
│ L │ R │ → [L] [R] │ TOP │ → [TOP]
│ │ │ ├─────────┤ [BOTTOM]
└────┴────┘ │ BOTTOM │
└─────────┘
Part 10: Inspecting an Array
When you receive an unfamiliar array, these attributes tell you everything about its structure. Notice they have no parentheses. They're descriptions of the object, not actions:
arr_big.shape # (rows, columns) — its dimensions
arr_big.size # total number of elements
arr_big.ndim # number of dimensions (1, 2, 3...)
arr_big.nbytes # total memory used, in bytes
arr_big.dtype # the data type of the elements (e.g. int64)
arr_big (a 12 × 10 array of integers)
┌────────────────────────────┐
│ .shape → (12, 10) │
│ .size → 120 │
│ .ndim → 2 │
│ .nbytes → 960 (120 × 8) │
│ .dtype → int64 │
└────────────────────────────┘
Part 11: Reshaping: Changing the Form Without Changing the Data
Reshaping rearranges the same numbers into a different grid. The total number of elements must stay the same.
arr_big.reshape(24, 5) # 120 elements → 24 rows × 5 cols (24×5 = 120 ✓)
The Magic -1
If you put -1 for one dimension, NumPy figures it out for you. If you have 120 elements and ask for 15 columns, NumPy works out you need 8 rows:
arr_big.reshape(-1, 15) # "15 columns, you figure out the rows" → 8 rows
120 elements, reshape(-1, 15)
"I want 15 columns. How many rows?"
120 ÷ 15 = 8 → NumPy fills in 8 automatically
Transpose — Flipping Rows and Columns
transpose() (or the shortcut .T) swaps rows and columns:
arr_big.transpose() # rows become columns, columns become rows
arr_big.T # identical, shorter
original transpose (.T)
[1 2 3] [1 4]
[4 5 6] → [2 5]
[3 6]
2 rows, 3 cols 3 rows, 2 cols
flatten vs ravel — A Critical Distinction
Both turn a multi-dimensional array into a flat 1D array, but they differ in a way that can cause real bugs if you don't understand it:
ravel()returns a view when possible. If the array is stored contiguously in memory (the common case), ravel shares the original data. For non-contiguous arrays (e.g. after a transpose), ravel has to make a copy instead. Either way, changing a ravel result may change the original.flatten()always makes a copy, regardless of memory layout. Changing the result never touches the original.
arr = np.array([[10, 20], [30, 40]])
rav = arr.ravel()
rav[0] = 99 # this ALSO changes arr! (shared memory)
flat = arr.flatten()
flat[1] = 88 # arr stays the same (independent copy)
ravel() → VIEW (window onto the same data)
┌──────────┐ change the view →
│ original │ ◄─────── change shows up in original too
└──────────┘
flatten() → COPY (separate data)
┌──────────┐ copy ┌──────────┐
│ original │ ──────► │ copy │ change here →
└──────────┘ └──────────┘ original untouched
Part 12: Shallow vs Deep Copy (Shared Memory)
This builds on the ravel/flatten idea and is one of NumPy's most important, and most surprising behaviors. NumPy avoids copying data whenever it can. Reshaping and slicing often hand you a new view of the same underlying data, not a fresh copy.
Watch what happens with a reshape:
a = np.arange(4) # [0 1 2 3]
b = a.reshape(2, 2) # b is a VIEW of a
a[0] = 10 # change a...
# ...and b changes too! They share the same data.
The same surprise happens with slicing:
mat1_new = mat1[0:6] # a slice — shares memory with mat1
mat1_new[2] = 25 # changing the slice ALSO changes mat1
a ──┐
├──► [ shared data block ] ← changing through either name
b ──┘ changes what BOTH see
This is a SHALLOW COPY: the data is reused, not duplicated.
Deep Copy — When You Want True Independence
If you want a genuinely separate array whose changes won't leak back, use .copy(). This is a deep copy: its own memory, fully independent.
original = np.array([1, 2, 3, 4])
deep_copy = original.copy()
deep_copy[0] = 77 # original stays [1, 2, 3, 4]
np.shares_memory(original, deep_copy) # False — proof they're separate
The takeaway: NumPy uses shallow copies (shared data) for cheap operations like reshape and slicing to stay fast. When you make permanent changes that should not affect the original, use .copy(). And whenever you're unsure whether two arrays share memory, np.shares_memory(a, b) will tell you.
Part 13: Broadcasting: Math Between Different-Shaped Arrays
Normally, doing math between two arrays requires them to be the same shape. Broadcasting is NumPy's clever system for stretching a smaller array to match a bigger one automatically, so the math just works.
The helper np.tile literally repeats an array to build bigger ones. To build the 4×3 matrix shown in the diagram below, we tile a column vector across 3 columns:
col = np.arange(0, 40, 10).reshape(4, 1) # column vector: [[0],[10],[20],[30]]
big_matrix = np.tile(col, (1, 3)) # repeat 3 times sideways → shape (4, 3)
But broadcasting does this stretching implicitly. Picture adding a single row to a full matrix:
Big matrix (4×3) Small row (1×3)
[ 0 0 0 ] [ 0 1 2 ]
[10 10 10 ] +
[20 20 20 ]
[30 30 30 ]
Broadcasting "stretches" the small row down to all 4 rows:
[ 0 0 0 ] [ 0 1 2 ] [ 0 1 2 ]
[10 10 10 ] + [ 0 1 2 ] = [10 11 12 ]
[20 20 20 ] [ 0 1 2 ] [20 21 22 ]
[30 30 30 ] [ 0 1 2 ] [30 31 32 ]
▲ replicated automatically
The single row [0, 1, 2] is broadcast across all four rows without you ever writing a loop or manually copying it. This is what makes NumPy code so concise. Operations that would need nested loops in plain Python become a single +.
Part 14: Getting Data In and Out of Files
So far we've built arrays by hand. In the real world, data lives in files. NumPy can read it in and write it back out.
Reading Text Files
There are two main loaders, and the choice depends on whether your data has gaps:
data = np.loadtxt('data.txt', delimiter="\t") # use when NO missing values
data = np.genfromtxt('data_miss.txt', delimiter="\t") # handles missing values
np.loadtxt is strict, which errors out if any value is missing. np.genfromtxt is forgiving , filling any gaps with nan. The rule of thumb:
Clean data, no gaps → np.loadtxt (fast, strict)
Real data with gaps → np.genfromtxt (fills blanks with NaN)
Try it on a file that actually has a gap and np.loadtxt stops cold instead of guessing:
data = np.loadtxt('data_miss.txt', delimiter="\t")
# ValueError: could not convert string '' to float64 at row 8, column 3.
Swap in np.genfromtxt on that exact same file and it loads without complaint, dropping a nan into the empty cell instead of raising:
data = np.genfromtxt('data_miss.txt', delimiter="\t") # the gap becomes nan, the rest loads fine
Reading CSVs with Mixed Types
CSV files often mix text and numbers (e.g. a customer's name and age). By default NumPy assumes everything is a float, which mangles text columns. Loading the Mall Customers file (which has a text "Genre" column) the naive way turns that whole column into nan, since "Male" and "Female" simply can't be parsed as numbers:
data = np.genfromtxt('Mall_Customers.csv', delimiter=",") # default dtype is float — the Genre column becomes nan
Forcing dtype='object' stops the column from being wiped out, but it creates a different headache: every single value, numbers included, comes back as a Python bytes object instead of a real number:
data = np.genfromtxt('Mall_Customers.csv', delimiter=",", dtype='object', skip_header=1)
data[0:5] # → array([[b'1', b'Male', b'19', b'15', b'39'], ...], dtype=object)
Those b'19'-style values look like numbers but are bytes underneath, so any arithmetic on them blows up:
# try dividing 4th col with 3rd col and we get an error
data[:, 3] / data[:, 2]
# TypeError: unsupported operand type(s) for /: 'bytes' and 'bytes'
The clean solution is to define a structured dtype that names each column and gives it the right type up front:
dt = np.dtype({'names': ["CustomerID", "Genre", "Age", "Annual_Income", "Spending_Score"],
'formats': [np.int16, 'U16', np.int16, np.int16, np.int16]})
data = np.genfromtxt('Mall_Customers.csv', delimiter=",", dtype=dt, skip_header=1)
Here skip_header=1 skips the header row, 'U16' means "text up to 16 characters," and each numeric column gets a proper integer type. Check data.ndim afterward and you'll see it's 1, not 2 — this is a 1D array of labeled records, one per row, not a grid. (In practice, this is exactly the kind of job where most people reach for Pandas instead, but it's good to know NumPy can do it — and the result drops straight into a DataFrame with pd.DataFrame(data) if you want to take it from there.)
Saving and Loading Arrays
To save arrays in NumPy's own fast binary format:
np.save('output.npy', data) # ONE array → .npy file
np.savez('outputs.npz', data) # MULTIPLE arrays → .npz file
a = np.load('output.npy') # load it back
.npy → a single array
.npz → several arrays in one file (named arr_0, arr_1, ...)
Part 15: Putting It All Together: A Real Case Study
Let's see why all of this matters. Imagine you're a data scientist at a major ride-sharing company (like Uber or Lyft). The city operations team hands you a batch of ride data from a busy Friday night and wants deep insights into driver efficiency, surge pricing dynamics, and rider satisfaction, all in service of optimizing the matching algorithm. Every technique below uses something from this guide.
Loading and Splitting the Columns
data = np.genfromtxt('ride_data.csv', delimiter=',', skip_header=1)
ride_ids = data[:, 0] # column slicing pulls out each variable
driver_ids = data[:, 1]
distances = data[:, 2] # distance in miles
durations = data[:, 3] # duration in minutes
fares = data[:, 4] # fare in USD
surges = data[:, 5] # surge multiplier (e.g. 1.2 = 20% extra)
d_ratings = data[:, 6] # driver rating
r_ratings = data[:, 7] # rider rating
Notice data[:, 0]. The slice "all rows, column 0." This is 2D indexing in action. Here's what each column actually holds:
ride_id— the unique ID given to every ride.driver_id— the unique ID of each driver.distance_miles— distance covered from origin to destination.duration_minutes— how long the ride took.fare_usd— the total fare of the ride, in US dollars.surge_multiplier— activates during high demand; a multiplier of 1.2 means the rider pays the normal fare × 1.2.driver_rating— the rating the rider gave the driver.rider_rating— the rating the driver gave the rider.
Revenue: GMV and the Platform's Cut
gmv = np.sum(fares) # Gross Merchandise Value — every dollar paid
platform_revenue = gmv * 0.25 # the company keeps a 25% commission
print(f"Total GMV: ${gmv:.2f} | Platform Cut: ${platform_revenue:.2f}")
# Total GMV: $503.50 | Platform Cut: $125.88
For this batch of rides, the platform collected $503.50 in total fares and kept $125.88 of it as commission. GMV (Gross Merchandise Value) is the total of every dollar that flowed through the platform before anyone takes their cut. The analogy: if you build an app connecting bakers to customers and someone buys a $100 cake, your revenue might be a $10 fee, but the GMV of that sale is the full $100. Big platforms (Uber, Airbnb, Amazon, eBay) live and die by GMV because it's the ultimate vanity metric for scale — a $1 billion GMV means the platform controls a massive chunk of that market. It also dictates the platform's ceiling: you can't grow revenue without either taking a bigger cut (which angers the supply side, e.g. drivers) or growing the overall GMV.
Speed Check (Element-wise Math + Masking)
speeds_mph = distances / (durations / 60) # element-wise: speed for every ride
speeding_rides = ride_ids[speeds_mph > 65] # masking: which rides exceeded 65 mph
print("Speeding Ride IDs:", speeding_rides.astype(int)) # Speeding Ride IDs: [16]
One line computes speed for every ride at once (element-wise division), and a boolean mask instantly isolates the suspicious ones — here, a single ride (ID 16) came back clocking over 65 mph, which is worth a manual look (bad GPS data, a highway trip misclassified, etc.).
Surge Revenue (Reverse-Engineering with Division)
base_fares = fares / surges # undo the surge to find the base fare
surge_extra_revenue = np.sum(fares - base_fares) # how much was pure surge premium
# Extra Revenue from Surge: $57.42
Surge pricing is a dynamic pricing model: the surge multiplier is the exact number the base fare gets multiplied by (1.5x, 2x, ...) during periods of high demand. Dividing the fare by the surge factor reverse-engineers what the ride would have cost without it, so the gap between the two is pure surge revenue — $57.42 of it here. It exists to solve a simple supply-and-demand problem, and it does two things at once:
- Incentivizes supply — a "2.0x surge in that area" notification pulls more drivers toward the high-demand zone.
- Controls demand — it discourages riders who don't urgently need a ride right now, freeing up cars for people willing to pay the premium (someone rushing to the airport, say).
Statistics: Median, Argmax, Percentiles
median_rating = np.median(d_ratings) # the typical driver rating → 4.7
max_fare_idx = np.argmax(fares) # POSITION of the single highest fare
print(f"Highest fare was ${fares[max_fare_idx]} on Ride ID {int(ride_ids[max_fare_idx])}")
# Highest fare was $85.0 on Ride ID 12
p90_fare = np.percentile(fares, 90) # the "high-value" threshold → $54.00
np.argmax returns the index of the largest value, so ride_ids[max_fare_idx] tells you exactly which ride earned the most (Ride 12, at $85). The 90th percentile is the dollar amount that separates the top 10% most expensive rides from the bottom 90% — here, $54.00 means 90% of rides cost less than that, and only the priciest tenth cost more. A ride can cross that line for several reasons: a long-distance airport run, a 3.0x surge, or a premium vehicle tier. Data scientists calculate it for a few concrete reasons:
- Operations — route the best, 5-star drivers to these high-value riders to protect the experience.
- Marketing — run a targeted campaign, e.g. "spend over $55 on a ride and get a free coffee voucher."
- Anomaly detection — a ride way past this threshold (say, $400) is worth a manual check in case the GPS glitched or the driver took a longer route than necessary.
Quick Ratios with np.mean on a Mask
short_trips_pct = np.mean(distances < 3.0) * 100 # % of rides under 3 miles → 25.0%
This is a neat trick: distances < 3.0 is a mask of True/False, and since True counts as 1, taking the mean of a mask gives you the proportion that satisfied the condition — a quarter of the night's rides here. Multiply by 100 for a percentage.
Per-Ride Calculations and Perfect Matches
rating_diff = np.abs(d_ratings - r_ratings) # disagreement per trip
dollars_per_minute = fares / durations # earning efficiency per ride
perfect_mask = (d_ratings == 5.0) & (r_ratings == 5.0) # both gave 5 stars
perfect_rides = ride_ids[perfect_mask]
print("Perfect Match Ride IDs:", perfect_rides.astype(int)) # Perfect Match Ride IDs: [10]
Three independent, per-ride calculations in one breath: rating_diff flags trips where driver and rider disagreed on quality, dollars_per_minute measures how efficiently each ride earned, and the perfect_mask boolean filter isolates the rides where both sides gave a flawless 5.0 — only Ride 10 managed that here.
np.where as If/Else: Driver Payouts with a Penalty
Here's the powerful second form of np.where, which works like an if/else across the whole array at once:
standard_payout = fares * 0.75
adjusted_payout = np.where(d_ratings < 4.0, # CONDITION
standard_payout - 3.00, # value if True (penalty)
standard_payout) # value if False
# Adjusted Payouts (First 5): [13.88 33.75 6.75 24. 9.38]
np.where(condition, value_if_true, value_if_false)
driver rating < 4.0 ?
├─ YES → payout − $3 penalty
└─ NO → normal payout
...applied to every driver simultaneously
Running Totals and Counting Uniques
cum_fares = np.cumsum(fares) # running total as the night progresses — ends at 403.5
unique_drivers, counts = np.unique(driver_ids, return_counts=True)
busiest_driver = unique_drivers[np.argmax(counts)] # most rides completed → Driver 101
np.cumsum gives a cumulative sum: fare after ride 1, after rides 1+2, and so on. np.unique with return_counts=True lists each distinct driver alongside how many rides they did, and np.argmax finds the busiest.
Sorting One Array by Another (argsort in action)
sorted_ride_ids = ride_ids[np.argsort(durations)] # ride IDs, shortest trip first
This is the classic argsort pattern: sort the durations, get the order of positions that would sort them, then apply that same order to the ride IDs array instead. It looks like a basic housekeeping task, but ordering rides by duration immediately surfaces two categories that matter a lot to the business: the shortest trips and the longest ones.
1. Analyzing "micro-trips" (the shortest durations). At the top of the sorted list — rides taking 2 to 5 minutes — you're looking at trips that are often inefficient for cars. A driver might spend 8 minutes fighting through city traffic just to pick someone up for a 2-minute ride; they earn very little and feel like their time was wasted. Once the platform can see the volume of these micro-trips, it can respond: introduce a "minimum fare" so drivers earn a fair amount regardless of trip length, or — if thousands of 3-minute rides cluster in one area — deploy a fleet of e-bikes or scooters there instead, nudging the app to suggest: "A car is $8 and takes 5 mins. A scooter is $2 and takes 4 mins."
2. Managing "deadheading" (the longest durations). At the bottom of the sorted list — 45+ minute rides — you're looking at rare but high-revenue outliers that create a logistical headache. A driver who takes a 60-minute ride out of the busy city center then has to drive 60 minutes back empty, since there's no demand where they dropped off. Driving empty like this is called "deadheading," and it costs the driver fuel and time with no fare to show for it. Spotting these long-duration rides is what led platforms to build features like a "Destination Filter" (matching a driver heading back to the city with a rider who wants to go the same way) and a "long trip premium" fee that compensates the driver for the empty return leg.
Reshaping into a 3D Tensor for Machine Learning
ratings_subset = np.vstack((d_ratings[:12], r_ratings[:12])) # stack into 2×12
tensor_3d = ratings_subset.reshape(2, 2, 6) # reshape to 2×2×6
print("3D Tensor Shape:", tensor_3d.shape) # 3D Tensor Shape: (2, 2, 6)
The ML team needs the first 12 driver and rider ratings reshaped into a 3D tensor for a neural network input. We vstack the two rating arrays together into a 2×12 grid, then reshape those same 24 values into a 2×2×6 cube — combining stacking and reshaping, both covered earlier, without changing a single value.
Quick Reference Summary
| Task | Code |
|---|---|
| Import NumPy | import numpy as np |
| Make array from list | np.array([1, 2, 3]) |
| Sequence | np.arange(start, end, step) |
| All ones / zeros | np.ones((r,c)) / np.zeros((r,c)) |
| Filled with a value | np.full((r,c), val) |
| Identity matrix | np.identity(n) |
| Random integers | np.random.randint(low, high, size) |
| Indexing / slicing | arr[2], arr[1:5], arr[::-1] |
| Negative index | arr[-1] (last element) |
| Masking (filter) | arr[arr > 25] |
| Combine conditions | arr[(a > 1) & (a < 9)] |
| Find positions | np.where(arr > 25) |
| If/else over array | np.where(cond, if_true, if_false) |
| Element-wise multiply | np.multiply(a, b) or a * b |
| Matrix multiply | np.dot(a, b) or a @ b |
| Rounding | np.ceil(), np.floor(), np.round(x, 2) |
| Percentile / median | np.percentile(a, 90) / np.median(a) |
| Check missing/infinite | np.isnan(a), np.isinf(a) |
| Any / all true | np.any(a) / np.all(a) |
| Aggregate by column | np.sum(a, axis=0) |
| Aggregate by row | np.sum(a, axis=1) |
| Sort (copy) | np.sort(a) |
| Sort (in-place) | a.sort() |
| Sort indices | np.argsort(a) |
| Index of max/min | np.argmax(a) / np.argmin(a) |
| Running total | np.cumsum(a) |
| Unique + counts | np.unique(a, return_counts=True) |
| Stack vertical / horizontal | np.vstack((a,b)) / np.hstack((a,b)) |
| Split | np.split(a, n, axis=) |
| Shape / size / dims | a.shape, a.size, a.ndim |
| Reshape (auto dim) | a.reshape(-1, cols) |
| Transpose | a.T |
| Flatten (copy) / ravel (view) | a.flatten() / a.ravel() |
| Deep copy | a.copy() |
| Check shared memory | np.shares_memory(a, b) |
| Read text / with gaps | np.loadtxt() / np.genfromtxt() |
| Save / load | np.save() / np.load() |
NumPy is the bedrock of the entire Python data science world. Once arrays feel natural: creating them, slicing them, filtering with masks, doing element-wise math, reshaping, and understanding when data is shared versus copied, everything built on top of it (Pandas, scikit-learn, deep learning frameworks) becomes far easier to reason about. The best way to make it stick is to grab a real CSV and try the case-study operations on it yourself: compute some sums by axis, filter rows with a mask, sort one column by another with argsort, and watch how a single line of NumPy replaces what would have been a tangle of loops.
NumPy is the bedrock of the entire Python data science world. Once arrays feel natural — creating them, slicing them, filtering with masks, doing element-wise math, reshaping, and understanding when data is shared versus copied — everything built on top of it becomes far easier to reason about. Grab a real CSV and try the case-study operations yourself: compute sums by axis, filter rows with a mask, sort one column by another with argsort. A single line of NumPy replaces what would have been a tangle of loops.