§12  NumPy
Python Data Science Series  ·  Article 12

NumPy: The Complete Beginner's Guide to Numerical Python

If Pandas is the workhorse of data analysis, NumPy is the engine sitting underneath it. NumPy (short for Numerical Python) is the foundational library for numerical computing in Python , which lets you do fast math on huge collections of numbers, and almost every other data science tool (Pandas, scikit-learn, TensorFlow) is built on top of it.

Why NumPy is 50× faster
Masking & Boolean indexing
Broadcasting explained visually
Views vs copies demystified
Full ride-sharing case study

This guide combines everything from the ground up: what NumPy is and why it's so fast, how to create and manipulate arrays, all the mathematical and statistical operations, how to read and write real data files, and finally a full real-world case study analyzing ride-sharing data the way a working data scientist would. Throughout, you'll find visual diagrams so you can actually see what's happening to the numbers, no plotting code, just pictures made of text.


Part 1: Why NumPy Exists

Before any syntax, it helps to understand the problem NumPy solves. Python lists are flexible, but slow for math. NumPy arrays are the opposite: a little stricter, but blazingly fast.

The Core Advantages

NumPy gives us four big wins: memory efficiency, speed, a huge library of ready-made functions, and ease of use. The speed comes from the fact that NumPy is written in the C language under the hood.

To use it, we import it with its near-universal nickname np:

python
import numpy as np

The "Square Every Number" Problem

Suppose you have a list of numbers and want to square each one. With a plain Python list, this doesn't work the way you'd hope:

python
list1 = [10, 20, 30, 45, 65, 87, 99]
list1 ** 2     # ERROR — you can't square a whole list at once

You'd have to write a loop:

python
twice = []
for val in list1:
    twice.append(val ** 2)

But with NumPy, you just... square the array. This is called element-wise operation, and it's the single biggest day-to-day convenience NumPy offers:

python
arr1 = np.array(list1)   # convert the list into a NumPy array
print(arr1 ** 2)         # squares every element at once — clean and simple
   Python list (manual loop)          NumPy array (element-wise)

   [10, 20, 30, 45]                   [10, 20, 30, 45]
     │   │   │   │                       │   │   │   │
   loop loop loop loop                  (all at once)
     ↓   ↓   ↓   ↓                       ↓   ↓   ↓   ↓
   [100,400,900,2025]                  [100,400,900,2025]

The Speed Advantage

The real story is speed. Using the %timeit tool to measure performance on a million numbers tells the tale:

python
l = range(1000000)
%timeit [i**2 for i in l]        # Python list: ~25.3 milliseconds per loop

l_arr = np.array(range(1000000))
%timeit l_arr**2                 # NumPy array: ~485 microseconds per loop

That's roughly 50x faster. Why such a dramatic difference?

  • NumPy arrays are densely packed in memory because every element is the same type.
  • NumPy uses SIMD instructions (Single Instruction, Multiple Data). One CPU instruction that operates on several numbers at once, like running four additions in the time it takes Python to run one.
  • The core operations are implemented in C, which runs far faster than interpreted Python.

Part 2: How NumPy Works Under the Hood

This is the most important concept for understanding NumPy rather than just using it. The key question: why is a NumPy array so much faster than a Python list?

The Analogy: A Scattered Library vs. a Packed Bookshelf

In a Python list, you can mix types: integers, strings, floats, all together. But there's a hidden cost. The actual values live scattered all over your computer's memory (RAM). The list itself only holds references (think of them as addresses, R1, R2, R3...) pointing to where each value really lives.

So reading one number is a two-step trip: the computer looks at the list to find the address, then travels to that address in RAM to fetch the actual value.

A C array (which is what NumPy actually is) stores all the values of the same type packed together in one continuous block of memory. Reading a value is a single step. You go straight to it by its position.

   PYTHON LIST  (references scattered in memory)

   list ──► [ R1 ][ R2 ][ R3 ][ R4 ]
              │     │     │     │
              ▼     ▼     ▼     ▼
            (10)  (20)  (30)  (40)      ← values scattered everywhere
            two steps to read each value


   NUMPY ARRAY  (values packed together — a C array)

   array ──► [ 10 ][ 20 ][ 30 ][ 40 ]   ← all values side by side
             one step to read each value

The trade-off: because everything is packed together by type, a NumPy array can only hold one data type (all integers, or all floats). It loses the flexibility of a Python list, but gains enormous speed. In short: NumPy arrays are not really Python lists. They are C arrays wearing a Python costume.


Part 3: Creating Arrays

From a List, and the Three "Dimensions"

The simplest way to make an array is from a list. And lists-within-lists make higher dimensions:

python
np.array([10, 20, 30])                       # 1D array
np.array([[10, 20], [30, 40], [60, 70]])     # 2D array

NumPy has special names for arrays of different dimensions, borrowed from math:

   1D = VECTOR          2D = MATRIX            3D+ = TENSOR
   (a single row)       (rows & columns)       (a stack of matrices)

   [10 20 30]           [[10 20]                  ┌─────────┐
                         [30 40]                 ┌─────────┐│
                         [60 70]]                │ [[..]]  ││
                                                 │ [[..]]  │┘
                                                 └─────────┘

Building Sequences with arange()

np.arange() is NumPy's version of Python's range() , which builds a sequence from a start to an end, stepping by some amount. The format is arange(start, end, step), and importantly, the end value is not included:

python
np.arange(1, 5)         # [1 2 3 4]      — stops before 5
np.arange(1, 5, 2)      # [1 3]          — step of 2
np.arange(1, 5, 0.5)    # [1. 1.5 2. ... 4.5]  — float steps allowed!

Unlike Python's range(), np.arange() happily accepts a floating-point step size, which is genuinely useful.

Ready-Made Array Builders

You'll often want an array pre-filled with a particular pattern. NumPy has a builder for each common case:

Function What it builds Mental picture
np.ones((2,3)) All ones a grid full of 1s
np.zeros((2,3)) All zeros a blank grid
np.full((2,3), 7) All the same value a grid full of 7s
np.identity(3) Identity matrix 1s on the diagonal, 0s elsewhere
np.eye(3) Like identity (more flexible) 1s on a diagonal
np.diag([1,2,3]) Values on the diagonal your numbers down the diagonal
np.random.random((2,2)) Random floats 0–1 a grid of random decimals
np.random.randint(10,51,size=15) Random integers in a range random whole numbers

The identity matrix is worth picturing, since it shows up constantly in math:

   np.identity(3)
   ┌           ┐
   │ 1   0   0 │
   │ 0   1   0 │   ← 1s run down the diagonal, 0s everywhere else
   │ 0   0   1 │
   └           ┘

Part 4: Indexing and Slicing

Once you have an array, you need to grab pieces of it. Let's use a random array of 15 numbers:

python
mat1 = np.random.randint(10, 51, size=15)

Basic Access

python
mat1[5]          # the element at position 5
mat1[16]         # ERROR — index out of bounds (there's no position 16)
mat1[3:10]       # positions 3 through 9
mat1[2:11:2]     # positions 2 to 10, stepping by 2
mat1[::-1]       # the whole array, reversed
mat1[[2,5,3,7]]  # fancy indexing — grab specific positions in any order

Negative Indexing

NumPy lets you count from the end using negative numbers. -1 is the last element:

python
mat1[-1]         # last element
mat1[-5:-1]      # from 5th-from-end up to (not including) last
mat1[-1:-5:-1]   # last four elements, in reverse
   Positive index:    0    1    2    3    4
   Array:           [ 42   17   33   25   48 ]
   Negative index:   -5   -4   -3   -2   -1
                      ▲                    ▲
                   first              mat1[-1] = 48

Part 5: Masking (Boolean / Fancy Indexing)

This is one of NumPy's superpowers and a concept you'll use every single day in data analysis: filtering an array by a condition.

When you write a comparison against an array, NumPy returns an array of True/False values. One per element, a boolean mask:

python
mat1 > 25       # [True, False, True, True, ...]   (a mask)

Think of the mask as a stencil laid over your data. Where the stencil says True, the value shows through; where it says False, it's filtered out:

   data:   [ 42   17   33   25   48   12 ]
   mask:   [  T    F    T    F    T    F ]   ← (data > 25)
            ───────────────────────────
   result: [ 42        33        48     ]   ← only the True values kept
python
mat1[mat1 > 25]                       # keeps only values greater than 25
mat1[(mat1 > 25) & (mat1 < 45)]       # combine conditions with &

When combining conditions, use & (and), | (or), and wrap each condition in its own parentheses. This is a strict rule in NumPy.

np.where — Finding Positions

Where masking returns the values, np.where returns the index positions where a condition is true:

python
np.where(mat1 > 25)                   # the positions of values > 25
np.where((mat1 > 25) & (mat1 < 45))   # positions matching both conditions

A second, very powerful form of np.where acts like an if/else for arrays. We'll use that in the case study at the end.


Part 6: Universal Functions (Math on Whole Arrays)

"Universal functions" (ufuncs) are functions that operate on every element of an array at once. They come in two flavors: those that take two arrays, and those that transform a single array.

Two-Array Math

Let's say we have two 3×3 arrays, arr3 and arr4:

python
np.add(arr3, arr4)        # same as arr3 + arr4
np.add(arr3, 10)          # adds 10 to every element
np.subtract(arr3, arr4)   # same as arr3 - arr4
np.multiply(arr3, arr4)   # ELEMENT-WISE multiply (not matrix multiply!)
np.divide(arr3, arr4)     # same as arr3 / arr4

There's an important distinction here. np.multiply multiplies matching positions. Matrix multiplication (the real linear-algebra operation) is a different function, np.dot:

   ELEMENT-WISE (np.multiply)        MATRIX MULTIPLY (np.dot)

   [a b]   [e f]   [a·e  b·f]        [a b]   [e f]   [ae+bg  af+bh]
   [c d] × [g h] = [c·g  d·h]        [c d] · [g h] = [ce+dg  cf+dh]

   each cell × matching cell         rows combined with columns
python
np.dot(arr3, arr4)        # matrix multiplication, same as arr3 @ arr4

Single-Array Transformations

These take one array (or number) and transform every value:

Function What it does Example result
np.log(10) natural log (base e) 2.302
np.log2(10) log base 2 3.321
np.log10(10) log base 10 1.0
np.exp(5) e raised to the power 148.4
np.sqrt(800) square root 28.28
np.square(80) the square 6400
np.abs(-10) absolute value 10
np.ceil(4.56) round up 5.0
np.floor(3.81) round down 3.0
np.round(12.6781, 2) round to N decimals 12.68

The ceiling/floor pair is easy to remember with a picture:

            ceil → 5.0      (always rounds UP to next integer)
              ▲
            4.56
              ▼
           floor → 4.0      (always rounds DOWN)

Percentiles

A percentile tells you the value below which a given percentage of data falls. The 25th percentile is the value that 25% of the data sits below:

python
np.percentile([150,120,30,45,60,78,81,93,25,33,110], 25)  # 25th percentile
np.percentile([...], 50)    # 50th percentile = the median
np.percentile([...], 100)   # 100th percentile = the maximum

Comparison and Validity Checks

These ufuncs compare arrays element-by-element or check for special values:

python
a = np.array([1, 2, 3])
b = np.array([1, 3, 2])

np.equal(a, b)        # [True, False, False]
np.not_equal(a, b)    # [False, True, True]
np.greater(a, b)      # [False, False, True]
np.less(a, b)         # [False, True, False]

NumPy has special values for "infinity" (np.inf) and "not a number / missing" (np.nan). You check for them with:

python
arr = np.array([1, 2, np.inf, 4, np.nan, -np.inf])
np.isfinite(arr)   # True only for normal, finite numbers
np.isinf(arr)      # True where the value is infinite
np.isnan(arr)      # True where the value is missing (NaN)

Finally, np.any and np.all collapse a whole array into a single yes/no:

  • np.any(arr) → True if at least one element is truthy (non-zero).
  • np.all(arr) → True only if every element is truthy.
   arr = [1, 2, 0, 4]
   np.any(arr) → True   (at least one non-zero)
   np.all(arr) → False  (one value is 0, so not ALL are truthy)

Part 7: Aggregation: and the All-Important axis

Aggregation functions summarize an array down to fewer numbers: np.sum, np.min, np.max, np.mean. On their own, they collapse the entire array to a single value:

python
np.sum(arr3)    # one number: the total of everything
np.mean(arr3)   # one number: the average of everything

But the real power appears with the axis parameter, which controls which direction you collapse. This is the concept beginners struggle with most, so here's the picture:

   axis=0  →  collapse DOWN the columns (one result per column)
   axis=1  →  collapse ACROSS the rows  (one result per row)

           col0 col1 col2
   row0  [  1    2    3  ]  ──► axis=1 sums across →  6
   row1  [  4    5    6  ]  ──► axis=1 sums across → 15
           │    │    │
           ▼    ▼    ▼
   axis=0  5    7    9      (sums down each column)
python
np.sum(arr3, axis=0)    # column totals  → [5, 7, 9]
np.sum(arr3, axis=1)    # row totals     → [6, 15]
np.max(arr3, axis=0)    # max of each column
np.mean(arr3, axis=1)   # average of each row

A simple way to remember it: axis=0 squashes the rows together (vertical), axis=1 squashes the columns together (horizontal).


Part 8: Sorting

1D Arrays — Copy vs. In-Place

There's a subtle but important difference between two ways of sorting:

python
np.sort(arr_unsorted)   # returns a SORTED COPY; original is untouched
arr_unsorted.sort()     # sorts IN-PLACE; modifies the original, returns None
   np.sort(a)              a.sort()
   ┌─────────────┐         ┌─────────────┐
   │ a unchanged │         │ a is now    │
   │ copy → new  │         │ sorted      │
   └─────────────┘         └─────────────┘
   "give me a sorted copy"  "sort yourself"

argsort — Sorting by Position

np.argsort doesn't return the sorted values; it returns the indices that would sort the array. This is incredibly useful when you want to reorder one array based on the order of another (we'll use exactly this in the case study):

python
np.argsort(arr_unsorted_2)   # e.g. [3, 0, 1, 2] — "put position 3 first, then 0..."

2D Arrays

For 2D arrays, sorting respects the axis direction:

python
arr_2d = np.array([[3, 10, 2],
                   [1,  5, 7],
                   [2,  7, 5]])

np.sort(arr_2d, axis=0)   # sorts within each column
np.sort(arr_2d, axis=1)   # sorts within each row

Part 9: Combining and Splitting Arrays

Stacking Arrays Together

There are several ways to glue arrays together. The two foundational ones are concatenation along an axis:

python
np.concatenate((arr3, arr4), axis=0)   # stack top-and-bottom (vertically)
np.concatenate((arr3, arr4), axis=1)   # stack side-by-side (horizontally)

NumPy provides friendlier named shortcuts that do the same thing:

   vstack / concatenate axis=0          hstack / concatenate axis=1
   (stack vertically)                   (stack horizontally)

   [A A A]                              [A A A] [B B B]
   [A A A]            +B  →             [A A A] [B B B]
   [B B B]   ← B added below            ▲ A and B side by side
   [B B B]
Function What it does
np.vstack((a, b)) stack vertically (same as concatenate axis=0)
np.hstack((a, b)) stack horizontally (same as concatenate axis=1)
np.column_stack((a, b)) stack 1D arrays as columns

Splitting Arrays Apart

np.split breaks one array into several pieces. The second argument tells it how:

  • An integer n splits the array into n equal parts.
  • A list of indices splits at those exact positions.
python
arr_big = np.random.randint(10, 51, size=(12, 10))

np.split(arr_big, 2, axis=1)   # split into 2 along columns
np.split(arr_big, 4, axis=0)   # split into 4 along rows
np.hsplit(arr_big, 2)          # horizontal split shortcut
np.vsplit(arr_big, 4)          # vertical split shortcut
   np.hsplit(arr, 2)            np.vsplit(arr, 2)
   ┌────┬────┐                  ┌─────────┐
   │ L  │ R  │   →  [L] [R]     │   TOP   │   →  [TOP]
   │    │    │                  ├─────────┤      [BOTTOM]
   └────┴────┘                  │ BOTTOM  │
                                └─────────┘

Part 10: Inspecting an Array

When you receive an unfamiliar array, these attributes tell you everything about its structure. Notice they have no parentheses. They're descriptions of the object, not actions:

python
arr_big.shape    # (rows, columns) — its dimensions
arr_big.size     # total number of elements
arr_big.ndim     # number of dimensions (1, 2, 3...)
arr_big.nbytes   # total memory used, in bytes
arr_big.dtype    # the data type of the elements (e.g. int64)
   arr_big  (a 12 × 10 array of integers)
   ┌────────────────────────────┐
   │ .shape  → (12, 10)          │
   │ .size   → 120               │
   │ .ndim   → 2                 │
   │ .nbytes → 960  (120 × 8)    │
   │ .dtype  → int64             │
   └────────────────────────────┘

Part 11: Reshaping: Changing the Form Without Changing the Data

Reshaping rearranges the same numbers into a different grid. The total number of elements must stay the same.

python
arr_big.reshape(24, 5)    # 120 elements → 24 rows × 5 cols (24×5 = 120 ✓)

The Magic -1

If you put -1 for one dimension, NumPy figures it out for you. If you have 120 elements and ask for 15 columns, NumPy works out you need 8 rows:

python
arr_big.reshape(-1, 15)   # "15 columns, you figure out the rows" → 8 rows
   120 elements, reshape(-1, 15)
   "I want 15 columns. How many rows?"
   120 ÷ 15 = 8  →  NumPy fills in 8 automatically

Transpose — Flipping Rows and Columns

transpose() (or the shortcut .T) swaps rows and columns:

python
arr_big.transpose()   # rows become columns, columns become rows
arr_big.T             # identical, shorter
   original          transpose (.T)
   [1 2 3]           [1 4]
   [4 5 6]    →      [2 5]
                     [3 6]
   2 rows, 3 cols    3 rows, 2 cols

flatten vs ravel — A Critical Distinction

Both turn a multi-dimensional array into a flat 1D array, but they differ in a way that can cause real bugs if you don't understand it:

  • ravel() returns a view when possible. If the array is stored contiguously in memory (the common case), ravel shares the original data. For non-contiguous arrays (e.g. after a transpose), ravel has to make a copy instead. Either way, changing a ravel result may change the original.
  • flatten() always makes a copy, regardless of memory layout. Changing the result never touches the original.
python
arr = np.array([[10, 20], [30, 40]])

rav = arr.ravel()
rav[0] = 99       # this ALSO changes arr! (shared memory)

flat = arr.flatten()
flat[1] = 88      # arr stays the same (independent copy)
   ravel()  → VIEW (window onto the same data)
   ┌──────────┐         change the view →
   │ original │ ◄─────── change shows up in original too
   └──────────┘

   flatten() → COPY (separate data)
   ┌──────────┐  copy   ┌──────────┐
   │ original │ ──────► │   copy   │  change here →
   └──────────┘         └──────────┘  original untouched

Part 12: Shallow vs Deep Copy (Shared Memory)

This builds on the ravel/flatten idea and is one of NumPy's most important, and most surprising behaviors. NumPy avoids copying data whenever it can. Reshaping and slicing often hand you a new view of the same underlying data, not a fresh copy.

Watch what happens with a reshape:

python
a = np.arange(4)      # [0 1 2 3]
b = a.reshape(2, 2)   # b is a VIEW of a

a[0] = 10             # change a...
# ...and b changes too! They share the same data.

The same surprise happens with slicing:

python
mat1_new = mat1[0:6]   # a slice — shares memory with mat1
mat1_new[2] = 25       # changing the slice ALSO changes mat1
   a ──┐
       ├──► [ shared data block ]   ← changing through either name
   b ──┘                              changes what BOTH see

   This is a SHALLOW COPY: the data is reused, not duplicated.

Deep Copy — When You Want True Independence

If you want a genuinely separate array whose changes won't leak back, use .copy(). This is a deep copy: its own memory, fully independent.

python
original = np.array([1, 2, 3, 4])
deep_copy = original.copy()

deep_copy[0] = 77       # original stays [1, 2, 3, 4]
np.shares_memory(original, deep_copy)   # False — proof they're separate

The takeaway: NumPy uses shallow copies (shared data) for cheap operations like reshape and slicing to stay fast. When you make permanent changes that should not affect the original, use .copy(). And whenever you're unsure whether two arrays share memory, np.shares_memory(a, b) will tell you.


Part 13: Broadcasting: Math Between Different-Shaped Arrays

Normally, doing math between two arrays requires them to be the same shape. Broadcasting is NumPy's clever system for stretching a smaller array to match a bigger one automatically, so the math just works.

The helper np.tile literally repeats an array to build bigger ones. To build the 4×3 matrix shown in the diagram below, we tile a column vector across 3 columns:

python
col = np.arange(0, 40, 10).reshape(4, 1)   # column vector: [[0],[10],[20],[30]]
big_matrix = np.tile(col, (1, 3))           # repeat 3 times sideways → shape (4, 3)

But broadcasting does this stretching implicitly. Picture adding a single row to a full matrix:

   Big matrix (4×3)         Small row (1×3)

   [ 0   0   0 ]            [ 0  1  2 ]
   [10  10  10 ]     +
   [20  20  20 ]
   [30  30  30 ]

   Broadcasting "stretches" the small row down to all 4 rows:

   [ 0   0   0 ]     [ 0  1  2 ]      [ 0   1   2 ]
   [10  10  10 ]  +  [ 0  1  2 ]   =  [10  11  12 ]
   [20  20  20 ]     [ 0  1  2 ]      [20  21  22 ]
   [30  30  30 ]     [ 0  1  2 ]      [30  31  32 ]
                     ▲ replicated automatically

The single row [0, 1, 2] is broadcast across all four rows without you ever writing a loop or manually copying it. This is what makes NumPy code so concise. Operations that would need nested loops in plain Python become a single +.


Part 14: Getting Data In and Out of Files

So far we've built arrays by hand. In the real world, data lives in files. NumPy can read it in and write it back out.

Reading Text Files

There are two main loaders, and the choice depends on whether your data has gaps:

python
data = np.loadtxt('data.txt', delimiter="\t")     # use when NO missing values
data = np.genfromtxt('data_miss.txt', delimiter="\t")  # handles missing values

np.loadtxt is strict, which errors out if any value is missing. np.genfromtxt is forgiving , filling any gaps with nan. The rule of thumb:

   Clean data, no gaps   →  np.loadtxt   (fast, strict)
   Real data with gaps   →  np.genfromtxt (fills blanks with NaN)

Try it on a file that actually has a gap and np.loadtxt stops cold instead of guessing:

python
data = np.loadtxt('data_miss.txt', delimiter="\t")
# ValueError: could not convert string '' to float64 at row 8, column 3.

Swap in np.genfromtxt on that exact same file and it loads without complaint, dropping a nan into the empty cell instead of raising:

python
data = np.genfromtxt('data_miss.txt', delimiter="\t")   # the gap becomes nan, the rest loads fine

Reading CSVs with Mixed Types

CSV files often mix text and numbers (e.g. a customer's name and age). By default NumPy assumes everything is a float, which mangles text columns. Loading the Mall Customers file (which has a text "Genre" column) the naive way turns that whole column into nan, since "Male" and "Female" simply can't be parsed as numbers:

python
data = np.genfromtxt('Mall_Customers.csv', delimiter=",")   # default dtype is float — the Genre column becomes nan

Forcing dtype='object' stops the column from being wiped out, but it creates a different headache: every single value, numbers included, comes back as a Python bytes object instead of a real number:

python
data = np.genfromtxt('Mall_Customers.csv', delimiter=",", dtype='object', skip_header=1)
data[0:5]   # → array([[b'1', b'Male', b'19', b'15', b'39'], ...], dtype=object)

Those b'19'-style values look like numbers but are bytes underneath, so any arithmetic on them blows up:

python
# try dividing 4th col with 3rd col and we get an error
data[:, 3] / data[:, 2]
# TypeError: unsupported operand type(s) for /: 'bytes' and 'bytes'

The clean solution is to define a structured dtype that names each column and gives it the right type up front:

python
dt = np.dtype({'names':   ["CustomerID", "Genre", "Age", "Annual_Income", "Spending_Score"],
               'formats': [np.int16,     'U16',   np.int16, np.int16,      np.int16]})

data = np.genfromtxt('Mall_Customers.csv', delimiter=",", dtype=dt, skip_header=1)

Here skip_header=1 skips the header row, 'U16' means "text up to 16 characters," and each numeric column gets a proper integer type. Check data.ndim afterward and you'll see it's 1, not 2 — this is a 1D array of labeled records, one per row, not a grid. (In practice, this is exactly the kind of job where most people reach for Pandas instead, but it's good to know NumPy can do it — and the result drops straight into a DataFrame with pd.DataFrame(data) if you want to take it from there.)

Saving and Loading Arrays

To save arrays in NumPy's own fast binary format:

python
np.save('output.npy', data)        # ONE array → .npy file
np.savez('outputs.npz', data)      # MULTIPLE arrays → .npz file
a = np.load('output.npy')          # load it back
   .npy  →  a single array
   .npz  →  several arrays in one file (named arr_0, arr_1, ...)

Part 15: Putting It All Together: A Real Case Study

Let's see why all of this matters. Imagine you're a data scientist at a major ride-sharing company (like Uber or Lyft). The city operations team hands you a batch of ride data from a busy Friday night and wants deep insights into driver efficiency, surge pricing dynamics, and rider satisfaction, all in service of optimizing the matching algorithm. Every technique below uses something from this guide.

Loading and Splitting the Columns

python
data = np.genfromtxt('ride_data.csv', delimiter=',', skip_header=1)

ride_ids   = data[:, 0]   # column slicing pulls out each variable
driver_ids = data[:, 1]
distances  = data[:, 2]   # distance in miles
durations  = data[:, 3]   # duration in minutes
fares      = data[:, 4]   # fare in USD
surges     = data[:, 5]   # surge multiplier (e.g. 1.2 = 20% extra)
d_ratings  = data[:, 6]   # driver rating
r_ratings  = data[:, 7]   # rider rating

Notice data[:, 0]. The slice "all rows, column 0." This is 2D indexing in action. Here's what each column actually holds:

  • ride_id — the unique ID given to every ride.
  • driver_id — the unique ID of each driver.
  • distance_miles — distance covered from origin to destination.
  • duration_minutes — how long the ride took.
  • fare_usd — the total fare of the ride, in US dollars.
  • surge_multiplier — activates during high demand; a multiplier of 1.2 means the rider pays the normal fare × 1.2.
  • driver_rating — the rating the rider gave the driver.
  • rider_rating — the rating the driver gave the rider.

Revenue: GMV and the Platform's Cut

python
gmv = np.sum(fares)                # Gross Merchandise Value — every dollar paid
platform_revenue = gmv * 0.25      # the company keeps a 25% commission
print(f"Total GMV: ${gmv:.2f} | Platform Cut: ${platform_revenue:.2f}")
# Total GMV: $503.50 | Platform Cut: $125.88

For this batch of rides, the platform collected $503.50 in total fares and kept $125.88 of it as commission. GMV (Gross Merchandise Value) is the total of every dollar that flowed through the platform before anyone takes their cut. The analogy: if you build an app connecting bakers to customers and someone buys a $100 cake, your revenue might be a $10 fee, but the GMV of that sale is the full $100. Big platforms (Uber, Airbnb, Amazon, eBay) live and die by GMV because it's the ultimate vanity metric for scale — a $1 billion GMV means the platform controls a massive chunk of that market. It also dictates the platform's ceiling: you can't grow revenue without either taking a bigger cut (which angers the supply side, e.g. drivers) or growing the overall GMV.

Speed Check (Element-wise Math + Masking)

python
speeds_mph = distances / (durations / 60)     # element-wise: speed for every ride
speeding_rides = ride_ids[speeds_mph > 65]    # masking: which rides exceeded 65 mph
print("Speeding Ride IDs:", speeding_rides.astype(int))   # Speeding Ride IDs: [16]

One line computes speed for every ride at once (element-wise division), and a boolean mask instantly isolates the suspicious ones — here, a single ride (ID 16) came back clocking over 65 mph, which is worth a manual look (bad GPS data, a highway trip misclassified, etc.).

Surge Revenue (Reverse-Engineering with Division)

python
base_fares = fares / surges                       # undo the surge to find the base fare
surge_extra_revenue = np.sum(fares - base_fares)   # how much was pure surge premium
# Extra Revenue from Surge: $57.42

Surge pricing is a dynamic pricing model: the surge multiplier is the exact number the base fare gets multiplied by (1.5x, 2x, ...) during periods of high demand. Dividing the fare by the surge factor reverse-engineers what the ride would have cost without it, so the gap between the two is pure surge revenue — $57.42 of it here. It exists to solve a simple supply-and-demand problem, and it does two things at once:

  • Incentivizes supply — a "2.0x surge in that area" notification pulls more drivers toward the high-demand zone.
  • Controls demand — it discourages riders who don't urgently need a ride right now, freeing up cars for people willing to pay the premium (someone rushing to the airport, say).

Statistics: Median, Argmax, Percentiles

python
median_rating = np.median(d_ratings)         # the typical driver rating → 4.7
max_fare_idx = np.argmax(fares)              # POSITION of the single highest fare
print(f"Highest fare was ${fares[max_fare_idx]} on Ride ID {int(ride_ids[max_fare_idx])}")
# Highest fare was $85.0 on Ride ID 12
p90_fare = np.percentile(fares, 90)          # the "high-value" threshold → $54.00

np.argmax returns the index of the largest value, so ride_ids[max_fare_idx] tells you exactly which ride earned the most (Ride 12, at $85). The 90th percentile is the dollar amount that separates the top 10% most expensive rides from the bottom 90% — here, $54.00 means 90% of rides cost less than that, and only the priciest tenth cost more. A ride can cross that line for several reasons: a long-distance airport run, a 3.0x surge, or a premium vehicle tier. Data scientists calculate it for a few concrete reasons:

  • Operations — route the best, 5-star drivers to these high-value riders to protect the experience.
  • Marketing — run a targeted campaign, e.g. "spend over $55 on a ride and get a free coffee voucher."
  • Anomaly detection — a ride way past this threshold (say, $400) is worth a manual check in case the GPS glitched or the driver took a longer route than necessary.

Quick Ratios with np.mean on a Mask

python
short_trips_pct = np.mean(distances < 3.0) * 100   # % of rides under 3 miles → 25.0%

This is a neat trick: distances < 3.0 is a mask of True/False, and since True counts as 1, taking the mean of a mask gives you the proportion that satisfied the condition — a quarter of the night's rides here. Multiply by 100 for a percentage.

Per-Ride Calculations and Perfect Matches

python
rating_diff = np.abs(d_ratings - r_ratings)        # disagreement per trip
dollars_per_minute = fares / durations             # earning efficiency per ride
perfect_mask = (d_ratings == 5.0) & (r_ratings == 5.0)   # both gave 5 stars
perfect_rides = ride_ids[perfect_mask]
print("Perfect Match Ride IDs:", perfect_rides.astype(int))   # Perfect Match Ride IDs: [10]

Three independent, per-ride calculations in one breath: rating_diff flags trips where driver and rider disagreed on quality, dollars_per_minute measures how efficiently each ride earned, and the perfect_mask boolean filter isolates the rides where both sides gave a flawless 5.0 — only Ride 10 managed that here.

np.where as If/Else: Driver Payouts with a Penalty

Here's the powerful second form of np.where, which works like an if/else across the whole array at once:

python
standard_payout = fares * 0.75
adjusted_payout = np.where(d_ratings < 4.0,        # CONDITION
                           standard_payout - 3.00,  # value if True (penalty)
                           standard_payout)         # value if False
# Adjusted Payouts (First 5): [13.88 33.75  6.75 24.    9.38]
   np.where(condition, value_if_true, value_if_false)

   driver rating < 4.0 ?
        ├─ YES → payout − $3 penalty
        └─ NO  → normal payout
   ...applied to every driver simultaneously

Running Totals and Counting Uniques

python
cum_fares = np.cumsum(fares)    # running total as the night progresses — ends at 403.5

unique_drivers, counts = np.unique(driver_ids, return_counts=True)
busiest_driver = unique_drivers[np.argmax(counts)]   # most rides completed → Driver 101

np.cumsum gives a cumulative sum: fare after ride 1, after rides 1+2, and so on. np.unique with return_counts=True lists each distinct driver alongside how many rides they did, and np.argmax finds the busiest.

Sorting One Array by Another (argsort in action)

python
sorted_ride_ids = ride_ids[np.argsort(durations)]   # ride IDs, shortest trip first

This is the classic argsort pattern: sort the durations, get the order of positions that would sort them, then apply that same order to the ride IDs array instead. It looks like a basic housekeeping task, but ordering rides by duration immediately surfaces two categories that matter a lot to the business: the shortest trips and the longest ones.

1. Analyzing "micro-trips" (the shortest durations). At the top of the sorted list — rides taking 2 to 5 minutes — you're looking at trips that are often inefficient for cars. A driver might spend 8 minutes fighting through city traffic just to pick someone up for a 2-minute ride; they earn very little and feel like their time was wasted. Once the platform can see the volume of these micro-trips, it can respond: introduce a "minimum fare" so drivers earn a fair amount regardless of trip length, or — if thousands of 3-minute rides cluster in one area — deploy a fleet of e-bikes or scooters there instead, nudging the app to suggest: "A car is $8 and takes 5 mins. A scooter is $2 and takes 4 mins."

2. Managing "deadheading" (the longest durations). At the bottom of the sorted list — 45+ minute rides — you're looking at rare but high-revenue outliers that create a logistical headache. A driver who takes a 60-minute ride out of the busy city center then has to drive 60 minutes back empty, since there's no demand where they dropped off. Driving empty like this is called "deadheading," and it costs the driver fuel and time with no fare to show for it. Spotting these long-duration rides is what led platforms to build features like a "Destination Filter" (matching a driver heading back to the city with a rider who wants to go the same way) and a "long trip premium" fee that compensates the driver for the empty return leg.

Reshaping into a 3D Tensor for Machine Learning

python
ratings_subset = np.vstack((d_ratings[:12], r_ratings[:12]))   # stack into 2×12
tensor_3d = ratings_subset.reshape(2, 2, 6)                    # reshape to 2×2×6
print("3D Tensor Shape:", tensor_3d.shape)   # 3D Tensor Shape: (2, 2, 6)

The ML team needs the first 12 driver and rider ratings reshaped into a 3D tensor for a neural network input. We vstack the two rating arrays together into a 2×12 grid, then reshape those same 24 values into a 2×2×6 cube — combining stacking and reshaping, both covered earlier, without changing a single value.


Quick Reference Summary

Task Code
Import NumPy import numpy as np
Make array from list np.array([1, 2, 3])
Sequence np.arange(start, end, step)
All ones / zeros np.ones((r,c)) / np.zeros((r,c))
Filled with a value np.full((r,c), val)
Identity matrix np.identity(n)
Random integers np.random.randint(low, high, size)
Indexing / slicing arr[2], arr[1:5], arr[::-1]
Negative index arr[-1] (last element)
Masking (filter) arr[arr > 25]
Combine conditions arr[(a > 1) & (a < 9)]
Find positions np.where(arr > 25)
If/else over array np.where(cond, if_true, if_false)
Element-wise multiply np.multiply(a, b) or a * b
Matrix multiply np.dot(a, b) or a @ b
Rounding np.ceil(), np.floor(), np.round(x, 2)
Percentile / median np.percentile(a, 90) / np.median(a)
Check missing/infinite np.isnan(a), np.isinf(a)
Any / all true np.any(a) / np.all(a)
Aggregate by column np.sum(a, axis=0)
Aggregate by row np.sum(a, axis=1)
Sort (copy) np.sort(a)
Sort (in-place) a.sort()
Sort indices np.argsort(a)
Index of max/min np.argmax(a) / np.argmin(a)
Running total np.cumsum(a)
Unique + counts np.unique(a, return_counts=True)
Stack vertical / horizontal np.vstack((a,b)) / np.hstack((a,b))
Split np.split(a, n, axis=)
Shape / size / dims a.shape, a.size, a.ndim
Reshape (auto dim) a.reshape(-1, cols)
Transpose a.T
Flatten (copy) / ravel (view) a.flatten() / a.ravel()
Deep copy a.copy()
Check shared memory np.shares_memory(a, b)
Read text / with gaps np.loadtxt() / np.genfromtxt()
Save / load np.save() / np.load()

NumPy is the bedrock of the entire Python data science world. Once arrays feel natural: creating them, slicing them, filtering with masks, doing element-wise math, reshaping, and understanding when data is shared versus copied, everything built on top of it (Pandas, scikit-learn, deep learning frameworks) becomes far easier to reason about. The best way to make it stick is to grab a real CSV and try the case-study operations on it yourself: compute some sums by axis, filter rows with a mask, sort one column by another with argsort, and watch how a single line of NumPy replaces what would have been a tangle of loops.

NumPy is the bedrock of the entire Python data science world. Once arrays feel natural — creating them, slicing them, filtering with masks, doing element-wise math, reshaping, and understanding when data is shared versus copied — everything built on top of it becomes far easier to reason about. Grab a real CSV and try the case-study operations yourself: compute sums by axis, filter rows with a mask, sort one column by another with argsort. A single line of NumPy replaces what would have been a tangle of loops.