Creating Sample Files for This Guide
Before diving in, it helps to create a few sample files to use throughout the examples, the same way you'd set up your materials before running an experiment.
import os
os.makedirs("file_handling_demo", exist_ok=True)
with open("file_handling_demo/learn.txt", "w", encoding="utf-8") as f:
f.write("Day 1: Started learning Python file handling!\n")
f.write("Day 2: Understood the 'with' statement. Life changing!\n")
f.write("Day 3: Read and wrote my first CSV file like a data scientist!\n")
print("Sample files created successfully!")
1. Opening a File
Before you can read or write a file, you must open it first, the same way you'd open a book before reading it.
Python's built-in open() function does this job:
file_object = open(filename, mode, encoding)
| Parameter | Description | Default |
|---|---|---|
filename |
Path to the file (absolute or relative) | Required |
mode |
How you want to open it ('r', 'w', 'a', etc.) |
'r' (read) |
encoding |
Character encoding to use | Platform-dependent |
The function returns a file object (also called a file handle).
Worth doing every time: always specify
encoding='utf-8'explicitly. The default encoding varies by operating system (Windows tends to usecp1252, Linux and Mac useutf-8). Being explicit makes your code behave the same way on every machine it runs on.
f = open("file_handling_demo/learn.txt", "r", encoding="utf-8")
print("File opened successfully!")
print("Type of file object:", type(f))
f.close()
File opened successfully!
Type of file object: <class '_io.TextIOWrapper'>
2. File Modes: Telling Python What You Want to Do
The mode is the most important argument after the filename. It tells Python what you plan to do with the file.
Analogy: think of it like booking a table at a restaurant. 'r' means you're just reading the menu, no ordering. 'w' means you're handed a brand new menu, and the old one gets thrown away. 'a' means you want to add new items to the existing menu without touching the rest. 'x' means you want to create a new menu, but only if one doesn't already exist.
Mode Quick Reference
| Mode | Full Name | Creates file? | Overwrites? | Can Read? | Can Write? |
|---|---|---|---|---|---|
'r' |
Read | ❌ | ❌ | ✅ | ❌ |
'w' |
Write | ✅ | ✅ | ❌ | ✅ |
'a' |
Append | ✅ | ❌ | ❌ | ✅ |
'x' |
Exclusive create | ✅ (fails if exists) | ❌ | ❌ | ✅ |
'r+' |
Read + Write | ❌ | ❌ | ✅ | ✅ |
'w+' |
Write + Read | ✅ | ✅ | ✅ | ✅ |
'a+' |
Append + Read | ✅ | ❌ | ✅ | ✅ |
Text vs. Binary
Add 'b' to any mode to open in binary mode, for images, videos, audio, or any file that isn't plain text:
- 'rb': read a binary file (like reading a photo)
- 'wb': write a binary file (like saving a downloaded image)
f = open("file_handling_demo/learn.txt", "r", encoding="utf-8")
print("'r' readable:", f.readable(), "| writable:", f.writable())
f.close()
f = open("file_handling_demo/temp_w.txt", "w", encoding="utf-8")
print("'w' readable:", f.readable(), "| writable:", f.writable())
f.close()
f = open("file_handling_demo/temp_a.txt", "a", encoding="utf-8")
print("'a' readable:", f.readable(), "| writable:", f.writable())
f.close()
f = open("file_handling_demo/learn.txt", "r+", encoding="utf-8")
print("'r+' readable:", f.readable(), "| writable:", f.writable())
f.close()
try:
f = open("file_handling_demo/brand_new.txt", "x", encoding="utf-8")
f.write("I am a brand new file!\n")
f.close()
print("'x' mode: file created successfully!")
except FileExistsError:
print("'x' mode: file already exists, creation failed (as expected)!")
'r' readable: True | writable: False
'w' readable: False | writable: True
'a' readable: False | writable: True
'r+' readable: True | writable: True
'x' mode: file created successfully!
3. Closing a File
When you open a file, the operating system allocates resources to manage it. If you forget to close it, memory is wasted, other programs may not be able to access the file, and data you wrote might not actually be saved yet due to internal buffering.
Analogy: it's like borrowing a book from a library. You're expected to return it, not leave it sitting on your desk forever.
There are three ways to close a file, and they're worth seeing side by side because the last one is genuinely better than the first two.
Method 1: close(), risky. If an error happens before this line runs, the file stays open.
f = open("file_handling_demo/learn.txt", "r", encoding="utf-8")
content = f.read()
f.close()
print("File closed:", f.closed)
File closed: True
Method 2: try/finally, safer. Guarantees closure even if an error occurs in between.
try:
f = open("file_handling_demo/learn.txt", "r", encoding="utf-8")
content = f.read()
finally:
f.close()
print("File closed:", f.closed)
File closed: True
Method 3: the with statement, best. Closes the file automatically the moment the block ends, no matter what happens inside it. This is the industry standard, and the one you should default to.
with open("file_handling_demo/learn.txt", "r", encoding="utf-8") as f:
content = f.read()
# File is automatically closed here, no matter what!
print("File closed:", f.closed)
File closed: True
Every example from this point on uses with.
4. Writing to a File
Python gives you two main methods to write data to a file.
write(string) writes a single string to the file and returns the number of characters written. It does not add a newline automatically; you have to include \n yourself. Think of it like typing on a typewriter: you control every character, including the line breaks.
writelines(list_of_strings) writes a list of strings all at once, which is more efficient when you have many lines. It also does not add newlines automatically; each string in the list needs its own \n already included. Think of it like stamping a whole stack of pages in one motion instead of one at a time.
with open("file_handling_demo/shopping_list.txt", "w", encoding="utf-8") as f:
f.write("My Shopping List:\n")
f.write("1. Milk\n")
f.write("2. Eggs\n")
f.write("3. Bread\n")
f.write("4. Coffee (the most important one)\n")
with open("file_handling_demo/shopping_list.txt", "r", encoding="utf-8") as f:
print(f.read())
My Shopping List:
1. Milk
2. Eggs
3. Bread
4. Coffee (the most important one)
students = [
"Alice's Score: 95\n",
"Bob's Score: 87\n",
"Carol's Score: 92\n",
"Dave's Score: 78\n",
"Eve's Score: 99\n",
]
with open("file_handling_demo/students.txt", "w", encoding="utf-8") as f:
f.write("Student Scores:\n")
f.writelines(students)
with open("file_handling_demo/students.txt", "r", encoding="utf-8") as f:
print(f.read())
Student Scores:
Alice's Score: 95
Bob's Score: 87
Carol's Score: 92
Dave's Score: 78
Eve's Score: 99
Switching to 'a' mode adds to the end of the file without erasing what's already there:
with open("file_handling_demo/shopping_list.txt", "a", encoding="utf-8") as f:
f.write("5. Chocolate\n")
f.write("6. Pizza\n")
with open("file_handling_demo/shopping_list.txt", "r", encoding="utf-8") as f:
print(f.read())
My Shopping List:
1. Milk
2. Eggs
3. Bread
4. Coffee (the most important one)
5. Chocolate
6. Pizza
Real-World Example: Writing Application Logs Like Google
Every major tech company maintains log files, a running record of everything that happens in their systems. Google processes billions of search requests daily, and every error, warning, or event gets written to a log file. Here's a simplified version of that pattern:
import datetime
def write_log(log_file, level, message):
timestamp = datetime.datetime.now().strftime("%Y-%m-%d %H:%M:%S")
entry = f"[{timestamp}] [{level:7s}] {message}\n"
with open(log_file, "a", encoding="utf-8") as f:
f.write(entry)
log_path = "file_handling_demo/app.log"
write_log(log_path, "INFO", "Server started on port 8080")
write_log(log_path, "INFO", "User 'alice@gmail.com' logged in")
write_log(log_path, "WARNING", "Response time exceeded 500ms for /api/search")
write_log(log_path, "ERROR", "Database connection timeout, retrying...")
write_log(log_path, "INFO", "Database reconnected successfully")
with open(log_path, "r", encoding="utf-8") as f:
print(f.read())
[2026-09-02 15:37:34] [INFO ] Server started on port 8080
[2026-09-02 15:37:34] [INFO ] User 'alice@gmail.com' logged in
[2026-09-02 15:37:34] [WARNING] Response time exceeded 500ms for /api/search
[2026-09-02 15:37:34] [ERROR ] Database connection timeout, retrying...
[2026-09-02 15:37:34] [INFO ] Database reconnected successfully
Each call to write_log opens the file in append mode, writes one line, and closes it again, exactly the pattern you'd want for a long-running server that logs events over hours or days.
5. Reading from a File
Python offers several ways to read file content, each suited to a different situation:
| Method | Returns | Use When |
|---|---|---|
read() |
Entire file as one string | File is small, you need everything |
read(n) |
Next n characters |
Processing in fixed-size chunks |
readline() |
Next line as a string | Processing line by line with fine control |
readlines() |
All lines as a list | You need every line available as a list |
for line in file |
Lines one at a time | Most common; memory-efficient |
read() loads the whole file into memory as one string:
with open("file_handling_demo/learn.txt", "r", encoding="utf-8") as f:
content = f.read()
print(content)
Day 1: Started learning Python file handling!
Day 2: Understood the 'with' statement. Life changing!
Day 3: Read and wrote my first CSV file like a data scientist!
read(n) reads exactly n characters, and each call picks up exactly where the last one left off:
with open("file_handling_demo/learn.txt", "r", encoding="utf-8") as f:
first_10 = f.read(10)
next_10 = f.read(10)
print(f"First 10 characters: {repr(first_10)}")
print(f"Next 10 characters: {repr(next_10)}")
First 10 characters: 'Day 1: Sta'
Next 10 characters: 'rted learn'
readline() reads one line at a time, including the trailing \n, and returns an empty string once it reaches the end of the file:
with open("file_handling_demo/learn.txt", "r", encoding="utf-8") as f:
line1 = f.readline()
line2 = f.readline()
line3 = f.readline()
line4 = f.readline() # Empty string = end of file
print("Line 1:", repr(line1))
print("Line 4 (EOF):", repr(line4))
Line 1: 'Day 1: Started learning Python file handling!\n'
Line 4 (EOF): ''
readlines() reads everything and hands it back as a list of lines:
with open("file_handling_demo/learn.txt", "r", encoding="utf-8") as f:
lines = f.readlines()
print("Type:", type(lines))
print("Number of lines:", len(lines))
print("First entry:", lines[0].strip())
print("Last entry:", lines[-1].strip())
Type: <class 'list'>
Number of lines: 3
First entry: Day 1: Started learning Python file handling!
Last entry: Day 3: Read and wrote my first CSV file like a data scientist!
Looping directly over the file object is the option you'll reach for most often, because it reads one line at a time without ever loading the whole file into memory:
with open("file_handling_demo/learn.txt", "r", encoding="utf-8") as f:
for line_num, line in enumerate(f, 1):
line = line.strip()
if line:
print(f" Entry {line_num}: {line}")
Entry 1: Day 1: Started learning Python file handling!
Entry 2: Day 2: Understood the 'with' statement. Life changing!
Entry 3: Day 3: Read and wrote my first CSV file like a data scientist!
Picture a 10 GB server log. read() would try to pull the entire thing into memory at once, which could crash the program. Looping line by line never holds more than one line in memory at a time, which is exactly why it's the default choice for anything that might be large.
Real-World Example: Reading User Data Like Netflix
Netflix analyzes watch history to power its recommendation engine, processing lines of history data stored in files. Here's a simplified version:
watch_history = [
"alice,Stranger Things,S01E01,45:30,2024-01-15\n",
"alice,Stranger Things,S01E02,50:22,2024-01-15\n",
"alice,The Crown,S01E01,58:10,2024-01-16\n",
"bob,Breaking Bad,S01E01,47:00,2024-01-15\n",
"bob,Breaking Bad,S01E02,48:30,2024-01-16\n",
"carol,Money Heist,S01E01,42:15,2024-01-17\n",
"alice,Breaking Bad,S01E01,47:00,2024-01-17\n",
]
with open("file_handling_demo/watch_history.txt", "w", encoding="utf-8") as f:
f.writelines(watch_history)
user_shows = {}
with open("file_handling_demo/watch_history.txt", "r", encoding="utf-8") as f:
for line in f:
parts = line.strip().split(",")
user, show = parts[0], parts[1]
if user not in user_shows:
user_shows[user] = set()
user_shows[user].add(show)
for user, shows in user_shows.items():
print(f" {user}: {', '.join(sorted(shows))}")
alice: Breaking Bad, Stranger Things, The Crown
bob: Breaking Bad
carol: Money Heist
This tiny script is the same basic shape as the first step of a real recommendation system: turn raw event logs into a per-user profile.
6. Navigating Inside a File: seek() and tell()
When you read from a file, Python tracks your position with an invisible cursor, the same way a text editor tracks the blinking cursor on screen. Every character you read moves it forward.
Analogy: think of a file like a cassette tape. tell() tells you which second of the tape you're currently at. seek() lets you fast-forward or rewind to any point instantly.
tell(): "Where is the cursor right now?" Returns the current byte position.seek(offset, whence): "Move the cursor to this position."
whence value |
Meaning |
|---|---|
0 (default) |
From the beginning of the file |
1 |
From the current position |
2 |
From the end of the file |
with open("file_handling_demo/learn.txt", "r", encoding="utf-8") as f:
print("Starting position:", f.tell()) # 0 at the beginning
first_line = f.readline()
print("After reading line 1:", f.tell()) # Moved forward!
second_line = f.readline()
print("After reading line 2:", f.tell())
f.seek(0) # Back to the very beginning
print("After seek(0):", f.tell())
line_again = f.readline()
print("Line 1 again:", repr(line_again))
f.seek(7) # Skip "Day 1: " (7 characters)
print("After seek(7):", repr(f.read(8)))
Starting position: 0
After reading line 1: 46
After reading line 2: 101
After seek(0): 0
Line 1 again: 'Day 1: Started learning Python file handling!\n'
After seek(7): 'Started '
7. File Object Properties
Once you open a file, the file object comes with properties that let you inspect its current state:
| Property / Method | Description |
|---|---|
.name |
The name or path of the file |
.mode |
The mode the file was opened in |
.closed |
True if the file is closed, False if open |
.readable() |
True if the file can be read |
.writable() |
True if the file can be written to |
.seekable() |
True if seek() is supported |
with open("file_handling_demo/learn.txt", "r+", encoding="utf-8") as f:
print(f" .name : {f.name}")
print(f" .mode : {f.mode}")
print(f" .closed : {f.closed}")
print(f" .readable() : {f.readable()}")
print(f" .writable() : {f.writable()}")
print(f" .seekable() : {f.seekable()}")
print(f"After 'with' block, .closed: {f.closed}")
.name : file_handling_demo/learn.txt
.mode : r+
.closed : False
.readable() : True
.writable() : True
.seekable() : True
After 'with' block, .closed: True
8. File and Directory Management: The os Module
The os module is Python's interface to your operating system. It lets you navigate the file system, create and delete directories, and rename or delete files, the same operations you'd normally do by clicking around in File Explorer or Finder, just from code instead.
Key Functions at a Glance
| Function | What It Does |
|---|---|
os.getcwd() |
Get the current working directory |
os.listdir(path) |
List files and folders |
os.mkdir(path) |
Create one directory |
os.makedirs(path) |
Create a directory plus all missing parent directories |
os.rmdir(path) |
Remove an empty directory |
os.rename(src, dst) |
Rename a file or directory |
os.remove(path) |
Delete a file |
os.path.exists(path) |
Check if a file or folder exists |
os.path.isfile(path) / isdir(path) |
Check whether it's a file or a directory |
os.path.join(...) |
Build paths safely, correct on every OS |
os.path.getsize(path) |
Get file size in bytes |
os.path.basename(path) / dirname(path) |
Get just the filename, or just the directory |
import os
os.mkdir("file_handling_demo/reports")
# makedirs() creates every missing intermediate directory in one call
os.makedirs("file_handling_demo/archive/2026/january", exist_ok=True)
exist_ok=True tells os.makedirs() to do nothing quietly if the directory already exists, instead of raising an error. By default, exist_ok is False, so calling os.makedirs() a second time on a folder that already exists raises FileExistsError:
try:
os.makedirs("file_handling_demo/reports", exist_ok=False)
except FileExistsError as e:
print(f"Confirmed: {e}")
Confirmed: [Errno 17] File exists: 'file_handling_demo/reports'
Renaming and deleting work the way you'd expect:
with open("file_handling_demo/old_name.txt", "w") as f:
f.write("I will be renamed!\n")
os.rename("file_handling_demo/old_name.txt", "file_handling_demo/new_name.txt")
print("Old name exists:", os.path.exists("file_handling_demo/old_name.txt"))
print("New name exists:", os.path.exists("file_handling_demo/new_name.txt"))
if os.path.exists("file_handling_demo/new_name.txt"):
os.remove("file_handling_demo/new_name.txt")
print("File removed safely!")
os.rmdir("file_handling_demo/reports")
print("Empty directory removed!")
Old name exists: False
New name exists: True
File removed safely!
Empty directory removed!
os.rmdir() only works on empty directories; if the folder still has files in it, it raises an error. (Section 13 covers shutil.rmtree(), which removes a directory and everything inside it in one call.)
os.path gives you a set of tools for inspecting and building paths without ever hardcoding a separator:
file_path = "file_handling_demo/learn.txt"
print(f" exists() : {os.path.exists(file_path)}")
print(f" isfile() : {os.path.isfile(file_path)}")
print(f" getsize() : {os.path.getsize(file_path)} bytes")
print(f" basename() : {os.path.basename(file_path)}")
print(f" dirname() : {os.path.dirname(file_path)}")
safe_path = os.path.join("file_handling_demo", "reports", "summary.txt")
print(f"os.path.join(): {safe_path}")
exists() : True
isfile() : True
getsize() : 164 bytes
basename() : learn.txt
dirname() : file_handling_demo
os.path.join(): file_handling_demo/reports/summary.txt
os.path.join() matters more than it looks: on Windows, path separators are \, while Mac and Linux use /. Building the string yourself with + bakes in an assumption about which OS the code will run on. os.path.join() always uses the right separator for whatever machine the code is actually running on.
9. Modern Path Handling: pathlib
os.path works, but Python 3.4 introduced pathlib, a more elegant, object-oriented way to work with paths. Instead of writing:
os.path.join("folder", "subfolder", "file.txt")
you write:
Path("folder") / "subfolder" / "file.txt"
The / operator joins path pieces, and a Path object comes with useful methods built directly onto it, rather than needing a separate function for everything.
from pathlib import Path
p = Path("file_handling_demo/learn.txt")
print(f" Path object : {p}")
print(f" .name : {p.name}") # filename with extension
print(f" .stem : {p.stem}") # filename without extension
print(f" .suffix : {p.suffix}") # just the extension
print(f" .parent : {p.parent}") # parent directory
print(f" .exists() : {p.exists()}")
print(f" .stat().st_size: {p.stat().st_size} bytes")
base = Path("file_handling_demo")
log_path = base / "logs" / "server.log"
print(f"Joined path: {log_path}")
print("All .txt files in file_handling_demo/:")
for txt_file in Path("file_handling_demo").glob("*.txt"):
print(f" {txt_file.name}")
Path object : file_handling_demo/learn.txt
.name : learn.txt
.stem : learn
.suffix : .txt
.parent : file_handling_demo
.exists() : True
.stat().st_size: 164 bytes
Joined path: file_handling_demo/logs/server.log
All .txt files in file_handling_demo/:
...
Path objects can also read and write files directly, without a separate open() call:
p = Path("file_handling_demo/pathlib_test.txt")
p.write_text("Hello from pathlib!\nNo open() needed!\n", encoding="utf-8")
content = p.read_text(encoding="utf-8")
print(content)
new_dir = Path("file_handling_demo/pathlib_dir")
new_dir.mkdir(exist_ok=True)
p.unlink() # Delete the file
new_dir.rmdir() # Remove the directory
print("Cleaned up!")
Hello from pathlib!
No open() needed!
Cleaned up!
os.path and pathlib do overlapping jobs. os.path is the older, function-based approach (os.path.join(a, b)), while pathlib treats a path as an object with its own behavior (a / b). Either works; most new code leans toward pathlib for exactly that readability.
10. Working with CSV Files
CSV (Comma-Separated Values) is one of the most common data formats there is. Export from a spreadsheet, download from a database, pull from an API, and there's a good chance it comes back as CSV.
Why not just use open() and split each line by commas? Because real CSV data has edge cases a plain .split(",") gets wrong: values that contain commas themselves, quoted strings, escaped characters, different delimiters. The csv module handles all of that correctly.
| Tool | Use Case |
|---|---|
csv.reader |
Read CSV rows as lists |
csv.writer |
Write lists as CSV rows |
csv.DictReader |
Read CSV rows as dictionaries (column names become keys) |
csv.DictWriter |
Write dictionaries as CSV rows |
import csv
employees = [
["Name", "Department", "Salary", "Years"],
["Alice Johnson", "Engineering", 95000, 5],
["Bob Smith", "Marketing", 72000, 3],
["Carol White", "Engineering", 105000, 8],
["Dave Brown", "HR", 65000, 2],
["Eve Davis, PhD", "Research", 115000, 10],
]
with open("file_handling_demo/employees.csv", "w", newline="", encoding="utf-8") as f:
writer = csv.writer(f)
writer.writerows(employees)
with open("file_handling_demo/employees.csv", "r", encoding="utf-8") as f:
print(f.read())
Name,Department,Salary,Years
Alice Johnson,Engineering,95000,5
Bob Smith,Marketing,72000,3
Carol White,Engineering,105000,8
Dave Brown,HR,65000,2
"Eve Davis, PhD",Research,115000,10
Notice "Eve Davis, PhD" got automatically wrapped in quotes in the output. That's csv.writer doing exactly the job described above: since that name contains a comma, quoting it is what keeps it from being misread as two separate columns when the file is read back.
with open("file_handling_demo/employees.csv", "r", encoding="utf-8") as f:
reader = csv.reader(f)
header = next(reader)
print(f"Columns: {header}")
for row in reader:
name, dept, salary, years = row
print(f" {name:<15} | {dept:<11} | ${int(salary):>7,} | {years} yrs")
Columns: ['Name', 'Department', 'Salary', 'Years']
Alice Johnson | Engineering | $ 95,000 | 5 yrs
Bob Smith | Marketing | $ 72,000 | 3 yrs
Carol White | Engineering | $105,000 | 8 yrs
Dave Brown | HR | $ 65,000 | 2 yrs
Eve Davis, PhD | Research | $115,000 | 10 yrs
csv.reader gives back plain lists, so you access fields by position (row[0], row[1]...). csv.DictReader reads the header row automatically and lets you access the same data by column name instead, which is far more readable in anything but a tiny script:
total_salary = 0
engineers = []
with open("file_handling_demo/employees.csv", "r", encoding="utf-8") as f:
reader = csv.DictReader(f)
for row in reader:
print(f" {row['Name']}: {row['Department']} (${int(row['Salary']):,})")
total_salary += int(row["Salary"])
if row["Department"] == "Engineering":
engineers.append(row["Name"])
print(f"\nTotal salary bill: ${total_salary:,}")
print(f"Engineers: {', '.join(engineers)}")
Alice Johnson: Engineering ($95,000)
Bob Smith: Marketing ($72,000)
Carol White: Engineering ($105,000)
Dave Brown: HR ($65,000)
Eve Davis, PhD: Research ($115,000)
Total salary bill: $452,000
Engineers: Alice Johnson, Carol White
csv.DictWriter is the mirror image, writing a list of dictionaries out as CSV rows:
new_employees = [
{"Name": "Frank Lee", "Department": "Engineering", "Salary": 98000, "Years": 4},
{"Name": "Grace Kim", "Department": "Design", "Salary": 82000, "Years": 6},
]
fieldnames = ["Name", "Department", "Salary", "Years"]
with open("file_handling_demo/new_employees.csv", "w", newline="", encoding="utf-8") as f:
writer = csv.DictWriter(f, fieldnames=fieldnames)
writer.writeheader()
writer.writerows(new_employees)
with open("file_handling_demo/new_employees.csv", "r", encoding="utf-8") as f:
print(f.read())
Name,Department,Salary,Years
Frank Lee,Engineering,98000,4
Grace Kim,Design,82000,6
Real-World Example: Processing Orders Like Amazon
Amazon processes millions of orders, often stored as CSV records. Here's a simplified order analytics pipeline:
import csv
from collections import defaultdict
orders = [
["order_id", "customer", "product", "category", "price", "quantity"],
["A001", "Alice", "Python Book", "Books", 29.99, 2],
["A002", "Bob", "Wireless Mouse", "Electronics", 45.00, 1],
["A003", "Alice", "Coffee Maker", "Appliances", 89.99, 1],
["A004", "Carol", "Python Book", "Books", 29.99, 3],
["A005", "Bob", "Keyboard", "Electronics", 75.00, 1],
["A006", "Dave", "Coffee Maker", "Appliances", 89.99, 2],
["A007", "Alice", "Headphones", "Electronics", 120.00, 1],
]
with open("file_handling_demo/orders.csv", "w", newline="", encoding="utf-8") as f:
csv.writer(f).writerows(orders)
revenue_by_category = defaultdict(float)
orders_by_customer = defaultdict(int)
with open("file_handling_demo/orders.csv", "r", encoding="utf-8") as f:
for row in csv.DictReader(f):
revenue = float(row["price"]) * int(row["quantity"])
revenue_by_category[row["category"]] += revenue
orders_by_customer[row["customer"]] += 1
print("Revenue by Category:")
for cat, rev in sorted(revenue_by_category.items(), key=lambda x: -x[1]):
print(f" {cat:<15}: ${rev:>8.2f}")
print("\nOrders by Customer:")
for customer, count in sorted(orders_by_customer.items()):
print(f" {customer:<10}: {count} orders")
Revenue by Category:
Appliances : $ 269.97
Electronics : $ 240.00
Books : $ 149.95
Orders by Customer:
Alice : 3 orders
Bob : 2 orders
Carol : 1 orders
Dave : 1 orders
11. Working with JSON Files
JSON (JavaScript Object Notation) is the language most web APIs, mobile app configs, and modern web services communicate in. It maps almost directly onto Python's own data types:
| JSON | Python |
|---|---|
object {} |
dict |
array [] |
list |
string "" |
str |
number |
int / float |
true / false |
True / False |
null |
None |
| Function | Direction | Use Case |
|---|---|---|
json.dump(obj, file) |
Python → File | Save Python data to a JSON file |
json.load(file) |
File → Python | Load JSON from a file into Python |
json.dumps(obj) |
Python → String | Convert Python to a JSON string |
json.loads(string) |
String → Python | Parse a JSON string into Python |
A memory trick worth keeping: dump and load work with files. dumps and loads work with strings (the trailing s is short for "string").
import json
user_profile = {
"username": "alice_wonder",
"email": "alice@example.com",
"age": 28,
"premium": True,
"watch_history": ["Stranger Things", "The Crown", "Breaking Bad"],
"settings": {"autoplay": True, "quality": "4K", "language": "English"},
"last_login": None
}
with open("file_handling_demo/user_profile.json", "w", encoding="utf-8") as f:
json.dump(user_profile, f, indent=4) # indent=4 makes it easy to read
with open("file_handling_demo/user_profile.json", "r", encoding="utf-8") as f:
print(f.read())
{
"username": "alice_wonder",
"email": "alice@example.com",
"age": 28,
"premium": true,
"watch_history": [
"Stranger Things",
"The Crown",
"Breaking Bad"
],
"settings": {
"autoplay": true,
"quality": "4K",
"language": "English"
},
"last_login": null
}
with open("file_handling_demo/user_profile.json", "r", encoding="utf-8") as f:
loaded_profile = json.load(f)
print(f"'premium' is type: {type(loaded_profile['premium'])}") # bool, not string
print(f"'age' is type: {type(loaded_profile['age'])}") # int, not string
'premium' is type: <class 'bool'>
'age' is type: <class 'int'>
Notice the types are preserved perfectly on the way back in: premium comes back as an actual bool, age as an actual int, not text that needs converting.
dumps and loads do the same conversions, but with a string instead of a file, which is exactly the shape you need for talking to a web API:
data = {"name": "Bob", "scores": [95, 87, 92], "passed": True}
json_string = json.dumps(data)
print(json_string)
pretty = json.dumps(data, indent=2, sort_keys=True)
print(pretty)
api_response = '{"status": "ok", "results": 42, "data": [1, 2, 3]}'
parsed = json.loads(api_response)
print(f"Status: {parsed['status']}")
{"name": "Bob", "scores": [95, 87, 92], "passed": true}
{
"name": "Bob",
"passed": true,
"scores": [
95,
87,
92
]
}
Status: ok
Real-World Example: App Settings Like Spotify
Every app stores user settings somewhere, and JSON is the near-universal format for that. Here's how a music app might save and load preferences:
import json
import os
SETTINGS_FILE = "file_handling_demo/spotify_settings.json"
DEFAULT_SETTINGS = {
"audio": {"quality": "high", "normalize": True, "equalizer": "flat"},
"ui": {"theme": "dark", "language": "en", "show_lyrics": True},
"social": {"share_activity": False, "follow_artists": True}
}
def load_settings():
if os.path.exists(SETTINGS_FILE):
with open(SETTINGS_FILE, "r", encoding="utf-8") as f:
return json.load(f)
return DEFAULT_SETTINGS.copy()
def save_settings(settings):
with open(SETTINGS_FILE, "w", encoding="utf-8") as f:
json.dump(settings, f, indent=4)
def update_setting(settings, section, key, value):
settings[section][key] = value
return settings
settings = load_settings()
print(f"Theme : {settings['ui']['theme']}")
settings = update_setting(settings, "ui", "theme", "light")
settings = update_setting(settings, "audio", "quality", "ultra")
save_settings(settings)
reloaded = load_settings()
print(f"After reload, theme: {reloaded['ui']['theme']}")
print(f"After reload, quality: {reloaded['audio']['quality']}")
Theme : dark
After reload, theme: light
After reload, quality: ultra
Settings persist across app restarts precisely because they're written to disk, not just held in a variable that disappears the moment the program exits.
12. Handling File Errors Gracefully
File operations can fail for many reasons. A good program handles these failures gracefully instead of crashing.
| Exception | When It Occurs |
|---|---|
FileNotFoundError |
File or directory does not exist |
PermissionError |
No permission to read or write the file |
FileExistsError |
File already exists (with 'x' mode) |
IsADirectoryError |
Expected a file but found a directory |
IOError / OSError |
General I/O errors |
UnicodeDecodeError |
The file's actual encoding doesn't match the one specified |
The golden rule: only catch exceptions you can meaningfully handle. Don't silently swallow errors just to avoid seeing a crash.
try:
with open("file_handling_demo/ghost_file.txt", "r") as f:
content = f.read()
except FileNotFoundError as e:
print(f"Caught: {e}")
Caught: [Errno 2] No such file or directory: 'file_handling_demo/ghost_file.txt'
UnicodeDecodeError is worth seeing directly, since the fix is easy to miss otherwise:
# Write a file using latin-1 encoding, then try to read it back as ascii
with open("file_handling_demo/encoding_test.txt", "w", encoding="latin-1") as f:
f.write("Café and naïve")
try:
with open("file_handling_demo/encoding_test.txt", "r", encoding="ascii") as f:
content = f.read()
except UnicodeDecodeError as e:
print(f"Caught: {type(e).__name__}")
# The fix: use the encoding the file was actually written with
with open("file_handling_demo/encoding_test.txt", "r", encoding="latin-1") as f:
print(f"Fixed. Content: {f.read()}")
Caught: UnicodeDecodeError
Fixed. Content: Café and naïve
The accented characters (é, ï) are exactly what trips this up: they exist in latin-1 but aren't valid under a strict ascii decode, so Python refuses to guess and raises an error instead of silently corrupting the text.
13. Copying and Moving Files with shutil
The os module can rename and delete individual files, but it doesn't copy files, and os.rmdir() refuses to remove a directory unless it's completely empty. That's where shutil (short for "shell utilities") comes in.
Analogy: if os is you personally carrying one box at a time, shutil is hiring movers who can pick up an entire room, furniture and all, in a single trip.
| Function | What It Does |
|---|---|
shutil.copy(src, dst) |
Copy a single file |
shutil.copy2(src, dst) |
Copy a file, and also preserve its metadata (timestamps) |
shutil.move(src, dst) |
Move or rename a file or folder |
shutil.copytree(src, dst) |
Copy an entire directory, including everything inside it |
shutil.rmtree(path) |
Delete a directory and everything inside it, even if it's not empty |
import shutil, os
# Copy a single file
shutil.copy("file_handling_demo/learn.txt", "file_handling_demo/learn_backup.txt")
print("Backup exists:", os.path.exists("file_handling_demo/learn_backup.txt"))
print("Original still exists:", os.path.exists("file_handling_demo/learn.txt"))
Backup exists: True
Original still exists: True
# Move that backup into its own folder
os.makedirs("file_handling_demo/backups", exist_ok=True)
shutil.move("file_handling_demo/learn_backup.txt", "file_handling_demo/backups/learn_backup.txt")
print("Moved to backups/:", os.path.exists("file_handling_demo/backups/learn_backup.txt"))
print("Old location gone:", not os.path.exists("file_handling_demo/learn_backup.txt"))
Moved to backups/: True
Old location gone: True
# Copy the entire backups/ folder, contents and all, in one call
shutil.copytree("file_handling_demo/backups", "file_handling_demo/backups_copy")
print("Copied folder contents:", os.listdir("file_handling_demo/backups_copy"))
# os.rmdir() refuses to remove a non-empty folder
try:
os.rmdir("file_handling_demo/backups_copy")
except OSError as e:
print(f"os.rmdir() fails on a non-empty folder: {e}")
# shutil.rmtree() removes it anyway, contents and all
shutil.rmtree("file_handling_demo/backups_copy")
print("Removed even though it had files in it:", not os.path.exists("file_handling_demo/backups_copy"))
Copied folder contents: ['learn_backup.txt']
os.rmdir() fails on a non-empty folder: [Errno 39] Directory not empty: 'file_handling_demo/backups_copy'
Removed even though it had files in it: True
Real-World Example: Auto-Organizing a Downloads Folder
A genuinely useful script: sort every file in a messy folder into subfolders by its extension, the same idea behind the "organize by type" feature in some file managers.
import shutil, os
def organize_by_extension(folder):
for filename in os.listdir(folder):
filepath = os.path.join(folder, filename)
if os.path.isfile(filepath):
ext = filename.split(".")[-1].upper()
dest_folder = os.path.join(folder, f"{ext}_files")
os.makedirs(dest_folder, exist_ok=True)
shutil.move(filepath, os.path.join(dest_folder, filename))
organize_by_extension("file_handling_demo/downloads")
for root, dirs, files in os.walk("file_handling_demo/downloads"):
for name in files:
print(os.path.join(root, name))
file_handling_demo/downloads/PDF_files/report.pdf
file_handling_demo/downloads/JPG_files/photo2.jpg
file_handling_demo/downloads/JPG_files/photo.jpg
file_handling_demo/downloads/TXT_files/notes.txt
Every photo, PDF, and text file ends up in its own labeled folder, using nothing but os.listdir() to see what's there and shutil.move() to relocate each one.
Your Python File Handling Cheat Sheet
Opening Files
with open("file.txt", "r", encoding="utf-8") as f:
...
Reading
f.read() # Entire file as a string
f.read(n) # Next n characters
f.readline() # One line at a time
f.readlines() # All lines as a list
for line in f: # Best for large files, memory-efficient
...
Writing
f.write("text") # Write a string (no automatic newline)
f.writelines(list) # Write a list of strings (no automatic newlines)
Navigation
f.tell() # Current cursor position (bytes)
f.seek(0) # Jump to the beginning
f.seek(0, 2) # Jump to the end
os Module
os.getcwd() # Current directory
os.listdir("path") # List directory contents
os.makedirs("a/b/c", exist_ok=True) # Create nested directories
os.rename("old", "new") # Rename a file or directory
os.remove("file.txt") # Delete a file
os.path.exists("path") # Check if a path exists
os.path.join("a", "b", "c") # Build a safe, cross-platform path
pathlib
from pathlib import Path
p = Path("folder") / "subfolder" / "file.txt"
p.read_text(encoding="utf-8")
p.write_text("content", encoding="utf-8")
p.exists(), p.is_file(), p.is_dir()
list(p.parent.glob("*.txt"))
shutil
import shutil
shutil.copy("src.txt", "dst.txt") # Copy one file
shutil.move("src.txt", "dst.txt") # Move or rename
shutil.copytree("src_dir", "dst_dir") # Copy an entire directory
shutil.rmtree("dir") # Delete a directory, even non-empty
CSV
import csv
with open("data.csv") as f:
for row in csv.reader(f): ... # rows as lists
for row in csv.DictReader(f): ... # rows as dicts
with open("data.csv", "w", newline="") as f:
csv.writer(f).writerows(data) # from lists
w = csv.DictWriter(f, fieldnames=[...])
w.writeheader(); w.writerows(data) # from dicts
JSON
import json
json.dump(data, file, indent=4) # Python → JSON file
json.load(file) # JSON file → Python
json.dumps(data) # Python → JSON string
json.loads(string) # JSON string → Python
A program that only holds data in memory forgets everything the instant it stops running. Files, and the tools built around them, open(), os, pathlib, csv, json, and shutil, are what let a program remember. Once reading, writing, and organizing files feels natural, working with real data, spreadsheets, API responses, log files, saved settings, stops feeling like a separate skill and starts feeling like a normal part of writing Python.
A program that only holds data in memory forgets everything the instant it stops running. Files, and the tools built around them, are what let a program remember. Once reading, writing, and organizing files feels natural, working with real data stops feeling like a separate skill and starts feeling like a normal part of writing Python.