§3  Sets & Dictionaries
Python Programming Series  ·  Article 3

Sets and Dictionaries: A Complete Beginner's Guide

Lists and tuples care about order and allow repeats. Sets and dictionaries flip that idea on its head in two different directions: a set cares about uniqueness above everything else, and a dictionary cares about looking things up by name instead of by position. Once you see what problem each one is built to solve, the syntax stops feeling arbitrary.

Sets: unique & unordered
Dictionaries: key-value lookup
2 outdated claims corrected
Real code from Facebook, Amazon, Stripe
Sets
Dictionaries

This guide covers every operation from the source notebook, verified by actually running the code, with two corrections to claims in the original notebook that no longer hold up (one about dictionary ordering, one about what has to be unique). Real examples come from Facebook, Netflix, Amazon, and Stripe.


Part 1: Sets: Unique, Unordered Collections

A set is a collection with one rule above all others: every value in it is unique. There's no way to store the same value twice, and there's no concept of position (first, second, index 3), only membership: is this value in the set, or isn't it?

Creating a Set

python
set1 = {10, 20, 30, 30, 40, 50, 50, 60, 70, 80, 90}
print(set1)
{70, 40, 10, 80, 50, 20, 90, 60, 30}

Notice two things happened automatically: the duplicate 30s and 50s collapsed down to one each, and the printed order has nothing to do with the order you typed the numbers in. Both are the defining features of a set.

   You typed:     { 10, 20, 30, 30, 40, 50, 50, 60, 70, 80, 90 }
                              ▼   ▼           ▼   ▼
                          duplicates silently dropped
                              ▼
   Set stores:    { 10, 20, 30, 40, 50, 60, 70, 80, 90 }   ← unique values only,
                                                                no guaranteed order

Why You Can't Index a Set

Since a set has no positions, asking for "the item at index 3" doesn't mean anything:

python
set1[3]
TypeError: 'set' object is not subscriptable

If you need to grab a specific item, that's usually a sign you wanted a list, not a set.

The Classic Use Case: Deduplicating a List

Because a set can't hold duplicates, converting a list into a set and back is the standard one-line trick for removing repeated values:

python
list1 = [10, 20, 30, 30, 40, 50, 50, 60, 70, 80, 90]
list1 = list(set(list1))
# list1 is now a list with only unique elements

The trade-off: you gain uniqueness but lose the original order (since it passed through a set on the way). If order matters, this isn't the right tool.

Adding Values

python
set1.add(100)               # adds ONE value
set1.update({89, 45, 77})   # adds MULTIPLE values at once

Removing Values

There are three ways to remove from a set, and they behave differently on missing items, which is the part worth remembering:

python
set1.pop()          # removes and returns an ARBITRARY item (you don't choose which)
set1.remove(89)      # removes a SPECIFIC item, by value
set1.discard(90)     # removes a SPECIFIC item, by value

remove and discard look identical until the item you ask for doesn't exist:

python
set1.remove(900)
KeyError: 900
python
set1.discard(900)   # no error at all: the set is simply left unchanged
   remove(900)                          discard(900)
   ┌─────────────────────┐              ┌─────────────────────┐
   │ Item not found?      │              │ Item not found?      │
   │ → raise KeyError     │              │ → do nothing, no error│
   └─────────────────────┘              └─────────────────────┘
   Use when you EXPECT              Use when "it might already
   the item to be there              be gone" is a normal case

Set Math: Union, Intersection, and Friends

This is where sets earn their keep. Every one of these mirrors an operation from math class:

python
set2 = {10, 20, 30, 40, 50, 60, 70}
set3 = {60, 70, 80, 90, 100, 110, 120}

set2.union(set3)                # everything in EITHER set, no duplicates
set2.intersection(set3)         # only what's in BOTH sets
set2.difference(set3)           # in set2 but NOT in set3
set3.difference(set2)           # in set3 but NOT in set2
set2.symmetric_difference(set3) # in one set or the other, but NOT both
set2.isdisjoint(set3)           # True if they share NOTHING in common
        set2                    set3
    ┌─────────────┐        ┌─────────────┐
    │  10  20  30 │        │  80  90 100 │
    │   ┌─────────┼────────┼──────┐      │
    │  40│  50    │  60 70 │      │ 110  │
    └────┼────────┘        └──────┼──120─┘
         │      intersection      │
         │      {60, 70}          │

   union              = everything in either circle
   intersection       = only the overlapping middle {60, 70}
   difference(set3)   = set2's circle, minus the overlap
   symmetric_diff     = both circles, MINUS the overlap

Two more check for relationships between sets rather than combining them:

python
set4 = {10, 30, 40, 50}

set4.issubset(set2)     # True if EVERY item in set4 is also in set2
set2.issuperset(set4)   # the reverse check: is set2 "big enough" to contain all of set4?

Real-World Examples

Facebook: Mutual Friends. Finding overlap between two friend lists is exactly what intersection is for, and since nobody can be "friends twice," the unique-values guarantee of a set fits naturally.

python
my_friends = {"user_101", "user_202", "user_303", "user_404"}
their_friends = {"user_303", "user_404", "user_505", "user_606"}

mutual_friends = my_friends.intersection(their_friends)
print(f"Mutual Friend IDs: {mutual_friends}")
Mutual Friend IDs: {'user_404', 'user_303'}

Netflix: Unique Genres Watched. Your watch history is a list with plenty of repeats (you probably watched more than one comedy). Converting it to a set instantly gives Netflix the distinct genres to show in a "Filter by Genre" row.

python
watched_genres = ["Sci-Fi", "Comedy", "Sci-Fi", "Drama", "Comedy", "Sci-Fi"]
unique_genres = set(watched_genres)
print(f"Your Genre Preferences: {unique_genres}")
Your Genre Preferences: {'Comedy', 'Drama', 'Sci-Fi'}

Instagram: Unique Story Viewers. If the same person watches your story ten times, the view count should still only go up by one. Adding the same ID to a set repeatedly has no effect after the first time.

python
viewers = {"alice_99", "bob_burns"}
viewers.add("bob_burns")       # already present, no change
viewers.add("charlie_brown")   # new, added

print(f"Total Unique Views: {len(viewers)}")
Total Unique Views: 3

Part 2: Dictionaries: Key-Value Pairs

A dictionary stores data as key-value pairs: instead of finding something by its position (like a list), you find it by its name. Look up 'name' and get back 'Shikhar', instantly, no matter how large the dictionary is.

Two corrections before we go further. The original notebook makes two claims about dictionaries that don't hold up when you actually test them, so it's worth clearing these up first:

  1. "Dictionaries are unordered." This was true before Python 3.7, but it hasn't been true for years. Since Python 3.7, dictionaries are guaranteed to preserve insertion order: the order you add keys in is the order you get back when you iterate. You can verify this yourself, and it's also why popitem() (covered below) removes a specific, predictable item rather than a random one; that would make no sense in a genuinely unordered structure.
  2. "Dictionaries only allow unique values." It's actually the keys that must be unique, not the values. Two different keys can happily point to the identical value ({'a': 1, 'b': 1} is completely valid). What Python won't allow is two identical keys; if you try, the second one silently overwrites the first rather than raising an error.

Creating a Dictionary

A dictionary's values can be absolutely anything, including lists, tuples, and even other dictionaries:

python
dict1 = {
    'name': 'Shikhar',
    'place': 'India',
    'roles': ['Data Scientist', 'Data Science Trainer'],
    'hobbies': ('Fitness', 'Football'),
    'education': {1: 'Masters', 2: 'Bachelors'}
}
   dict1
   ┌────────────┬──────────────────────────────────────┐
   │ KEY        │ VALUE                                │
   ├────────────┼──────────────────────────────────────┤
   │ 'name'     │ 'Shikhar'                             │
   │ 'place'    │ 'India'                                │
   │ 'roles'    │ ['Data Scientist', 'Data Science Tr..] │  ← a list
   │ 'hobbies'  │ ('Fitness', 'Football')                │  ← a tuple
   │ 'education'│ {1: 'Masters', 2: 'Bachelors'}          │  ← another dict!
   └────────────┴──────────────────────────────────────┘

Accessing Values

python
dict1['roles']            # access by key, in square brackets
dict1.get('roles')        # the safer alternative, see below
dict1['roles'][1]         # double indexing: get the list, THEN index into it
dict1['education'][2]     # same idea: get the nested dict, THEN look up key 2

dict1['roles'] and dict1.get('roles') return the same thing when the key exists, but they disagree the moment it doesn't: square brackets raise a KeyError and stop your program, while .get() quietly returns None (or a fallback value you specify: dict1.get('missing_key', 'default')). Use .get() whenever a missing key is a normal possibility rather than a bug.

Inserting Values

python
dict1['new_key'] = 45                  # add a single new key-value pair
dict1.update({'new_key2': (12, 43)})   # add one or more pairs at once

If the key you assign to already exists, this same syntax updates it instead of creating a duplicate; that's the practical result of "keys must be unique" in action.

Deleting Values

python
dict1.popitem()        # removes and returns the LAST-inserted pair
dict1.pop('new_key')   # removes a SPECIFIC pair, by key, and returns its value
del dict1['place']     # removes a SPECIFIC pair, by key (no return value)

That popitem() behavior, always removing the most recently added pair, is only possible because Python's dictionaries track insertion order. It's a small piece of evidence for the correction above.

Dictionary Methods

python
dict1_copy = dict1.copy()   # a shallow copy (same caveat as list.copy(), see Article 2)
dict1.keys()                 # every key, as a view
dict1.values()                # every value, as a view
dict1.items()                  # every (key, value) pair, as a view of tuples

.keys(), .values(), and .items() don't return plain lists; they return special "view" objects that stay live if the dictionary changes later. In practice you'll usually loop over them or wrap them in list() if you need an actual list.

Real-World Examples

Amazon: Instant Product Lookup. With millions of products, Amazon can't afford to scan a list top to bottom every time you click something. A dictionary keyed by SKU turns "find this product" into a single instant lookup.

python
amazon_catalog = {
    "B07XJ8C8F5": {"name": "Echo Dot", "price": 49.99, "stock": 120},
    "B08N5WRWJ5": {"name": "MacBook Air", "price": 999.00, "stock": 45},
}

sku = "B07XJ8C8F5"
if sku in amazon_catalog:
    item = amazon_catalog[sku]
    print(f"Product: {item['name']} | Price: ${item['price']}")
Product: Echo Dot | Price: $49.99

YouTube: Video Metadata. The alphanumeric ID in a YouTube URL is the key that maps straight to that video's title, views, and stats, which is why clicking a link loads the right video instantly.

python
youtube_videos = {
    "dQw4w9WgXcQ": {"title": "Never Gonna Give You Up", "views": 1200000000},
    "y6120O5EUPU": {"title": "Python Tutorial", "views": 500000}
}

video_id = "dQw4w9WgXcQ"
print(f"Watching: {youtube_videos[video_id]['title']}")
Watching: Never Gonna Give You Up

Google: DNS Lookup. Computers route traffic by IP address, not by domain name, so somewhere a lookup table maps google.com to its actual address. That's a dictionary at planet scale.

python
dns_table = {
    "google.com": "142.250.190.46",
    "facebook.com": "157.240.22.35",
    "amazon.com": "54.239.28.85"
}

url = "google.com"
print(f"Connecting to {url} at IP {dns_table[url]}...")
Connecting to google.com at IP 142.250.190.46...

Stripe: Friendly Error Messages. A failed payment comes back from Stripe as a technical code like card_declined. A dictionary translates that code into something a customer can actually read.

python
error_messages = {
    "card_declined": "Your card was declined by the bank.",
    "expired_card": "The card has expired. Please use a different one.",
    "incorrect_cvc": "The security code (CVC) is incorrect."
}

code_from_api = "expired_card"
print(f"Error for User: {error_messages.get(code_from_api, 'Unknown Error')}")
Error for User: The card has expired. Please use a different one.

Notice the .get() here comes with a fallback value, 'Unknown Error'. If Stripe ever returns a code that isn't in this dictionary, the program shows a sensible default instead of crashing.


Choosing the Right One

   What are you actually trying to store?
   │
   ├── A collection where duplicates should be IMPOSSIBLE
   │   and you mainly care about "is X in here?"     →  SET
   │
   └── Data you look up by NAME, not by position
       ("give me the value for this key")             →  DICTIONARY

A useful gut check: if you find yourself writing for item in my_list: if item == target, searching a list end to end for a value, that's often a sign a set or a dictionary would do the same job far faster and more clearly.


Quick Reference Summary

Task Set Dictionary
Create {1, 2, 3} {'key': 'value'}
Access an item Not applicable (no positions) d['key'] or d.get('key')
Add one item .add(x) d['key'] = value
Add multiple items .update({a, b, c}) d.update({...})
Remove (errors if missing) .remove(x) del d['key']
Remove (silent if missing) .discard(x) d.pop('key', default)
Remove and return .pop() (arbitrary item) .pop('key') (specific) or .popitem() (last pair)
Copy .copy() .copy() (shallow)
Combine with another .union() .update()
Common elements .intersection() n/a
Elements only in this one .difference() n/a
Elements not shared by either .symmetric_difference() n/a
Relationship checks .issubset() .issuperset() .isdisjoint() n/a
List the keys / values n/a .keys() / .values() / .items()
Guaranteed unique Every value Every key (not values)
Ordered? No Yes, insertion order (Python 3.7+)

Sets and dictionaries solve two very different problems that both happen to rhyme: sets guarantee "no duplicates, ever," and dictionaries turn "search through everything" into "look it up directly." Between lists (ordered, allows duplicates), tuples and strings (ordered, fixed), sets (unordered, unique), and dictionaries (ordered, key-based), you now have a tool suited to almost any shape of data you'll run into.

The shape of the problem tells you which one to reach for. Need to guarantee no duplicates and don't care about order? That's a set. Need to look something up by name instead of scanning for it? That's a dictionary. Both trade the simplicity of a list for a specific, powerful guarantee.