If you've built a few machine learning projects already, a classifier in a notebook, maybe a Kaggle entry, something you're proud enough of to put on GitHub, you've done the hard intellectual part. You know how to clean data, engineer features, train a model, and squeeze out a good accuracy score.
Here's the question this article sits with: what happens the day after that?
To answer your friend when they ask how to actually use your model, you'd need to figure out how to save it so it survives after the notebook closes, and how someone else's computer runs your code without your exact setup. You'd need a way for a website or app to actually ask your model a question and get an answer back. And you'd need an answer for what happens when you improve the model next month (email everyone the new file?), or when the data six months from now doesn't look like the data you trained on.
None of that is a machine learning question. It's an engineering question, and the discipline built to answer it is called MLOps. This article explains what that means, why it exists, and how the rest of this series, starting with Git and GitHub next, fits together.
What this article covers
| # | Topic |
|---|---|
| 1 | The moment every ML learner hits a wall |
| 2 | What MLOps actually is |
| 3 | Why we need it, and the problems it solves |
| 4 | A real example of what happens without MLOps |
| 5 | A real example of what mature MLOps looks like |
| 6 | MLOps vs. DevOps vs. data engineering |
| 7 | The machine learning lifecycle: a loop, not a line |
| 8 | The core pillars of MLOps |
| 9 | A day in the life, before and after MLOps |
| 10 | Common myths about MLOps |
| 11 | The roadmap ahead |
Part 1The moment every ML learner hits a wall
Picture a familiar scene. You spend a weekend on a customer churn model: load the data, clean it up, try a few algorithms, land on a random forest that hits 89% accuracy. You feel good about it. You show a friend, and they ask the obvious question.
"Cool, so how do I actually use it?"
And there's the wall.
Your model lives in a notebook, on your laptop, inside a Python session that vanishes the moment you close it. None of the questions that follow are about machine learning. Your 89% accuracy isn't in doubt. What's in doubt is everything around the model, and that's what MLOps is about.
This gap isn't a minor inconvenience. It's the biggest reason machine learning projects never end up mattering. A commonly cited figure puts it at roughly 87%: for every ten ML projects that get built, eight or nine never make it past the prototype stage into anything a real user or business relies on. That specific number traces back to a single 2019 conference panel, not a rigorous study, so it's worth seeing it alongside a couple of independent estimates rather than treating it as one precise, agreed-upon measurement.
| Source | Year | Finding |
|---|---|---|
| VentureBeat | 2019 | 87% of data science projects never reach production (from a Transform 2019 conference panel) |
| Gartner | 2022 | Only 54% of AI models move from pilot to production |
| McKinsey | 2023 | Only 15% of businesses' ML projects ever succeed |
Different organizations, different years, different definitions of "success," so these three numbers shouldn't be read as three measurements of the same thing. They don't need to agree precisely to make the same point: independently, across different methodologies, the honest answer to "how many ML projects actually make it" keeps landing somewhere between "most don't" and "the large majority don't," even as cloud infrastructure and tooling have improved substantially since 2019. The bottleneck was never really the modeling.
MLOps is the answer to that bottleneck.
Part 2What is MLOps, really?
MLOps stands for machine learning operations. It's a set of practices, tools, and habits that bring the discipline of software engineering, reliability, automation, testing, versioning, to the messier, constantly changing world of machine learning.
The name borrows from DevOps (development plus operations), the discipline that changed how regular software gets built and shipped over the last fifteen years. DevOps solved a familiar problem: developers wrote code that worked on their machine, then handed it to an operations team who had to somehow run it reliably in the real world, and chaos followed. DevOps fixed that by making development and operations work as one continuous process instead of two disconnected teams.
MLOps takes that same idea and extends it to cover the parts of machine learning that DevOps was never built to handle.
The one difference that matters
Here's the idea the rest of this article, and honestly the rest of this series, hangs on: in traditional software, only the code changes. In machine learning, the code, the data, and the model all change, independently of each other, and any one of them can quietly break the system.
A traditional web app doesn't spontaneously get worse over time just by sitting there. A machine learning model can. Customer behavior shifts, a market crashes, a competitor changes the game, a pandemic rewrites what "normal" data even looks like, and your model, which hasn't changed one line of code, starts making confidently wrong predictions. Nothing crashed. No error was thrown. It just quietly stopped being right. We'll see this happen to a real company in Part 4.
Building a car vs. running an airline. Here's a way to feel the difference, not just understand it.
Training a good model is like a skilled engineer building one excellent prototype car. Given enough time, talent, and a controlled test track, a good engineer can build a car that performs beautifully. That's the data science part, and it's genuinely hard work.
Running that model in production is more like running a commercial airline. An airline doesn't build one great plane and hope for the best. It runs pre-flight checklists, schedules maintenance, keeps redundant systems on standby, has air traffic control watching every flight in real time, and monitors weather closely enough to ground flights before a storm arrives. Every aircraft carries a black box recorder, because when something does go wrong, someone needs to know exactly what happened. None of this is about whether the plane can fly. It's about making sure it flies safely, every day, for years, as conditions around it keep changing.
MLOps is the checklist, the maintenance schedule, the air traffic control, and the black box recorder for machine learning. It's not really about building a better model. It's about keeping the good model you already built flying safely long after you've stopped watching it.
Part 3Why do we need MLOps? The problems it solves
Let's get specific. Here are the recurring pains MLOps exists to fix, the same pains that add up to the 87% failure rate from Part 1.
| Problem | What it looks like in practice |
|---|---|
| Irreproducibility | "The model in production was trained six weeks ago. Nobody can say exactly which code, which data, and which hyperparameters produced it." |
| The handoff gap | The data scientist who built the model speaks Jupyter and pandas. The engineer who has to run it speaks servers and uptime. Neither fully understands the other's world, and the model gets stuck in the middle. |
| "Works on my machine" | The model runs perfectly on the data scientist's laptop and immediately breaks on the server: different Python version, missing system library, a package that silently upgraded. |
| Manual, error-prone deployment | Shipping a new model version means someone manually copying a .pkl file onto a server at 11pm and hoping nothing else changed. |
| Silent model decay | The model was 92% accurate at launch. Nobody re-measures it against new data, so nobody notices it's now effectively guessing. |
| No audit trail | A regulator, a customer, or your own CEO asks why the model made this exact decision for this exact customer, and there's no record good enough to answer. |
| Duplicated effort | Every new ML project reinvents its own ad-hoc way to save models, serve predictions, and track experiments, because there's no shared, repeatable system. |
Every tool on this roadmap, Git, DVC, Docker, Flask, FastAPI, MLflow, Airflow, GitHub Actions, Terraform, Kubernetes, Prometheus, exists to close one or more of these specific gaps. None of it is complexity for its own sake. These are practical answers to problems that have already cost real companies real money. Let's look at two of those companies.
Part 4What happens without MLOps
In 2018, the real-estate company Zillow launched a business called Zillow Offers. The idea: use a machine learning pricing model, built on the company's "Zestimate" technology, to make instant cash offers on homes, buy them directly from sellers, and resell them for a modest profit. No negotiation, no waiting around, just the algorithm pricing homes at scale.
For a while, it worked reasonably well. Then the world changed faster than the model could keep up.
When the pandemic hit, the U.S. housing market didn't just move, it convulsed. Prices swung unpredictably, inventory patterns broke from anything in the historical record, and the statistical relationships the model had learned from years of past housing data stopped holding true. This is a textbook case of concept drift: the real-world relationship between a home's features and its true market price had shifted, but the model kept scoring homes as though the old, stable market still applied.
Making things worse, reports indicate Zillow's own management pushed the algorithm to make more aggressive offers to hit growth targets, turning up the dial on a model that was already working off outdated assumptions, instead of having an independent system checking whether those assumptions still held.
By November 2021, the damage was done. Zillow disclosed a write-down of more than $500 million on homes it had bought for prices that no longer made sense, and shut Zillow Offers down entirely, laying off roughly a quarter of the unit's workforce.
This wasn't a story about a bad algorithm. Zillow had skilled data scientists and enormous amounts of data. It was a story about missing operations around the model: no automated system was comparing the model's predictions against reality, nothing was flagging that the input data no longer resembled the training data, no early-warning dashboard existed to catch the problem before it became a half-billion-dollar one. Monitoring, drift detection, and automated retraining triggers are core pieces of MLOps, and we'll build toward all three later in this series, especially the observability tools in Phase 4 of the roadmap in Part 11.
Part 5What mature MLOps looks like
Now the other side of the coin. Uber runs machine learning at a scale that's hard to picture: trip time estimates, driver-rider matching, fraud detection, UberEats delivery predictions, and dozens more, running continuously, worldwide.
Uber didn't get there by having each team build its own deployment process from scratch. Starting in 2016, the company built an internal platform called Michelangelo, meant to take a model from a data scientist's experiment to a reliable, monitored production service, repeatably, for any team, without reinventing the process each time.
The reason for building it is worth sitting with. Internally, Uber found that machine learning's impact across the company was limited by how much engineering effort it took to turn any single local model into something running reliably in production. The modeling itself wasn't the bottleneck. The road from notebook to production was.
Today, Michelangelo manages roughly 400 active ML projects and generates more than 20,000 model training jobs a month, covering the full lifecycle: managing data, training, evaluating, deploying, serving live predictions, and continuously monitoring how those predictions hold up in the real world.
Michelangelo didn't make Uber's models smarter. It turned the path from idea to reliable production into a solved, reusable capability, the opposite of what happened at Zillow. That's MLOps at its best: not a single clever tool, but a set of shared, engineered systems that make going to production safely the default outcome of an ML project instead of a stroke of luck.
Part 6MLOps vs. DevOps vs. data engineering
MLOps didn't appear out of nowhere. It sits at the intersection of three existing disciplines, borrowing something from each.
Here's a direct comparison, because the differences from plain DevOps are exactly where beginners get tripped up.
| Aspect | DevOps (traditional software) | MLOps (machine learning) |
|---|---|---|
| What gets versioned | Code | Code, plus data, plus trained models |
| What "testing" means | Does the code behave correctly? | Does the code work, and is the model's accuracy, fairness, and performance acceptable? |
| Why quality degrades | Bugs introduced by new code changes | Bugs from code changes, plus silent decay from the real world changing (drift) even when no code changed |
| The core pipeline | CI/CD (continuous integration / continuous deployment) | CI/CD plus CT (continuous training: retraining models as new data arrives) |
| What "rollback" means | Revert to a previous code version | Revert the code, and the model, and possibly the data snapshot |
| Who's involved | Developers and operations engineers | Data scientists, ML engineers, and operations engineers, sharing one workflow |
MLOps also overlaps heavily with data engineering, the discipline of building reliable pipelines that move and transform data, since a model is only as trustworthy as the data feeding it. In practice, MLOps absorbs the parts of data engineering concerned with getting clean, versioned, monitored data into a training and serving pipeline.
Part 7The machine learning lifecycle: a loop, not a line
Here's a mental model shift that changes a lot: traditional software has a release cycle. You ship version 1.0, then work toward 2.0. Machine learning has a lifecycle loop that never really stops, because the moment a model is deployed, the real world starts drifting away from the data it was trained on, and the whole cycle has to run again.
Notice step 6 has two arrows leaving it: one loops back to step 1 (retrain, because the world moved), and one just continues serving (because things still look healthy). A production ML system is never really finished. It's a loop that has to keep spinning correctly, indefinitely, which is why operations has to be a real part of the discipline and not an afterthought.
Part 8The core pillars of MLOps
Everything MLOps does can be sorted into a small number of foundational pillars. Think of these as the load-bearing columns holding up a reliable production ML system. Remove any one of them, and the whole structure gets shaky.
Versioning means code, data, and models are all tracked, all the time, so you can always answer what exactly produced this prediction. Articles 1 through 3 of this series cover Git and DVC, the tools that make this possible.
Automation, often shortened to CI/CD/CT, means building, testing, training, and deploying happen through repeatable, automated pipelines instead of manual, error-prone steps. GitHub Actions and orchestration tools like Airflow cover this later in the series.
Reproducibility means that given the same code, the same data, and the same parameters, you get the exact same model every time, on any machine, not just yours. That's what Docker guarantees, and what DVC and Git make possible together.
Monitoring and observability means someone, or something automated, is always watching how the model performs against fresh, real-world data, so decay gets caught in days instead of in a half-billion-dollar write-down. This is the pillar that was missing at Zillow, and it's what Prometheus, Grafana, and Evidently AI will give you later in the series.
Part 9A day in the life, before and after MLOps
Let's make this concrete with one running example: a subscription business's customer churn model. Here's how the same situations play out with and without MLOps.
| Situation | Without MLOps | With MLOps |
|---|---|---|
| New data arrives at month-end | A data scientist manually reruns their notebook, eyeballs the new accuracy, and, if it looks fine, manually copies a new model file onto the server. | An automated pipeline retrains the model on schedule, evaluates it against the current production model, and only promotes it if it's actually better. |
| A bug shows up in production | Nobody can say for certain which code, which data, and which model version is even currently live. | Git, DVC, and a model registry together tell you the exact code commit, data version, and model version running right now. |
| Comparing two experiments | Accuracy numbers live in scattered notebook print statements, Slack messages, and someone's memory. | A tool like MLflow shows every run's parameters, metrics, and artifacts side by side, searchable and permanent. |
| The model quietly gets worse | Nobody notices until a customer complaint, a bad business quarter, or an executive asks an uncomfortable question. | Automated drift detection and dashboards flag the degradation within days, before it becomes expensive. |
| The data scientist who built it leaves | The model becomes an unmaintainable black box that nobody dares touch. | Everything, code, data, environment, model, is versioned and documented, so any team member can pick it up. |
| Traffic suddenly spikes | The single server hosting the model falls over, and predictions stop entirely. | Container orchestration (Kubernetes) automatically scales up more instances to absorb the load. |
Every row in the "With MLOps" column maps to a tool or practice on this series' roadmap, which is the whole point of Part 11.
Part 10Common myths about MLOps
Before the roadmap, a few misconceptions worth clearing up.
One myth: MLOps is just DevOps with a new name. It isn't, not quite. MLOps borrows heavily from DevOps, but it has to solve problems DevOps was never designed for. Code doesn't spontaneously get worse by sitting still. Data does, quietly, through drift. That difference is why data and model versioning, continuous training, and drift monitoring exist as entirely new categories of tooling that DevOps never needed.
Another myth: you need Kubernetes, a cloud budget, and a platform team to "do" MLOps. You don't. MLOps is a set of practices, not a specific stack. Someone running Git, DVC, and Docker on a laptop is already practicing real MLOps principles, versioning and reproducibility, just at a smaller scale than Uber's. This series builds up to Kubernetes and cloud infrastructure later because you'll understand why you need them by the time you get there, not because they're required on day one.
A third myth: MLOps only matters once you're "in production." It doesn't. The habits that make production reliable, versioning data alongside code, writing reproducible training scripts, tracking experiments properly, are much easier to build from day one than to retrofit later. Waiting until "it's serious" to start caring about MLOps is exactly how projects end up in the 87% that never make it.
Part 11The roadmap ahead
Here's how the tools in this series map onto everything you've just read, organized into four phases that mirror the pillars from Part 8 and the lifecycle loop from Part 7.
This four-phase shape isn't specific to this series. Google Cloud publishes its own MLOps maturity model with almost the identical progression, in three levels instead of four: Level 0 is a fully manual process, someone runs every step by hand. Level 1 adds ML pipeline automation, the training pipeline itself runs on a schedule and retrains automatically. Level 2 adds full CI/CD for the pipeline, so changes to the training code get tested, built, and deployed the same automated way changes to a web app would be. Phases 1 and 2 of this roadmap build toward Google's Level 1; Phase 3's GitHub Actions work is what actually gets you to Level 2. The point isn't that the numbers match exactly, they don't need to, it's that "manual, then an automated pipeline, then a fully automated pipeline" is how the industry itself describes MLOps maturity, not a framework invented for this series.
Let's slow down and go tool by tool before zooming back out to the phase level.
Tool by tool: the problem each one actually solves
Every tool on this roadmap exists because of one specific, recurring problem, not because it's trendy. To make this concrete, we'll follow one running example the whole way through: the customer churn model from Part 9, as it moves from a data scientist's laptop to a real, monitored production system.
1. Git and DVC
Regular version control tracks code well, but it was never built to handle datasets and trained model files: large binary blobs that bloat a repository and can't be meaningfully diffed. Without a way to version data and models alongside code, nobody can answer the question that matters most after something goes wrong: what data, what code, and what model produced this exact prediction?
Here's how that plays out. Three data scientists each retrain the churn model on slightly different data pulls over a month. Six weeks later, leadership asks the team to reproduce the exact model that flagged a batch of high-value customers as "high churn risk" last quarter, because that flag triggered a costly retention campaign and finance wants to audit it. Nobody can say for certain which data snapshot and which code version produced that model. With Git tracking the code and DVC tracking the data, linked by a small pointer file Git can version, the team runs git checkout to the right commit and dvc checkout to pull the matching data, and reproduces the exact model, byte for byte, months later.
2. Docker
A model that trains and predicts correctly on a data scientist's laptop can behave differently, or simply crash, the moment it moves to another machine, because of differences in Python version, library versions, or missing system dependencies. This is the classic "works on my machine" problem.
The churn model was built and tested on a laptop running Python 3.11 and scikit-learn 1.4. Handed off to the company's production server, which runs Python 3.9 with an older scikit-learn, predictions come out subtly different, and a required system library is missing entirely, crashing the service. Docker packages the exact Python version, exact package versions, and any system dependencies into one portable container. The same container that ran correctly on the laptop runs identically on staging and production, and this whole category of bug disappears.
3. Flask, then FastAPI
A trained model saved as a .pkl file can only be used by loading it inside a Python script. Real consumers of a prediction (a customer support dashboard, a mobile app, another internal service) are rarely written in Python and shouldn't need to know anything about pickle files or scikit-learn.
The support team's dashboard, built in a completely different stack, needs to show a live "churn risk" badge next to each customer's name. Wrapping the model in a FastAPI endpoint (POST /predict) turns it into a universal service: the dashboard sends a customer's data as plain JSON over HTTP and gets a churn probability back, with no knowledge of Python, pickle files, or how the model works internally.
4. MLflow, experiment tracking
Finding the best-performing version of a model usually means running many experiments across different algorithms, features, and hyperparameters. Without a system to track them, results end up scattered across notebook print statements, an outdated spreadsheet, and half-remembered Slack messages, and it becomes nearly impossible to say with confidence which configuration was actually best.
Over two weeks, the team runs 40 variations of the churn model: logistic regression, random forests, gradient boosting, each with different features and hyperparameters. MLflow logs every run's parameters, metrics, and output artifacts to one searchable place automatically. Instead of digging through notebooks, the team opens the MLflow UI, sorts by F1 score, and finds that run 27, a random forest with 200 trees and max depth 8, outperformed everything else, with the exact code and data version that produced it already attached.
5. MLflow, model registry
Even after finding the best model, there's usually no formal, safe way to declare "this is the model serving real customers right now" versus "this is a candidate we're still testing." Teams fall back on renaming files by hand, something like model_v2_FINAL_use_this_one.pkl, and that system breaks down the moment more than one person touches it.
The team promotes run 27 from staging to production inside the MLflow Model Registry with a single action. The live prediction API always loads whichever model is currently tagged production, so releasing a better model later is a clean stage transition, not a manual file swap at 11pm that someone might get wrong.
6. Apache Airflow or Prefect
A production model needs fresh data on a regular schedule to stay accurate, which usually means running several scripts in a strict order: extract, clean, train, evaluate. Relying on a person to remember to run these manually, in order, every single week, is exactly the kind of fragile process that breaks the moment that person is on vacation or just forgets.
The churn model needs retraining every Monday on the past week's customer activity. An Airflow DAG defines "extract, then preprocess, then train, then evaluate" as a dependency graph that runs automatically at 2am every Monday, retries any step that fails, and alerts the team the moment something breaks, with no human needing to remember any of it.
7. GitHub Actions
A developer can fix a bug in the code and push it to Git, then forget to rebuild and redeploy the Docker image that actually runs in production. The "fixed" code sits safely in the repository while the live system keeps running the old, buggy container, and this mismatch can go unnoticed for weeks.
A data scientist fixes a preprocessing bug in the churn model's feature pipeline and pushes the change. A GitHub Actions workflow triggers automatically: it runs the unit tests, builds a fresh Docker image if they pass, and pushes the new image to the container registry, so the code in Git and the container running in production can't quietly drift apart again.
8. Terraform
Cloud infrastructure (servers, networking rules, storage) often gets set up by clicking through a cloud provider's console by hand. That works until someone needs to recreate the setup later: a new staging environment, a disaster recovery scenario, or simply because the person who originally clicked through it has left the company and nobody remembers the settings.
The engineer who set up the churn model API's server, load balancer, and storage bucket on AWS leaves the company. Six months later, the team needs an identical staging environment to test a new model version safely, and nobody can recall the exact configuration. With Terraform, that infrastructure was described as code from the start, so recreating it identically is one command, terraform apply, not a two-hour guessing game through a console.
9. Kubernetes
A model API running on a single server handles everyday traffic fine, but has no way to automatically absorb a sudden spike in demand, and a server that falls over during a spike takes down service for every user, not just the extra ones.
Every Monday morning, the sales team runs a bulk churn-risk check against thousands of customers ahead of their weekly planning meeting. On a single server, that spike overwhelms the churn API and it stops responding for everyone, including customer support agents relying on it in real time. Kubernetes detects the load increase, starts additional copies of the API to absorb it, then scales back down once the spike passes, with no 3am page and no manual server provisioning.
10. Prometheus, Grafana, and Evidently AI
This is the exact gap that sank Zillow Offers back in Part 4: a model's real-world performance can degrade slowly and silently as the data it sees drifts away from what it was trained on, and without active monitoring, nobody notices until the damage shows up in a quarterly report.
Three months after launch, the churn model's accuracy quietly drops because a pricing change shifted customer behavior in ways the training data never saw. Prometheus continuously collects operational metrics like request latency and volume. Grafana turns them into a live dashboard. Evidently AI compares the live customer data hitting the model against the original training distribution, and flags the drift within days of it starting, not three months later in a boardroom.
Ten tools, ten separate scenarios. It's worth stepping back and looking at the same churn model moving through all of them as one continuous journey, rather than ten disconnected stories, since that's what actually happens to it in a real pipeline.
Read left to right, then right to left: the notebook becomes versioned, then portable, then reachable, then measured, then formally promoted, then automated, then continuously shipped, then able to absorb load, then watched. Nothing on this path is optional if the goal is a system real users can depend on, and nothing on it requires all ten tools on day one, that's exactly what the four phases below are for.
Now, zooming back out to the phase level: here's what each of these phases builds toward as a whole.
Phase 1 — environment, versioning, and packaging (days 1–7)
This phase builds the versioning and reproducibility pillars from Part 8. It runs across Articles 1 through 7, right after this one.
- Git and DVC (Articles 1–3) version your code, your data, and your trained models, so you can always answer what exactly produced this.
- Docker (Articles 4–5) packages your code and every dependency into one portable unit, killing the "works on my machine" problem for good.
- Flask, then FastAPI (Articles 6–7) wrap your trained model in an API so anything, a website, an app, another service, can ask it for a prediction.
By the end of Phase 1, the gap from Part 1 closes: your model won't be trapped in a notebook anymore. It'll be reproducible, portable, and reachable over a network.
Phase 2 — experiment tracking and model registries (days 8–14)
This phase brings order to the experimentation side of the lifecycle loop, steps 3 and 4 from Part 7. Train ten variations of a model without it, and you'll likely end up tracking results in scattered print statements or a messy spreadsheet.
- MLflow tracking logs every experiment's parameters, metrics, and output artifacts automatically, so comparing runs stops being guesswork.
- The MLflow model registry gives each trained model a managed lifecycle, moving it formally through stages like staging, production, and archived, instead of "the latest file someone happened to copy onto the server."
This phase directly answers the "comparing two experiments" and "which model is live" rows from Part 9's table.
Phase 3 — orchestration and CI/CD (days 15–22)
This is where the automation pillar takes over, turning the manual steps from Part 9's "without MLOps" column into something that runs on its own.
- Apache Airflow or Prefect let you write DAGs, directed acyclic graphs, that run your pipeline's steps in the correct order (extract data, preprocess, train, evaluate) on a schedule, without anyone kicking each one off by hand.
- GitHub Actions automatically tests, builds, and packages your Docker image the moment new code is pushed, then pushes the result to a container registry, closing the loop between "I changed the code" and "the new version is ready to deploy."
This phase is what makes continuous training (CT), the ML-specific addition to CI/CD from Part 6's comparison table, actually real instead of theoretical.
Phase 4 — deployment, scaling, and monitoring (days 23–30)
The final phase builds the monitoring and observability pillar, the exact one that was missing at Zillow in Part 4.
- Terraform provisions the cloud infrastructure your API needs through code, instead of clicking through a cloud console by hand and being unable to reliably repeat what you clicked.
- Kubernetes manages your containerized API at scale, so a traffic spike gets absorbed automatically instead of taking your service down.
- Prometheus, Grafana, and Evidently AI expose real operational metrics like latency and request volume, visualize them, and, critically, compare incoming data against your original training data to catch drift before it becomes a Zillow-sized problem.
By the end of Phase 4, every row in Part 9's "With MLOps" column is a system you'll have actually built, not just a concept you've read about.
Summary
Before you move on to Article 1, here's what should be locked in. A working notebook and a reliable production system are two different achievements, and closing that gap is what this series is for. MLOps takes DevOps principles and extends them to handle the fact that in ML, code, data, and models all change independently, and any one of them can quietly break things. Skipping it costs real money: irreproducibility, handoff friction, "works on my machine," manual deployment, silent model decay, and missing audit trails added up to Zillow's $500 million write-down. Done well, it looks like Uber's Michelangelo platform, which turned "get a model to production reliably" into a solved, reusable capability across 400-plus projects.
The whole thing rests on four pillars: versioning, automation, reproducibility, and monitoring. And the lifecycle itself is a loop, not a line. Production ML is never really finished, it just keeps spinning, which is why operations has to be part of the discipline from the start.
From here, the roadmap runs in four phases: foundational versioning first, then experiment tracking, then automation, then deployment at scale with real monitoring.
Sources and further reading
- VentureBeat, "Why do 87% of data science projects never make it into production?" (2019). The original source of the widely cited failure rate.
- Stanford Graduate School of Business, "Flip Flop: Why Zillow's Algorithmic Home Buying Venture Imploded." Background on the Zillow Offers shutdown.
- Uber Engineering Blog, "Meet Michelangelo: Uber's Machine Learning Platform" and "From Predictive to Generative: How Michelangelo Accelerates Uber's AI Journey." Background on Uber's internal ML platform.
You'll start building the first pillar, versioning, with your own hands in Article 1. Everything in this article was the why. Starting with Article 1, it's all how.