A Week's Record of Trying to Predict Denver Bus Delays with Naive Bayes

Last Tuesday night, the Route 15 bus running down Colfax Avenue was late again by 20 minutes. I was sitting on the bench at the stop, watching the estimated arrival time on Google Maps keep getting pushed back. That's when a thought popped into my head: if I collected data on this delay pattern, couldn't I predict in advance whether the bus would be late or not?

I had just learned Naive Bayes in my graduate class. After hearing that it's good for classification problems, simple to implement, and works reasonably well even with limited data, I was itching to try it on something. And since the bus I ride every day is late every day, I figured there couldn't be a more perfect project.

Scraping Together the Data

RTD (Regional Transportation District) is the agency that operates public transit in the greater Denver area, and it provides real-time bus location data through a GTFS-realtime feed. Just figuring that out was the first struggle. After poking around the RTD website for a while, I finally found the developer page and applied for an API key. It took two days to get approved.

As soon as the API key arrived, I wrote a script in Python. It requested the feed every 30 seconds and recorded the difference between the actual and scheduled arrival times for each stop on the Route 15 line. I ran it on my laptop. I was too lazy to set up something like a cron job, so I just left a terminal window open and kept it running. On the first day, I closed my laptop lid before going to sleep and discovered the next morning that the script had died. I repeated this mistake three times over three days.

In the end, my plan to collect a week's worth of data actually resulted in roughly five and a half days of incomplete data. The early morning hours were almost entirely empty, and the weekend data only covered part of Saturday morning. Still, I managed to collect a decent amount of data for weekday rush hours, a single CSV file with a few thousand rows.

Laptop screen showing a spreadsheet of transit delay data next to a coffee mug on a desk

I Built the Model, and It Blew Up

After deliberating over what to use as features, I went with the simple route: day of the week, time of day (split into morning rush/midday/evening rush/night), and stop number. For labels, I did binary classification, 0 or 1 for whether the delay was five minutes or more. I imported scikit-learn's GaussianNB, called fit, then predict. The code itself was surprisingly short, fewer than ten lines.

The accuracy on the training set came out to about 0.72. At first, I thought that was pretty good. But when I ran it on the test set, the precision was dismal. The model was great at correctly labeling non-delayed cases as "no delay," but the answer to the question I actually wanted to know ("Will the bus be late today?"), the rate at which it correctly identified delays as delays, was rock bottom. Since the vast majority of the data was "no delay," the model could just always guess "not late" and still get high accuracy.

Class imbalance. It was a concept I'd learned in class, but encountering it in my own data felt completely different. When I read about it in a textbook, I just thought "okay, makes sense" and moved on. But staring at data I'd spent a week collecting from a bench, it felt pretty frustrating.

I tried oversampling. I tried SMOTE too. Accuracy dropped and recall went up a little, but it was still far from what you'd call practical. More importantly, it was Thursday night when I realized I was fundamentally missing something.

What I Had Overlooked

That night, waiting for the bus at the stop, a thought suddenly hit me. Why is the bus late? The day of the week? The time of day? That's not all. It's late when it snows. It's late when there's an accident on I-25 and traffic spills onto Colfax. It's late on evenings when the Broncos have a home game and the surrounding roads are gridlocked. It's late when there's a construction zone. Not a single one of these was reflected in my model.

You might say I could just add weather data. True, but that wasn't simply a matter of plugging in one more API. The effect of weather on road conditions, the way that effect spreads through traffic flow with a time lag, the different road structures along each route: to incorporate all of that as features, I would have needed a deep understanding of Denver's transportation system and road conditions. It was an entirely different kind of knowledge from writing good code.

I spread out the RTD route map and traced the path of the Route 15 bus stop by stop. The East Colfax section toward Aurora has closely spaced traffic lights, while West Colfax on the other end is relatively open. I realized for the first time, looking at the map, that delay patterns could be completely different from one segment to another on the same route. My model did include stop number as a feature, but it captured nothing about what road that stop was on or what the surrounding environment was like.

Overhead view of a transit route map with handwritten annotations on a cluttered desk

In the end, the most important variables for this project weren't in my CSV file. It wasn't a code problem. It was a data design problem, and the data design problem ultimately stemmed from a lack of domain knowledge.

What I Learned from the Bench

On Friday, I shelved the idea of improving the model. Instead, I just visualized the data I'd collected that week. When I plotted the distribution of delays by time of day using matplotlib, I could clearly see that delays were concentrated between 7:30 and 8:30 in the morning. There was a smaller peak around 5 PM as well. This was something I could have figured out from a single histogram, no ML model needed. I finally understood in my gut why people say you should look at your data before running a complex classifier.

You could call this a failed project, but for me it was a week that taught me more than ten class assignments. Implementing an algorithm is just a small part of the whole process. Knowing what data to collect, in what context, and how to interpret it (in other words, understanding the problem) is where the real weight lies. If I try something like this again, before opening a code editor, I think I'd first ask a bus driver, "Where do you usually get stuck in traffic?"

The Route 15 bus is still late. But now, when I'm standing in front of that delay, I find myself asking slightly different questions.

Comments