Tuesday, April 5, 2022

Why I Call Myself a Data Scientist

​This week, I am at the INFORMS Business Analytics conference, one of the two conferences I attend regularly as an Operations Research PhD and enthusiast​. In fact, INFORMS conferences ​are the only ones ​I have attended at all since graduation.​ In my work, I identify myself as a data scientist. What is interesting about these two facts is that INFORMS is not even on the map when it comes to data science.


What I find most valuable about my background in Operations Research is that by the end of your PhD for sure, and likely after a Masters, you have internalized one key lesson: The problem is always up for discussion. Unfortunately, you don't receive that lesson explicitly. Instead, what you get is a series of courses focused on "reformulating" problems. As an example, you learn that linear problems are easiest mathematically, and so you use your training to rewrite problems as linear subproblems. After spending 2-6 years rewriting problems it becomes crystal clear that the first way you think of writing down a problem is unlikely to be the best.


This mindset puts you in a place to succeed as a data scientist (and also as a consultant). Traditionally, the best data scientists are able to take a business problem and understand how to leverage ML and tools from analytics to solve that problem. Data science training programs focus on teaching you primarily how to use ML algorithms and code. This puts Operations Researchers in an odd position as they have so many of the hard-to-find skills on the business side that make exceptional data scientists, but often are lacking the ML skills that are considered "table stakes" for these roles.


As a result of all this, I call myself a data scientist. It is expedient and people generally hand me the right kinds of problems when I market myself that way. At some point they catch on or I warn them that not all data scientists approach problems the same way I do. Depending on the setting I go so far as to explain Operations Research and why they should consider hiring OR professionals for their data science needs. However, what I really wish is that the mindset of OR became the common framework for all data scientists. By articulating the notion that the problem is always up for discussion, you start to realize how much of your value comes not from just solving the problem you were asked to solve, but from getting to the why and how along the way.

Saturday, December 26, 2020

My bathroom non-remodel

In graduate school I took a class called “Systems Engineering” somewhat on a whim since it was being taught by the former secretary of the navy and I like systems. In preparing to now teach the same course in the spring quarter for DU, I have been reflecting on the guiding principles of the discipline and how I can convey those to my students.

I feel “mindset” is the most distinctive feature of many of my favorite disciplines (and what makes me a good consultant). What sets Operations Research apart is the perspective that the problem is up for discussion as well, not just the solution. Similarly, lean engineering is a way of understanding and improving production systems. I see systems engineering as also primarily a perspective and set of tools: one focused on how people can design large systems that succeed.

Today’s illustration however is not a large system. Namely, it is a home improvement project I was contemplating but hesitant to move forward with. When we first bought our house 3 years ago, I was suspicious of the shower bench in the master bathroom. It seemed to have serious mold issues, and as a former Michigander I am incredibly suspicious of mold. A month later, I had patched in new tile around the bench and felt confident for the short term but knew within a few years I would want to do the whole shower properly.

After finishing my last project this summer, I have been debating how urgent the master bathroom project is. It seems like the next major project I should tackle, but should I tackle it now? With three kids home and work to do, it seemed like an obvious “no.” But still, the molding grout and caulk shouted for something to be done.

And then I realized I was making a classic mistake I learned about in systems engineering. I considered plenty of “revolutionary” alternatives (should I redo the whole bathroom, or just the shower?), but had forgotten to include an “evolutionary” alternative (leave things mostly the same, but re-caulk). One of the important lessons in Systems Engineering is to contemplate your alternatives carefully. If you are not mindful of the alternatives, you often end up with a sub-optimal solution. Like a large remodel in the middle of a pandemic.

I am happy to report that my shower is once-again mold free. And with just a couple hours invested!

Thursday, November 7, 2019

Computers like to cheat

One year ago I got to hear Janelle Shane speak about her blog "AI weirdness." Her illustrations helped me start to understand sort of the... logic? of more sophisticated machine learning algorithms like neural nets.

This week, her book on the same topic was published and I got to hear her speak again. First off, I highly recommend both the blog and the book. Her justification for writing "You Look Like a Thing and I Love You" is that while we have many examples in science fiction of super-smart AIs like C3PO and Ultron, we don't actually have examples of AI as we have it today. Machines with brains maybe as powerful as a worm, but who are being trusted to screen applicants and drive cars.

The part that I continue to find most interesting in her presentations is the subject of my blog. She explains that while sometimes you don't have enough data to train an algorithm, much more often the problem is that you asked the computer to solve the wrong problem. You wanted a computer to caption images, but neglected to mention that "I'm not sure" is an acceptable answer. Or you asked a computer to make unbiased hiring decisions, but fed it discriminatory examples. These may sound like isolated examples, but Janelle's presentation helps you to understand that computers are always going to take advantage of the smallest oversights in your problem statement to win at the wrong task.

This intuitive idea that the computer will be trying to "cheat" any way it can is helpful as we navigate hype around AI and when we should actually trust it. Can this image recognition software tell the difference between dogs and wolves? Or has it just learned that wolves are often photographed in snow. Should we trust an algorithm because it has a lot of training data? Or might that data have important holes on issues we care about.

None of this is to say ML and AI don't have a place in the world today. But it does help us as individuals in the modern era understand how our lives may be changing for both good and bad as more decisions are handed over to computers.

I'm also going to throw out her Ted talk from last week, if you're not quite ready to commit to a whole book.

Friday, September 7, 2018

Not that multiverse

I recently started reading the book "Fooled by Randomness" by Nassim Taleb. So far it is not a book I would recommend to most people (the person who suggested I read it said he usually recommends people start with his most recent book, Antifragile). The author covers very interesting content, but not in a way that is easy to follow or digest. This is the first of probably (hopefully?) a series of posts trying to translate the subject of Taleb's book to an easier to digest format.

While I lived in Ann Arbor during graduate school there was a turn I had to drive about once a month. The unfortunate thing about this turn was that it was a left turn immediately after taking a left at a light. It was so close that I had to make a decision: either move into the middle lane of the road, which was a left turn lane for traffic coming the opposite direction, or remain in the line of traffic, and wait for any oncoming traffic to be clear.

After a few times of taking the turn, I wondered which of the two not-great options I should choose going forward. It seemed to me that I could either risk a low likelihood of a head-on-collision, or a relatively higher likelihood of being rear-ended in the other lane. I settled on staying in my lane and risking being rear-ended because of how much more destructive head-on collisions are.

A few years after making the decision, I made the left turn, and waited for the oncoming traffic to clear as usual. The person who was driving behind me saw the brake lights and stopped. Unfortunately, the person behind them didn't and bumped the middle car into mine. It was fairly minor damage all around, but it is easy to wonder given what happened if I actually made the right choice.

One of the messages from Taleb is that there is complexity in judging the quality of a decision based on random outcomes. For the person who bought a lottery ticket and won, it was a good decision. However, we should advise each person not to buy lottery tickets because in most versions of the universe, the individual you are talking to does not win. 

This notion of "most versions of the universe" is a useful one when talking about randomness since it lets you still give weight to things that didn't happen. And while it can be a good idea to update your estimates of probabilities as you get new information, the fundamentals before an event are the same as they are after. As an example, after being rear-ended I did conclude that maybe I should be a bit more aggressive in taking my turn between oncoming traffic. But the fundamentals of my decision didn't change because of it.

Tuesday, April 17, 2018

Soft vs. Hard constraints

​Last week, at a meeting to prepare for an on-site kickoff with a client, I was asked if I had any real-life examples of the "squishy rules"​ I wanted to discuss with the customer. At first nothing was coming to mind, but my airline helpfully solved that problem for me on my way to the kickoff.

After I had started my first flight, my second flight was cancelled. I found myself in the customer service line behind several other people also trying to figure out how to satisfy their constraints and priorities in the best way possible (scheduled meetings the next day, no private jets, how far were they willing to drive a rental car). What struck me was how much those constraints and priorities varied among the 4 people ahead of me in line. Some people were fine with getting in the next night, others (like me) were willing to give up anything except being on time the next day.

Now, you may have noticed above that I combined constraints and priorities into a single list. When I had booked my flight, I chose to fly to the actual city I was headed to. Once that flight was cancelled, I had a choice to make. What used to be two hard constraints now gave me zero "feasible solutions" -- I could either miss one day of the one-and-a-half day kickoff, or I needed to fly to a different city. Now, some very creative people find themselves in this situation and will fly to some other middle city and then to their destination. But my airline either didn't or couldn't suggest those options, and if you had asked me before the cancellation if I would consider a 3-leg trip, I would have given a flat no. So if I had no possible solutions, what could I do?

Well, this happens a lot. People will often list their preferences as needs until pressed. And as long as there is a feasible solution, it doesn't have to become obvious. One of the people ahead of me in line chose not to give up any of their hard constraints, which meant there were still no options available. It was obvious that something had to give unless the goal had changed -- nevermind, I didn't need to go to that city after all. But knowing which of your rules to turn "squishy" is the key to still achieving your goal.

In my case, I flew to a neighboring city instead. In fact, my boss had flown directly to my alternate city and planned from the start to drive the remaining distance -- he had never made flying to the final city a constraint. As a result of this experience I also finally bought some plane tickets for the summer I had been putting off for weeks. I am now flying to the 2-hour-away airport for less than half the price of the tickets to the actual city.

Have you ever realized you were overconstraining your problem? Which constraints turned out to be a lot squishier than you realized?

Saturday, February 3, 2018

​You can't inspect in quality

​
This is just a short post on applying industrial engineering principles to daily life.

At some point in my education, someone told me that it is impossible to inspect in quality. At the time it made sense from what I knew about inspections: people are bad at noticing rare events. 

Since then I have found a semi common application at home where attempting to inspect in quality is both tempting and a bad idea... Cleaning up broken glass. I have no idea how other people do it, but the system I have found to avoid the unpleasant outcome of stepping on glass is to clean extremely thoroughly twice, and then to conduct my first inspection. If I find any glass, I assume there are several more pieces I missed and do another cleaning pass.

Do you have any tips to speed up this process? Thoughts of other scenarios where it is tempting to try to "inspect in" quality? Leave your thoughts in the comments!

Saturday, October 14, 2017

Solving the “real” problem

In undergrad when I learned about the field of operations research I assumed people would write down their objective and constraints, get the optimal solution, and then do whatever the model told them to. Eventually I took a class from an adjunct professor my first year of grad school who explained that the hardest part of working in OR was convincing people to implement the output of the model. Basically "decision makers" (aka, people who did not know math) would not believe the output of the model, and so we had to design things so they could follow all the steps in our analysis.

I internalized that people would have reasons not to believe the model, but for a long time continued to believe it was mostly because of mistakes people made. You would build them a beautiful model, and then they would see the solution and realize that they forgot to give you important constraints. Or they would see the result and just determine it too weird and insist on a sub-optimal solution which looked more like what they had been doing. Over time I developed a more complete list of why people would not trust a model, but I still fundamentally thought of the models as right.

At some point though, that changed. I stopped thinking of people as the problem. I started this blog under the premise that not solving the right problem (type 3 error) was avoidable, but took a careful study to get to the problem you should solve. Even now, I have continued to find it challenging to really talk about that mental shift. In fact, this particular blog post has been sitting in purgatory since July while I was figuring out just the right way to convey the distinction.

But yesterday, while reading the HBR article “Are you solving the right problems?” the author described reframing a problem as not simply redefining the “real” problem, but instead recognizing that there is a better problem to solve. Realizing that you could be solving a better problem is not a simple process. It often requires attempting to solve other problems first. Even the notion of a “better” problem is not straightforward. It may have to do with the intractability of your current problem, or the realization that your first solution does not achieve what you thought it would.


If you are reading this and have a problem that could use some reframing, feel free to reach out to me or leave a comment here. Oftentimes, just explaining the situation to someone a bit further from the problem is all it takes to shift your context.