Skip to main content

Posts

Showing posts with the label Machine Learning

A practical advice about building models

One of the most practical pieces of advice I recently learned about building models is counterintuitive. It suggests that we should not immediately jump into training models on the data. Instead, we should first try to create heuristic rules for the prediction problem at hand. For example, if we are trying to predict whether a customer will buy the latest edition of the iPhone or not, a simple heuristic rule would be that customers with an annual income greater than $80,000 USD and a history of purchasing Apple products would have a higher probability of buying the new iPhone. You could write a simple SQL query to test out such heuristic rules on your training and holdout sets and evaluate their effectiveness. This approach could sometimes help you create better features, identify inherent target leakage issues, and provide a baseline that you could aim to beat with the models.

Solving Customer Churn with a hammer!

Learning when data should take a back seat and give way to domain knowledge is a valuable skill. Suppose you built a machine learning model on the data of your customers to predict churn risk. Now that you have a risk score for each customer, what do you do next? Do you filter the top n% based on the risk and send them a coupon with a discount in the hopes that it will prevent churn? But what if price is not the factor driving churn in many of these customers? Customers might have been treated poorly by customer service, which drove them away from your company's product.  Or there might have been an indirect competitor's product or service that removes the need for your company's product altogether (this happened to companies like Blockbuster and Kodak in the past!) There could be a myriad of factors, but you get the point! Dashboards and models cannot guide any company's strategic actions directly. If companies try to use them without additional context, more often tha...

A random problem of decision tree

There's something you need to know about if you are building decision tree models using Python's famous scikit-learn package. The algorithm it uses for building the models is deterministic (it produces consistent results across multiple executions if the inputs don't change). Despite this nature, scikit-learn provides a 'Random state' hyperparameter to the decision tree's class. This hyperparameter is only needed when an algorithm is not deterministic, as fixing the random state to a constant integer value arrests the randomness. So, the random state must be a redundant parameter when building decision trees, right? Not really. The decision tree algorithm could use the value of the random state passed to it for making a 'decision' in the below three cases: i) If you set the max_features hyperparameter to an integer value lesser than the total number of features. It means the algorithm needs to decide which random subset of features to use at each node to...

Kryptonite of the correlations

  It is easy to get lost in the world of correlations! But make no mistake, you have valuable information about your data by looking at the correlations between the variables. And if you’re planning to build a linear model using supervised machine learning, the presence of correlated variables will exacerbate the model’s accuracy. Hence, looking at the correlations between continuous variables is almost a non-negotiable task, and one of the ways to do it is by calculating the Pearson correlation coefficient for these variables. Pearson coefficient is the default mode for many, including me, for testing correlations. But Pearson coefficient has found its kryptonite in the form of the non-linear relationship between the variables. Pearson coefficient makes sense only when there is a linear relationship between the variables and is not very useful when the variables have a non-linear relationship. But it is high time we adopt other measures of correlation in addition to the Pearson co...

If you had data of the entire world at your feet, would you still build a deep learning model on it?

  Suppose you solved a crucial business problem by building a complex deep neural network model to generate predictions on a new set of business data. Deploying and maintaining this model costs your business a fortune, but the decision-makers consider it a small price to pay for a greater good. It is all well and good. But could you have optimized the cost by opting for a simple regression model instead of a deep neural network? Of course, you tried regression as a baseline model, and you went for a deep neural network only after you noticed that it outperformed the regression model by miles. But if you had to suddenly build this model again with an enormous volume of training data, would you notice any difference in the performance metrics of the baseline and the complex models? Or, for that matter, would an increase in the volume of training data improve the performance of a baseline model and make it on par with a complex model? Two Microsoft researchers, Michele Banko and Eric ...