Nhi Luong
  • Home
  • Data Science Projects
    • Korean Drama Analysis
    • Airline Reviews Analysis
    • Explore tmcn R Package
    • Making USA State Maps!
    • Web Scraping Application
  • Statistics Projects
    • Asian American Quality of Life
    • Spotify Song Characteristics
  • Machine Learning Projects
    • Handwritten Digits Classification
    • Recipes Origin Prediction

On this page

  • Introduction
    • Wrangle dataset
  • EDA
    • Summary Statistics
    • Graphs
  • Model building
    • Random Forest
    • LASSO Regression Model
  • Model refinement
    • Random Forest: Top 10 most important variables
    • LASSO: Top 10 most important variables
  • Conclusion

Recipe Origin Prediction

Applying Random Forest and LASSO Models

Author

James Le, Nhi Luong, Henry Black

Published

December 8, 2025

Introduction

In the culmination of our work this semester, we were hungry for a dataset that would let us apply our newly-learn modeling skills in a meaningful and interesting way. The dataset we selected to satisfy that hunger was, perhaps fittingly, one centered entirely around food. Specifically, we worked with a collection of 2,218 recipes from Allrecipes.com, covering dishes posted between 2009 and 2025 and curated by Brian Mubia. Each observation includes a substantial number of variables: country of origin, nutritional information, serving sizes, cooking and preparation times, user-generated ratings and reviews and most significantly, a list of the ingredients used in the recipe. Our primary goal is to develop a model that predicts the continent of origin for each recipe based on the information provided in the recipe entry.

Before beginning any modeling, we first worked on some data-wrangling to prepare the recipe dataset for analysis. Although the original dataset was already relatively clean, we introduced an important new variable continent, which grouped each recipe’s country of origin into a broader regional category. After reviewing the distribution of these categories, we restricted our analysis to the three continents with the highest counts: Asia, Europe, and North America. This step allowed us to focus on the groups with the most data and avoid instability from sparsely represented regions. We then removed variables that were irrelevant to our modeling goal or redundant by definition, such as name, author, url, and the separate prep_time and cook_time variables, since total_time already captured their combined effect. Finally, we converted the new continent variable into a factor, dropped observations containing missing values, and retained only the predictors that would contribute to classification. This wrangling stage ensured that the dataset was structured, consistent, and aligned with the questions we aimed to answer.

Wrangle dataset

  • The cuisines dataset is already clean so there was not any cleaning step involved. Our group did create a new variables called continent that groups different countries under specific continent.
The number of recipies in the dataset from each Continent
continent n
Europe 699
Asia 679
North America 527
South America 168
Oceania 65
Multiple Regions 61
NA 19
  • To create the data set for model building, we chose three continents with the highest number of counts and they are Asia, Europe, and North America.

  • In this wrangling step, we filtered the three continents, turned continent into factor and deselected variables that we do not want to be in the model. The original data set included more variables that we decided to discard such as cooking_time and prep_time since total_timeis the sum of both of them. We also get rid of name, author, url and date_published as they are irrelevant to our question and also can overfit our model. Finally, we used drop_na() to take out the NAs.

EDA

Summary Statistics

Characteristic N = 1,7731
continent
    Asia 624 (35%)
    Europe 646 (36%)
    North America 503 (28%)
calories 332 (201, 485)
fat 15 (8, 26)
carbs 26 (12, 45)
protein 14 (4, 27)
avg_rating 4.60 (4.40, 4.80)
total_ratings 26 (6, 94)
reviews 23 (6, 80)
total_time 60 (35, 120)
servings 6 (4, 10)
1 n (%); Median (Q1, Q3)
  • As mentioned above, the only new variables we associated with the model is continent, our response varibles which have three level: Asia, Europe and North America due to their relevancy.

  • Calories, fat, carbs and protein are numeric variables that possess the nutritional value, hence their namesake, per serving. Avg_rating is the mean value of all ratings on a specific recipie out of 5 with total_ratings indicate the total number of rating in thousands unless they are below 1000, in which case the exact number is displayed. Similarly, reviews is the total number of reviews per recipe in the thousands unless below 1000. Total_time is the sum of both cook_time and prep_time in minutes and serving is the number of servings. Due to practical reason, ingredients are not included in the summary table but it is included in the model post-tokenization in which each word of the ingredients can be used to predict the response variables.

  • Once we finished the tokenization of the ingredients, we added the 100 most frequently occurring variables to our data for help with the training and regularization.

Graphs

Distribution of average recipe ratings across continents. Violin and boxplots show the spread and central tendency of average ratings for recipes originating from Asia, Europe, and North America. The medians of ratings for three continents are a little over 4.5. All three distributions are skewed left.

Distribution of servings across continents. Violin and boxplots show the spread and central tendency of average ratings for recipes originating from Asia, Europe, and North America. The medians of ratings for three continents are around 5. All three distributions are skewed right
  • We can see that all continents have a lot of similar average rating in the range of 4 to 5 stars and similar number of servings as represented in the box and violin plot. However, Europe and Asia have some very low rating compared to North America and the number of servings from European dishes can be way higher than the other twos

Model building

Random Forest

  • We first fit a Random Fores model optimizing mtry value using bootstrap. There are 10 levels of mtry and the graph below shows the accuracy of each mtry value. The best mtry we picked based on the one standard rule is 4.

Accuracy vs mtry
The accuracy of the Random Forest with the optimized ‘mtry’ of 4 is 0.644. This shows 64.4% of the continents are classified correctly by the model.
.metric .estimator .estimate
accuracy multiclass 0.6463964

Table: The confusion matrix shows North America has the most missclassification, mostly being missclassified with Europe.

  • For the random forest, we decided to used 100 trees after trial and error of using larger number of trees (200, 500, 1000) and got very similar accuracy. Since the data is not unbalanced, accuracy gave a pretty good view on the performance of the model. We also tuned mtry based on the one standard error rule using boostraping and got the best mtry of 4 for our tree. When used on our testing dataset, the model yielded 65% accuracy and our confusion matrix showed that North America is the continent with the most misclassification. This makes intuitive sense as we have the fewest number of observations from North America, so the model is then less likely to select North America than it is the other continents.

LASSO Regression Model

  • We also fit a LASSO regression model optimizing the penalty using bootstrap. There are 10 levels of penalty and the graph below shows the accuracy of each penalty value. The best penalty we picked based on the one standard rule is 0.00599.

Accuracy vs penalty plot
The accuracy of the LASSO regression model with optimized ‘penalty’ of 0.00599 is 0.667. This shows 66.7% of the continents are classified correctly by the LASSO model.
.metric .estimator .estimate
accuracy multiclass 0.6666667

Table: The confusion matrix from the LASSO model shows North America also has the most missclassification, mostly being missclassified with Europe.

  • Our model perform with 67% accuracy and our confusion matrix shows that North America has the most misclassification.

Model refinement

  • We have put together two tables below to show the Top 10 most important variables for each model using vip function.

Random Forest: Top 10 most important variables

Random Forest’s most important variables

LASSO: Top 10 most important variables

Lasso regression’s most important variables
  • In the Random Forest, the top 5 most important indicators are soy, ginger, butter, sauce, and pepper. In the LASSO model, the top 5 most important indicators are oil, soy, ginger, paste and water. These indicators are all under variable ingredients, and we saw that there is no known relationships in the problem, so we decided to not exclude any variables from the model. Although it is worth to noted that the tokenized approach is imperfect as it’s using individual words and not the conventional meaning of the word ingredients. Such as the indicators soy and sauce can be a separate concepts or they are the same thing: soy sauce, or the indicators all, purpose, flour can be all purpose flour. We also need to taken them with a grain of salt because the model evaluated their important on their count (i.e. how many times they are written out) instead of an more intuitive way of basing them on how much are they used per recipe (e.g. 1/2 table spoon of oil).

Conclusion

In summary, our analysis demonstrates that both the Random Forest and LASSO models are reasonably effective at predicting the continent of origin for a recipe using nutritional information, preparation details, and tokenized ingredient text. While the overall accuracies of 0.644 for Random Forest and 0.667 for LASSO leave room for improvement, they show that significant patterns exist within the data. Across both models, different ingredients emerged as the most influential predictors, with recurring tokens such as soy and ginger reflecting the distinctive flavor profiles of certain cuisines from different continents. At the same time, some surprising indicators and odd tokens revealed the limitations of relying on single-word ingredient features, suggesting that more advanced text-processing methods could strengthen future models. Despite these challenges, our results highlight the potential of combining structured nutritional data with textual information to classify recipes by geographic origin, offering a promising foundation for more refined approaches moving forward.

Summary Table

Model Accuracy Top 5 Variables
Random Forest 0.653 soy, ginger, butter, sauce, pepper
LASSO 0.667 oil, soy, ginger, paste, water

LASSO Model Coefficients

  • We pulled out the top 5 most positive and negative coefficients in each class (Asia, North America, and Europe) to see how they are affecting the predicted classification.
For Asia, the most positive coefficient is 28.5. For Europe, it is 13 and for North America, it is 10.6. These coefficients have the most positive contribution to the probability of predicting each continent.
Class Term Estimate Penalty
Asia tfidf_ingredients_ginger 28.480741 0.0059948
Asia tfidf_ingredients_soy 25.021863 0.0059948
Asia tfidf_ingredients_oil 23.943768 0.0059948
Asia tfidf_ingredients_paste 20.188929 0.0059948
Asia tfidf_ingredients_rice 13.523074 0.0059948
Europe tfidf_ingredients_all 13.031080 0.0059948
Europe tfidf_ingredients_onion 10.555620 0.0059948
Europe tfidf_ingredients_olive 9.312119 0.0059948
Europe tfidf_ingredients_butter 7.950544 0.0059948
Europe tfidf_ingredients_sugar 7.032319 0.0059948
North America tfidf_ingredients_pepper 10.565244 0.0059948
North America tfidf_ingredients_bell 10.103820 0.0059948
North America tfidf_ingredients_green 9.291270 0.0059948
North America tfidf_ingredients_juice 6.460613 0.0059948
North America tfidf_ingredients_brown 5.646521 0.0059948
For Asia, the most negative coefficient is -8.95. For Europe, it is -11.2 and for North America, it is -8.29. These coefficients have the most negative contribution to the probability of predicting each continent.
Class Term Estimate Penalty
Asia tfidf_ingredients_cheese -8.945128 0.0059948
Asia tfidf_ingredients_dried -8.507417 0.0059948
Asia tfidf_ingredients_tomato -8.373700 0.0059948
Asia tfidf_ingredients_salt -8.133772 0.0059948
Asia tfidf_ingredients_all -7.428275 0.0059948
Europe tfidf_ingredients_cayenne -11.236615 0.0059948
Europe tfidf_ingredients_garlic -10.589621 0.0059948
Europe tfidf_ingredients_green -7.856457 0.0059948
Europe tfidf_ingredients_brown -7.570824 0.0059948
Europe tfidf_ingredients_powder -6.359980 0.0059948
North America tfidf_ingredients_grated -8.292466 0.0059948
North America tfidf_ingredients_lemon -7.543970 0.0059948
North America tfidf_ingredients_rice -7.458780 0.0059948
North America tfidf_ingredients_red -6.255168 0.0059948
North America tfidf_ingredients_onions -5.739782 0.0059948
  • For all continents, the top 5 highest and lowest coefficients all come from the same variable: ingredients. Each tokenized indicator shows how diverse the ingredients each continent tends to use.
  • Asia: Highest coefficients are within expectation with ginger and soy as these two ingredients also appear in the Random Forest vip plot. While the lowest coefficients are quite surprising with salt having a very negative impact.
  • Europe: Highest coefficients included olive which can be both the vegetable and the oil which is not surprising. The lowest coefficient is from ingredient cayenne, which is a type of chili pepper.
  • North America: Highest coefficients included another mysterious ingredients: juice, brown and green, with bell pepper as an essential ingredients which does make sense upon reflection. The most negative coefficient for North America is -8.292 and the term is tfidf_ingredients_grated. This is not an ingredient but rather a cooking method in the ingredients variable.