| continent | n |
|---|---|
| Europe | 699 |
| Asia | 679 |
| North America | 527 |
| South America | 168 |
| Oceania | 65 |
| Multiple Regions | 61 |
| NA | 19 |
Recipe Origin Prediction
Applying Random Forest and LASSO Models
Introduction
In the culmination of our work this semester, we were hungry for a dataset that would let us apply our newly-learn modeling skills in a meaningful and interesting way. The dataset we selected to satisfy that hunger was, perhaps fittingly, one centered entirely around food. Specifically, we worked with a collection of 2,218 recipes from Allrecipes.com, covering dishes posted between 2009 and 2025 and curated by Brian Mubia. Each observation includes a substantial number of variables: country of origin, nutritional information, serving sizes, cooking and preparation times, user-generated ratings and reviews and most significantly, a list of the ingredients used in the recipe. Our primary goal is to develop a model that predicts the continent of origin for each recipe based on the information provided in the recipe entry.
Before beginning any modeling, we first worked on some data-wrangling to prepare the recipe dataset for analysis. Although the original dataset was already relatively clean, we introduced an important new variable continent, which grouped each recipe’s country of origin into a broader regional category. After reviewing the distribution of these categories, we restricted our analysis to the three continents with the highest counts: Asia, Europe, and North America. This step allowed us to focus on the groups with the most data and avoid instability from sparsely represented regions. We then removed variables that were irrelevant to our modeling goal or redundant by definition, such as name, author, url, and the separate prep_time and cook_time variables, since total_time already captured their combined effect. Finally, we converted the new continent variable into a factor, dropped observations containing missing values, and retained only the predictors that would contribute to classification. This wrangling stage ensured that the dataset was structured, consistent, and aligned with the questions we aimed to answer.
Wrangle dataset
- The
cuisinesdataset is already clean so there was not any cleaning step involved. Our group did create a new variables calledcontinentthat groups different countries under specific continent.
To create the data set for model building, we chose three continents with the highest number of counts and they are Asia, Europe, and North America.
In this wrangling step, we filtered the three continents, turned
continentinto factor and deselected variables that we do not want to be in the model. The original data set included more variables that we decided to discard such ascooking_timeandprep_timesincetotal_timeis the sum of both of them. We also get rid ofname,author,urlanddate_publishedas they are irrelevant to our question and also can overfit our model. Finally, we useddrop_na()to take out the NAs.
EDA
Summary Statistics
| Characteristic | N = 1,7731 |
|---|---|
| continent | |
| Asia | 624 (35%) |
| Europe | 646 (36%) |
| North America | 503 (28%) |
| calories | 332 (201, 485) |
| fat | 15 (8, 26) |
| carbs | 26 (12, 45) |
| protein | 14 (4, 27) |
| avg_rating | 4.60 (4.40, 4.80) |
| total_ratings | 26 (6, 94) |
| reviews | 23 (6, 80) |
| total_time | 60 (35, 120) |
| servings | 6 (4, 10) |
| 1 n (%); Median (Q1, Q3) | |
As mentioned above, the only new variables we associated with the model is continent, our response varibles which have three level:
Asia,EuropeandNorth Americadue to their relevancy.Calories,fat,carbsandproteinare numeric variables that possess the nutritional value, hence their namesake, per serving.Avg_ratingis the mean value of all ratings on a specific recipie out of 5 withtotal_ratingsindicate the total number of rating in thousands unless they are below 1000, in which case the exact number is displayed. Similarly,reviewsis the total number of reviews per recipe in the thousands unless below 1000.Total_timeis the sum of bothcook_timeandprep_timein minutes andservingis the number of servings. Due to practical reason,ingredientsare not included in the summary table but it is included in the model post-tokenization in which each word of the ingredients can be used to predict the response variables.Once we finished the tokenization of the ingredients, we added the 100 most frequently occurring variables to our data for help with the training and regularization.
Graphs
- We can see that all continents have a lot of similar average rating in the range of 4 to 5 stars and similar number of servings as represented in the box and violin plot. However, Europe and Asia have some very low rating compared to North America and the number of servings from European dishes can be way higher than the other twos
Model building
Random Forest
- We first fit a Random Fores model optimizing
mtryvalue using bootstrap. There are 10 levels ofmtryand the graph below shows the accuracy of each mtry value. The best mtry we picked based on the one standard rule is 4.
| .metric | .estimator | .estimate |
|---|---|---|
| accuracy | multiclass | 0.6463964 |
Table: The confusion matrix shows North America has the most missclassification, mostly being missclassified with Europe.
- For the random forest, we decided to used 100 trees after trial and error of using larger number of trees (200, 500, 1000) and got very similar accuracy. Since the data is not unbalanced, accuracy gave a pretty good view on the performance of the model. We also tuned mtry based on the one standard error rule using boostraping and got the best mtry of 4 for our tree. When used on our testing dataset, the model yielded 65% accuracy and our confusion matrix showed that
North Americais the continent with the most misclassification. This makes intuitive sense as we have the fewest number of observations from North America, so the model is then less likely to select North America than it is the other continents.
LASSO Regression Model
- We also fit a LASSO regression model optimizing the
penaltyusing bootstrap. There are 10 levels of penalty and the graph below shows the accuracy of each penalty value. The best penalty we picked based on the one standard rule is 0.00599.
| .metric | .estimator | .estimate |
|---|---|---|
| accuracy | multiclass | 0.6666667 |
Table: The confusion matrix from the LASSO model shows North America also has the most missclassification, mostly being missclassified with Europe.
- Our model perform with 67% accuracy and our confusion matrix shows that
North Americahas the most misclassification.
Model refinement
- We have put together two tables below to show the Top 10 most important variables for each model using
vipfunction.
Random Forest: Top 10 most important variables
LASSO: Top 10 most important variables
- In the Random Forest, the top 5 most important indicators are soy, ginger, butter, sauce, and pepper. In the LASSO model, the top 5 most important indicators are oil, soy, ginger, paste and water. These indicators are all under variable
ingredients, and we saw that there is no known relationships in the problem, so we decided to not exclude any variables from the model. Although it is worth to noted that the tokenized approach is imperfect as it’s using individual words and not the conventional meaning of the word ingredients. Such as the indicators soy and sauce can be a separate concepts or they are the same thing: soy sauce, or the indicators all, purpose, flour can be all purpose flour. We also need to taken them with a grain of salt because the model evaluated their important on their count (i.e. how many times they are written out) instead of an more intuitive way of basing them on how much are they used per recipe (e.g. 1/2 table spoon of oil).
Conclusion
In summary, our analysis demonstrates that both the Random Forest and LASSO models are reasonably effective at predicting the continent of origin for a recipe using nutritional information, preparation details, and tokenized ingredient text. While the overall accuracies of 0.644 for Random Forest and 0.667 for LASSO leave room for improvement, they show that significant patterns exist within the data. Across both models, different ingredients emerged as the most influential predictors, with recurring tokens such as soy and ginger reflecting the distinctive flavor profiles of certain cuisines from different continents. At the same time, some surprising indicators and odd tokens revealed the limitations of relying on single-word ingredient features, suggesting that more advanced text-processing methods could strengthen future models. Despite these challenges, our results highlight the potential of combining structured nutritional data with textual information to classify recipes by geographic origin, offering a promising foundation for more refined approaches moving forward.
Summary Table
| Model | Accuracy | Top 5 Variables |
|---|---|---|
| Random Forest | 0.653 | soy, ginger, butter, sauce, pepper |
| LASSO | 0.667 | oil, soy, ginger, paste, water |
LASSO Model Coefficients
- We pulled out the top 5 most positive and negative coefficients in each class (Asia, North America, and Europe) to see how they are affecting the predicted classification.
| Class | Term | Estimate | Penalty |
|---|---|---|---|
| Asia | tfidf_ingredients_ginger | 28.480741 | 0.0059948 |
| Asia | tfidf_ingredients_soy | 25.021863 | 0.0059948 |
| Asia | tfidf_ingredients_oil | 23.943768 | 0.0059948 |
| Asia | tfidf_ingredients_paste | 20.188929 | 0.0059948 |
| Asia | tfidf_ingredients_rice | 13.523074 | 0.0059948 |
| Europe | tfidf_ingredients_all | 13.031080 | 0.0059948 |
| Europe | tfidf_ingredients_onion | 10.555620 | 0.0059948 |
| Europe | tfidf_ingredients_olive | 9.312119 | 0.0059948 |
| Europe | tfidf_ingredients_butter | 7.950544 | 0.0059948 |
| Europe | tfidf_ingredients_sugar | 7.032319 | 0.0059948 |
| North America | tfidf_ingredients_pepper | 10.565244 | 0.0059948 |
| North America | tfidf_ingredients_bell | 10.103820 | 0.0059948 |
| North America | tfidf_ingredients_green | 9.291270 | 0.0059948 |
| North America | tfidf_ingredients_juice | 6.460613 | 0.0059948 |
| North America | tfidf_ingredients_brown | 5.646521 | 0.0059948 |
| Class | Term | Estimate | Penalty |
|---|---|---|---|
| Asia | tfidf_ingredients_cheese | -8.945128 | 0.0059948 |
| Asia | tfidf_ingredients_dried | -8.507417 | 0.0059948 |
| Asia | tfidf_ingredients_tomato | -8.373700 | 0.0059948 |
| Asia | tfidf_ingredients_salt | -8.133772 | 0.0059948 |
| Asia | tfidf_ingredients_all | -7.428275 | 0.0059948 |
| Europe | tfidf_ingredients_cayenne | -11.236615 | 0.0059948 |
| Europe | tfidf_ingredients_garlic | -10.589621 | 0.0059948 |
| Europe | tfidf_ingredients_green | -7.856457 | 0.0059948 |
| Europe | tfidf_ingredients_brown | -7.570824 | 0.0059948 |
| Europe | tfidf_ingredients_powder | -6.359980 | 0.0059948 |
| North America | tfidf_ingredients_grated | -8.292466 | 0.0059948 |
| North America | tfidf_ingredients_lemon | -7.543970 | 0.0059948 |
| North America | tfidf_ingredients_rice | -7.458780 | 0.0059948 |
| North America | tfidf_ingredients_red | -6.255168 | 0.0059948 |
| North America | tfidf_ingredients_onions | -5.739782 | 0.0059948 |
- For all continents, the top 5 highest and lowest coefficients all come from the same variable:
ingredients. Each tokenized indicator shows how diverse the ingredients each continent tends to use. Asia: Highest coefficients are within expectation with ginger and soy as these two ingredients also appear in the Random Forestvipplot. While the lowest coefficients are quite surprising with salt having a very negative impact.Europe: Highest coefficients included olive which can be both the vegetable and the oil which is not surprising. The lowest coefficient is from ingredient cayenne, which is a type of chili pepper.North America: Highest coefficients included another mysterious ingredients: juice, brown and green, with bell pepper as an essential ingredients which does make sense upon reflection. The most negative coefficient forNorth Americais -8.292 and the term istfidf_ingredients_grated. This is not an ingredient but rather a cooking method in theingredientsvariable.