---
title: "Machine Learning and Sensory Chemistry: Deciphering the Pleasure of Drinking a Beer"
description: "The relationship between the chemical composition of a beer and its consumer appreciation has been an elusive field for food science, due to the non-linear nature of perception and its complex interactions."
url: https://www.thebeertimes.com/en/machine-learning-sensory-chemistry-deciphering-pleasure-drinking-beer/
date: 2025-10-28
modified: 2026-07-01
author: "Carlos Uhart M."
image: https://www.thebeertimes.com/wp-content/uploads/2025/10/Machine-learning-y-quimica-sensorial.jpg
categories: ["Technology"]
tags: ["Big Data", "Machine Learning", "Sensory Analysis", "Technology"]
type: post
lang: en
---

# Machine Learning and Sensory Chemistry: Deciphering the Pleasure of Drinking a Beer

Beer is one of the most complex fermented beverages, with a flavor profile that arises from the interaction of thousands of chemical compounds.

![Machine learning and sensory chemistry](https://www.thebeertimes.com/wp-content/uploads/2025/10/Machine-learning-y-quimica-sensorial.jpg)*Machine learning and sensory chemistry*

Traditionally, the relationship between the chemical composition of a beer and its sensory perception or consumer appreciation has been an elusive field for food science, due to the non-linear nature of perception and the complex interactions between aromatic compounds.

A pioneering study entitled “Predicting and improving complex beer flavor through machine learning,” published in Nature Communications, addresses this challenge by combining large-scale chemical analyses, trained sensory panels, and online consumer reviews with advanced machine learning models to predict and improve beer flavor and appreciation.

## From Chemistry to Perception

To generate an unprecedented dataset, the researchers selected 250 commercial Belgian beers spanning 22 different styles, from Lagers and Blonds to Stouts and [Lambic sour beers](https://www.thebeertimes.com/ocho-mitos-las-cervezas-lambic-refutados/).

For each beer, 226 different chemical properties were measured, including fermentation parameters such as ethanol and glycerol, hop bitter acids such as [iso-alpha acids](https://www.thebeertimes.com/en/the-science-of-beer-aftertaste-physiology-and-chemistry-of-persistent-flavor/), a wide range of fruity esters produced by yeast, sulfur compounds, organic acids, and hop-derived terpenoids responsible for herbal and citrus aromas.

This exhaustive chemical characterization was achieved through analytical techniques such as gas chromatography coupled with mass spectrometry and flame photometry detectors, as well as enzymatic and near-infrared analyses.

In parallel, a trained sensory panel of 16 [tasters who evaluated each beer](https://www.thebeertimes.com/en/how-to-taste-beer-learn-to-appreciate-beer-like-a-true-expert/) for 50 sensory attributes was assembled, rating the intensity of malt, hop, ester, acidity, bitterness, and mouthfeel sensations, among others.

To complement this expert panel data and obtain a measure of general consumer appreciation, more than 180,000 public reviews from the [RateBeer](https://www.thebeertimes.com/ratebeer-cierra-sus-puertas-definitivamente-el-fin-de-una-era-en-la-cerveza-artesanal/) platform were collected and processed.

This included not only numerical scores for aroma, flavor, and overall quality, but also the development of automated text analysis tools to extract mentions of specific sensory attributes from review texts, allowing the creation of a flavor profile based on the wisdom of the crowd.

![Chemical correlation vs descriptors](https://www.thebeertimes.com/wp-content/uploads/2025/10/Correlacion-quimica-vs-descriptores-300x295.jpg)*Chemical correlation vs. descriptors*

## The Predictive Power of Machine Learning

With these massive, multidimensional datasets, the team trained and compared 10 different machine learning models to predict sensory responses and appreciation from chemical profiles.

The evaluated models ranged from conventional statistical techniques such as linear regression and partial least squares regression, to more advanced decision tree-based algorithms such as Random Forest, XGBoost, and Gradient Boosting, as well as an artificial neural network and a support vector machine.

The results were revealing. Tree-based models, particularly the Gradient Boosting Regressor, significantly outperformed traditional linear models, which tend to suffer from overfitting when faced with a large number of correlated predictors.

The Gradient Boosting model demonstrated superior performance, with R² determination coefficient values of up to 0.75 for some sensory attributes.

Interestingly, all models predicted RateBeer data more accurately than trained panel data, which is likely because the large volume of reviews averages out the sensory variability inherent in individuals.

Furthermore, the models were generally better at predicting taste than aroma, possibly because attributes such as bitterness have a more direct relationship with specific chemical compounds such as iso-alpha acids.

## Chemical Drivers of Appreciation

One of the most innovative aspects of the study was the dissection of the best models to identify which chemical compounds are the primary drivers of consumer appreciation.

Two approaches were used: first, impurity-based feature importance, and the SHAP method, which quantifies the contribution of each feature to the model’s predictions for each individual sample.

The analysis identified a set of compounds as critical for predicting high appreciation.

Ethanol and ethyl acetate — an ester that imparts fruity and sometimes solvent-like notes at high concentrations — consistently emerged as the two most important predictors.

This underscores the fundamental role of ethanol not only as a flavor compound, but also as a modulator of the release of volatile aromatic compounds from the beer matrix.

Other unexpected compounds that appeared in the top 15 were phenylethyl acetate, often associated with beer aging, and methanethiol, a sulfur compound that smells like rotting cabbage at high concentrations.

This suggests that, while high concentrations of these compounds are undesirable, modest levels could contribute positively to flavor complexity.

Proteins, components that influence the mouthfeel and body of beer, were also identified as a highly important factor, as was lactic acid, characteristic of highly appreciated sour beers.

It is crucial to note that many of these key compounds would not have been identified through conventional Spearman correlation analysis, which only captures linear relationships.

For example, the correlation of lactic acid with appreciation was low, but the Gradient Boosting model was able to capture its importance by recognizing that sour beers form a distinct group with high appreciation — a non-linear relationship that simple correlation overlooks.

## Experimental Validation

To validate the model’s predictions, the researchers conducted practical beer improvement experiments.

They took a commercial Blond beer and a non-alcoholic one, and added to them a combination of the compounds identified as most important (ethyl acetate, ethyl hexanoate, isoamyl acetate, ethyl phenylacetate, glycerol, and lactic acid), raising their concentrations to the 95th percentile within their style.

In blind sensory tests, the “enhanced” versions were significantly preferred over the originals by the trained panel. Tasters noted an increase in the intensity of [ester flavors](https://www.thebeertimes.com/esteres-vs-fenoles-en-la-cerveza-cual-es-la-diferencia/), sweetness, body, and alcohol perception, resulting in greater overall appreciation.

Even in the non-alcoholic beer, the addition of the compound mixture (excluding ethanol) significantly improved its acceptance, demonstrating the model’s potential to optimize low-alcohol products.

## Frequently Asked Questions (FAQ)

### Why Did Ethanol and Ethyl Acetate Emerge as the Two Most Important Chemical Predictors of Consumer Appreciation?

Answer: Ethanol, beyond being psychoactive, is fundamental because it acts as a chemical modulator. It significantly influences how other volatile aromatic compounds are released from the beer matrix, altering perception. Ethyl acetate, for its part, imparts fruity notes (sherry or solvent-like in excess) and is key to complexity. Its high importance suggests that consumers associate these compounds with perceived intensity and body — indicators of quality in many Belgian styles.

### If Machine Learning Models Predict RateBeer Appreciation Better Than Trained Panels, What Is the Value of Continuing to Use Expert Sensory Panels?

Answer: The value of trained panels lies in the precision and objectivity of describing specific attributes (e.g., intensity of “lactic acidity”). RateBeer data, while easier to predict due to its large volume averaging individual variability, only reflects overall appreciation. Trained panels are indispensable for product optimization (e.g., if the goal is to reduce solvent aroma, the panel identifies the precise intensity — which is less clear in a general review).

### How Can Craft Brewers Use the Identification of Methanethiol and Phenylethyl Acetate as Appreciation Drivers?

Both compounds are usually associated with defects or aging, so their appearance in the top 15 suggests a concept of “threshold complexity.” Methanethiol (sulfurous aroma) could be an indicator of positive yeast reactions or terroir at extremely low, non-offensive levels. Phenylethyl acetate (rose/honey) at modest levels enhances aroma. Brewers could focus on controlling and maintaining these compounds at very low and precise concentrations to contribute subtle complexity without crossing the defect threshold.

### Why Were Decision Tree-Based Models More Effective Than Traditional Linear Regression in This Study?

Answer: Linear models can only capture direct linear relationships (e.g., more iso-alpha acids = more bitterness). The sensory profile of beer is inherently non-linear and complex (e.g., the interaction between acidity and sweetness). Tree-based models are superior because they can model non-linear relationships, thresholds, and interaction effects between multiple chemical compounds, allowing them to capture the true complexity of human perception — such as the high appreciation of sour beers that simple linear regression did not recognize.

### How Is the “Percentile 95” Concept Applied in Experimental Validation and Why Is It Important for Flavor Improvement?

The “95th percentile” refers to the highest concentration of a chemical compound found in 95% of beers of that style. By adding the key compounds up to that upper limit (P95), the researchers were testing the hypothesis that complexity and intensity increase appreciation within an acceptable range for consumers. The improvement in acceptance demonstrated that flavor optimization is achieved by pushing chemical profiles toward the upper limits of their natural concentration spectrum.

## Reference

Schreurs, M., Piampongsant, S., Roncoroni, M., et al. (2024). *Big data and machine learning reveal the chemical drivers of beer flavor and appreciation*. *Nature Communications*, 15, 2368. [https://doi.org/10.1038/s41467-024-46346-0](https://doi.org/10.1038/s41467-024-46346-0)

## Recommended

- [Researchers present the first sensory wheel with standardized lexicon for brewing malt](https://www.thebeertimes.com/lexico-sensorial-y-analisis-de-compuestos-aromaticos-en-malta-cervecera/)
- [Aeration and oxygenation when brewing beer: why are they so important?](https://www.thebeertimes.com/aireacion-y-oxigenacion-al-elaborar-cerveza-por-que-son-tan-importantes/)
