Modeling Feeding Tolerance, Gut Microbiome, and Intestinal Immune Markers in a Clinical Trial on Infant Formula

The project

This project is a clinical study carried out at the Nestlé Clinical Research Unit and published in BMC Nutrition (2026). It evaluated an infant formula built around a whey protein concentrate co-enriched in α-lactalbumin, milk fat globule membrane (MFGM), and Sn-2 palmitate, designed to bring the protein and lipid profile of the formula closer to that of human milk. The study followed a single-arm, open-label design in which formula-fed infants were compared to a parallel breastfed reference group over six weeks. The primary objective was to test the non-inferiority of formula feeding in terms of gastrointestinal tolerance, measured with the validated IGSQ-13 questionnaire, while the secondary objectives covered the gut microbiome (Bifidobacteria abundance and diversity), fecal short-chain fatty acids, and a panel of immune, inflammation, and intestinal-barrier markers.

From a statistical point of view, the data combined several difficulties within a single analysis: a non-inferiority hypothesis, repeated measures at baseline and end of study, and compositional microbiome data with a large proportion of zeros, each calling for a different modeling approach.

My contributions

I was responsible for implementing the statistical analysis and reporting the results, working in R and in close collaboration with the study scientists. I built the set of models behind the published results: ANCOVA and robust ANCOVA for the primary IGSQ non-inferiority endpoint and the continuous secondary endpoints, zero-inflated negative binomial models for bacterial relative abundances to handle the excess of zeros typical of microbiome data, Poisson regression for stool frequency, Shannon diversity together with PCoA and PERMANOVA for community structure, and Spearman correlation heatmaps with Benjamini-Hochberg correction to study the associations between microbiota and fecal metabolites. I also produced the corresponding tables and figures and wrote the statistical report that the scientific team relied on.

Beyond the modeling itself, I took care of the data curation. I identified and reported several irregularities in the dataset, including missing values linked to participant dropouts and duplicated measurements that turned out to be a data-entry error at the analyzing laboratory, and I recommended changes to some of the models to resolve convergence issues. The report I produced was used to write the final paper (published in BMC Nutrition, a Springer Nature journal), on which I am listed as a co-author and which I also helped review before submission.