Longitudinal Multi-Batch Characterisation of Beer Wort Fermentation Foam Using Image-Based Semantic Segmentation
Anca Sipos, Otto Ketney, Mariana-Liliana Păcală, Romulus IagarauImage-derived descriptors of fermentation foam may complement conventional brewing-process measurements, but they require evaluation at the level of independent fermentation batches. We derived image-based estimates of foam height and foam-free exposed area from hourly images of 12 beer wort fermentations conducted at initial extract levels of 12, 15 and 18 °P, and analysed 1147 hourly observations together with synchronised pH, temperature and dissolved-oxygen records. Foam was segmented with convolutional encoder–decoder networks trained separately for each fermentation, using 223 annotated frames overall (19.4% of the 1147 acquired frames; 14–24 frames per batch), all acquired within the first 72 h. Because hourly frames were nested within fermentation batches, observation-level correlations were treated as exploratory. In a batch-aware sensitivity analysis based on 70 batch-by-window means, linear mixed-effects models with batch as a random intercept identified an effect of operational time window on foam height (F(5, 53.10) = 72.17; p < 0.001) and on foam-free exposed area (F(5, 62.00) = 80.57; p < 0.001), whereas an effect of initial extract level was not detected (p = 0.520 and p = 0.404, respectively). Mean window-specific foam height was highest at 43.37 ± 12.80 mm in the >10–18 h window, and pH declined monotonically from 4.69 to 3.94. Within each fermentation, the annotated frames were partitioned at frame level into a training and a validation subset only, with no held-out test subset and no pooled model; the class-wise evaluation recorded for nine of these per-batch models gave a mean intersection-over-union of 0.190 ± 0.077 across the classes each model used (six in eight cases, five in one) and 0.371 ± 0.226 for the foam class, so segmentation accuracy on fermentations unseen during training was not established. The 307 observations recorded beyond 72 h (26.8% of the dataset) lie outside the annotated temporal domain. The findings support the proof-of-concept feasibility of longitudinal, image-based foam characterisation under controlled laboratory conditions but do not establish absolute measurement accuracy, generalisation to fermentations unseen during training, or readiness for process-control deployment. The semantic-segmentation component was used as an image-analysis instrument for extracting longitudinal foam descriptors rather than as a newly proposed segmentation architecture. In a targeted experiment on a fixed batch-level partition, with three complete fermentations withheld for testing, replacing class-weighted cross-entropy with an asymmetric Tversky loss raised foam recall on sparse-foam frames from 0.074 to 0.318 without any change to the network architecture, whereas the tested attention-gated configuration, which adds skip connections and attention gates together, did not improve foam segmentation under weighted cross-entropy. Accordingly, the methodological contribution of the study resides in the longitudinal multi-batch acquisition design, integration of image-derived descriptors with synchronised physicochemical measurements, and batch-aware analysis of repeated fermentation observations. The present data do not establish architectural superiority, segmentation generalisation to unseen fermentations, or a technical advance in semantic-segmentation methodology.