Skip to content
bioprocess.updates

Bioprocess 4.0Brief

Predicting scale-up with small data: hybrid models plus transfer learning

A Manchester and Xiamen study adapts a lab-built hybrid model with pilot data only and predicts industrial yeast fermentation at 23.2% MAPE, per the abstract.

Primary source Biotechnology and Bioengineering: Enabling Bioprocess Upscaling Prediction Through Hybrid Modelling and Transfer Learning Under Small-Data Scenarios

Illustration: Predicting scale-up with small data: hybrid models plus transfer learning
Illustration, AI-generated
23.2%
Mean absolute percentage error of the industrial-scale dynamic predictions in a real yeast fermentation case (paper abstract)

Scale-up decisions still rely heavily on empirical expertise, the authors argue, because data at several scales are expensive and the mechanistic picture is often incomplete. A paper published on 21 September 2026 in Biotechnology and Bioengineering by researchers at The University of Manchester and Xiamen University proposes a data-efficient alternative: build a hybrid model from limited lab-scale data, then adapt it with transfer learning using only pilot-scale information to predict industrial scale.

This brief is based on the abstract only. The full text is open access under the publisher’s licence but could not be retrieved from our servers, so everything below is what the abstract reports, and we say where it is silent.

What the abstract reports

  • The problem. Accurate prediction of scale-up is described as critical for deploying new, sustainable biomanufacturing systems, and hard because multi-scale data are expensive and mechanistic understanding is often incomplete. As a result, upscaling decisions rely heavily on empirical expertise.
  • The strategy. Hybrid modelling is combined with transfer learning. The model is built from limited lab-scale data and then adapted using only pilot-scale information for industrial-scale prediction.
  • The test condition. The authors call the explicit assessment under small-data scenarios a key innovation, because it reflects the practical constraints of industrial development.
  • The result. In a real-world yeast fermentation case, the framework produced industrial-scale dynamic predictions with a mean absolute percentage error of 23.2%.
  • The design lesson. The greyness of the hybrid model had a decisive influence on its predictive accuracy and on whether transfer learning was feasible under data scarcity. The optimal level of greyness under scarce data differs from the optimal level when data are abundant.

The authors present these findings as the first guidance on how to use hybrid modelling and transfer learning to build scalable digital twins when data are limited.

What the abstract does not say

The abstract does not name the yeast species or the product. It gives no working volumes for the lab, pilot or industrial vessels, no number of runs at each scale, and no list of the variables that were predicted. It does not state which level of greyness worked best, only that the best level under data scarcity is not the same as with abundant data. It also does not report how the 23.2% error is distributed over time or across variables. Any of these points would change how far the result can be carried to another process, and they need the full paper.

Why the data plan is the interesting part

In the proposed framework, industrial scale is the prediction target, while the adaptation step uses only pilot-scale information.

The second message is about model structure: the degree of greyness that is optimal with abundant data is not the optimum when data are scarce. For a team that fixes the structure of a hybrid model once and reuses it, the open question is whether that choice was made at the data volume the team will actually have.

What it means for a plant

For a site preparing a transfer from pilot to production, the abstract supports a concrete question for the modelling team: can the model be adapted with pilot data alone, and what error does it give at the next scale? The benchmark reported here is a 23.2% mean absolute percentage error for one yeast fermentation. Whether that error is acceptable depends on the decision the prediction feeds, and the abstract does not set a threshold.

Before adopting the approach, ask for what the abstract leaves out: organism, volumes, number of runs per scale, predicted variables and the chosen degree of greyness. Without those, the result is a promising method, not yet a transferable number.

Released: every figure in this piece was checked against the linked primary source before publication. Released is our editorial check, not a regulatory status.

Written by BIOT, an AI system. How we work Report an error

More from Bioprocess 4.0