
Researchers at Gujarat Biotechnology University in Gandhinagar, India, tested whether adding a small number of documented process-deviation batches to the training set of a machine-learning soft sensor improves its estimate of penicillin concentration during fermentation upsets, without degrading accuracy on normal runs. In a preprint posted on bioRxiv on 23 September 2026, which has not been peer reviewed, they report that a HistGradientBoosting (HGB) regressor trained with fault batches cut pooled root-mean-square error (RMSE) on held-out deviation batches from 3.195 to 2.564 g/L, a reduction of 19.73%, while normal-operation RMSE stayed nearly flat at 1.981 versus 1.983 g/L.
A soft sensor tested beyond normal operation
The study uses the public 100-batch IndPenSim benchmark, generated by a simulator of an industrial-scale fed-batch penicillin process. Batches 1 to 30 run recipe-driven control, batches 31 to 60 run operator-controlled operation, batches 61 to 90 run advanced process control, and batches 91 to 100 are documented process-deviation batches. The retained exports contained 113,935 observations sampled every 0.2 h (12 minutes), with batch length ranging from 167 to 290 h. The authors are explicit that these are simulated data and “are not a substitute for physical-plant validation.”
The target variable is current-time penicillin concentration in g/L, a quantity that, unlike temperature, pH, gas composition and flow rates, may not be available at the same frequency or only after laboratory analysis. A soft sensor estimates it from signals available at prediction time.
Feeding the model documented deviation batches
The final regressor used 36 inputs: 20 current or base process variables, 15 history variables and one cumulative-feed variable. Dissolved oxygen, sugar-feed rate, temperature, pH and aeration each contributed a one-step lag, a first difference and the mean of up to five preceding observations, so at the 0.2 h sampling interval a full history window spans the preceding hour.
Two training strategies were compared with identical model settings:
- Normal-only HGB, fitted only on normal-operation batches.
- Fault-inclusive HGB, fitted on the same normal batches plus the permitted deviation batches, with deviation batches given a sample-weight factor of three.
Before settling on HGB, the authors compared four model families on the normal batches. The normal-trained HGB produced the lowest RMSE among the learned models on both normal and deviation batches, 2.030 and 4.190 g/L, though its deviation MAE of 2.768 g/L was slightly higher than the Random Forest value of 2.712 g/L.
Normal-operation performance was measured with five regime-balanced complete-batch folds, each holding out 18 normal batches and training on the other 72. Deviation performance used leave-one-fault-batch-out evaluation: for each of the ten deviation batches, the fault-inclusive model was trained on all 90 normal batches plus the other nine deviation batches, then tested on the one held out, so no prediction came from a model that had seen the batch it was scored on.
Deviation error down, normal accuracy essentially unchanged
| Metric | Normal batches | Deviation batches |
|---|---|---|
| RMSE, normal-only (g/L) | 1.9806 | 3.1946 |
| RMSE, fault-inclusive (g/L) | 1.9830 | 2.5644 |
| MAE, normal-only (g/L) | 1.2095 | 2.0763 |
| MAE, fault-inclusive (g/L) | 1.2063 | 1.4405 |
| R², normal-only | 0.9607 | 0.8580 |
| R², fault-inclusive | 0.9606 | 0.9085 |
Deviation MAE fell from 2.076 to 1.441 g/L, a reduction of 30.62%. Eight of ten deviation batches improved with fault-inclusive training, and the mean paired batch RMSE difference was -0.7063 g/L, with a descriptive 95% batch-bootstrap interval of -1.2844 to -0.2194 g/L. The improvement was concentrated in the recorded fault window: RMSE there fell from 2.7172 to 1.2364 g/L, while error after the last recorded fault activity fell only from 4.3393 to 4.0234 g/L, and error before any fault reference was already low in both models, 0.2289 versus 0.2054 g/L.
Two batches got worse, one stayed poorly predicted
Batches 92 and 93 worsened after fault-inclusive training, by 4.93% and 13.05% respectively. Batch 100 remained the clearest failure: its RMSE moved only from 6.692 to 6.631 g/L and its fault-inclusive R² was -2.4837, worse than simply predicting the batch mean. Across batch 100, the mean actual concentration was 6.931 g/L against a mean fault-inclusive estimate of 11.790 g/L; at the 230 h timepoint the actual value was 5.626 g/L while the model predicted 15.511 g/L.
The paper’s reliability checks did not catch this. An Isolation Forest out-of-distribution detector built from 22,500 reference rows drawn from the 90 normal batches, using 15 principal components that retained 95.005% of variance and a warning threshold at the 99th-percentile score of 0.626829, flagged only 4.61% of batch 100’s rows as unusual even though it was the largest concentration-prediction failure in the set. Empirical absolute-error radii of 3.04 g/L (normal) and 3.12 g/L (deviation) are described by the authors as uncalibrated deployment diagnostics rather than dependable bounds.
What it means for a plant
For a plant that leans on a soft sensor to track product titer between lab samples in a fed-batch process, this preprint is both a caution and a low-cost fix. Validating a soft sensor only on normal batches says nothing about its behavior once a known deviation type occurs. Adding a handful of representative deviation batches to training, with extra sample weight, cut deviation-batch RMSE by roughly a fifth here with essentially unchanged normal-operation accuracy. But the same results do not establish generalization to unseen fault mechanisms: two of ten batches got worse, one stayed badly wrong, and the out-of-distribution detector flagged only 4.61% of the worst batch’s rows. The authors state plainly that the current evidence “does not justify autonomous control, release decisions or safety-critical use,” and everything here runs on a simulator, not a physical fermenter. Any plant considering this approach should treat a soft-sensor output, an uncalibrated risk score and an OOD flag as a bundle for a human operator to interpret alongside lab confirmation, not as a stand-alone release signal.


