
Chromatography is where purity is won and where money is spent. Resin type, ligand, pH, conductivity, gradient shape and step order multiply into more conditions than any laboratory can test. That is why “AI for protein purification” is an attractive headline. Three open-access papers from groups in Vienna, Aachen and Edinburgh set out what the models in use actually predict, what data they need first, and where they break. Two of them come from the same Vienna and Aachen collaboration.
Three kinds of model, three kinds of cost
Bernau and co-authors separate the field cleanly. Data-driven models, from multiple linear regression on a design of experiments to neural networks and support vector regression, are fast and accurate inside the parameter space they were trained on, need little experimental effort and simple analysis, but cannot extrapolate beyond it. Mechanistic models describe mass transport and sorption from physicochemical principles, so they can extrapolate outside the tested conditions, but they need calibration data and computing power, and their complexity keeps them mostly in late-stage process characterisation. Hybrid models sit in between: mechanistic structure with data-driven components, better process understanding, and lower demands on data quantity and quality than a purely data-driven model.
The review is explicit about what mechanistic modelling buys. It reports a case in which the experimental effort to characterise a cation exchange step for a monoclonal antibody was cut by about 75% compared with traditional laboratory-based process characterisation, because the model could predict a priori how a change in protein surface charge would affect separation. It is equally explicit about the entry ticket.
The entry ticket: measurements before models
Before any prediction, the column and the system have to be measured:
- Porosity. An inaccurately determined column porosity distorts the equilibrium constant fitted to a peak. Interparticle porosities of around 12% have been measured experimentally, below the theoretical limit of around 26% for densely packed spheres, because packing at linear flow rates of up to 7 m/h compresses the beads.
- Dead volumes. For an AKTA pure 25 L with an XK 16/20 column, the injection and column valves contributed 22.7 x 10-5 L against 6.3 x 10-5 L for the tubing, and the column peripheral liquid volume (frits and connectors) reached 76.6 x 10-5 L. The ratio of that peripheral volume to bed volume runs from 0.767 at a 1 mL bed to 0.024 at the maximum 31 mL bed.
- Isotherm parameters. Static batch binding is simple but gives capacities around 1.5-fold higher than dynamic determination, so it is used for screening rather than calibration. Frontal analysis often needs 50 to 200 mg of pure protein per run on a 1 mL column, while inverse fitting of an elution peak needs around 0.6 mg, with purity above 50% in the authors’ hands.
Descriptors: the part that is genuinely hard
The Vienna and Aachen group’s second paper reviews the descriptors that turn a protein structure into the handful of numbers a model can use, and is candid about their limits. Descriptors average: two proteins with different surface charge distributions can share the same dipole moment, and two proteins with very different domain architecture can share the same enclosing ellipsoid, and so the same eccentricity. Patch descriptors need arbitrary thresholds, and using several thresholds inflates the number of collinear descriptors, which makes selection harder and overfitting easier.
Even the arithmetic is tool-dependent. The authors report electric dipole moments for catalase ranging from 371 to 812 Debye using one commercial package with different hydrogen addition and energy minimisation settings, against 121 Debye from an open-access server. Of the 24 hydrophobicity scales implemented in ProtScale, 19 are strongly connected with an average Pearson correlation of 0.81, and the remaining five are inverted relative to the first group.
Two effects that matter in a real feed are rarely captured in the models. Host cell proteins that bind to the product and co-purify with it, known as hitchhiking, would require protein-protein interactions in the model; docking calculations can take up to 11 hours, and several hundred host cell proteins may be present. And for large colloids such as virus-like particles of around 100 nm, almost all pores of common resins are inaccessible, which is why dynamic binding capacity can be unexpectedly low.
What data-driven models are good at
A third paper shows where the purely data-driven route pays. A group in Edinburgh, with co-authors at BOKU Vienna and at the column manufacturer Repligen, took around 25,000 quality control runs of pre-packed columns collected over about 10 years, removed incomplete records and averaged repeats of identical column types, leaving 546 independent entries. An extreme gradient boosting model predicted reduced plate height and peak asymmetry with 90% and 93% predictive capability, a mean absolute percentage error of 10% and 7% on the test set.
The interesting result is the ranking. Resin backbone accounted for 68.4% of the importance for plate height and 77.0% for asymmetry, with functional mode second at around 15%, ahead of particle size, column diameter and column length. Those two are qualitative attributes that first-principles rate models do not take into account by design, as the authors point out.
What it means for a plant
The practical question is not whether to use AI in purification but which model repays its calibration. A data-driven model is the cheap option when the question stays inside a known design space, as in resin and column selection from historical quality data. A mechanistic or hybrid model is what survives a change of scale or resin, and its cost is measured in porosity determinations, dead volume corrections and gradient runs, not in software licences. Three consequences for a downstream group:
- Characterise the system before the protein. An inaccurate porosity or an uncorrected dead volume distorts the fitted isotherm parameters and can travel with the model to the next scale.
- Check which descriptor tool and which settings produced the numbers in a model you inherit. The same protein can differ several-fold between packages.
- Expect the model to miss hitchhiking host cell proteins and size-excluded large particles. Those stay experimental questions.
Sources
- Bernau CR, Knoedler M, Emonts J, Jaepel RC, Buyel JF, The use of predictive models to develop chromatography-based purification processes, Frontiers in Bioengineering and Biotechnology 10:1009102, 12 October 2022.
- Emonts J, Buyel JF, An overview of descriptors to capture protein properties: tools and perspectives in the context of QSAR modeling, Computational and Structural Biotechnology Journal 21:3234-3247, 24 May 2023.
- Jiang Q, Seth S, Scharl T, Schroeder T, Jungbauer A, Dimartino S, Prediction of the performance of pre-packed purification columns through machine learning, Journal of Separation Science 45:1445-1457, 20 March 2022.


