
- July-August 2026 Online-only Peer-reviewed Articles
- Volume 41
- Issue wp8
Uniform Design-Augmented Process Data for Multivariate Raman Calibration in Bioprocess Monitoring
Key Takeaways
- Hybrid UD-plus-process calibration expanded concentration coverage while retaining process realism, supporting lifecycle-managed Raman/PAT model robustness under QbD expectations with fewer process runs.
- UF/DF monitoring achieved low RMSEP for mAb (0.10 mg/mL in UF1; 1.57 mg/mL in DF) and maintained near-zero predictions for absent excipients during UF1.
A hybrid Uniform Design and process-data calibration strategy improves Raman-based PLS model development for real-time monitoring of ultrafiltration/diafiltration and in vitro transcription steps in biopharmaceutical manufacturing.
Nimesh Khadka and Lin Zhang, Thermo Fisher Scientific, Tewksbury, Massachusetts, USA
Abstract
Process analytical technology (PAT) increasingly relies on multivariate spectroscopic models for real-time monitoring, yet model development is often constrained by the limited availability of process-representative calibration data. In this study, we evaluated a hybrid calibration strategy that integrates Uniform Design (UD) laboratory datasets with limited in-process Raman data for ultrafiltration/diafiltration (UF/DF) and in vitro transcription (IVT) applications. Models were developed across relevant analyte concentration ranges and assessed using grouped internal cross-validation and external testing on independent process runs, with analyte-specific prediction errors reported. The results indicate that UD-assisted hybrid calibration can improve calibration coverage and support robust Raman-based monitoring while reducing experimental burden. Overall, these findings support UD-enabled hybrid dataset construction as a promising and potentially resource-efficient strategy for spectroscopic model development in the specific bioprocess systems examined here.
Introduction
Within the framework of PAT implemented under Quality by Design (QbD) principles, spectroscopic tools coupled with chemometric models are increasingly deployed to enable tighter process control in bioprocessing.1 Such model-based analysis must demonstrate fitness for purpose, robustness to process variability, and sustained performance across the analytical procedure lifecycle. From both regulatory and scientific perspectives, multivariate calibration models are no longer regarded as static constructs developed at a single point in time, but rather as lifecycle-managed analytical systems that must remain reliable under changing operating conditions, raw material variability, and dynamic process behavior.2,3 Because multivariate analytical procedures are lifecycle-managed under PAT/QbD expectations, calibration strategies must balance controlled experimental variation with representative process variability.
A widely adopted approach for chemometric model development combines Design of Experiments (DoE)-generated data with process data. This hybrid strategy integrates orthogonality and controlled factor variation of DoE with complex process variations. DoE datasets reduce multicollinearity by varying factors independently and enable calibration data to span a broad concentration range of the target analyte, which may be impractical or infeasible to obtain directly from routine process operation. The benefit of expanding the calibration range can be illustrated using simple linear regression, where the variances of the estimated intercept and slope are given by
with σ2 denoting the residual variance and Σi=1n(xi − x̄)2 representing the spread of predictor values.4,5 These relationships show that increasing predictor spread reduces parameter variance and improves model stability, an effect readily achieved through DoE. However, DoE datasets alone often lack the structured noise, temporal dependencies, nonlinear interactions, and operating-condition effects present in real bioprocesses. As a result, models developed solely from DoE data may exhibit bias or slope error when applied to process data. In contrast, process data capture correlated concentration changes, matrix effects, instrumental drift, and dynamic system behavior, enabling models to better reflect actual operating conditions. Incorporating both data types therefore enhances model robustness and generalizability while reducing sensitivity to systematic and random variation.
Efficient design strategies are required to adequately cover the relevant design space while minimizing experimental burden. UD is a space-filling DoE approach in which sampling points are uniformly distributed across the design domain, ensuring representative coverage without reliance on predefined linear or non-linear model assumptions.6,7 Such uniform coverage promotes accurate estimation of the predictor’s mean and maximizes the effective spread of predictor values, which is consistent with equations [1] and [2], reduces variance in both slope and intercept estimates. By avoiding localized clustering or imbalanced spread, UD limits sensitivity to sampling bias and improves robustness of the resulting calibration. These characteristics make UD particularly well suited for multivariate chemometric modeling, where underlying relationships may be linear, nonlinear, or partially unknown. Other key advantages of UD includes efficient exploration of high-dimensional spaces, accommodation of multiple-factor levels, and minimal bias toward any specific model structure.7
In multivariate methods such as Partial Least Squares (PLS), the regression vector is expressed as
where estimator stability depends on the conditioning of the predictor covariance matrix XTX. Poorly designed calibration datasets, characterized by clustered samples or insufficient coverage of the experimental domain, can lead to unstable latent variable extraction, increased estimation variance, and reduced predictive robustness for the PLS model.8 Also, previous studies showed that UD was better than center composite design and orthogonal array design for non-linear systems.7 Thus, we hypothesized that UD would improve conditioning, stabilize latent variables, and lower prediction variance for the PLS model.
In this study, we evaluate the effectiveness of UD in accelerating the PLS model development for Raman-based bioprocessing applications with the performance acceptable for process monitoring and control. Raman applications for these bioprocesses, including downstream UF/DF operations and IVT, have been previously demonstrated and serve as benchmarks for comparison in this work.9–11
Materials and Methods
The calibration samples for sucrose, histidine, arginine, and protein quantification models for UF/DF process were generated using publicly available UD tables.12 A single sample was prepared per UD level. Sucrose, histidine, and arginine were evaluated across fifteen concentration levels spanning 0 - 200 mg/mL, 0 - 15 mg/mL, and 0 - 50 mg/mL, respectively. Raman spectra were acquired using MarqMetrix™ All-in-One Process Raman Analyzer with DoE solutions flowed through a flow cell at 100 mL/min to match UF/DF operating conditions (450 mW laser power, 3 s exposure, three spectral averages). Inline Raman data in UF/DF process were collected across three stages (UF1, DF, and UF2), with five data points per stage. During UF1, protein concentrations ranged from 0 - 33 mg/mL in Tris buffer. During DF, protein concentrations ranged from 33 - 36 mg/mL, with corresponding histidine, arginine, and sucrose concentrations ranging from 0 - 5 mg/mL, 0 - 7 mg/mL, and 0 - 120 mg/mL, respectively. During UF2, protein concentrations increased from 36 to 155 mg/mL, while histidine and arginine remained approximately constant (5.0 - 5.2 mg/mL and 7.0 - 7.1 mg/mL, respectively), and sucrose decreased from 120 to 108 mg/mL due to volume-exclusion effects. Reference measurements for histidine and arginine were obtained using HPLC, while protein concentrations were quantified using Thermo Scientific™ NanoDrop™ Ultra spectrophotometers and fluorometers. HPLC and protein quantification analysis were performed in triplicate, and the averaged reference value was assigned to the replicate Raman measurements for each UD or process sample.
For IVT reaction monitoring, ATP, GTP, CTP, CleanCap (m7G(5′)ppp(5′)(2′-O-methyl-A)pG), and N1-methyl-Ψ-UTP were selected as factors for the DoE. The design space comprised fifteen concentration levels for each analyte, spanning 0 mM (lower control limit, LCL) to 10 mM (upper control limit, UCL). Raman spectra were acquired using an immersible Raman probe with same acquisition parameters as described above. A single IVT process was conducted, with samples collected at 0, 10, 30, 40, and 50 min. HPLC reference measurements were performed in triplicate and averaged, and the average value was assigned to the replicate Raman measurements.
DoE and process samples were merged to develop calibration models. PLS models for histidine, arginine, and sucrose were developed using Raman spectra in the range of 800 - 3240 cm−1. The protein PLS model was developed using spectral region 1570 to 1750 cm−1 and 3100 - 3240 cm−1. Spectral preprocessing was applied sequentially as follows: (i) region selection (ii) normalization to the water band using infinity norm calculated over the 3100 - 3240 cm−1 region; (iii) Savitzky-Golay first-derivative filtering (order 2, window width 13); and (iv) mean centering of both spectral and concentration datasets. All replicate spectrum and corresponding reference measurements for UD or process data were labeled as a single data group.
Model optimization was performed using leave-one-out cross-validation implemented in a grouped fashion, such that all replicates associated with one group were excluded together in each fold, preventing information leakage between training and cross validation sets. All preprocessing steps and latent-variable (LV) selection were nested within each cross-validation fold. The optimal number of LVs was selected by minimizing the root mean square error of calibration (RMSEC) and cross-validation (RMSECV), while maintaining an RMSECV-to-RMSEC ratio close to unity. PLS models for IVT components were developed using the same workflow, using spectral regions of 500 - 3240 cm−1. The final models with all parameters locked were evaluated on independent process runs that were not included in training. The overall workflow is summarized in Figure 1. For UF/DF validation, five samples were collected from each stage, and reference measurements for protein, histidine, and arginine were compared with corresponding inline Raman predictions. Model performance was assessed using the root mean square error of prediction (RMSEP). For IVT validation, samples were collected at 0, 10, 30, 40, and 60 minutes, with analyte concentrations determined by HPLC. The performance of each model was evaluated across four independent IVT runs using RMSEP.
Figure 1. Model development workflow.
Data management and chemometric analyses were performed using Python and SOLO 9.5 (Eigenvector Research, Inc., Manson, WA, USA).
Results and Discussion
The UD tables used to generate experimental conditions are represented by Un(ns) where n is the number of experimental runs and s is the number of factors.12 The UD table was constructed based on centered L2 discrepancies minimization.6 For sucrose, histidine, and arginine we used U15(153) table as shown in Figure 2a. Each column of the UD table corresponds to a factor, and each row represents an experimental run.
The principal component analysis (PCA) of the DoE levels is shown in Figure 2b. The three principal components (PC1, PC2, and PC3) each explained approximately 33% of the total variance, demonstrating that variability was evenly distributed across three orthogonal latent directions rather than being dominated by a single component. Loading vectors for sucrose (red), histidine (green), and arginine (blue) were overlaid to visualize the contribution of each factor to the PCA space. The direction of each vector was determined by the sign of its PCA loadings and indicated the direction of increasing factor levels in the latent variable space, while the vector length, determined by the magnitude of the loadings, reflected the extent to which each factor was represented by the retained principal components. The similar lengths of the vectors indicated that all three factors contributed comparably to overall variability, with no single factor disproportionately influencing the design space. The angles between vectors, calculated via their dot products, were close to 90°, and the resulting near orthogonal orientation indicated low correlation among the experimental factors within the explored design space. Together, the near-orthogonality of the loading vectors, their comparable magnitudes, and the uniform dispersion of samples relative to these vectors were consistent with the premise of UD design, namely uniform distribution of sample with minimum redundancy.
Figure 2. UD for DoE generation. (a) UD table showing factor levels for sucrose, histidine, and arginine; (b) three-dimensional representation illustrating wide distribution of samples in the design space; (c) mapped concentrations after scaling UD levels to experimental ranges (LCL and UCL).
To convert the UD levels into concentrations, we mapped the discrete levels to the concentration domain using a linear scaling transformation as:
where uij is the UD table entry and q is the total number of levels. This formulation ensured that the lowest and highest UD levels correspond exactly to the lower and upper bounds of the experimental range, respectively, with intermediate levels evenly distributed between these limits as shown in Figure 2c.
Raman spectrum of the DoE mixtures containing sucrose, histidine, and arginine are shown in Figure 3a, while the augmented dataset with inclusion of UF/DF process data is presented in Figure 3b. Relative to the DoE data, the process spectra had elevated baseline contributions and additional spectral features arising from process related components. Notably, a distinct protein amide I Raman band appeared in the 1600 - 1700 cm−1 region, which is clearly resolved in the second derivative spectra shown in Figure 3c. This band primarily originates from C=O stretching vibrations of the peptide backbone and contains information associated with protein secondary structure.13 Inspection of Figure 3c showed minimal spectral interference in the amide I region from water, sucrose, histidine, arginine, or other UF/DF components.
Figure 3. Raman spectral datasets used for model development. (a) spectra of DoE mixtures containing sucrose, histidine, and arginine; color coded by histidine concentration (b) augmented dataset including UF/DF process spectra showing additional baseline and matrix contributions; color coded by protein concentration (c) second-derivative spectra highlighting the protein amide I region (~1600 - 1700 cm−1) with minimal interference from other components. The spectra are color coded by protein concentration.
Protein was excluded from the DoE design based on prior knowledge that the amide I Raman band is spectrally distinct in this UF/DF system. When distinct spectral features are present, multivariate methods such as PLS can more effectively isolate analyte specific information. This observation is consistent with our previous work where we demonstrated that a classical least squares model based solely on the amide I region and water band achieved performance comparable to PLS due to minimum interferences.14 Excluding protein from the DoE also provided practical advantages such as base dataset transferability across different protein systems and avoidance of potential protein instability under DoE conditions.
Although protein information was derived exclusively from process data, inclusion of protein free DoE samples was critical for improving model specificity. These samples served as negative controls, ensuring minimal covariance between excipient related spectral variation and protein concentration. In general, robust PLS models are obtained when latent variables capture both positive covariance associated with the target analyte and near zero or negative covariance associated with non-target variation. In this case, the DoE data constrained the PLS weights to emphasize protein-specific spectral features and reduced the impact of multicollinearity.
In contrast, sucrose, histidine, and arginine exhibited significant spectral overlap with each other and with the protein, complicating model development. Histidine and arginine are also present as amino acids within the protein backbone, leading to highly similar Raman signatures in their free and protein-bound forms. As a result, the measured Raman spectrum can be considered a composite of protein-associated signals and residual contributions from excipients and other UF/DF components. Distinct protein features, most notably the amide I band, enable partial deconvolution of the spectrum into protein-associated and residual components. The residual signal can then be leveraged to predict concentrations of sucrose, histidine, and arginine. Although PLS does not perform this separation explicitly, it captures these relationships through its latent variable structure.
The calibration and cross-validation statistics for the developed PLS models are summarized in Table 1. All models showed strong calibration performance, with high R2 values for cross validation (0.96 - 0.99) and low RMSECV. The protein model required fewer latent variables, reflecting the presence of distinct amide I region, whereas the other analytes required more latent variables due to high multicollinearity.
Table 1. Summary of PLS calibration model parameters and performance metrics for each analyte, including spectral regions used, concentration ranges, number of latent variables (LVs), root mean square error of cross-validation (RMSECV), and coefficient of determination (R² CV).
Analyte
Spectral region (cm−1)
Concentration range (mg/mL)
LVs
RMSECV (mg/mL)
R² CV
Histidine
800 - 3240
0 - 15
5
0.42
0.97
Arginine
800 - 3240
0 - 50
5
0.45
0.98
Sucrose
800 - 3240
0 - 200
5
2.40
0.96
Protein
1570 - 1750 and 3100 - 3240
0 - 155
2
0.72
0.99
Next, the performance of the developed hybrid models was evaluated using an independent UF/DF process. As shown in Figure 4, the ultrafiltration step (UF1) started with 10.52 mg/mL monoclonal antibody (mAb) in Tris buffer and concentrated the protein stepwise to 33 mg/mL. During this stage, protein predictions showed strong agreement with reference measurements in five samples across range of 0 to 33 mg/mL, with a root mean square error of prediction (RMSEP) of 0.10 mg/mL. As expected, histidine, arginine, and sucrose were absent during UF1, and their predicted concentrations remained near zero with RMSEP of 0.21, 0.27, and 1.86 mg/mL respectively. During the subsequent DF steps, protein remained relatively constant while Tris buffer was progressively exchanged with histidine, arginine, and sucrose, reaching steady-state target concentrations (95 mg/mL sucrose, 4.0 mg/mL histidine, and 7.0 mg/mL arginine) within approximately 80 minutes. These trends were well captured by Raman predictions. RMSEP values for histidine, arginine, sucrose, and protein, calculated from five pooled DF samples collected at the beginning, midpoint, and end of DF were 0.32, 0.25, 2.33, and 1.57 mg/mL, respectively. These prediction errors were comparable to values reported in the literature from process-data-based models.9,15 Previous studies have reported correlations between nonspecific background effects and Raman-based protein predictions during UF/DF processing.9 No such nonspecific background correlation was observed in this study. Analysis of the loadings, variable importance on projection (VIP) score and selectivity ratio indicated that the model was specific for Raman bands associated with proteins as discussed elsewhere.16,17
Figure 4. Performance of the hybrid Raman models applied to an independent UF/DF process. Predicted concentrations of mAb, sucrose, histidine, and arginine are shown over time. During UF1, protein concentration increases while other components remain near zero. During DF, protein concentration remains relatively constant, while histidine, arginine, and sucrose increase and reach steady-state values consistent with buffer exchange dynamics.
Following successful validation of the UD-based DoE approach for the UF/DF process, UD was evaluated for monitoring IVT reactions. The calibration statistics and predictive performance of the developed PLS models are summarized in the insets of the correlation plots shown in Figure 5(a-e). The average coefficients of determination (R2) greater than 0.98 showed strong linear agreement across both calibration and test datasets. RMSEP values ranged from 0.29 to 0.44 mM (Figure 5f) demonstrating good quantitative accuracy across the concentration range studied. Overall, these results demonstrated that the hybrid UD-based approach enabled multicomponent prediction in an IVT reaction with accuracy comparable to values reported in literature.10
Figure 5. Performance of Raman based PLS models for IVT reaction monitoring. Predicted versus measured concentrations are shown for (a) ATP, (b) GTP, (c) CTP, (d) modified UTP, and (e) CleanCap, with calibration and test datasets indicated. (f) The RMSEP for each analyte is listed.
In addition, we have previously demonstrated a UD-based strategy for glucose monitoring in bioreactor processes.18 In that study, mixtures containing glucose, lactate, and glutamine were prepared using UD table (U24(243)), and Raman-based calibration models were developed for glucose. When applied to an independent bioreactor run, the model successfully predicted the trend for glucose consumption, although a prediction bias was observed. This bias was reduced by augmenting the calibration dataset with simulated baseline variations. The averaged RMSEP for two bioreactor runs (n = 90 samples) was approximately 0.5 g/L, which was close to the reported values in the literature for models developed using process-only datasets.19
Conclusion
This study indicates that UD can provide an efficient and practical approach for constructing calibration datasets for multivariate modeling in biomanufacturing workflows. When UD datasets were combined with limited in-process data, the resulting hybrid datasets appeared to improve design-space coverage and model stability. The case studies presented here demonstrate the applicability of this approach to Raman-based monitoring in UF/DF and IVT processes. While these results support UD-assisted hybrid calibration as a promising strategy for the specific systems examined, broader claims regarding transferability, scalability, and suitability for advanced process control will require further benchmarking across a larger number of independent runs, more diverse batch conditions, and direct comparison with alternative sample-selection strategies.
References
(1) Food and Drug Administration. PAT — A Framework for Innovative Pharmaceutical Development, Manufacturing, and Quality Assurance, 2004.
(2) Food and Drug Administration. Q14 Analytical Procedure Development, 2024.
(3) Food and Drug Administration. Q2(R2) Validation of Analytical Procedures, 2024.
(4) Draper, N. R.; Smith, H. Fitting a Straight Line by Least Squares. In Applied Regression Analysis; John Wiley & Sons, Ltd, 1998; pp 15–46. DOI:
(5) Beebe, K. R.; Pell, R. J.; Seasholtz, M. B. Chemometrics: A Practical Guide; John Wiley & Sons: New York, 1998.
(6) Fang, K.-T.; Lin, D. K. J.; Winker, P.; Zhang, Y. Uniform Design: Theory and Application. Technometrics 2000, 42 (3), 237–248. DOI:
(7) Zhang, L.; Liang, Y.-Z.; Jiang, J.-H.; Yu, R.-Q.; Fang, K.-T. Uniform Design Applied to Nonlinear Multivariate Calibration by ANN. Anal. Chim. Acta 1998, 370 (1), 65–77. DOI:
(8) Martens, H.; Næs, T. Multivariate Calibration; John Wiley & Sons, Ltd: Chichester, UK, 1991.
(9) Rolinger, L.; Hubbuch, J.; Rüdt, M. Monitoring of Ultra- and Diafiltration Processes by Kalman-Filtered Raman Measurements. Anal. Bioanal. Chem. 2023, 415 (5), 841–854. DOI:
(10) Fulber, J. P. C.; Xu, X.; Radi, H. E.; Grollier, K.; Kamen, A. A. Real-Time Monitoring of In Vitro Transcription for Production of mRNA Using Raman Spectroscopy. Biotechnol. Bioeng. 2026. DOI:
(11) Khadka, N. Quantification of Adenine Residues in the PolyA Tail of Oligonucleotides Using Raman Spectroscopy. Spectroscopy 2025, 40 (5), 32–42. DOI:
(12) Ma, C.-X. Uniform Design Based on CL2.
(13) Peters, J.; Park, E.; Kalyanaraman, R.; Luczak, A.; Ganesh, V. Protein Secondary Structure Determination Using Drop Coat Deposition Confocal Raman Spectroscopy. Spectroscopy 2016, 31 (10), 31–39.
(14) Khadka, N.; Pleitt, K.; Nolasco, M. A Single Concentration-Based Amide I Model for Real-Time Protein Quantification and Quality Assessment in Ultrafiltration/Diafiltration Using Raman Spectroscopy. Presented at IFPAC, Baltimore, MD, March 2026.
(15) Passno, C.; Cowan, C.; Tustian, A. Use of Raman Spectroscopy in Downstream Purification. US 12,398,176 B2, August 26, 2025.
(16) Nolasco, M.; Pleitt, K.; Khadka, N. Raman-Based Accurate Protein Quantification in a Matrix that Interferes with UV-Vis Measurement; Application Note AN1163; Thermo Fisher Scientific: Waltham, MA, 2024.
(17) Khadka, N.; Pleitt, K.; Rivera, M. N. Raman-Based Quality Monitoring of Biopharmaceutical Production Processes. WO 2025/250834 A1, December 4, 2025.
(18) Khainovski, N.; Zhang, L.; Khadka, N. System and Method for Spectroscopic Determination of a Chemometric Model from Sample Scans. WO 2025/024537 A9, February 12, 2026.
(19) Abu-Absi, N. R.; Kenty, B. M.; Cuellar, M. E.; Borys, M. C.; Sakhamuri, S.; Strachan, D. J.; Hausladen, M. C.; Li, Z. J. Real Time Monitoring of Multiple Parameters in Mammalian Cell Culture Bioreactors Using an In-Line Raman Spectroscopy Probe. Biotechnol. Bioeng. 2011, 108 (5), 1215–1221. DOI:
About the Authors
Nimesh Khadka, Ph.D. is a Senior Products Specialist at Thermo Fisher Scientific, specializing in advanced spectroscopic technologies and AI/ML-driven chemometric solutions for product applications, product marketing, product development, and new product introduction (NPI). He develops and deploys innovative Raman, FT-IR, and other vibrational spectroscopy solutions to address complex analytical challenges across bioprocessing, pharmaceuticals, chemicals, and industrial manufacturing. His expertise includes Process Analytical Technology (PAT), process development, chemometric modeling, machine learning, multivariate data analysis, and digital process analytics for real-time monitoring, process optimization, and automation. Dr. Khadka is passionate about translating cutting-edge science into practical, data-driven solutions and collaborates with customers worldwide to implement advanced spectroscopic technologies that accelerate innovation, improve product quality, enhance process understanding, and enable intelligent decision-making. He holds a Ph.D. in Analytical Biochemistry from Utah State University and a B.S. in Physical Biochemistry from Pokhara University.
Lin Zhang, Ph.D., is an algorithm development manager at Thermo Fisher Scientific in Tewksbury, Massachusetts. His expertise includes chemometrics, multivariate calibration, and spectroscopic method development for process analytical technology and portable analytical instruments. He holds a Ph.D. in Analytical Chemistry and an M.S. in Computer Science from Ohio University, as well as M.S. and B.S. degrees in Analytical Chemistry from Hunan University.




