Columns | Column: Chemometrics in Spectroscopy

This is Chemometrics in Spectroscopy Column Number 251. In Column Number 250, we left you with a puzzle. We had just shown that a little-known piece of two-hundred-year-old mathematics, Lagrange's method of undetermined multipliers, could force an MLR calibration to ignore one particular, carefully chosen direction of repack-induced spectral change. It worked, but it only handled one direction at a time, and we admitted we had no idea whether the trick could be extended to PCR or PLS. That bothered us. So in this installment we go looking for the more general version of the idea, and we find it sitting in some very good recent work by Neal Gallagher and Nathanial Watson on what they call clutter suppression (a framework built for a completely different problem, target detection in hyperspectral imaging, that turns out to fit ours almost perfectly). Put the two together and you get an algorithm that whitens the calibration spectra with a clutter covariance matrix estimated from repack replicates, then hands the whitened data to the same Lagrangian machine we built last time. The payoff: the correction now generalizes to several, non-proportional directions of diffuse-reflection variability at once, and, because it is just a preprocessing step, it works ahead of PLS and PCR as well as MLR, which finally answers the question we could not answer before. This approach may very well be an answer to repack variation in repeated solid or slurry sample measurements using diffuse reflection. In this column we change gears with our writing tone and format by taking a "chemometry" approach, rather than a strict tutorial one. In future columns we hope to "unpack" this information in our more typical extended and tutorial manner. Let's explore solving this repack variation problem together.

Gold Anniversary Candles: Our 250th Article on Statistics, Chemometrics, and AI in over 40 years! ©  Valerii Evlakhov -chronicles-stock.adobe.com

In their milestone 250th column, Howard Mark and Jerome Workman, Jr. describe a mathematically rigorous algorithm that minimizes or eliminates sampling repack variation in near-infrared spectroscopy. The method separates systematic spectral changes caused by sample rearrangement from true compositional information, enabling more robust calibration models and significantly improving analytical repeatability for powdered and heterogeneous solid samples.

Hand holding a glowing AI sphere symbolizing the power and potential of artificial intelligence. | Image Credit: © lucegrafiar - stock.adobe.com.

This “Chemometrics in Spectroscopy” column traces the historical and technical development of these methods, emphasizing their application in calibrating spectrophotometers for predicting measured sample chemical or physical properties—particularly in near-infrared (NIR), infrared (IR), Raman, and atomic spectroscopy—and explores how AI and deep learning are reshaping the spectroscopic landscape.

Big data concept. | Image Credit: © your123 - stock.adobe.com.

This column is the continuation of our previous column that describes and explains some algorithms and data transforms beyond those most commonly used. We present and discuss algorithms that are rarely, if ever, seen or used in practice, despite that they have been proposed and described in the literature.

Abstract graphic world map illustration on blue background, big data and networking concept. 3D Rendering | Image Credit: © Pixels Hunter - stock.adobe.com.

In this column and its successor, we describe and explain some algorithms and data transforms beyond those commonly used. We present and discuss algorithms that are rarely, if ever, used in practice, despite having been described in the literature. These comprise algorithms used in conjunction with continuous spectra, as well as those used with discrete spectra.

Spectroscopy

A newly discovered effect can introduce large errors in many multivariate spectroscopic calibration results. The CLS algorithm can be used to explain this effect. Having found this new effect that can introduce large errors in calibration results, an investigation of the effects of this phenomenon to calibrations using principal component regression (PCR) and partial least squares (PLS) is examined.

Spectroscopy

As we have previously discussed, the most time consuming and bothersome issue associated with calibration modeling and the routine use of multivariate models for quantitative analysis in spectroscopy are the constant intercept (bias) or slope adjustments. These adjustments must be routinely performed for every product and each constituent model. For transfer and maintenance of multivariate calibrations this procedure must be continuously implemented to maintain calibration prediction accuracy over time. Sample composition, reference values, within and between instrument drift, and operator differences may be the cause of variation over time. When calibration transfer is attempted using instruments of somewhat different vintage or design type the problem is amplified. In this discussion of the problem we continue to delve into the issues causing prediction error, bias and slope changes for quantitative calibrations using spectroscopy.