Predicting DNA/RNA Extraction Yields for NGS Using Machine Learning-Based Analysis of H&E-Stained Slides
Authors
Dmitrii Ivchenkov, Vladimir Kushnarev, Anna Belozerova, Anna Bejanyan, Kirill Kriukov, Ekaterina Postovalova, Alexander Bagaev, Alexander Sarachakov
- BostonGene, Corp., Waltham, USA
Abstract
Introduction:
Ensuring high-quality nucleic acid (NA) extractions from tumors while conserving valuable tissue specimens is pivotal to the success of NGS for molecular tumor profiling. Traditionally, pathologists visually assess H&E-stained formalin-fixed paraffin-embedded (FFPE) slides to determine tumor purity and tissue quantities needed for optimal DNA/RNA yields. Since this method can be highly subjective, we developed an ML-based tool to enhance the precision of predicting NA yields from H&E FFPE slides and as such, reduce pre-analytical failure rates of NGS workflows in clinical practice.
Methods:
We used 2,823 H&E-stained FFPE slides of diverse tumor origin from an internal cohort. Model training employed a 5-fold cross-validation on 2,399 samples, with the remaining 424 slides reserved for final model testing. All samples passed quality assessment by pathologists, and samples with adequate NA yields were sequenced. The model architecture comprised distinct convolutional neural networks for precise tissue and cellular segmentation to generate features, including tissue area and cell density. Then, these features were fed into RNA/DNA Yield Regression Models employing a Langmuir isotherm model and ResNet neural network to refine NA yield prediction. Accordingly, our tool determined the optimal slide counts needed to achieve the target NA yield (10 ng/µl), with yields < 2 ng/µl considered as failures. R2 and mean absolute error (MAE) were calculated to assess model performance.
Results:
Model testing on 424 test samples showed R²=0.21 (MAE=1.2 µg) and R²=0.15 (MAE=1.2 µg) for DNA and RNA extraction, respectively. These metrics indicate that our model could yield meaningful predictions of sample quantities needed for extraction yields that meet the target NA concentration. Specifically, it reliably predicted the number of slides needed per sample up to 10 slides, with approximately 1% failure rate. Moreover, since the failure rate increased when the slide count exceeded 10, one may use our model to identify samples that are unlikely to produce meaningful NA yields if they require >10 slides.
Conclusions:
Our findings suggest that our model for ML-based analysis of H&E slides is reliable for samples requiring up to 10 slides for successful NA extraction. This model is beneficial to clinical laboratories because it enhances NA extraction outcomes by reliably predicting the yield, potential problems, and failure analytics, contributing to considerable resource savings. This enables laboratories to improve their cost-efficiency in processing complex samples with limited quantities in order to achieve desirable analytic outcomes in tumor profiling, leading to more precise diagnostic outcomes.
Ensuring high-quality nucleic acid (NA) extractions from tumors while conserving valuable tissue specimens is pivotal to the success of NGS for molecular tumor profiling. Traditionally, pathologists visually assess H&E-stained formalin-fixed paraffin-embedded (FFPE) slides to determine tumor purity and tissue quantities needed for optimal DNA/RNA yields. Since this method can be highly subjective, we developed an ML-based tool to enhance the precision of predicting NA yields from H&E FFPE slides and as such, reduce pre-analytical failure rates of NGS workflows in clinical practice.
Methods:
We used 2,823 H&E-stained FFPE slides of diverse tumor origin from an internal cohort. Model training employed a 5-fold cross-validation on 2,399 samples, with the remaining 424 slides reserved for final model testing. All samples passed quality assessment by pathologists, and samples with adequate NA yields were sequenced. The model architecture comprised distinct convolutional neural networks for precise tissue and cellular segmentation to generate features, including tissue area and cell density. Then, these features were fed into RNA/DNA Yield Regression Models employing a Langmuir isotherm model and ResNet neural network to refine NA yield prediction. Accordingly, our tool determined the optimal slide counts needed to achieve the target NA yield (10 ng/µl), with yields < 2 ng/µl considered as failures. R2 and mean absolute error (MAE) were calculated to assess model performance.
Results:
Model testing on 424 test samples showed R²=0.21 (MAE=1.2 µg) and R²=0.15 (MAE=1.2 µg) for DNA and RNA extraction, respectively. These metrics indicate that our model could yield meaningful predictions of sample quantities needed for extraction yields that meet the target NA concentration. Specifically, it reliably predicted the number of slides needed per sample up to 10 slides, with approximately 1% failure rate. Moreover, since the failure rate increased when the slide count exceeded 10, one may use our model to identify samples that are unlikely to produce meaningful NA yields if they require >10 slides.
Conclusions:
Our findings suggest that our model for ML-based analysis of H&E slides is reliable for samples requiring up to 10 slides for successful NA extraction. This model is beneficial to clinical laboratories because it enhances NA extraction outcomes by reliably predicting the yield, potential problems, and failure analytics, contributing to considerable resource savings. This enables laboratories to improve their cost-efficiency in processing complex samples with limited quantities in order to achieve desirable analytic outcomes in tumor profiling, leading to more precise diagnostic outcomes.
Latest publications