A novel machine learning classifier for SCLC transcriptional subtypes
Authors
Simon Heeke¹, Ivan Valiev², Mikhail Kleimenov², Anastasia Sobol², Evgeny Barykin², Mikhail Slizen², Anastasia Evdokimova², Vladimir Kushnarev², Artem Kosmin², Anna Butusova², Francesca Paradiso², Patrick Clayton², Sheila T. Yong², Alexander Bagaev², Lauren Byers¹, John V. Heymach¹, Carl M. Gay¹
- 1 The University of Texas MD Anderson Cancer Center, Houston, TX, USA
- 2 BostonGene Corporation, Waltham, MA, USA
Abstract
Introduction:
Small cell lung cancer (SCLC) exhibits significant molecular and clinical heterogeneity, leading to suboptimal clinical outcomes. To improve patient stratification and treatment selection, Gay et al. (2021) developed a transcriptional classification system that identified four distinct subtypes: SCLC-A (ASCL1), N (NEUROD1), P (POU2F3), and I (Inflamed). In particular, SCLC-I showed higher frequency of an inflamed gene signature than SCLC-A, N, and P, along with longer duration of response to immunotherapy when added to chemotherapy (e.g., IMpower133 and CASPIAN). However, the SCLC subtype clustering methods used in the original analyses by Gay et al. are not ideal for patient selection and stratification. While Heeke et al. (2024) created a bioinformatics tool for assigning individual patient samples to specific SCLC transcriptional subtypes, its implementation in a clinically applicable assay is still challenging. To address this unmet clinical need, we enhanced the bioinformatics classification approach by creating a machine learning SCLC transcriptional classifier for assigning individual patient tumor samples to subtypes A, N, P, and I.
Methods:
In total, 167 RNA-seq samples from George et al. (2015) and Heeke et al. (2024) were used as ground truth and split into training (n=125) and tuning (n=42) sets. Our classifier consisted of three hierarchy-arranged gradient boosting binary models, each trained to distinguish between specific SCLC transcriptomic patterns. The first model distinguished between SCLC-P and non-P (SCLC-A, I, and N), the second between SCLC-I and A plus N, and the third between SCLC-A and N. Obtained probabilities of these six pre-subtypes (two from each model) were harmonized to generate four final probabilities, one for each SCLC subtype, the sum of which equaled one. The classifier was evaluated on the tuning set and then tested on a metacohort of 271 publicly available SCLC samples. The expression profile of classified samples was compared to samples of corresponding subtypes in the tuning set. Gene expression-based tumor microenvironment (TME) classification (Immune-enriched fibrotic [IE/F], Immune-enriched non-fibrotic [IE], Immune desert [ID], and Fibrotic [F]; Bagaev et al. 2021) was also applied to the SCLC metacohort.
Results:
The SCLC classifier achieved 0.86 sensitivity and 0.95 specificity on the tuning set. For the evaluation samples, the classifier revealed gene expression patterns that were consistent with those in the tuning set. SCLC subtypes were also associated with specific TME subtypes (Chi-squared, p = 1.14e-07). SCLC-I correlated with IE (Pearson correlation, p = 1.5e-15) and IE/F TMEs (p = 0.004), while SCLC-A and N correlated with the ID TME (p = 0.005 and 0.001 respectively).
Conclusion:
The proposed transcriptomic tool for subtyping and target expression will be applied to stratify patients across treatment arms in an upcoming clinical trial.
Small cell lung cancer (SCLC) exhibits significant molecular and clinical heterogeneity, leading to suboptimal clinical outcomes. To improve patient stratification and treatment selection, Gay et al. (2021) developed a transcriptional classification system that identified four distinct subtypes: SCLC-A (ASCL1), N (NEUROD1), P (POU2F3), and I (Inflamed). In particular, SCLC-I showed higher frequency of an inflamed gene signature than SCLC-A, N, and P, along with longer duration of response to immunotherapy when added to chemotherapy (e.g., IMpower133 and CASPIAN). However, the SCLC subtype clustering methods used in the original analyses by Gay et al. are not ideal for patient selection and stratification. While Heeke et al. (2024) created a bioinformatics tool for assigning individual patient samples to specific SCLC transcriptional subtypes, its implementation in a clinically applicable assay is still challenging. To address this unmet clinical need, we enhanced the bioinformatics classification approach by creating a machine learning SCLC transcriptional classifier for assigning individual patient tumor samples to subtypes A, N, P, and I.
Methods:
In total, 167 RNA-seq samples from George et al. (2015) and Heeke et al. (2024) were used as ground truth and split into training (n=125) and tuning (n=42) sets. Our classifier consisted of three hierarchy-arranged gradient boosting binary models, each trained to distinguish between specific SCLC transcriptomic patterns. The first model distinguished between SCLC-P and non-P (SCLC-A, I, and N), the second between SCLC-I and A plus N, and the third between SCLC-A and N. Obtained probabilities of these six pre-subtypes (two from each model) were harmonized to generate four final probabilities, one for each SCLC subtype, the sum of which equaled one. The classifier was evaluated on the tuning set and then tested on a metacohort of 271 publicly available SCLC samples. The expression profile of classified samples was compared to samples of corresponding subtypes in the tuning set. Gene expression-based tumor microenvironment (TME) classification (Immune-enriched fibrotic [IE/F], Immune-enriched non-fibrotic [IE], Immune desert [ID], and Fibrotic [F]; Bagaev et al. 2021) was also applied to the SCLC metacohort.
Results:
The SCLC classifier achieved 0.86 sensitivity and 0.95 specificity on the tuning set. For the evaluation samples, the classifier revealed gene expression patterns that were consistent with those in the tuning set. SCLC subtypes were also associated with specific TME subtypes (Chi-squared, p = 1.14e-07). SCLC-I correlated with IE (Pearson correlation, p = 1.5e-15) and IE/F TMEs (p = 0.004), while SCLC-A and N correlated with the ID TME (p = 0.005 and 0.001 respectively).
Conclusion:
The proposed transcriptomic tool for subtyping and target expression will be applied to stratify patients across treatment arms in an upcoming clinical trial.
Latest publications