Enhancing germline variant calling: Adjustment with tumor samples to filter low-confidence variants from clonal hematopoiesis
Authors
Maria Shumilova, Danil Stupichev, Gleb Khegai, Irina Bulusheva, Sevda Nuralieva, Miryam Spektor, Daria Melikhova, Georgy Sagaradze, Michael Goldberg, Alexander Bagaev
- BostonGene Corp., Waltham, USA
Abstract
Background
High-quality germline variant calling is crucial for accurate genetic reports; yet, certain factors can lead to false identifications. Clonal hematopoiesis (CH), caused by somatic mutations in blood cells, can mimic germline variants, while pseudogenes and mosaic mutations generate low variant allele frequency (VAF) signals that can be misinterpreted. Current methods for addressing these artifacts use only non-tumor samples, limiting their use in clinical settings for generating genetic reports for tumor samples. To expand the capability of germline variant calling, we developed a filtering method to identify and eliminate low-confidence variants from CH, pseudogenes, and mosaicism within both non-tumor and tumor samples.
Methods
We analyzed whole-exome sequencing (WES) data from 2,078 patients (55.2% females and 44.8% males; mean age 58.7 years, range: 6–99) with 175 unique oncological diagnoses. From each patient, a tumor biopsy and non-tumor samples, including whole blood (71.4%), saliva (25.3%), and buccal swabs (3.3%), were obtained. Strelka v2.9.10 was used for variant calling; precise VAF values were calculated using SAMtools mpileup v1.10. The VAF values were compared between tumor and non-tumor samples. Copy number alterations (CNAs) detected with Sequenza v2.1.2 were used to evaluate expected VAF values in tumor samples adjusted for gene dosage. Mappability metrics were estimated using bigwig reference files for hg38 with various read lengths (75bp, 100bp, 150bp, 200bp) and pyBigWig v0.3.22. The binomial distribution model was calculated based on CNAs and sample purity to determine expected VAF range and identify low-confidence variants. Variants with VAFs outside the 99.9% confidence interval were flagged as potential artifacts.
Results
Overall, 3,591 unique low-confident variants were found in 1,428 patients, with about 29% being associated with low mappability (≤0.95). In total, 64 unique variants were found in 42 CH-associated genes in 121 patients (5.8%), including actionable genes ATM, BRCA1, TP53, and RB1. The mean VAF for CH mutations was 30% (standard deviation, SD 5%) in non-tumor samples, compared to 6% (SD 3%) in tumor samples. The mean patient age for filtered-out variants was 61, consistent with higher CH probability in older individuals. In non-tumor samples, CH artifacts were found mostly in whole blood (78%). Breast cancers and sarcomas were the main diagnoses with observed CH artifacts.
Conclusion
Our filtering strategy integrates WES data from both tumor and non-tumor samples, including the copy number alteration (CNA) profiles to improve the accuracy of germline reports. This unique method reduces CH-associated errors and other artifacts, consequently streamlining quality control of genetic reports for cancer patients to guide their therapeutic decision-making.
High-quality germline variant calling is crucial for accurate genetic reports; yet, certain factors can lead to false identifications. Clonal hematopoiesis (CH), caused by somatic mutations in blood cells, can mimic germline variants, while pseudogenes and mosaic mutations generate low variant allele frequency (VAF) signals that can be misinterpreted. Current methods for addressing these artifacts use only non-tumor samples, limiting their use in clinical settings for generating genetic reports for tumor samples. To expand the capability of germline variant calling, we developed a filtering method to identify and eliminate low-confidence variants from CH, pseudogenes, and mosaicism within both non-tumor and tumor samples.
Methods
We analyzed whole-exome sequencing (WES) data from 2,078 patients (55.2% females and 44.8% males; mean age 58.7 years, range: 6–99) with 175 unique oncological diagnoses. From each patient, a tumor biopsy and non-tumor samples, including whole blood (71.4%), saliva (25.3%), and buccal swabs (3.3%), were obtained. Strelka v2.9.10 was used for variant calling; precise VAF values were calculated using SAMtools mpileup v1.10. The VAF values were compared between tumor and non-tumor samples. Copy number alterations (CNAs) detected with Sequenza v2.1.2 were used to evaluate expected VAF values in tumor samples adjusted for gene dosage. Mappability metrics were estimated using bigwig reference files for hg38 with various read lengths (75bp, 100bp, 150bp, 200bp) and pyBigWig v0.3.22. The binomial distribution model was calculated based on CNAs and sample purity to determine expected VAF range and identify low-confidence variants. Variants with VAFs outside the 99.9% confidence interval were flagged as potential artifacts.
Results
Overall, 3,591 unique low-confident variants were found in 1,428 patients, with about 29% being associated with low mappability (≤0.95). In total, 64 unique variants were found in 42 CH-associated genes in 121 patients (5.8%), including actionable genes ATM, BRCA1, TP53, and RB1. The mean VAF for CH mutations was 30% (standard deviation, SD 5%) in non-tumor samples, compared to 6% (SD 3%) in tumor samples. The mean patient age for filtered-out variants was 61, consistent with higher CH probability in older individuals. In non-tumor samples, CH artifacts were found mostly in whole blood (78%). Breast cancers and sarcomas were the main diagnoses with observed CH artifacts.
Conclusion
Our filtering strategy integrates WES data from both tumor and non-tumor samples, including the copy number alteration (CNA) profiles to improve the accuracy of germline reports. This unique method reduces CH-associated errors and other artifacts, consequently streamlining quality control of genetic reports for cancer patients to guide their therapeutic decision-making.
Latest publications