Structural Variant benchmarking frameworks: Parameterization, matching logic, and evaluation assumptions presented through HG002 and NA12878

Wait 5 sec.

by Gamze Maden, Mehmet Baysan, Nizamettin AydınStructural Variants (SVs) constitute a significant yet complicated class of genomic changes. Due to their imprecise breakpoints, variable size, genomic context and representations across SV callers complicate their accurate detection. Several SV callers have been developed over the years; however, the process of evaluating the outputs of SV callers is inherently more complicated than single nucleotide polymorphisms (SNPs). Unlike SNPs where there is simply a binary representation of whether the nucleotide at a specific position exists, SV analysis must clarify partially overlapping events caused by imprecise breakpoints. This ambiguity has motivated the development of specialized frameworks to support the comparison of SVs. Each of these frameworks implement different assumptions about breakpoint resolution, size concordance or sequence similarity. In this study, we described the underlying strategy of SV comparison modules of three frameworks: Truvari, EvalSVcallers, and SVbenchmark. The primary focus of our study is to reveal how the implemented parameters of each framework influence matching behavior as well as their evaluation outcomes. The conceptual and algorithmic differences of these frameworks are further illustrated through an empirical example based on the query sets of five callers (Manta, Delly, Lumpy, GRIDSS, Wham) for the HG002 and NA12878 reference samples to stress how distinct benchmarking strategies translate into framework-dependent evaluation outcomes. Thus, the reported metrics demonstrate effects of such differences on reproducibility and interpretation rather than prioritizing performance ranking.