Hi Nelson,
I'm using SNPGenie to analyze pooled NGS data and have two related questions about coverage.
Background:
My sample inclusion criteria: ≥80% of sites with >200x depth
In practice, it's nearly impossible to achieve 100% coverage — some sites are always under-covered or not covered at all
Questions:
Does incomplete or uneven coverage affect the accuracy of SNPGenie's estimates (e.g., π, dN/dS)? If so, what's the recommended way to handle under-covered sites?
I'm using lofreq for SNP calling, which explicitly states that users don't need to filter by depth since it handles this internally. However, this means the output doesn't tell me which sites were effectively analyzed. How should I determine the denominator for per-site diversity calculations? Do I need to independently calculate coverage from BAM files to define valid sites?
Thanks!
Hi Nelson,
I'm using SNPGenie to analyze pooled NGS data and have two related questions about coverage.
Background:
My sample inclusion criteria: ≥80% of sites with >200x depth
In practice, it's nearly impossible to achieve 100% coverage — some sites are always under-covered or not covered at all
Questions:
Does incomplete or uneven coverage affect the accuracy of SNPGenie's estimates (e.g., π, dN/dS)? If so, what's the recommended way to handle under-covered sites?
I'm using lofreq for SNP calling, which explicitly states that users don't need to filter by depth since it handles this internally. However, this means the output doesn't tell me which sites were effectively analyzed. How should I determine the denominator for per-site diversity calculations? Do I need to independently calculate coverage from BAM files to define valid sites?
Thanks!