UCL Discovery
UCL home » Library Services » Electronic resources » UCL Discovery

SWAMP: Sliding Window Alignment Masker for PAML.

Harrison, PW; Jordan, GE; Montgomery, SH; (2014) SWAMP: Sliding Window Alignment Masker for PAML. Evol Bioinform Online , 10 197 - 204. 10.4137/EBO.S18193. Green open access

[thumbnail of f_4531-EBO-SWAMP-Sliding-Window-Alignment-Masker-for-PAML.pdf_6055.pdf] PDF
f_4531-EBO-SWAMP-Sliding-Window-Alignment-Masker-for-PAML.pdf_6055.pdf

Download (1MB)

Abstract

With the greater availability of genetic data, large genome-wide scans for positive selection increasingly incorporate data from a range of sources. These data sets may be derived from different sequencing methods, each of which has potential sources of error. Sequencing errors, compounded by alignment errors, greatly increase the number of false positives in tests for adaptive evolution. Genome-wide analyses often fail to fully address these issues or to provide sufficient detail on postalignment masking/filtering. Here, we introduce a Sliding Window Alignment Masker for Phylogenetic Analysis by Maximum Likelihood (SWAMP) that scans multiple-sequence alignments for short regions enriched with unreasonably high rates of nonsynonymous substitutions caused, for example, by sequence or alignment errors. SWAMP prevents their inclusion in downstream evolutionary analyses and therefore increases the reliability of downstream analyses. It is able to effectively mask short stretches of erroneous sequence, particularly prevalent in low-coverage genomes, which may not be detected by existing methods based on filtering by sitewise conservation or alignment confidence. SWAMP offers a flexible masking approach, and the user can apply different masking regimens to specific branches or sequences in the phylogeny allowing the stringency of masking to vary according to branch length, expected divergence levels, or assembly quality. We exemplify SWAMPs effectiveness on a dataset of 6,379 protein-coding genes from primate species, including data of variable quality. Full reporting of the software parameters will further improve the reproducibility of genome-wide analyses, as well as reduce false-positive rates.

Type: Article
Title: SWAMP: Sliding Window Alignment Masker for PAML.
Location: New Zealand
Open access status: An open access version is available from UCL Discovery
DOI: 10.4137/EBO.S18193
Publisher version: http://dx.doi.org/10.4137/EBO.S18193
Language: English
Additional information: © the authors, publisher and licensee Libertas Academica Limited. This is an open-access article distributed under the terms of the Creative Commons CCCC-BY-NCNC 3.0 License.
Keywords: PAML, adaptive evolution, genome evolution, molecular evolution, phylogenetics, sequence analysis
UCL classification: UCL
UCL > Provost and Vice Provost Offices > School of Life and Medical Sciences
UCL > Provost and Vice Provost Offices > School of Life and Medical Sciences > Faculty of Life Sciences
URI: https://discovery.ucl.ac.uk/id/eprint/1460143
Downloads since deposit
173Downloads
Download activity - last month
Download activity - last 12 months
Downloads by country - last 12 months

Archive Staff Only

View Item View Item