Protein MSA
Immutable multiple sequence alignment object, consisting of two or more ProteinSeqRecords with equal lengths. The data can then be regarded as a matrix of letters, with well defined columns. ProteinMSA also supports all read-only collection operations on the list of ProteinSeqRecords.
Please note while the Kotlin "x..y" range operator is supported, and works very similarly to Python's slice "x:y", y is inclusive here in Kotlin, and exclusive in Python.
Creating a ProteinMSA object: You will typically use Bio.AlignIO to read in alignments from files as ProteinMSA objects. You can also use Bio.Align to align sequences of uneven length and generate a MultipleSeqAlignment object.
You can also create a MultipleSeqAlignment object directly, with argument:
seqs - List of sequence records, required (type: ImmutableList
or List )
@throws IllegalStateException if seqs has less than two elements, or if the sequence records in seqs are not all of the same length.
from Bio.Seq import Seq from Bio.Seq import ProteinSeqRecord from Bio.Seq import ProteinMSA val record_1 = ProteinSeqRecord(Seq("MHQA"), "1") val record_2 = ProteinSeqRecord(Seq("MHQ-"), "2") val alignment = ProteinMSA(listOf(record_1, record_2))
Functions
Function to create a new ProteinSeq containing the sequence with all of the gaps for a given sampleIdx index. This allows for retrieval of sequence out of the ProteinSeq Note: this is 0 based.
Function to create a new ProteinSeq removing the gaps from the sequence for a given sampleIdx index. This allows for retrieval of sequence out of the ProteinSeq Note: This is 0 based.
Returns the number of sequences in the alignment.
Returns the ProteinMSA at the specified index idx with idx starting at zero. Negative indices start from the end of the sampleList, i.e. -1 is the last sample
Function to filter the MSA by sample index based on the provided filterLambda. This will return another ProteinMSA. Note this will not work with negative indices
Function to return a NucMSA given a collection of both indices and sample names. Returns a subset of the NucSeqRecords in the NucMSA as a new NucMSA based on the provided indices and sample names. Note this will work with both positive and negative indices.
Returns a subset of the ProteinSeqRecords in the ProteinMSA as a List, based on the sample IntRange given. Kotlin range operator is "..". Indices start at zero. Note Kotlin IntRange are inclusive end, while Python slices exclusive end. Negative slices "-3..-1" start from the last base (i.e. would return the last three bases).
Function to filter the MSA by sample index based on the provided filterLambda. This will return another ProteinMSA.
Function to filter the ProteinMSA sites using a lambda function. This will return another ProteinMSA object Note this will not work with negative indices
Function to filter down a ProteinMSA at a single site. This will return another ProteinMSA
Function to filter out the ProteinMSA based on a collection of siteIndices This Collection will first be Sorted and Duplicates removed so the resulting ProteinMSA's ProteinSeqs will be in the correct order. Note: This will work with negative indices.
Function to slice the ProteinMSA by siteRange. This will return another ProteinMSA and does support negative indices
Returns a string summary of the ProteinMSA.