Package-level declarations
Read-only sequence object for DNA, RNA, Protein Sequences and Multiple Sequence Alignments.
Sequence
Sequence objects are immutable. This prevents you from doing mySeq[5] = "A" for example, but does allow Seq objects to be used in multi-threaded applications. While String are used to initially create these objects, their sequences are elements of enum classes NUC and AminoAcid. NUC represents the full IUPAC encoding of DNA and RNA. These classes support information on ambiguity, masses, complements, etc.
The Seq object provides a number of string-like methods (such as ProteinSeq.count, NucSeq.indexOf, find, split), which are alphabet aware where appropriate. Please note while the Kotlin "x..y" range operator is supported, and works very similarly to Python's slice "x:y". y is inclusive here in Kotlin, and exclusive in Python.
Unlike BioPython, BioKotlin has two subclasses NucSeq and ProteinSeq that enforces type safe usages (i.e. no adding of DNA + Protein). Note DNA and RNA are stored in the same classes and backed data structure - only the view is different (T for DNA, U or RNA). A DNA sequence can be used to search RNA, and vice versa.
The Seq object provides a few methods, but most of the methods are part of NucSeq and ProteinSeq such as complement, reverse_complement, transcribe, back_transcribe and translate (which are not applicable to sequences with a protein alphabet).
Create a NucSeq object with a string for the sequence, and optional NucSet NUC.DNA,NUC.RNA,NUC.AmbiguousDNA, NUC.AmbiguousRNA.
import biokotlin.seq.*
val dna = NucSeq("GCTA") //inferred DNA
val rna = NucSeq("GCUA") //inferred RNA
val rnaSpecified = NucSeq("GCACCCCC", NUC.RNA)
println(dna.transcribe() == rna) //true
println(dna) //GCTA
println(dna.repr()) //NucSeqByte('GCTA',[A,C,G,T])A protein sequence is simply:
import biokotlin.seq.*
val protein = ProteinSeq("TGWR")You will typically use Bio.SeqIO to read in sequences from files as SeqRecord objects, whose sequence will be exposed as a Seq object via the seq property.
MultipleSequenceAlignment
Multiple Sequence Alignments objects are Immutable representations of both Nucleotide and Protein Sequence MSAs.
Most of the filtering functions for both NucMSA and ProteinMSA will return other MSA objects. Both implementations support filtering the MSA for both sites and samples by id and by a provided lambda. Also sample filtering is supported by id.
Both of these classes have a gappedSequence(sampleIndex) and nonGappedSequence(sampleIndex) which return the corresponding Seq objects as loaded or with gaps removed for a given sample index.
To create a simple NucMSA:
import biokotlin.seq.*
val nucRecords = mutableListOf(
NucSeqRecord(NucSeq("AACCACGTTTAA"), id="ID001"),
NucSeqRecord(NucSeq("CACCA-GTGGGT"), id="ID002"),
NucSeqRecord(NucSeq("CACCACGTT-GC"), id="ID003"))
val nucMSA = NucMSA(nucRecords)
val msaFiltered = nucMSA.sites(3 .. 6).samples(setOf(0,2))
val gappedFirstNucSeq = msaFiltered.gappedSequence(0)
val unGappedSecondNucSeq = msaFiltered.nonGappedSequence(1)Similary to create a simple ProteinMSA:
import biokotlin.seq.*
val proteinRecords = mutableListOf(
ProteinSeqRecord(ProteinSeq("MHQAIFIYQIGYP*LKSGYIQSIRSPEYDNW-"), id="ID001"),
ProteinSeqRecord(ProteinSeq("MH--IFIYQIGYAYLKSGYIQSIRSPEY-NW*"), id="ID002"),
ProteinSeqRecord(ProteinSeq("MHQAIFIYQIGYPYLKSGYIQSIRSPEYDNW*"), id="ID003")
)
val proteinMSA = ProteinMSA(proteinRecords)
val msaFiltered = proteinMSA.samples(setOf("ID001", "ID003")).sites(-7 until -2)
val gappedSecondProteinSeq = msaFiltered.nonGappedSequence(1)
val unGappedFirstProteinSeq = msaFiltered.gappedSequence(0)Types
Immutable multiple sequence alignment object, consisting of two or more SeqRecords with equal lengths. The data can then be regarded as a matrix of letters, with well defined columns.
Immutable multiple sequence alignment object, consisting of two or more NucSeqRecords with equal lengths. The data can then be regarded as a matrix of letters, with well defined columns. NucMSA also supports all read-only collection operations on the list of NucSeqRecords.
A NucSeqRecord consists of a NucSeq and several optional annotations.
Immutable multiple sequence alignment object, consisting of two or more ProteinSeqRecords with equal lengths. The data can then be regarded as a matrix of letters, with well defined columns. ProteinMSA also supports all read-only collection operations on the list of ProteinSeqRecords.
Main data structure for working with Protein sequences
A ProteinSeqRecord consists of a ProteinSeq and several optional annotations.
Properties
Functions
Preferred method for creating a Protein sequence
Creates a random ProteinSeq of the specified length
Create a Seq from a String, could be DNA or RNA. This functions provides compatibility with BioPython, but the preferred use is to use either NucSeq or ProteinSeq, as the Seq is less clear. Unlike Biopython Seq will not convert Protein String to Protein - use ProteinSeq