Package-level declarations

Read-only sequence object for DNA, RNA, Protein Sequences and Multiple Sequence Alignments.

Sequence

Sequence objects are immutable. This prevents you from doing mySeq[5] = "A" for example, but does allow Seq objects to be used in multi-threaded applications. While String are used to initially create these objects, their sequences are elements of enum classes NUC and AminoAcid. NUC represents the full IUPAC encoding of DNA and RNA. These classes support information on ambiguity, masses, complements, etc.

The Seq object provides a number of string-like methods (such as ProteinSeq.count, NucSeq.indexOf, find, split), which are alphabet aware where appropriate. Please note while the Kotlin "x..y" range operator is supported, and works very similarly to Python's slice "x:y". y is inclusive here in Kotlin, and exclusive in Python.

Unlike BioPython, BioKotlin has two subclasses NucSeq and ProteinSeq that enforces type safe usages (i.e. no adding of DNA + Protein). Note DNA and RNA are stored in the same classes and backed data structure - only the view is different (T for DNA, U or RNA). A DNA sequence can be used to search RNA, and vice versa.

The Seq object provides a few methods, but most of the methods are part of NucSeq and ProteinSeq such as complement, reverse_complement, transcribe, back_transcribe and translate (which are not applicable to sequences with a protein alphabet).

Create a NucSeq object with a string for the sequence, and optional NucSet NUC.DNA,NUC.RNA,NUC.AmbiguousDNA, NUC.AmbiguousRNA.

import biokotlin.seq.*
val dna = NucSeq("GCTA") //inferred DNA
val rna = NucSeq("GCUA") //inferred RNA
val rnaSpecified = NucSeq("GCACCCCC", NUC.RNA)

println(dna.transcribe() == rna) //true
println(dna) //GCTA
println(dna.repr()) //NucSeqByte('GCTA',[A,C,G,T])

A protein sequence is simply:

import biokotlin.seq.*
val protein = ProteinSeq("TGWR")

You will typically use Bio.SeqIO to read in sequences from files as SeqRecord objects, whose sequence will be exposed as a Seq object via the seq property.

MultipleSequenceAlignment

Multiple Sequence Alignments objects are Immutable representations of both Nucleotide and Protein Sequence MSAs.

Most of the filtering functions for both NucMSA and ProteinMSA will return other MSA objects. Both implementations support filtering the MSA for both sites and samples by id and by a provided lambda. Also sample filtering is supported by id.

Both of these classes have a gappedSequence(sampleIndex) and nonGappedSequence(sampleIndex) which return the corresponding Seq objects as loaded or with gaps removed for a given sample index.

To create a simple NucMSA:

import biokotlin.seq.*
val nucRecords = mutableListOf(
NucSeqRecord(NucSeq("AACCACGTTTAA"), id="ID001"),
NucSeqRecord(NucSeq("CACCA-GTGGGT"), id="ID002"),
NucSeqRecord(NucSeq("CACCACGTT-GC"), id="ID003"))

val nucMSA = NucMSA(nucRecords)

val msaFiltered = nucMSA.sites(3 .. 6).samples(setOf(0,2))

val gappedFirstNucSeq = msaFiltered.gappedSequence(0)
val unGappedSecondNucSeq = msaFiltered.nonGappedSequence(1)

Similary to create a simple ProteinMSA:

import biokotlin.seq.*
val proteinRecords = mutableListOf(
ProteinSeqRecord(ProteinSeq("MHQAIFIYQIGYP*LKSGYIQSIRSPEYDNW-"), id="ID001"),
ProteinSeqRecord(ProteinSeq("MH--IFIYQIGYAYLKSGYIQSIRSPEY-NW*"), id="ID002"),
ProteinSeqRecord(ProteinSeq("MHQAIFIYQIGYPYLKSGYIQSIRSPEYDNW*"), id="ID003")
)
val proteinMSA = ProteinMSA(proteinRecords)

val msaFiltered = proteinMSA.samples(setOf("ID001", "ID003")).sites(-7 until -2)

val gappedSecondProteinSeq = msaFiltered.nonGappedSequence(1)
val unGappedFirstProteinSeq = msaFiltered.gappedSequence(0)

Types

Link copied to clipboard

Definition of all amino acids, one char, three letter, and weights

Link copied to clipboard
enum BioSet : Enum<BioSet>
Link copied to clipboard

Immutable multiple sequence alignment object, consisting of two or more SeqRecords with equal lengths. The data can then be regarded as a matrix of letters, with well defined columns.

Link copied to clipboard
enum NUC : Enum<NUC>

Definition of DNA and RNA Nucleotides, IUPAC ambiguity, and nucleotide properties

Link copied to clipboard
class NucMSA(sequences: ImmutableList<NucSeqRecord>) : MultipleSeqAlignment, Collection<NucSeqRecord>

Immutable multiple sequence alignment object, consisting of two or more NucSeqRecords with equal lengths. The data can then be regarded as a matrix of letters, with well defined columns. NucMSA also supports all read-only collection operations on the list of NucSeqRecords.

Link copied to clipboard
interface NucSeq : Seq

Main data structure for working with DNA and RNA sequences

Link copied to clipboard
class NucSeqRecord(val sequence: NucSeq, val id: String, val name: String? = null, val description: String? = null, val annotations: ImmutableMap<String, String> = ImmutableMap.of(), val letterAnnotations: ImmutableMap<String, Array<out Any>> = ImmutableMap.of()) : SeqRecord, NucSeq

A NucSeqRecord consists of a NucSeq and several optional annotations.

Link copied to clipboard
typealias NucSet = ImmutableSet<NUC>

Define nucleotide sets for DNA, RNA, and associated ambiguity sets

Link copied to clipboard

Immutable multiple sequence alignment object, consisting of two or more ProteinSeqRecords with equal lengths. The data can then be regarded as a matrix of letters, with well defined columns. ProteinMSA also supports all read-only collection operations on the list of ProteinSeqRecords.

Link copied to clipboard
interface ProteinSeq : Seq

Main data structure for working with Protein sequences

Link copied to clipboard
class ProteinSeqRecord(val sequence: ProteinSeq, val id: String, val name: String? = null, val description: String? = null, val annotations: ImmutableMap<String, String> = ImmutableMap.of(), val letterAnnotations: ImmutableMap<String, Array<out Any>> = ImmutableMap.of()) : SeqRecord, ProteinSeq

A ProteinSeqRecord consists of a ProteinSeq and several optional annotations.

Link copied to clipboard
interface Seq

Basic data structure for biological sequences - Nucleotide or Protein

Link copied to clipboard
sealed class SeqRecord

A SeqRecord consists of a Seq and several optional annotations.

Properties

Link copied to clipboard
val atom_weights: <Error class: unknown class>

For Center of Mass Calculation. Taken from http://www.chem.qmul.ac.uk/iupac/AtWt/ & PyMol

Functions

Link copied to clipboard
operator fun NUC.div(nuc: NUC): NUC
Link copied to clipboard
fun NucSeq(vararg seq: String, convertStates: Boolean = true): NucSeq

Preferred method for creating a DNA or RNA sequence.

fun NucSeq(seq: String, preferredNucSet: NucSet, convertStates: Boolean = true): NucSeq

Create a NucSeq with a specified NucSet

Link copied to clipboard
fun NucSeqRecord(sequence: NucSeq, id: String, name: String? = null, description: String? = null, annotations: Map<String, String>, letterAnnotations: ImmutableMap<String, Array<out Any>> = ImmutableMap.of()): NucSeqRecord
Link copied to clipboard
operator fun NUC.plus(nuc: NUC): NucSeq
Link copied to clipboard

Preferred method for creating a Protein sequence

Link copied to clipboard
fun ProteinSeqRecord(sequence: ProteinSeq, id: String, name: String? = null, description: String? = null, annotations: Map<String, String>, letterAnnotations: ImmutableMap<String, Array<out Any>> = ImmutableMap.of()): ProteinSeqRecord
Link copied to clipboard
fun RandomNucSeq(length: Int, nucSet: NucSet = NUC.DNA, seed: Int = 0): NucSeq

Creates a random NucSeq of the specified length

Link copied to clipboard
fun RandomProteinSeq(length: Int, seed: Int = 0): ProteinSeq

Creates a random ProteinSeq of the specified length

Link copied to clipboard
fun Seq(seq: String): NucSeq

Create a Seq from a String, could be DNA or RNA. This functions provides compatibility with BioPython, but the preferred use is to use either NucSeq or ProteinSeq, as the Seq is less clear. Unlike Biopython Seq will not convert Protein String to Protein - use ProteinSeq