Skip to content

BioKotlin

BioKotlin is a high-performance bioinformatics library that brings the power and speed of compiled programming languages to scripting and big data environments.

It supports nucleotide and protein sequence manipulation, fast sequence IO, k-mer analysis, genomic ranges, and GFF feature trees. Because it runs on the JVM, it interoperates with the wider genomics ecosystem: HTSJDK, GATK, BioJava, and TASSEL.

Get started Browse the tutorials

A quick look

Sequences are immutable and type-safe. DNA and RNA share one backing store, and the alphabet is inferred from the sequence unless you name it explicitly.

import biokotlin.seq.*

fun main() {
    //sampleStart
    val dna = NucSeq("GCAGAT")         // DNA inferred from the 'T' amino acid
    println(dna.complement())          // CGTCTA
    println(dna.reverse_complement())  // ATCTGC
    println(dna.transcribe())          // GCAGAU
    println(dna.translate())           // AD
    //sampleEnd
}

Protein sequences are a separate type, so the compiler stops you from, say, adding DNA to a peptide. The AminoAcid enum carries the properties you usually have to look up.

import biokotlin.seq.*

fun main() {
    //sampleStart
    val protein = ProteinSeq("GCAGAT") + ProteinSeq("ARSQRS")
    println(protein) // GCAGATARSQRS
    println("Gly count: ${protein.count(AminoAcid.G)}")

    var mass = 0.0
    for (i in 0 until protein.size()) mass += protein[i].weight
    println("Mass: $mass daltons")
    //sampleEnd
}

Why Kotlin?

Kotlin is a high-performance language that runs on the Java Virtual Machine. It is fully interoperable with Java, but its syntax is closer to Python and other functional languages, and it has a number of features designed for scripting and domain-specific languages.

Because Kotlin is compiled, it can be many times faster than scripting languages for the same work - up to two orders of magnitude on the benchmarks we track. BioKotlin also stores DNA with two bits per base pair, which saves four- to eight-fold on RAM with only modest performance losses.

Where BioKotlin can, it copies BioPython's beautifully designed syntax; see the BioPython comparison.

Kotlin for data science

With GraalVM making JVM languages and other scripting languages (Python, R) interoperable, Kotlin works with all the most popular data science environments.

  • Kotlin in Jupyter supports a notebook environment alongside Python's numpy and R's dplyr and ggplot.
  • GraalVM is a polyglot environment supporting FastR, JVM, Python, and C languages.

What's in the library

Package What it does
biokotlin.seq Immutable DNA, RNA, and protein sequences, sequence records, and multiple sequence alignments
biokotlin.seqIO Fast readers and writers for FASTA, FASTQ, and GVCF
biokotlin.featureTree Parses GFF3 into an immutable gene → transcript → exon/CDS tree
biokotlin.genome Genomic ranges, GFF data frames, and MAF coverage, identity, and GVCF conversion
biokotlin.kmer Two-bit encoded k-mers (up to 32 bp) with counting and set operations
biokotlin.data NCBI genetic code and codon tables

The API reference documents all of them.

Where to go next

  • Getting started - install BioKotlin in Jupyter, a Kotlin script, or a Gradle project.
  • Tutorials - worked examples, generated from runnable notebooks in the repository.
  • Contributing - we welcome new contributors.