BioKotlin¶
BioKotlin is a high-performance bioinformatics library that brings the power and speed of compiled programming languages to scripting and big data environments.
It supports nucleotide and protein sequence manipulation, fast sequence IO, k-mer analysis, genomic ranges, and GFF feature trees. Because it runs on the JVM, it interoperates with the wider genomics ecosystem: HTSJDK, GATK, BioJava, and TASSEL.
Get started Browse the tutorials
A quick look¶
Sequences are immutable and type-safe. DNA and RNA share one backing store, and the alphabet is inferred from the sequence unless you name it explicitly.
import biokotlin.seq.*
fun main() {
//sampleStart
val dna = NucSeq("GCAGAT") // DNA inferred from the 'T' amino acid
println(dna.complement()) // CGTCTA
println(dna.reverse_complement()) // ATCTGC
println(dna.transcribe()) // GCAGAU
println(dna.translate()) // AD
//sampleEnd
}
Protein sequences are a separate type, so the compiler stops you from, say,
adding DNA to a peptide. The AminoAcid enum carries the properties you
usually have to look up.
import biokotlin.seq.*
fun main() {
//sampleStart
val protein = ProteinSeq("GCAGAT") + ProteinSeq("ARSQRS")
println(protein) // GCAGATARSQRS
println("Gly count: ${protein.count(AminoAcid.G)}")
var mass = 0.0
for (i in 0 until protein.size()) mass += protein[i].weight
println("Mass: $mass daltons")
//sampleEnd
}
Why Kotlin?¶
Kotlin is a high-performance language that runs on the Java Virtual Machine. It is fully interoperable with Java, but its syntax is closer to Python and other functional languages, and it has a number of features designed for scripting and domain-specific languages.
Because Kotlin is compiled, it can be many times faster than scripting languages for the same work - up to two orders of magnitude on the benchmarks we track. BioKotlin also stores DNA with two bits per base pair, which saves four- to eight-fold on RAM with only modest performance losses.
Where BioKotlin can, it copies BioPython's beautifully designed syntax; see the BioPython comparison.
Kotlin for data science¶
With GraalVM making JVM languages and other scripting languages (Python, R) interoperable, Kotlin works with all the most popular data science environments.
- Kotlin in Jupyter supports a notebook environment alongside Python's numpy and R's dplyr and ggplot.
- GraalVM is a polyglot environment supporting FastR, JVM, Python, and C languages.
What's in the library¶
| Package | What it does |
|---|---|
biokotlin.seq |
Immutable DNA, RNA, and protein sequences, sequence records, and multiple sequence alignments |
biokotlin.seqIO |
Fast readers and writers for FASTA, FASTQ, and GVCF |
biokotlin.featureTree |
Parses GFF3 into an immutable gene → transcript → exon/CDS tree |
biokotlin.genome |
Genomic ranges, GFF data frames, and MAF coverage, identity, and GVCF conversion |
biokotlin.kmer |
Two-bit encoded k-mers (up to 32 bp) with counting and set operations |
biokotlin.data |
NCBI genetic code and codon tables |
The API reference documents all of them.
Where to go next¶
- Getting started - install BioKotlin in Jupyter, a Kotlin script, or a Gradle project.
- Tutorials - worked examples, generated from runnable notebooks in the repository.
- Contributing - we welcome new contributors.