Package-level declarations

Types

Link copied to clipboard
abstract class AbstractKmerSet(val kmerSize: Int, val bothStrands: Boolean, val stepSize: Int, val keepMinOnly: Boolean)

Represents some collection of kmers generated from nucleic acid sequence(s).

Link copied to clipboard
abstract class AbstractSparseKmerSet(val kmerSize: Int, val bothStrands: Boolean, val stepSize: Int, val keepMinOnly: Boolean) : AbstractKmerSet

Represents some AbstractKmerSet where most of the unique kmers of kmerSize are not represented. The maximum number of kmers cannot exceed MAX_VALUE of Int.

Link copied to clipboard
data class DistanceMatrices(val seqIDs: List<String>, val h1Count: DistanceMatrix, val hManyCount: DistanceMatrix, val copyNumberCount: DistanceMatrix, val copyNumberDifference: DistanceMatrix)

Set of distance matrices for samples seqIDs using the following distance measures (all normalized by average length of the sequences). h1Count: Number of kmers that have a max Hamming distance value of 1 hManyCount: Number of kmers that have a max Hamming distance value of 2 or more copyNumberCount: Number of shared kmers that have different copy numbers between the two sequences copyNumberDifference: Of the kmers that appear in both sequences, the total difference in copy number

Link copied to clipboard

Two-dimensional array of doubles used to store kmer distance values.

Link copied to clipboard
data class HammingCounts(val h1KmerCount: Int, val hManyKmerCount: Int, val copyNumberKmerCount: Int, val copyNumberKmerDifference: Int)

h1KmerCount: number of kmers that had a max Hamming distance value of 1 hManyKmerCount: number of kmers that had a max Hamming distance value of 2 copyNumberCount: of the kmers that appeared in both sets, how many had a different number of counts. copyNumberDifference: of the kmers that appeared in both sets, the total difference in their counts

Link copied to clipboard
value class Kmer(val encoding: Long) : Comparable<Kmer>

2-bit encoded representation of a short nucleotide sequence (32 bases or fewer) Encoding: A=0, C=1, T/U=2, G=3 For kmers smaller than 32 bp, leftmost digits are padded with 0 Ambiguous bases (eg. N) are not allowed

Link copied to clipboard
class KmerBigSet(val kmerSize: Int = 21, val bothStrands: Boolean = true, val stepSize: Int = 1, val keepMinOnly: Boolean = false) : AbstractKmerSet

A set of kmers capable of storing up to 134217728 * (2^31 − 1) unique kmers. Includes a count value for each kmer, up to 128. Recommended use is for dense tables where most possible kmers of length kmerSize are present or where the number of unique kmers exceeds MAX_VALUE of Int.

Link copied to clipboard
class KmerIO(filename: String, isCompressed: Boolean = true) : Iterator<Pair<Kmer, Int>>

This class reads kmers from a text file (which may be LZ4 compressed) and can be used to construct KmerSet, KmerBigSet, and KmerMultiSet objects. It can also iterate through a kmer text file without constructing a set.

Link copied to clipboard
class KmerMultiSet(val kmerSize: Int = 21, val bothStrands: Boolean = true, val stepSize: Int = 1, val keepMinOnly: Boolean = false) : AbstractSparseKmerSet

A set of kmers with associated counts.

Link copied to clipboard
class KmerSet(val kmerSize: Int = 21, val bothStrands: Boolean = true, val stepSize: Int = 1, val keepMinOnly: Boolean = false) : AbstractSparseKmerSet

A set of kmers. Does not include any information about counts.

Functions

Link copied to clipboard
fun getGenomicKmerSet(fastaFile: String, kmerSize: Int): KmerMultiSet

Given a file fastaFile, returns a map of all kmers of size kmerSize in the genome. This assumes a "compact" representation (collapses forward and reverse complement) and includes both strands.

Link copied to clipboard

Compares the kmers of two KmerMultiSets, set1 and set2.

fun getHammingCounts(set1: KmerMultiSet, set2: KmerMultiSet, hashmap1: Long2ObjectOpenHashMap<LongOpenHashSet>, hashmap2: Long2ObjectOpenHashMap<LongOpenHashSet>): HammingCounts

Compares the kmers of two KmerMultiSets, set1 and set2. Also takes one hashmap for each set, hashmap1 and hashmap2.

Link copied to clipboard
fun getKmerConservationSet(fastaFiles: List<String>, kmerSize: Int): KmerBigSet

From the fastaFiles, return a KmerBigSet of kmers of length kmerSize where the value for any kmer key represents the number of samples (files) where that kmer was present. Each file may contain multiple fasta sequences (e.g. chromosomes).

Link copied to clipboard
fun Kmer(seq: NucSeq): Kmer

Create a Kmer from a NucSeq Sequence may be fewer than 32 bases long, left will be padded with A's to equal 32 bases

fun Kmer(seq: String): Kmer

Create a Kmer from a string Sequence may be fewer than 32 bases long, left will be padded with A's to equal 32 bases

Link copied to clipboard
fun kmerDistanceMatrix(fastaFile: String, kmerSize: Int = 21): DistanceMatrices

Calculates distance between sequences in fastaFile based on differences in kmer sets. The length of the kmers is kmerSize. Returns distance measures described in DistanceMatrices.

Link copied to clipboard
fun printMatrixToFile(seqIDs: List<String>, matrix: DistanceMatrix, fileName: String)

Prints a square matrix matrix with row and column names seqIDs to file fileName.

Link copied to clipboard
fun writeKmerSet(set: KmerBigSet, filename: String, zip: Boolean = true)

Writes set to file filename. By default, text file is zipped using LZ4 compression algorithm but zip may be set to false to write a plain text file.

fun writeKmerSet(set: KmerMultiSet, filename: String, zip: Boolean = true, sortedByEncoding: Boolean = false)
fun writeKmerSet(set: KmerSet, filename: String, zip: Boolean = true, sortedByEncoding: Boolean = false)

Writes set to file filename. By default, text file is zipped using LZ4 compression algorithm but zip may be set to false to write a plain text file. By default, order of kmers is not guaranteed but sortedByEncoding may be set to true to guarantee kmers are sorted numerically by their Long representation.