Package-level declarations
Types
Represents some collection of kmers generated from nucleic acid sequence(s).
Represents some AbstractKmerSet where most of the unique kmers of kmerSize are not represented. The maximum number of kmers cannot exceed MAX_VALUE of Int.
Set of distance matrices for samples seqIDs using the following distance measures (all normalized by average length of the sequences). h1Count: Number of kmers that have a max Hamming distance value of 1 hManyCount: Number of kmers that have a max Hamming distance value of 2 or more copyNumberCount: Number of shared kmers that have different copy numbers between the two sequences copyNumberDifference: Of the kmers that appear in both sequences, the total difference in copy number
Two-dimensional array of doubles used to store kmer distance values.
h1KmerCount: number of kmers that had a max Hamming distance value of 1 hManyKmerCount: number of kmers that had a max Hamming distance value of 2 copyNumberCount: of the kmers that appeared in both sets, how many had a different number of counts. copyNumberDifference: of the kmers that appeared in both sets, the total difference in their counts
2-bit encoded representation of a short nucleotide sequence (32 bases or fewer) Encoding: A=0, C=1, T/U=2, G=3 For kmers smaller than 32 bp, leftmost digits are padded with 0 Ambiguous bases (eg. N) are not allowed
A set of kmers capable of storing up to 134217728 * (2^31 − 1) unique kmers. Includes a count value for each kmer, up to 128. Recommended use is for dense tables where most possible kmers of length kmerSize are present or where the number of unique kmers exceeds MAX_VALUE of Int.
This class reads kmers from a text file (which may be LZ4 compressed) and can be used to construct KmerSet, KmerBigSet, and KmerMultiSet objects. It can also iterate through a kmer text file without constructing a set.
A set of kmers with associated counts.
Functions
From the fastaFiles, return a KmerBigSet of kmers of length kmerSize where the value for any kmer key represents the number of samples (files) where that kmer was present. Each file may contain multiple fasta sequences (e.g. chromosomes).
Writes set to file filename. By default, text file is zipped using LZ4 compression algorithm but zip may be set to false to write a plain text file.
Writes set to file filename. By default, text file is zipped using LZ4 compression algorithm but zip may be set to false to write a plain text file. By default, order of kmers is not guaranteed but sortedByEncoding may be set to true to guarantee kmers are sorted numerically by their Long representation.