Package-level declarations

Types

Link copied to clipboard
data class AlignmentBlock(val chromName: String, val start: Int, val size: Int, val strand: String, val chrSize: Int, val alignment: String)

This class takes a UCSC MAF file, a reference fasta, a sample name and an output file name. It creates a gvcf file from the MAF and reference, writing the data to the output file.

Link copied to clipboard
data class AssemblyVariantInfo(var chr: String, var startPos: Int, var endPos: Int, var genotype: String, var refAllele: String, var altAllele: String, var isVariant: Boolean, var alleleDepths: IntArray = intArrayOf(), var asmChrom: String = "", var asmStart: Int = -1, var asmEnd: Int = -1, var asmStrand: String = "", var isMissing: Boolean = false)
Link copied to clipboard
data class BKSamRecord(val queryName: String, val flag: Int, val queryLength: Int, val strand: String, val targetName: String, val targetStart: Int, val targetEnd: Int, val mapQ: Int, val NM: Int, val numM: Int, val numEQ: Int, val numX: Int, val numI: Int, val numD: Int, val numH: Int, val numS: Int, val sequence: String) : SAMDataFrame

A BioKotlin SAM Record, which is the data class used SAMDataFrame

Link copied to clipboard
data class ChromStats(val contig: String, val numRegionBPs: Int, val percentCov: Double, val percentId: Double)

This file holds methods used to process MAF Files with the intent of creating Wiggle or BED formatted files for IGV viewing

Link copied to clipboard
class GenomicFeatures(val gffFile: String, val refFasta: String? = null)

The GenomicFeatures class processes data from a GFF formatted file. Internally it stores the GFF features in Kotlin dataframes. THis allows quicker access than using an internal database. The Kotlin dataframe object also allows for filtering based on columns, metrics for columns and rows.

Link copied to clipboard

This class has methods which create coverage and identity counts for a list of mafFiles.

Link copied to clipboard
data class MAFRecord(val score: Double, val refRecord: AlignmentBlock, val altRecord: AlignmentBlock)
Link copied to clipboard
class MAFToGVCF
Link copied to clipboard
Link copied to clipboard
data class Position(val contig: String, val position: Int) : Comparable<Position>

Data class to represent a position on a contig. Position is 1-based

Link copied to clipboard
data class PositionRange(val contig: String, val start: Int, val end: Int) : Comparable<PositionRange>
Link copied to clipboard
interface SAMDataFrame

SAM or BAM file parsed into a dataframe with the CIGAR string summarized into counts

Link copied to clipboard
data class SampleGamete(val name: String, val gameteId: Int = 0) : Comparable<SampleGamete>

A SampleGamete is a combination of a sample name and a gamete id.

Link copied to clipboard
data class SeqPosition(val seqRecord: SeqRecord?, val site: Int) : Comparable<SeqPosition>

Class defines SeqPosition as an optional seqRecord and a site. All sites must be non-zero positive. Sites are physical positions. The SeqPosition class is used when defining an SRange.

Link copied to clipboard
Link copied to clipboard

Factory to create comparators for SeqPositionRanges - comparing on SeqRecord and Ranges More types can be added based on biologists specifications for sorting - these are just a few examples.

Link copied to clipboard

Sorting for the SeqRecord object

Link copied to clipboard
Link copied to clipboard
data class SRangeDataRow(val ID: String, val start: Int, val end: Int, val range: IntRange)

Transform the set of SRanges into a DataFrame with the SeqRange ID, start, end and IntRange columns

Link copied to clipboard
typealias SRangeSet = Set<SRange>

Properties

Link copied to clipboard
Link copied to clipboard
@get:JvmName(name = "geneDataRow_biotype")
val ColumnsContainer<GenomicFeatures.geneDataRow>.biotype: DataColumn<String>
@get:JvmName(name = "transcriptDataRow_biotype")
val ColumnsContainer<GenomicFeatures.transcriptDataRow>.biotype: DataColumn<String>
@get:JvmName(name = "geneDataRow_biotype")
val DataRow<GenomicFeatures.geneDataRow>.biotype: String
@get:JvmName(name = "transcriptDataRow_biotype")
val DataRow<GenomicFeatures.transcriptDataRow>.biotype: String
Link copied to clipboard
@get:JvmName(name = "featureRangeDataRow_data")
val ColumnsContainer<GenomicFeatures.featureRangeDataRow>.data: DataColumn<String>
@get:JvmName(name = "featureRangeDataRow_data")
val DataRow<GenomicFeatures.featureRangeDataRow>.data: String
Link copied to clipboard
@get:JvmName(name = "cdsDataRow_end")
val ColumnsContainer<GenomicFeatures.cdsDataRow>.end: DataColumn<Int>
@get:JvmName(name = "exonDataRow_end")
val ColumnsContainer<GenomicFeatures.exonDataRow>.end: DataColumn<Int>
@get:JvmName(name = "featureRangeDataRow_end")
val ColumnsContainer<GenomicFeatures.featureRangeDataRow>.end: DataColumn<Int>
@get:JvmName(name = "fivePrimeDataRow_end")
val ColumnsContainer<GenomicFeatures.fivePrimeDataRow>.end: DataColumn<Int>
@get:JvmName(name = "geneDataRow_end")
val ColumnsContainer<GenomicFeatures.geneDataRow>.end: DataColumn<Int>
@get:JvmName(name = "threePrimeDataRow_end")
val ColumnsContainer<GenomicFeatures.threePrimeDataRow>.end: DataColumn<Int>
@get:JvmName(name = "transcriptDataRow_end")
val ColumnsContainer<GenomicFeatures.transcriptDataRow>.end: DataColumn<Int>
@get:JvmName(name = "cdsDataRow_end")
val DataRow<GenomicFeatures.cdsDataRow>.end: Int
@get:JvmName(name = "exonDataRow_end")
val DataRow<GenomicFeatures.exonDataRow>.end: Int
@get:JvmName(name = "featureRangeDataRow_end")
val DataRow<GenomicFeatures.featureRangeDataRow>.end: Int
@get:JvmName(name = "fivePrimeDataRow_end")
val DataRow<GenomicFeatures.fivePrimeDataRow>.end: Int
@get:JvmName(name = "geneDataRow_end")
val DataRow<GenomicFeatures.geneDataRow>.end: Int
@get:JvmName(name = "threePrimeDataRow_end")
val DataRow<GenomicFeatures.threePrimeDataRow>.end: Int
@get:JvmName(name = "transcriptDataRow_end")
val DataRow<GenomicFeatures.transcriptDataRow>.end: Int
Link copied to clipboard
@get:JvmName(name = "SAMDataFrame_flag")
val ColumnsContainer<SAMDataFrame>.flag: DataColumn<Int>
@get:JvmName(name = "SAMDataFrame_flag")
val DataRow<SAMDataFrame>.flag: Int
Link copied to clipboard
@get:JvmName(name = "chromDataRow_length")
val ColumnsContainer<GenomicFeatures.chromDataRow>.length: DataColumn<Int>
@get:JvmName(name = "chromDataRow_length")
val DataRow<GenomicFeatures.chromDataRow>.length: Int
Link copied to clipboard
@get:JvmName(name = "geneDataRow_logic_name")
val ColumnsContainer<GenomicFeatures.geneDataRow>.logic_name: DataColumn<String>
@get:JvmName(name = "geneDataRow_logic_name")
val DataRow<GenomicFeatures.geneDataRow>.logic_name: String
Link copied to clipboard
@get:JvmName(name = "SAMDataFrame_mapQ")
val ColumnsContainer<SAMDataFrame>.mapQ: DataColumn<Int>
@get:JvmName(name = "SAMDataFrame_mapQ")
val DataRow<SAMDataFrame>.mapQ: Int
Link copied to clipboard
Link copied to clipboard
@get:JvmName(name = "cdsDataRow_name")
val ColumnsContainer<GenomicFeatures.cdsDataRow>.name: DataColumn<String>
@get:JvmName(name = "exonDataRow_name")
val ColumnsContainer<GenomicFeatures.exonDataRow>.name: DataColumn<String>
@get:JvmName(name = "geneDataRow_name")
val ColumnsContainer<GenomicFeatures.geneDataRow>.name: DataColumn<String>
@get:JvmName(name = "transcriptDataRow_name")
val ColumnsContainer<GenomicFeatures.transcriptDataRow>.name: DataColumn<String>
@get:JvmName(name = "cdsDataRow_name")
val DataRow<GenomicFeatures.cdsDataRow>.name: String
@get:JvmName(name = "exonDataRow_name")
val DataRow<GenomicFeatures.exonDataRow>.name: String
@get:JvmName(name = "geneDataRow_name")
val DataRow<GenomicFeatures.geneDataRow>.name: String
@get:JvmName(name = "transcriptDataRow_name")
val DataRow<GenomicFeatures.transcriptDataRow>.name: String
Link copied to clipboard
@get:JvmName(name = "SAMDataFrame_NM")
val ColumnsContainer<SAMDataFrame>.NM: DataColumn<Int>
@get:JvmName(name = "SAMDataFrame_NM")
val DataRow<SAMDataFrame>.NM: Int
Link copied to clipboard
@get:JvmName(name = "SAMDataFrame_numD")
val ColumnsContainer<SAMDataFrame>.numD: DataColumn<Int>
@get:JvmName(name = "SAMDataFrame_numD")
val DataRow<SAMDataFrame>.numD: Int
Link copied to clipboard
@get:JvmName(name = "SAMDataFrame_numEQ")
val ColumnsContainer<SAMDataFrame>.numEQ: DataColumn<Int>
@get:JvmName(name = "SAMDataFrame_numEQ")
val DataRow<SAMDataFrame>.numEQ: Int
Link copied to clipboard
@get:JvmName(name = "SAMDataFrame_numH")
val ColumnsContainer<SAMDataFrame>.numH: DataColumn<Int>
@get:JvmName(name = "SAMDataFrame_numH")
val DataRow<SAMDataFrame>.numH: Int
Link copied to clipboard
@get:JvmName(name = "SAMDataFrame_numI")
val ColumnsContainer<SAMDataFrame>.numI: DataColumn<Int>
@get:JvmName(name = "SAMDataFrame_numI")
val DataRow<SAMDataFrame>.numI: Int
Link copied to clipboard
@get:JvmName(name = "SAMDataFrame_numM")
val ColumnsContainer<SAMDataFrame>.numM: DataColumn<Int>
@get:JvmName(name = "SAMDataFrame_numM")
val DataRow<SAMDataFrame>.numM: Int
Link copied to clipboard
@get:JvmName(name = "SAMDataFrame_numS")
val ColumnsContainer<SAMDataFrame>.numS: DataColumn<Int>
@get:JvmName(name = "SAMDataFrame_numS")
val DataRow<SAMDataFrame>.numS: Int
Link copied to clipboard
@get:JvmName(name = "SAMDataFrame_numX")
val ColumnsContainer<SAMDataFrame>.numX: DataColumn<Int>
@get:JvmName(name = "SAMDataFrame_numX")
val DataRow<SAMDataFrame>.numX: Int
Link copied to clipboard
@get:JvmName(name = "cdsDataRow_phase")
val ColumnsContainer<GenomicFeatures.cdsDataRow>.phase: DataColumn<Int>
@get:JvmName(name = "cdsDataRow_phase")
val DataRow<GenomicFeatures.cdsDataRow>.phase: Int
Link copied to clipboard
@get:JvmName(name = "SAMDataFrame_queryLength")
val ColumnsContainer<SAMDataFrame>.queryLength: DataColumn<Int>
@get:JvmName(name = "SAMDataFrame_queryLength")
val DataRow<SAMDataFrame>.queryLength: Int
Link copied to clipboard
@get:JvmName(name = "SAMDataFrame_queryName")
val ColumnsContainer<SAMDataFrame>.queryName: DataColumn<String>
@get:JvmName(name = "SAMDataFrame_queryName")
val DataRow<SAMDataFrame>.queryName: String
Link copied to clipboard
@get:JvmName(name = "exonDataRow_rank")
val ColumnsContainer<GenomicFeatures.exonDataRow>.rank: DataColumn<Int>
@get:JvmName(name = "exonDataRow_rank")
val DataRow<GenomicFeatures.exonDataRow>.rank: Int
Link copied to clipboard
Link copied to clipboard
@get:JvmName(name = "cdsDataRow_seqid")
val ColumnsContainer<GenomicFeatures.cdsDataRow>.seqid: DataColumn<String>
@get:JvmName(name = "chromDataRow_seqid")
val ColumnsContainer<GenomicFeatures.chromDataRow>.seqid: DataColumn<String>
@get:JvmName(name = "exonDataRow_seqid")
val ColumnsContainer<GenomicFeatures.exonDataRow>.seqid: DataColumn<String>
@get:JvmName(name = "featureRangeDataRow_seqid")
val ColumnsContainer<GenomicFeatures.featureRangeDataRow>.seqid: DataColumn<String>
@get:JvmName(name = "fivePrimeDataRow_seqid")
val ColumnsContainer<GenomicFeatures.fivePrimeDataRow>.seqid: DataColumn<String>
@get:JvmName(name = "geneDataRow_seqid")
val ColumnsContainer<GenomicFeatures.geneDataRow>.seqid: DataColumn<String>
@get:JvmName(name = "threePrimeDataRow_seqid")
val ColumnsContainer<GenomicFeatures.threePrimeDataRow>.seqid: DataColumn<String>
@get:JvmName(name = "transcriptDataRow_seqid")
val ColumnsContainer<GenomicFeatures.transcriptDataRow>.seqid: DataColumn<String>
@get:JvmName(name = "cdsDataRow_seqid")
val DataRow<GenomicFeatures.cdsDataRow>.seqid: String
@get:JvmName(name = "chromDataRow_seqid")
val DataRow<GenomicFeatures.chromDataRow>.seqid: String
@get:JvmName(name = "exonDataRow_seqid")
val DataRow<GenomicFeatures.exonDataRow>.seqid: String
@get:JvmName(name = "featureRangeDataRow_seqid")
val DataRow<GenomicFeatures.featureRangeDataRow>.seqid: String
@get:JvmName(name = "fivePrimeDataRow_seqid")
val DataRow<GenomicFeatures.fivePrimeDataRow>.seqid: String
@get:JvmName(name = "geneDataRow_seqid")
val DataRow<GenomicFeatures.geneDataRow>.seqid: String
@get:JvmName(name = "threePrimeDataRow_seqid")
val DataRow<GenomicFeatures.threePrimeDataRow>.seqid: String
@get:JvmName(name = "transcriptDataRow_seqid")
val DataRow<GenomicFeatures.transcriptDataRow>.seqid: String
Link copied to clipboard
@get:JvmName(name = "SAMDataFrame_sequence")
val ColumnsContainer<SAMDataFrame>.sequence: DataColumn<String>
@get:JvmName(name = "SAMDataFrame_sequence")
val DataRow<SAMDataFrame>.sequence: String
Link copied to clipboard
@get:JvmName(name = "cdsDataRow_start")
val ColumnsContainer<GenomicFeatures.cdsDataRow>.start: DataColumn<Int>
@get:JvmName(name = "exonDataRow_start")
val ColumnsContainer<GenomicFeatures.exonDataRow>.start: DataColumn<Int>
@get:JvmName(name = "featureRangeDataRow_start")
val ColumnsContainer<GenomicFeatures.featureRangeDataRow>.start: DataColumn<Int>
@get:JvmName(name = "fivePrimeDataRow_start")
val ColumnsContainer<GenomicFeatures.fivePrimeDataRow>.start: DataColumn<Int>
@get:JvmName(name = "geneDataRow_start")
val ColumnsContainer<GenomicFeatures.geneDataRow>.start: DataColumn<Int>
@get:JvmName(name = "threePrimeDataRow_start")
val ColumnsContainer<GenomicFeatures.threePrimeDataRow>.start: DataColumn<Int>
@get:JvmName(name = "transcriptDataRow_start")
val ColumnsContainer<GenomicFeatures.transcriptDataRow>.start: DataColumn<Int>
@get:JvmName(name = "cdsDataRow_start")
val DataRow<GenomicFeatures.cdsDataRow>.start: Int
@get:JvmName(name = "exonDataRow_start")
val DataRow<GenomicFeatures.exonDataRow>.start: Int
@get:JvmName(name = "featureRangeDataRow_start")
val DataRow<GenomicFeatures.featureRangeDataRow>.start: Int
@get:JvmName(name = "fivePrimeDataRow_start")
val DataRow<GenomicFeatures.fivePrimeDataRow>.start: Int
@get:JvmName(name = "geneDataRow_start")
val DataRow<GenomicFeatures.geneDataRow>.start: Int
@get:JvmName(name = "threePrimeDataRow_start")
val DataRow<GenomicFeatures.threePrimeDataRow>.start: Int
@get:JvmName(name = "transcriptDataRow_start")
val DataRow<GenomicFeatures.transcriptDataRow>.start: Int
Link copied to clipboard
@get:JvmName(name = "cdsDataRow_strand")
val ColumnsContainer<GenomicFeatures.cdsDataRow>.strand: DataColumn<String>
@get:JvmName(name = "exonDataRow_strand")
val ColumnsContainer<GenomicFeatures.exonDataRow>.strand: DataColumn<String>
@get:JvmName(name = "featureRangeDataRow_strand")
val ColumnsContainer<GenomicFeatures.featureRangeDataRow>.strand: DataColumn<String>
@get:JvmName(name = "fivePrimeDataRow_strand")
val ColumnsContainer<GenomicFeatures.fivePrimeDataRow>.strand: DataColumn<String>
@get:JvmName(name = "geneDataRow_strand")
val ColumnsContainer<GenomicFeatures.geneDataRow>.strand: DataColumn<String>
@get:JvmName(name = "threePrimeDataRow_strand")
val ColumnsContainer<GenomicFeatures.threePrimeDataRow>.strand: DataColumn<String>
@get:JvmName(name = "transcriptDataRow_strand")
val ColumnsContainer<GenomicFeatures.transcriptDataRow>.strand: DataColumn<String>
@get:JvmName(name = "SAMDataFrame_strand")
val ColumnsContainer<SAMDataFrame>.strand: DataColumn<String>
@get:JvmName(name = "cdsDataRow_strand")
val DataRow<GenomicFeatures.cdsDataRow>.strand: String
@get:JvmName(name = "exonDataRow_strand")
val DataRow<GenomicFeatures.exonDataRow>.strand: String
@get:JvmName(name = "featureRangeDataRow_strand")
val DataRow<GenomicFeatures.featureRangeDataRow>.strand: String
@get:JvmName(name = "fivePrimeDataRow_strand")
val DataRow<GenomicFeatures.fivePrimeDataRow>.strand: String
@get:JvmName(name = "geneDataRow_strand")
val DataRow<GenomicFeatures.geneDataRow>.strand: String
@get:JvmName(name = "threePrimeDataRow_strand")
val DataRow<GenomicFeatures.threePrimeDataRow>.strand: String
@get:JvmName(name = "transcriptDataRow_strand")
val DataRow<GenomicFeatures.transcriptDataRow>.strand: String
@get:JvmName(name = "SAMDataFrame_strand")
val DataRow<SAMDataFrame>.strand: String
Link copied to clipboard
@get:JvmName(name = "SAMDataFrame_targetEnd")
val ColumnsContainer<SAMDataFrame>.targetEnd: DataColumn<Int>
@get:JvmName(name = "SAMDataFrame_targetEnd")
val DataRow<SAMDataFrame>.targetEnd: Int
Link copied to clipboard
@get:JvmName(name = "SAMDataFrame_targetName")
val ColumnsContainer<SAMDataFrame>.targetName: DataColumn<String>
@get:JvmName(name = "SAMDataFrame_targetName")
val DataRow<SAMDataFrame>.targetName: String
Link copied to clipboard
@get:JvmName(name = "SAMDataFrame_targetStart")
val ColumnsContainer<SAMDataFrame>.targetStart: DataColumn<Int>
@get:JvmName(name = "SAMDataFrame_targetStart")
val DataRow<SAMDataFrame>.targetStart: Int
Link copied to clipboard
@get:JvmName(name = "cdsDataRow_transcript")
val ColumnsContainer<GenomicFeatures.cdsDataRow>.transcript: DataColumn<String>
@get:JvmName(name = "exonDataRow_transcript")
val ColumnsContainer<GenomicFeatures.exonDataRow>.transcript: DataColumn<String>
@get:JvmName(name = "fivePrimeDataRow_transcript")
val ColumnsContainer<GenomicFeatures.fivePrimeDataRow>.transcript: DataColumn<String>
@get:JvmName(name = "threePrimeDataRow_transcript")
val ColumnsContainer<GenomicFeatures.threePrimeDataRow>.transcript: DataColumn<String>
@get:JvmName(name = "cdsDataRow_transcript")
val DataRow<GenomicFeatures.cdsDataRow>.transcript: String
@get:JvmName(name = "exonDataRow_transcript")
val DataRow<GenomicFeatures.exonDataRow>.transcript: String
@get:JvmName(name = "fivePrimeDataRow_transcript")
val DataRow<GenomicFeatures.fivePrimeDataRow>.transcript: String
@get:JvmName(name = "threePrimeDataRow_transcript")
val DataRow<GenomicFeatures.threePrimeDataRow>.transcript: String
Link copied to clipboard
@get:JvmName(name = "featureRangeDataRow_type")
val ColumnsContainer<GenomicFeatures.featureRangeDataRow>.type: DataColumn<String>
@get:JvmName(name = "featureRangeDataRow_type")
val DataRow<GenomicFeatures.featureRangeDataRow>.type: String

Functions

Link copied to clipboard

target is a mutable list of maf records that are assumed to be sorted by chromosome and start. If there are reference positions present in source that are absent from target then the sequence for those positions will be added to target. MAF start positions are 0-based numbers

Link copied to clipboard
Link copied to clipboard
fun bedfileToSRangeSet(bedfile: String, fasta: String): SRangeSet

Takes a genome fasta and a bedFile of ranges, creates a set of SRanges

Link copied to clipboard
fun calculateCoverageAndIdentity(alignments: List<String>, coverageCnt: IntArray, identityCnt: IntArray, startStop: ClosedRange<Int>)

This function will walk the ref and query sequences, checking for bp coverage, and bp matches. The counts at each overlapping position in the coverageArray and identityArray will be updated The alignments pulled are those that overlap the user requested coordinates so it is necessary to calculate the start and end positions on the alignment that should be processed. This function processes a single alignment block. It may have multiple query sequences, but they are all to the same contig and coordinates on that contig

Link copied to clipboard
fun coalescingSetOf(comparator: Comparator<SRange> = SeqRangeSort.by(SeqRangeSort.numberThenAlphaSort,leftEdge), ranges: List<SRange>): SRangeSet

Coalescing Sets: When added to set, ranges that overlap or are embedded will be merged. This does not merge adjacent ranges (ie, 14..29 and 30..35 are not merged, but 14..29 and 29..31 are merged) User must supply a comparator - either their own or one defined in SeqRangeSort

Link copied to clipboard
fun coalescingsetOf(comparator: Comparator<SRange> = SeqRangeSort.by(SeqRangeSort.numberThenAlphaSort,leftEdge), vararg ranges: SRange): SRangeSet

Coalescing Sets: When added to set, ranges that overlap or are embedded will be merged. This does not merge adjacent ranges (ie, 14..29 and 30..35 are not merged, but 14..29 and 29..31 are merged) User must supply a comparator - either their own or one defined in SeqRangeSort

Link copied to clipboard
fun compareId(first: String?, second: String?, numFirst: Boolean): Int
Link copied to clipboard
fun complement(boundaryRange: SRange, intersectingRanges: Set<SRange>): SRangeSet

This method takes an SRange and a set of ranges that intersect it. It splits the "boundaryRange" into a set of ranges, none of which contain positions that overlap positions from the "intersectingRanges" list.

Link copied to clipboard
fun SRangeSet.complement(boundaryRange: SRange): SRangeSet

Find the complement for a set of ranges. The boundaryRange parameter sets the upper and lower limits for the complemented ranges.

Link copied to clipboard

function to bgzip and create tabix index on a gvcf file

Link copied to clipboard
fun convertSAMToDataFrame(inputFile: String, includeSequence: Boolean = true, filter: BKSamRecord.() -> Boolean = { true }): DataFrame<SAMDataFrame>

Convert a SAM or BAM file and convert to a in memory DataFrame, where the CIGAR strings are parsed into counts. This can produce a very large file if there is lots of sequences. To keep things managable, consider setting includeSequence to false or filter the reads. e.g. convertSAMToDataFrame("path/myFile.sam",false){NM<2 && numM 100} - would not include the sequence, and retain only sequences with less than two mismatches and matched for more than 100 bp.

Link copied to clipboard
fun createBedFileFromCoverageIdentity(coverage: IntArray, identity: IntArray, contig: String, refStart: Int, minCoverage: Int, minIdentity: Int, outputFile: String)
Link copied to clipboard
fun createShuffledSubRangeList(targetLen: Int, ranges: Set<SRange>): List<SRange>

The function creates a list of sub-ranges from a given set of SRanges. It takes a set or SRanges, creates subsets of length "targetLen". These are created using a window of 1. This method does NOT change the sequence in the record. It assumes the ranges represent an interval on the sequence. It then shuffles the list prior to returning. This method is used by applications (e.g. pairedInterval() to find negative peaks) which want a list that doesn't prioritize a specific section of a range.

Link copied to clipboard
fun createWiggleFilesFromCoverageIdentity(coverage: IntArray, identity: IntArray, contig: String, outputDir: String)
Link copied to clipboard
Link copied to clipboard

This function extracts a subblock from a MAF record. The subblock is defined by the indices passed as paramters. The indices are the start and end indices of the subblock in the alignment.

Link copied to clipboard
fun extractSubMafRecord(start: Int, end: Int, mafRecord: MAFRecord): MAFRecord?

This function takes a MafRecord, a start and an end position. From the MafRecord it extracts the part of the ref and alt blocks that fall within the range start,end. It returns a new MafRecord. If the MafRecord does not fall within the range start,end then it returns null.

Link copied to clipboard

Takes an input fasta, returns a Map of where NucSeq is the Biokotlin data structure for DNA and RNA sequences.

Link copied to clipboard
fun findGaps(target: List<MAFRecord>, source: List<MAFRecord>): List<Range<Int>>?

Method takes 2 lists of Maf Records, returns the gaps indicating which positions appear in the source that are not in the target. The calling method then augments the target with sequence from the gaps indicated by the findGaps method.

Link copied to clipboard

Given 2 sets of SRanges, return an SRangeSet of the intersecting positions from the 2 input SRange Sets.

Link copied to clipboard
fun findIntersectingSRanges(query: SRange, searchSpace: Set<SRange>): SRangeSet

Given a single SRange, find peaks that overlap For intersecting ranges, you can stop after a while if you have a sorted set and only want those that are intersecting.

Link copied to clipboard
fun findNegativePeaks(positive: NucSeq, rangeList: List<SRange>, pairingFunc: (NucSeq, NucSeq) -> Boolean, count: Int): Set<SRange>

Find peaks with criteria matching that specified by the user-supplied pairing function. This method assumes the sequence in the NucSeqRecord hasn't changed from the original, and the ranges indicate which subsection of the sequence to pull.

Link copied to clipboard
fun findPair(positive: NucSeq, negativeSpace: Set<SRange>, pairingFunc: (NucSeq, NucSeq) -> Boolean, count: Int = 1): Set<SRange>
fun findPair(positive: SRange, negativeSpace: Set<SRange>, pairingFunc: (NucSeq, NucSeq) -> Boolean, count: Int = 1): Set<SRange>
Link copied to clipboard
Link copied to clipboard
fun SRange.flankBoth(count: Int): Set<SRange>

Flank both ends of the range by the specified amount The lower end (left edge) will not drop below 1. The upper end (right edge) will not exceed max size of sequence in the SeqRecord

Link copied to clipboard

subtract "count" bps from the left (lower) end of each range

fun SRange.flankLeft(count: Int): SRange?

Flank the lower end of the range if it isn't already at 1

Link copied to clipboard

Add "count" bps to the right (upper) end of each range

fun SRange.flankRight(count: Int): SRange?

Flank the upper end of the range if it isn't already at max

Link copied to clipboard
fun getCoverageIdentityPercentForMAF(mafFile: String, region: String = "all"): DataFrame<ChromStats>?

The getCoverageIdentityPercentForMAF takes a single UCSC MAF formatted file and for each contig represented, calculates the coverage and identity percentages as relates to the REF aligned against.

Link copied to clipboard

This function splits a MAF file into a Set of Maf lines, where each entry in the set is a list representing an individual maf block. It is intended for programs that want to process each MAF block in some manner, perhaps looking for data from the "e", "i" or "q" lines. No lines in the maf block are filtered.

Link copied to clipboard

This function takes 2 sets of Kotlin IntRange and returns a Set of overlapping positions. It is only called internally from findIntersectingPosition().

Link copied to clipboard
fun indexOfNonGapCharacters(seq: String, start: Int, end: Int): IntArray

This function returns the indices of the start and end of the non-gap characters in a sequence. The gap is defined as a dash '-'. The start and end parameteres determine where in the sequence to make the search for the non-gap chahacters. An IntArray is returned with the start and end indices of the non-gap characters.

Link copied to clipboard

Find the intersecting positions from 2 SRangeSets

Link copied to clipboard

This function identifies the intersecting SRanges, not specific positions on these ranges.

Link copied to clipboard
fun DataFrame<SAMDataFrame>.mapped(): DataFrame<SAMDataFrame>
Link copied to clipboard
fun DataFrame<SAMDataFrame>.mappedProperPair(): DataFrame<SAMDataFrame>
Link copied to clipboard
fun DataFrame<SAMDataFrame>.mateMapped(): DataFrame<SAMDataFrame>
Link copied to clipboard
fun DataFrame<SAMDataFrame>.mateUnmapped(): DataFrame<SAMDataFrame>
Link copied to clipboard
fun SRangeSet.merge(count: Int, comparator: Comparator<SRange> = SeqRangeSort.by(SeqRangeSort.numberThenAlphaSort,leftEdge)): SRangeSet

Merge will merge overlapping and embedded ranges, and other ranges where distance between them is "count" or less base pairss. It will not merge adjacent/non-overlapping ranges A comparator is necessary as we can't merge until the ranges are sorted. SRange is Kotlin Set which is immutable, but not necessarily sorted. Merging of ranges requires that the upper endpoint SeqRecord of the first range matches the lower endpoint SeqRecord of the next range.

Link copied to clipboard
fun mergeWiggleFiles(file1: String, file2: String, contig: String, outputFile: String)
Link copied to clipboard
fun nonCoalescingSetOf(comparator: Comparator<SRange> = SeqRangeSort.by(SeqRangeSort.numberThenAlphaSort,leftEdge), vararg ranges: SRange): SRangeSet

Method takes a comma separated list of SRanges and adds them to a sorted set. Overlapping intervals are NOT coalesced. User must supply a comparator - either their own or one defined in SeqRangeSort

fun nonCoalescingSetOf(comparator: Comparator<SRange> = SeqRangeSort.by(SeqRangeSort.numberThenAlphaSort,leftEdge), ranges: List<SRange>): SRangeSet

Method takes a List and adds them to a sorted set. Overlapping intervals are NOT coalesced. User must supply a comparator - either their own or one defined in SeqRangeSort

Link copied to clipboard
fun DataFrame<SAMDataFrame>.notPrimaryAlignment(): DataFrame<SAMDataFrame>
Link copied to clipboard
fun overlaps(peak: SRange, search: SRange): Boolean

This function uses DeMorgan's law to determine if ranges overlap. Return: boolean indicating if the ranges overlapped.

Link copied to clipboard
fun DataFrame<SAMDataFrame>.paired(): DataFrame<SAMDataFrame>
Link copied to clipboard
fun SRange.pairedInterval(searchSpace: Set<SRange>, pairingFunc: (NucSeq, NucSeq) -> Boolean, count: Int = 1): Set<SRange>
Link copied to clipboard
fun parseIdSite(idSite: String): Pair<String, Int>

separate the SeqRecord:id from the site. This returns a site value minus the commas. This is necessary for .toInt() We could instead return 2 Strings, leaving the commas in place, and let the conversion to Int take place elsewhere.

Link copied to clipboard
Link copied to clipboard
fun DataFrame<SAMDataFrame>.primaryAlignment(): DataFrame<SAMDataFrame>
Link copied to clipboard
fun DataFrame<SAMDataFrame>.propAligned(): DataFrame<SAMDataFrame>
Link copied to clipboard
fun DataFrame<SAMDataFrame>.propAlignIdentical(): DataFrame<SAMDataFrame>
Link copied to clipboard
fun <Error class: unknown class><SRange>.range(siteSort: ClosedRange<Int>): <Error class: unknown class><SRange>
Link copied to clipboard
fun rangesFromPositions(positions: List<Int>): List<Range<Int>>?
Link copied to clipboard

function to read alignment block from the MAF file.

Link copied to clipboard

Return the sequence represented by this SRange. Return null of the SRange does not have a SeqRecord.

Link copied to clipboard

Shift each range in the set by "count" bps. Can be positive or negative number

fun SRange.shift(count: Int): SRange

Shift the given range by "count" positions. If the number is negative, it is a left shift. If the number is positive, shift right This will not shift into an adjacent contig of the genome

Link copied to clipboard
fun srangeIDMatch(peak: SRange, searchSpace: Set<SRange>)
Link copied to clipboard

Find intersections between 2 SRanges. Returns null if they don't intersect, returns the intersecting values if there is an overlap.

Link copied to clipboard
fun SRangeSet.subtract(removeRanges: Set<SRange>): SRangeSet

For a set of SRanges, create new ranges based on removing any portion of the old range which overlaps with any of the ranges in the "removeRanges" set.

fun SRange.subtract(removeRanges: Set<SRange>): SRangeSet

For a given range, create new ranges based on removing any portion of the old range which overlaps with any of the ranges in the "removeRanges" set. Return a new set of ranges

Link copied to clipboard
Link copied to clipboard
fun DataFrame<SAMDataFrame>.unmapped(): DataFrame<SAMDataFrame>