Package-level declarations
Types
This class takes a UCSC MAF file, a reference fasta, a sample name and an output file name. It creates a gvcf file from the MAF and reference, writing the data to the output file.
A BioKotlin SAM Record, which is the data class used SAMDataFrame
This file holds methods used to process MAF Files with the intent of creating Wiggle or BED formatted files for IGV viewing
The GenomicFeatures class processes data from a GFF formatted file. Internally it stores the GFF features in Kotlin dataframes. THis allows quicker access than using an internal database. The Kotlin dataframe object also allows for filtering based on columns, metrics for columns and rows.
This class has methods which create coverage and identity counts for a list of mafFiles.
SAM or BAM file parsed into a dataframe with the CIGAR string summarized into counts
A SampleGamete is a combination of a sample name and a gamete id.
Class defines SeqPosition as an optional seqRecord and a site. All sites must be non-zero positive. Sites are physical positions. The SeqPosition class is used when defining an SRange.
Factory to create comparators for SeqPositionRanges - comparing on SeqRecord and Ranges More types can be added based on biologists specifications for sorting - these are just a few examples.
Sorting for the SeqRecord object
Transform the set of SRanges into a DataFrame with the SeqRange ID, start, end and IntRange columns
Properties
Functions
target is a mutable list of maf records that are assumed to be sorted by chromosome and start. If there are reference positions present in source that are absent from target then the sequence for those positions will be added to target. MAF start positions are 0-based numbers
Takes a genome fasta and a bedFile of ranges, creates a set of SRanges
This function will walk the ref and query sequences, checking for bp coverage, and bp matches. The counts at each overlapping position in the coverageArray and identityArray will be updated The alignments pulled are those that overlap the user requested coordinates so it is necessary to calculate the start and end positions on the alignment that should be processed. This function processes a single alignment block. It may have multiple query sequences, but they are all to the same contig and coordinates on that contig
Coalescing Sets: When added to set, ranges that overlap or are embedded will be merged. This does not merge adjacent ranges (ie, 14..29 and 30..35 are not merged, but 14..29 and 29..31 are merged) User must supply a comparator - either their own or one defined in SeqRangeSort
Coalescing Sets: When added to set, ranges that overlap or are embedded will be merged. This does not merge adjacent ranges (ie, 14..29 and 30..35 are not merged, but 14..29 and 29..31 are merged) User must supply a comparator - either their own or one defined in SeqRangeSort
This method takes an SRange and a set of ranges that intersect it. It splits the "boundaryRange" into a set of ranges, none of which contain positions that overlap positions from the "intersectingRanges" list.
Find the complement for a set of ranges. The boundaryRange parameter sets the upper and lower limits for the complemented ranges.
function to bgzip and create tabix index on a gvcf file
Convert a SAM or BAM file and convert to a in memory DataFrame, where the CIGAR strings are parsed into counts. This can produce a very large file if there is lots of sequences. To keep things managable, consider setting includeSequence to false or filter the reads. e.g. convertSAMToDataFrame("path/myFile.sam",false){NM<2 && numM 100} - would not include the sequence, and retain only sequences with less than two mismatches and matched for more than 100 bp.
The function creates a list of sub-ranges from a given set of SRanges. It takes a set or SRanges, creates subsets of length "targetLen". These are created using a window of 1. This method does NOT change the sequence in the record. It assumes the ranges represent an interval on the sequence. It then shuffles the list prior to returning. This method is used by applications (e.g. pairedInterval() to find negative peaks) which want a list that doesn't prioritize a specific section of a range.
This function extracts a subblock from a MAF record. The subblock is defined by the indices passed as paramters. The indices are the start and end indices of the subblock in the alignment.
This function takes a MafRecord, a start and an end position. From the MafRecord it extracts the part of the ref and alt blocks that fall within the range start,end. It returns a new MafRecord. If the MafRecord does not fall within the range start,end then it returns null.
Takes an input fasta, returns a Map of
Method takes 2 lists of Maf Records, returns the gaps indicating which positions appear in the source that are not in the target. The calling method then augments the target with sequence from the gaps indicated by the findGaps method.
Given 2 sets of SRanges, return an SRangeSet of the intersecting positions from the 2 input SRange Sets.
Given a single SRange, find peaks that overlap For intersecting ranges, you can stop after a while if you have a sorted set and only want those that are intersecting.
Find peaks with criteria matching that specified by the user-supplied pairing function. This method assumes the sequence in the NucSeqRecord hasn't changed from the original, and the ranges indicate which subsection of the sequence to pull.
Add "count" bps to the right (upper) end of each range
Flank the upper end of the range if it isn't already at max
The getCoverageIdentityPercentForMAF takes a single UCSC MAF formatted file and for each contig represented, calculates the coverage and identity percentages as relates to the REF aligned against.
This function splits a MAF file into a Set of Maf lines, where each entry in the set is a list
This function returns the indices of the start and end of the non-gap characters in a sequence. The gap is defined as a dash '-'. The start and end parameteres determine where in the sequence to make the search for the non-gap chahacters. An IntArray is returned with the start and end indices of the non-gap characters.
This function identifies the intersecting SRanges, not specific positions on these ranges.
Merge will merge overlapping and embedded ranges, and other ranges where distance between them is "count" or less base pairss. It will not merge adjacent/non-overlapping ranges A comparator is necessary as we can't merge until the ranges are sorted. SRange is Kotlin Set which is immutable, but not necessarily sorted. Merging of ranges requires that the upper endpoint SeqRecord of the first range matches the lower endpoint SeqRecord of the next range.
Method takes a comma separated list of SRanges and adds them to a sorted set. Overlapping intervals are NOT coalesced. User must supply a comparator - either their own or one defined in SeqRangeSort
Method takes a List
separate the SeqRecord:id from the site. This returns a site value minus the commas. This is necessary for .toInt() We could instead return 2 Strings, leaving the commas in place, and let the conversion to Int take place elsewhere.
function to read alignment block from the MAF file.
Shift each range in the set by "count" bps. Can be positive or negative number
Shift the given range by "count" positions. If the number is negative, it is a left shift. If the number is positive, shift right This will not shift into an adjacent contig of the genome
Find intersections between 2 SRanges. Returns null if they don't intersect, returns the intersecting values if there is an overlap.
For a set of SRanges, create new ranges based on removing any portion of the old range which overlaps with any of the ranges in the "removeRanges" set.
For a given range, create new ranges based on removing any portion of the old range which overlaps with any of the ranges in the "removeRanges" set. Return a new set of ranges