Genomic features¶
Open this tutorial as a notebook
GenomicFeatures reads a GFF3 file into Kotlin
DataFrame objects, one per feature type.
You can query them with the DataFrame API or with the convenience functions on
the class itself.
This is a different view of the same data from the
feature tree, which models the gene → transcript
→ exon hierarchy as a tree. Reach for GenomicFeatures when you want
tabular, column-oriented access.
Loading a GFF¶
The file used here is a short excerpt of the B73 maize reference annotation that ships with the repository.
val gff = "resources/b73_shortened.gff"
val myGF = GenomicFeatures(gff)
// Everything the class can do:
myGF.help()
readGffToLists: number of file lines read: 22
readGffToLists: end, total lines processed: 22
Size of chromList = 1
Size of GeneList = 2
Size of cdsList = 7
Size of exonList = 4
Size of transcriptList = 2
Available Biokotlin GenomicFeature functions:
cds() - returns a DataFrame of GFF CDS entries
chromosomes() - returns a DataFrame of GFF chromosome entries
columnNames(type:String) - returns a list of DataFrame columns for the specified feature.
The "type" parameter must be one of the following:
CDS, chromosome, exon, gene, three_prime_UTR, five_prime_UTR or mRNA
exons() - returns a DataFrame of GFF exon entries
genes() - returns a DataFrame of GFF gene entries
threePrimeUTRs() - returns a DataFrame of GFF three_prime_UTR entries
fivePrimeUTRs() - returns a DataFrame of GFF five_prime_UTR entries
transcripts() - returns a Da
taFrame of GFF mRNA entries
featuresByRange( chr:String, range:IntRange, features:String = "ALL")
- returns a DataFrame containing all feature entries filtered by chromosome, range and requested feature types.
The "features" parameter must be a comma-separated list containing 1 or more features from the list below:
CDS, chromosome, exon, gene, three_prime_UTR, five_prime_UTR or mRNA
If the features parameter is not specified, entries for all are returned.
featuresWithTranscript(searchTranscript:String)
- returns a DataFrame containing all GFF entries relating to the specified transcript
sequenceForChrRange(chr:String, positions:IntRange) - returns the sequence as a String for the specified chr/range.
If fasta is not present, or chr doesn't exist in the fasta, null is returned
help() - returns a list of functions that may be run against a GenomicFeatures object
Feature tables¶
Each feature type has an accessor returning a DataFrame: genes,
transcripts, exons, cds, chromosomes, fivePrimeUTRs, and
threePrimeUTRs. columnNames reports the columns of any of them.
CDS column names:
name,seqid,start,end,strand,phase,transcript
Select a few columns from the CDS table:
seqid start end phase
0 chr1 34722 35318 0
1 chr1 36037 36174 0
2 chr1 36259 36504 0
3 chr1 36600 36713 0
4 chr1 36822 37004 0
5 chr1 37416 37633 0
6 chr1 38021 38366 1
Print the exons in full:
name seqid start end strand rank transcript
0 Zm00001eb000010_T001.exon.1 chr1 34617 35318 + 1 Zm00001eb000010_T001
1 Zm00001eb000010_T001.exon.2 chr1 36037 36174 + 2 Zm00001eb000010_T001
2 Zm00001eb000020_T002.exon.1 chr1 41263 41588 - 8 Zm00001eb000020_T002
3 Zm00001eb000020_T002.exon.2 chr1 41742 42375 - 7 Zm00001eb000020_T002
Querying by range¶
featuresByRange returns every feature overlapping a chromosome interval.
Coordinates are 1-based and inclusive, as in the GFF itself.
// Restrict to particular feature types with a comma-separated list.
myGF.featuresByRange("chr1", 34617..36200, "three_prime_UTR,five_prime_UTR").print()
seqid start end strand type data
0 chr1 34617 34721 + five_prime_UTR transcript:Zm00001eb000010_T001
// Leave the types off and you get everything overlapping the interval.
myGF.featuresByRange("chr1", 34617..36200).print()
seqid start end strand type data
0 chr1 34617 35318 + exon rank:1;transcript:Zm00001eb000010_T001
1 chr1 34617 34721 + five_prime_UTR transcript:Zm00001eb000010_T001
2 chr1 34617 40204 + gene name:Zm00001eb000010;biotype:protein_...
3 chr1 34617 40204 + transcript name:Zm00001eb000010_T001;biotype:pro...
4 chr1 34722 35318 + cds phase:0;transcript:Zm00001eb000010_T001
5 chr1 36037 36174 + exon rank:2;transcript:Zm00001eb000010_T001
6 chr1 36037 36174 + cds phase:0;transcript:Zm00001eb000010_T001
Querying by transcript¶
featuresWithTranscript collects every feature belonging to one transcript.
seqid start end strand type data
0 chr1 34617 35318 + exon rank:1;transcript:Zm00001eb000010_T001
1 chr1 34617 34721 + five_prime_UTR transcript:Zm00001eb000010_T001
2 chr1 34617 40204 + transcript name:Zm00001eb000010_T001;biotype:pro...
3 chr1 34722 35318 + cds phase:0;transcript:Zm00001eb000010_T001
4 chr1 36037 36174 + exon rank:2;transcript:Zm00001eb000010_T001
5 chr1 36037 36174 + cds phase:0;transcript:Zm00001eb000010_T001
6 chr1 36259 36504 + cds phase:0;transcript:Zm00001eb000010_T001
7 chr1 36600 36713 + cds phase:0;transcript:Zm00001eb000010_T001
8 chr1 36822 37004 + cds phase:0;transcript:Zm00001eb000010_T001
9 chr1 37416 37633 + cds phase:0;transcript:Zm00001eb000010_T001
10 chr1 38021 38366 + cds phase:1;transcript:Zm00001eb000010_T001
11 chr1 38367 38482 + three_prime_UTR transcript:Zm00001eb000010_T001
12 chr1 38571 39618 + three_prime_UTR transcript:Zm00001eb000010_T001
Filtering with the DataFrame API¶
The tables are ordinary DataFrames, so the full Kotlin DataFrame vocabulary is available when the built-in queries are not enough.
name seqid start end strand phase transcript
0 Zm00001eb000010_P001 chr1 36037 36174 + 0 Zm00001eb000010_T001
1 Zm00001eb000010_P001 chr1 36259 36504 + 0 Zm00001eb000010_T001
2 Zm00001eb000010_P001 chr1 36600 36713 + 0 Zm00001eb000010_T001
3 Zm00001eb000010_P001 chr1 36822 37004 + 0 Zm00001eb000010_T001
Adding sequence¶
GenomicFeatures(gff, fasta) also loads the reference FASTA, which enables
sequenceForChrRange(chr, positions) and lets you pull the sequence behind
any feature:
val myGFwithSeq = GenomicFeatures(gff, referenceFasta)
println(myGFwithSeq.sequenceForChrRange("chr1", 34617..34676))
The reference FASTA for this annotation is too large to ship with the repository, so that step is not run here.