Skip to content

Genomic features

Open this tutorial as a notebook

GenomicFeatures reads a GFF3 file into Kotlin DataFrame objects, one per feature type. You can query them with the DataFrame API or with the convenience functions on the class itself.

This is a different view of the same data from the feature tree, which models the gene → transcript → exon hierarchy as a tree. Reach for GenomicFeatures when you want tabular, column-oriented access.

%use biokotlin

import biokotlin.genome.*
import org.jetbrains.kotlinx.dataframe.api.*

Loading a GFF

The file used here is a short excerpt of the B73 maize reference annotation that ships with the repository.

val gff = "resources/b73_shortened.gff"
val myGF = GenomicFeatures(gff)

// Everything the class can do:
myGF.help()
readGffToLists: number of file lines read: 22
readGffToLists: end, total lines processed: 22
Size of chromList = 1
Size of GeneList = 2
Size of cdsList = 7
Size of exonList = 4
Size of transcriptList = 2

Available Biokotlin GenomicFeature functions:
  cds() - returns a DataFrame of GFF CDS entries
  chromosomes() - returns a DataFrame of GFF chromosome entries
  columnNames(type:String) - returns a list of DataFrame columns for the specified feature.
      The "type" parameter must be one of the following:
      CDS, chromosome, exon, gene, three_prime_UTR, five_prime_UTR or mRNA
  exons() - returns a DataFrame of GFF exon entries
  genes() - returns a DataFrame of GFF gene entries
  threePrimeUTRs() - returns a DataFrame of GFF three_prime_UTR entries
  fivePrimeUTRs() - returns a DataFrame of GFF five_prime_UTR entries
  transcripts() - returns a Da
taFrame of GFF mRNA entries
  featuresByRange( chr:String, range:IntRange, features:String = "ALL")
      - returns a DataFrame containing all feature entries filtered by chromosome, range and requested feature types.
      The "features" parameter must be a comma-separated list containing 1 or more features from the list below:
        CDS, chromosome, exon, gene, three_prime_UTR, five_prime_UTR or mRNA
      If the features parameter is not specified, entries for all are returned.
  featuresWithTranscript(searchTranscript:String)
      - returns a DataFrame containing all GFF entries relating to the specified transcript
  sequenceForChrRange(chr:String, positions:IntRange) - returns the sequence as a String for the specified chr/range.
      If fasta is not present, or chr doesn't exist in the fasta, null is returned
  help() - returns a list of functions that may be run against a GenomicFeatures object

Feature tables

Each feature type has an accessor returning a DataFrame: genes, transcripts, exons, cds, chromosomes, fivePrimeUTRs, and threePrimeUTRs. columnNames reports the columns of any of them.

println("CDS column names:")
println(myGF.columnNames("CDS"))
CDS column names:
name,seqid,start,end,strand,phase,transcript

Select a few columns from the CDS table:

myGF.cds().select { seqid and start and end and phase }.print()
   seqid start   end phase
 0  chr1 34722 35318     0
 1  chr1 36037 36174     0
 2  chr1 36259 36504     0
 3  chr1 36600 36713     0
 4  chr1 36822 37004     0
 5  chr1 37416 37633     0
 6  chr1 38021 38366     1

Print the exons in full:

myGF.exons().print()
                          name seqid start   end strand rank           transcript
 0 Zm00001eb000010_T001.exon.1  chr1 34617 35318      +    1 Zm00001eb000010_T001
 1 Zm00001eb000010_T001.exon.2  chr1 36037 36174      +    2 Zm00001eb000010_T001
 2 Zm00001eb000020_T002.exon.1  chr1 41263 41588      -    8 Zm00001eb000020_T002
 3 Zm00001eb000020_T002.exon.2  chr1 41742 42375      -    7 Zm00001eb000020_T002

Querying by range

featuresByRange returns every feature overlapping a chromosome interval. Coordinates are 1-based and inclusive, as in the GFF itself.

// Restrict to particular feature types with a comma-separated list.
myGF.featuresByRange("chr1", 34617..36200, "three_prime_UTR,five_prime_UTR").print()
   seqid start   end strand           type                            data
 0  chr1 34617 34721      + five_prime_UTR transcript:Zm00001eb000010_T001
// Leave the types off and you get everything overlapping the interval.
myGF.featuresByRange("chr1", 34617..36200).print()
   seqid start   end strand           type                                     data
 0  chr1 34617 35318      +           exon   rank:1;transcript:Zm00001eb000010_T001
 1  chr1 34617 34721      + five_prime_UTR          transcript:Zm00001eb000010_T001
 2  chr1 34617 40204      +           gene name:Zm00001eb000010;biotype:protein_...
 3  chr1 34617 40204      +     transcript name:Zm00001eb000010_T001;biotype:pro...
 4  chr1 34722 35318      +            cds  phase:0;transcript:Zm00001eb000010_T001
 5  chr1 36037 36174      +           exon   rank:2;transcript:Zm00001eb000010_T001
 6  chr1 36037 36174      +            cds  phase:0;transcript:Zm00001eb000010_T001

Querying by transcript

featuresWithTranscript collects every feature belonging to one transcript.

myGF.featuresWithTranscript("Zm00001eb000010_T001").print()
    seqid start   end strand            type                                     data
  0  chr1 34617 35318      +            exon   rank:1;transcript:Zm00001eb000010_T001
  1  chr1 34617 34721      +  five_prime_UTR          transcript:Zm00001eb000010_T001
  2  chr1 34617 40204      +      transcript name:Zm00001eb000010_T001;biotype:pro...
  3  chr1 34722 35318      +             cds  phase:0;transcript:Zm00001eb000010_T001
  4  chr1 36037 36174      +            exon   rank:2;transcript:Zm00001eb000010_T001
  5  chr1 36037 36174      +             cds  phase:0;transcript:Zm00001eb000010_T001
  6  chr1 36259 36504      +             cds  phase:0;transcript:Zm00001eb000010_T001
  7  chr1 36600 36713      +             cds  phase:0;transcript:Zm00001eb000010_T001
  8  chr1 36822 37004      +             cds  phase:0;transcript:Zm00001eb000010_T001
  9  chr1 37416 37633      +             cds  phase:0;transcript:Zm00001eb000010_T001
 10  chr1 38021 38366      +             cds  phase:1;transcript:Zm00001eb000010_T001
 11  chr1 38367 38482      + three_prime_UTR          transcript:Zm00001eb000010_T001
 12  chr1 38571 39618      + three_prime_UTR          transcript:Zm00001eb000010_T001

Filtering with the DataFrame API

The tables are ordinary DataFrames, so the full Kotlin DataFrame vocabulary is available when the built-in queries are not enough.

myGF.cds()
    .filter { seqid == "chr1" && start <= 37000 && end >= 36000 }
    .print()
                   name seqid start   end strand phase           transcript
 0 Zm00001eb000010_P001  chr1 36037 36174      +     0 Zm00001eb000010_T001
 1 Zm00001eb000010_P001  chr1 36259 36504      +     0 Zm00001eb000010_T001
 2 Zm00001eb000010_P001  chr1 36600 36713      +     0 Zm00001eb000010_T001
 3 Zm00001eb000010_P001  chr1 36822 37004      +     0 Zm00001eb000010_T001

Adding sequence

GenomicFeatures(gff, fasta) also loads the reference FASTA, which enables sequenceForChrRange(chr, positions) and lets you pull the sequence behind any feature:

val myGFwithSeq = GenomicFeatures(gff, referenceFasta)
println(myGFwithSeq.sequenceForChrRange("chr1", 34617..34676))

The reference FASTA for this annotation is too large to ship with the repository, so that step is not run here.