Data Set Details

Data Sets / Genome or genetic data, such as gene sequences

Integrase-On-Demand-Pipeline Data Set

Total Size: 1 GB

Files needed to run the Integrase-On-Demand-Pipeline, a program designed to provide users with a list of putative attachment site and integrase pairs for a prokaryotic genome of interest.

  1. isles.pkl: Serialized python-object file, containing a dictionary of attachment site sequences and reference genomic island information extracted from the Genomic island database
  2. ints.gff: Gene format file containing annotations for all integrases referenced in isles.pkl. The source genome, gene coordinates, integrase name, protein IDs and amino acid sequence included.
  3. reps.msh: Binary file containing 1000 128-bit MurmurHash3 hashes for >80,000 genomes

Contributors

McClain, Hannah; Williams, Kelly; Torrance, Ellis; Mageeney, Catherine

Data Files

Transfer data using Globus

Source

10.1093/nar/gkag106

Fields

  1. isles.pkl: When loaded, the fields contains information about genomic islands, grouped by identical attachment site sequences and organized by species. Structure:
    • attachment site sequence (nucleotide sequences with lengths ranging from 22 to 276 base pairs)
      • Genome Taxonomy Database species name
        • Genomic island (GI) information per island in the species utilizing the above attachment site sequence: integrase name, GI identifier ([9-digit GCA].[gene number].[gene identifier]), genome scaffold(s) and coordinates, direction and attachment site coordinates, cross-over site (7-bp or ID-block) coordinates, attachment site sequence in the mobile element, source program(s), support value, island type.
  2. ints.gff: Standard gff format, the following fields are tab separated:
    1. Scaffold ID
    2. Source program of annotation
    3. Feature (CDS, tRNA, etc.)
    4. Start coordinate
    5. End coordinate
    6. Score
    7. Direction (+/-)
    8. Frame (0, 1, or 2)
    9. Attributes column: protein ID (9-digit GCA, gene number), GCA, integrase name, and amino acid sequence
  1. reps.msh: Binary file containing, 9-digit GCA, length of each original genome, and 1000 128-bit MurmurHash3 (non-cryptographic hash function) hashes for each of the 85,205 genomes

Sample Data Set

isles.pkl (after loading python object):

"TAGCGCACTTgcatggggtgcaaggggtcgagtgttcgaatcactccgtcccgaccaTATATTTCAA": {
"Pseudomonas_E__alloputida": [
[
"Y-SXTArm_02675",
"000007565.32.P",
"AE015451.2/2817039-2848838",
"+;2817034-2817080;2848834-2848880",
"42-49",
"TAAAATAAGCgcatggggtgcaaggggtcgagtgttcgaatcactccgtcccgaccaAAAAACCCTA",
"Islander,TIGER",
"7",
"other"
]

ints.gff:

GG666849.1 Prodigal:2.6 CDS 224532 225896 . + 0 ID=000003135_00192;gca=000003135;name=Y-Tn916Arm_00192;seq=MVEETKRPLRRRRFGCILERKGSTGDVTSIEARYISPINGQRVSKRFAPGRRGDAEDWLETERSIVDLHRRGMMTWIPPRDRDGNTLTPKLTFGVFADGYVRRHRRKDGAEIAGSTLRNLRNDIKHLKEAFGDVKLAELTEELVTEWYYGPHPNGEWQFRSECIRLKMLLREACAPGSKGAPPLLAENPFTLPIPPEPEAGSSDIPPVTPDELYHIYNAMPGYTRLSVYLAACAGGMRIGEVCGLMDTDFDLENKVLMIRRSVSHGADDLGPSRIGRLKTKGSRRTVPIPDMLIPLIRMHLEDRPDQSNHMFFQAKRGEILCQNTLRNHFMKARKAAGRPDLQFRTLRVTHATRLMLDGSSLKETMDALGHVREETTLRHYLRAVPEHQREAAERTAAYLLSADPSLAIASLPPAPAGLDADGDAGDMAALALLLGQVSTLMAGIAARNAQPAG

 

reps.msh is binary

 

Supplemental Information

isles.pkl uses the built-in python module, pickle, to save python-objects. To read to contents of the file, please use python 3.8+ (the software associated with this data set was written in 3.12)

To read reps.msh the external program MASH (v2.3) (https://github.com/marbl/Mash) needs to be available as a system-wide executable.

Software Needed

The GitHub page with software using this dataset can be found here: https://github.com/sandialabs/Integrase-On-Demand

Coverage

Spatial Coverage

N/A

Temporal Coverage

N/A

Top