Data Set Details
Integrase-On-Demand-Pipeline Data Set
Total Size: 1 GB
Files needed to run the Integrase-On-Demand-Pipeline, a program designed to provide users with a list of putative attachment site and integrase pairs for a prokaryotic genome of interest.
- isles.pkl: Serialized python-object file, containing a dictionary of attachment site sequences and reference genomic island information extracted from the Genomic island database
- ints.gff: Gene format file containing annotations for all integrases referenced in isles.pkl. The source genome, gene coordinates, integrase name, protein IDs and amino acid sequence included.
- reps.msh: Binary file containing 1000 128-bit MurmurHash3 hashes for >80,000 genomes
Contributors
McClain, Hannah; Williams, Kelly; Torrance, Ellis; Mageeney, Catherine
Data Files
Source
10.1093/nar/gkag106
Fields
- isles.pkl: When loaded, the fields contains information about genomic islands, grouped by identical attachment site sequences and organized by species. Structure:
- attachment site sequence (nucleotide sequences with lengths ranging from 22 to 276 base pairs)
- Genome Taxonomy Database species name
- Genomic island (GI) information per island in the species utilizing the above attachment site sequence: integrase name, GI identifier ([9-digit GCA].[gene number].[gene identifier]), genome scaffold(s) and coordinates, direction and attachment site coordinates, cross-over site (7-bp or ID-block) coordinates, attachment site sequence in the mobile element, source program(s), support value, island type.
- Genome Taxonomy Database species name
- attachment site sequence (nucleotide sequences with lengths ranging from 22 to 276 base pairs)
- ints.gff: Standard gff format, the following fields are tab separated:
-
- Scaffold ID
- Source program of annotation
- Feature (CDS, tRNA, etc.)
- Start coordinate
- End coordinate
- Score
- Direction (+/-)
- Frame (0, 1, or 2)
- Attributes column: protein ID (9-digit GCA, gene number), GCA, integrase name, and amino acid sequence
- reps.msh: Binary file containing, 9-digit GCA, length of each original genome, and 1000 128-bit MurmurHash3 (non-cryptographic hash function) hashes for each of the 85,205 genomes
Sample Data Set
isles.pkl (after loading python object):
"TAGCGCACTTgcatggggtgcaaggggtcgagtgttcgaatcactccgtcccgaccaTATATTTCAA": {
"Pseudomonas_E__alloputida": [
[
"Y-SXTArm_02675",
"000007565.32.P",
"AE015451.2/2817039-2848838",
"+;2817034-2817080;2848834-2848880",
"42-49",
"TAAAATAAGCgcatggggtgcaaggggtcgagtgttcgaatcactccgtcccgaccaAAAAACCCTA",
"Islander,TIGER",
"7",
"other"
]
ints.gff:
GG666849.1 Prodigal:2.6 CDS 224532 225896 . + 0 ID=000003135_00192;gca=000003135;name=Y-Tn916Arm_00192;seq=MVEETKRPLRRRRFGCILERKGSTGDVTSIEARYISPINGQRVSKRFAPGRRGDAEDWLETERSIVDLHRRGMMTWIPPRDRDGNTLTPKLTFGVFADGYVRRHRRKDGAEIAGSTLRNLRNDIKHLKEAFGDVKLAELTEELVTEWYYGPHPNGEWQFRSECIRLKMLLREACAPGSKGAPPLLAENPFTLPIPPEPEAGSSDIPPVTPDELYHIYNAMPGYTRLSVYLAACAGGMRIGEVCGLMDTDFDLENKVLMIRRSVSHGADDLGPSRIGRLKTKGSRRTVPIPDMLIPLIRMHLEDRPDQSNHMFFQAKRGEILCQNTLRNHFMKARKAAGRPDLQFRTLRVTHATRLMLDGSSLKETMDALGHVREETTLRHYLRAVPEHQREAAERTAAYLLSADPSLAIASLPPAPAGLDADGDAGDMAALALLLGQVSTLMAGIAARNAQPAG
reps.msh is binary
Supplemental Information
isles.pkl uses the built-in python module, pickle, to save python-objects. To read to contents of the file, please use python 3.8+ (the software associated with this data set was written in 3.12)
To read reps.msh the external program MASH (v2.3) (https://github.com/marbl/Mash) needs to be available as a system-wide executable.
Software Needed
The GitHub page with software using this dataset can be found here: https://github.com/sandialabs/Integrase-On-Demand
Coverage
Spatial Coverage
N/A
Temporal Coverage
N/A