Skip to content

Discrepancies Between Genome Sequences from BED Coordinates and tfrecords in Basenji Dataset #37

Description

@yangzhao1230

I have identified discrepancies between the sequences extracted from the hg38.ml.fa reference genome using BED file coordinates and those stored in the tfrecords within the Basenji dataset hosted on Google Cloud (https://console.cloud.google.com/storage/browser/basenji_barnyard?pageState=(%22StorageObjectListTable%22:(%22f%22:%22%255B%255D%22). This inconsistency is concerning as it affects the reliability of our data used for genomic analyses.

Detailed Observations:

  1. I retrieved sequences using the sequences.bed file from the hg38.ml.fa reference genome.

  2. Upon comparing these sequences with those recorded in the tfrecords dataset, I discovered mismatches. For example, in the validation set for human data, only two samples had sequences that matched their coordinates.

Questions:

  1. What could be causing these discrepancies between the retrieved sequences and those in the tfrecords?

  2. How is the hg38.ml.fa genome processed, and how does it differ from the standard reference genome?

Any assistance or guidance on how to verify the consistency of the data in the Google Cloud repository would be greatly appreciated.

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions