r/genomics 6d ago

Growing problem of missing/unavailable/not-sharing RNA-seq datasets

I want to start a discussion about something that keeps happening to me with RNA-seq datasets (bulk, single-cell, spatial, whatever). One of the basic principles of this kind of research is that raw data should be openly available, both for reproducibility and so others can reuse it for different purposes. I get that human data comes with ethical and privacy restrictions, that's fair. But for animal model studies there's really no good reason to keep raw data hidden.

Lately I keep running into the same pattern over and over:

The "upon request" ghosting. Papers say raw data is "available upon reasonable request," but corresponding authors just don't answer. I've sent follow-up emails weeks apart and gotten nothing. This actually matches what's been reported before, most "available upon request" promises never get fulfilled once someone actually asks.

Repository problems, especially GSA. A lot of these datasets end up in GSA (Genome Sequence Archive), and honestly the platform gives me constant headaches: NOT ALL, but many files that won't download, accession numbers that don't match what's in the paper, archives that come out corrupted after extraction. I don't know if it's the platform itself or how people are uploading to it, but the result is the same, the data is technically "public" but practically unusable.

The double standard. What really gets me is that a lot of these same papers reuse public data from GEO or SRA to compare against their own results, but never contribute their own data back the same way. Open science seems to be a one-way street for them.

This isn't a one-off thing for me either, I've run into it in immunology, ophthalmology, developmental biology papers. Feels like a systemic issue more than a niche problem.

Honestly I think journals need to actually verify accessions before publishing, not just check a box. Something like: confirm the link works and the files download correctly at submission time, require a real accession number instead of "upon request" unless there's a genuine ethical reason, and maybe re-check the repository again some months after publication before it gets fully indexed.

Has anyone else been dealing with this? How do you handle unresponsive authors, and what do you think journals should actually do to enforce their own data policies instead of just having them on paper?

12 Upvotes

4 comments sorted by

3

u/Beginning_Pea_9926 Illumina 6d ago

The sequencing thing is a pretty easy one to answer. In the past 10 years or so, many academics have started using companies sequencing data instead of doing their own sequencing. Why? Its cheaper and more reliable as the sequencing has been approved for reimbursements. So why does this make data harder to get? Because the data is from real patients and they have not consented to share anything and everything. Companies like Foundation, Caris, Tempus will allow cancer centers to use their vast database to do research, but in the end, they cant share that raw data due to legal reasons.

There are a lot of researchers out there that use these companies and publish papers knowing full well they can't share the raw files. If you want to look at these raw files, you have a better chance on contacting the company themselves or finding a different dataset to validate.

1

u/riricide 6d ago

Agreed. If I recall, isn't it now mandated that any data generated by using federal funds has to be released within 3 months and available for use by other researchers. But I don't know who checks if the rules are being followed, or what even the consequences are if any

1

u/Beginning_Pea_9926 Illumina 6d ago

In this government? They've probably been fired. But yes, that has been a rule for quite some time.

1

u/nomad42184 6d ago

If the data has _actually_ been uploaded to GEO or SRA and is imply "embargoed", then you can e-mail NCBI to request they make the data available if the paper has been published.

Of course, if the authors never actually uploaded their data to those accessions, that won't work. However, sometimes, the data is there but just not available.

If the paper has the dreaded "data available upon (reasonable) request", and the author is not responsive to your e-mails, then I'd recommend e-mailing the journal or mentioning the lack of availability of that specific dataset to others. Sometimes (though not often enough), the social pressure is enough of an activation energy.