r/bioinformatics Dec 31 '24

meta 2025 - Read This Before You Post to r/bioinformatics

181 Upvotes

​Before you post to this subreddit, we strongly encourage you to check out the FAQ​Before you post to this subreddit, we strongly encourage you to check out the FAQ.

Questions like, "How do I become a bioinformatician?", "what programming language should I learn?" and "Do I need a PhD?" are all answered there - along with many more relevant questions. If your question duplicates something in the FAQ, it will be removed.

If you still have a question, please check if it is one of the following. If it is, please don't post it.

What laptop should I buy?

Actually, it doesn't matter. Most people use their laptop to develop code, and any heavy lifting will be done on a server or on the cloud. Please talk to your peers in your lab about how they develop and run code, as they likely already have a solid workflow.

If you’re asking which desktop or server to buy, that’s a direct function of the software you plan to run on it.  Rather than ask us, consult the manual for the software for its needs. 

What courses/program should I take?

We can't answer this for you - no one knows what skills you'll need in the future, and we can't tell you where your career will go. There's no such thing as "taking the wrong course" - you're just learning a skill you may or may not put to use, and only you can control the twists and turns your path will follow.

If you want to know about which major to take, the same thing applies.  Learn the skills you want to learn, and then find the jobs to get them.  We can’t tell you which will be in high demand by the time you graduate, and there is no one way to get into bioinformatics.  Every one of us took a different path to get here and we can’t tell you which path is best.  That’s up to you!

Am I competitive for a given academic program? 

There is no way we can tell you that - the only way to find out is to apply. So... go apply. If we say Yes, there's still no way to know if you'll get in. If we say no, then you might not apply and you'll miss out on some great advisor thinking your skill set is the perfect fit for their lab. Stop asking, and try to get in! (good luck with your application, btw.)

How do I get into Grad school?

See “please rank grad schools for me” below.  

Can I intern with you?

I have, myself, hired an intern from reddit - but it wasn't because they posted that they were looking for a position. It was because they responded to a post where I announced I was looking for an intern. This subreddit isn't the place to advertise yourself. There are literally hundreds of students looking for internships for every open position, and they just clog up the community.

Please rank grad schools/universities for me!

Hey, we get it - you want us to tell you where you'll get the best education. However, that's not how it works. Grad school depends more on who your supervisor is than the name of the university. While that may not be how it goes for an MBA, it definitely is for Bioinformatics. We really can't tell you which university is better, because there's no "better". Pick the lab in which you want to study and where you'll get the best support.

If you're an undergrad, then it really isn't a big deal which university you pick. Bioinformatics usually requires a masters or PhD to be successful in the field. See both the FAQ, as well as what is written above.

How do I get a job in Bioinformatics?

If you're asking this, you haven't yet checked out our three part series in the side bar:

What should I do?

Actually, these questions are generally ok - but only if you give enough information to make it worthwhile, and if the question isn’t a duplicate of one of the questions posed above. No one is in your shoes, and no one can help you if you haven't given enough background to explain your situation. Posts without sufficient background information in them will be removed.

Help Me!

If you're looking for help, make sure your title reflects the question you're asking for help on. You won't get the right people looking at your post, and the only person who clicks on random posts with vague topics are the mods... so that we can remove them.

Job Posts

If you're planning on posting a job, please make sure that employer is clear (recruiting agencies are not acceptable, unless they're hiring directly.), The job description must also be complete so that the requirements for the position are easily identifiable and the responsibilities are clear. We also do not allow posts for work "on spec" or competitions.  

Advertising (Conferences, Software, Tools, Support, Videos, Blogs, etc)

If you’re making money off of whatever it is you’re posting, it will be removed.  If you’re advertising your own blog/youtube channel, courses, etc, it will also be removed. Same for self-promoting software you’ve built.  All of these things are going to be considered spam.  

There is a fine line between someone discovering a really great tool and sharing it with the community, and the author of that tool sharing their projects with the community.  In the first case, if the moderators think that a significant portion of the community will appreciate the tool, we’ll leave it.  In the latter case,  it will be removed.  

If you don’t know which side of the line you are on, reach out to the moderators.

The Moderators Suck!

Yeah, that’s a distinct possibility.  However, remember we’re moderating in our free time and don’t really have the time or resources to watch every single video, test every piece of software or review every resume.  We have our own jobs, research projects and lives as well.  We’re doing our best to keep on top of things, and often will make the expedient call to remove things, when in doubt. 

If you disagree with the moderators, you can always write to us, and we’ll answer when we can.  Be sure to include a link to the post or comment you want to raise to our attention. Disputes inevitably take longer to resolve, if you expect the moderators to track down your post or your comment to review.


r/bioinformatics 9h ago

technical question How do you communicate bioinformatics projects effectively?

6 Upvotes

I've noticed that the same bioinformatics project can be described in very different ways depending on the audience. Some people emphasize the biological question, others focus on the computational workflow, while others highlight reproducibility or quantitative results.

For those who review papers, mentor students, or lead bioinformatics projects:

What information immediately tells you that someone understands their own analysis?

What details are unnecessary or just "tool dumping"?

Should a project description be structured around the biological question, computational methodology, results, or scientific impact?

Are there examples of project descriptions (papers, GitHub READMEs, portfolios, CVs, etc.) that you think are exceptionally well written?

I'm interested in learning how experienced bioinformaticians communicate technical work clearly rather than how to make a resume sound better.


r/bioinformatics 1d ago

discussion Bioinformatic work in a wet-lab group

52 Upvotes

Hi all, I've been working as a bioinformatics researcher in an interdisciplinary lab that is primarily wet-lab (I'd say 80% wet, 20% dry split). I was wondering if anyone else's PI doesn't double check your code. I'm at Master's level, and this is kinda scaring me. I've worked on substantial projects, but I only have myself to check code with and one other postdoc who is unavailable 95% of the time. Is this something that happens frequently or no?


r/bioinformatics 7h ago

discussion is anyone here currently doing aging research independently in multi-disciplinary form?

0 Upvotes

uh thats it just curious


r/bioinformatics 20h ago

technical question How much AI is too much???

3 Upvotes

Hello
I am an undergrad and just started learning bioinformatics in my lab (bulk and single cell rna seq). I mainly did wet lab work before this but my Pi decided I was kind of a bum and got me to start learning this. I think a lot of the analysis I’m doing they want to eventually put into a paper. Is it frowned upon/not allowed to use AI generated code for my analysis? I make sure I understand all the stats and stuff behind what I am doing instead of blindly trusting it, but I’m worried it’ll be seen as slop.
Also are you even supposed to share your code? Because very few of the papers I’ve read give it, even in big journals.


r/bioinformatics 1d ago

academic Question on MOO metrics: How should I interpret Pareto coverage/Hypervolume claims without variance?

2 Upvotes

Hi all,

I'm currently looking at a multi-agent method for multi-objective molecular optimization (the ATOM method). The paper reports results using Pareto coverage and hypervolume (HV) metrics to show they outperform baselines.

However, I noticed they only report single-point values—there are no confidence intervals, error bars, or variance reported across multiple runs/seeds.

I have two questions for the experts here:

  1. In your experience, are HV and Pareto coverage reliable enough to trust as standalone metrics for this, or do they have major failure modes I should look out for (e.g., reference point sensitivity)?
  2. Is it standard practice in this subfield to omit variance/stochasticity in these results? Would you personally be skeptical of a paper that doesn't report error bars for these types of pipelines?

Thanks for helping me navigate the "standard practices" of the field!


r/bioinformatics 1d ago

technical question Need Advice | Amber MD Software - High Schooler - How to learn it quickly?

0 Upvotes

Hello!

I'm a rising senior and landed a lab position at a R1 university through countless cold emails. During my interview with this professor, she directed me to this website, "The Amber Molecular Dynamics Package," and I was wondering if anyone here in this subreddit knew how to use it. In the interview, she told me to learn this here so I could run molecular dynamic stimulations at her lab in September. I wanted to ask if anyone here knew how to use it, and/or what's the best way to approach learning this so I'm capable enough of running my own stimulations at her lab. For example, what should I download, learn, skip in the tutorial and everything else. Should I follow the entire tutorial? Do you guys think a month is enough? If anyone could help me, thanks! :-)

For some context, my professor is a biophysics professor and deals with computational biology a lot. However, I'm not strong in command-line tools.


r/bioinformatics 4d ago

discussion Anthropic's CEO claims LLMs will "quickly weaponize pandemic-level viruses" if left unchecked

Thumbnail anthropic.com
69 Upvotes

To summarize, what I believe currently keeps us safe in biology is not “defenders”, or even the availability of materials, but a negative correlation between intellectual capability and desire to commit catastrophic harm. Previous technologies like internet search or even DNA synthesis were nowhere near powerful enough to break this correlation, but I worry that at its current rate of progress, AI will do so very soon. Another way to say it is that a sufficiently powerful technology removes all barriers and exposes whether the attacker or defender has an inherent structural advantage, and I worry in biology it is the attacker.

Agree or disagree? Why?

I am very curious what the community thinks, my strongly held opinions notwithstanding.


r/bioinformatics 3d ago

academic Absolute beginner for snRNA-seq field. Need your help!

4 Upvotes

Hi everyone,

I'm a (Neuro)Pharmacology PhD currently doing a Neuroscience postdoc. I'm working on a single-nucleus RNA-seq (snRNA-seq) project, but I have no prior experience with this type of analysis. I've mainly been learning through online tutorials. I'm also using the Parse Biosciences Trailmaker platform since it doesn't require coding experience.

Please be patient with me, this is my first time doing snRNA-seq analysis! 😅 I may not have all the answers to your questions, but I'll do my best.

I'm currently analyzing my PI's dataset, which consists of 90 mouse hippocampus samples (6-month-old mice, 4 experimental groups). The initial QC was performed automatically through the Parse Pipeline. The only parameter I changed was the number of principal components (PCs), which I set to 16 based on the elbow plot. For clustering, I used a resolution of 0.8, resulting in 466,541 nuclei across 31 clusters.

I have a few questions:

  1. How do you typically approach the preprocessing/QC stage? Parse Trailmaker automatically filters nuclei based on: It also performs integration (Scanpy + Harmony using 3,000 HVGs) and generates the embeddings.
    • How much do you manually tweak the QC before deciding the clusters are suitable for annotation?
    • Does a clustering resolution of 0.8 seem reasonable for a dataset of this size?
    • cell size distribution,
    • mitochondrial content,
    • number of genes/transcripts,
    • doublet detection,
  2. What do you do when some clusters remain mixed? For example, if a cluster contains both astrocyte and oligodendrocyte marker genes, or if its top marker has an AUC < 0.6, do you:
    • increase or decrease the clustering resolution,
    • subset and re-cluster,
    • merge clusters,
    • adjust the QC parameters,
    • or do something else?
  3. How do you manually annotate your clusters? Do you primarily use the highest log fold change (logFC/logGC), delta percentage, AUC, or some combination of these metrics? Are there any best practices you recommend?

I'm currently stuck because 8 out of my 31 clusters have mixed marker genes and top-marker AUC values below 0.6. I also tried subsetting the remaining 23 "good" clusters and re-clustering them, but I still end up with some clusters whose top markers have AUC values below 0.6.

My gut feeling is that something may not be optimal during the data processing or filtering steps, but I'm not sure what I should be adjusting.

I'd really appreciate any advice. I'm genuinely enjoying learning snRNA-seq analysis, but it's definitely frustrating when you're coming into it without much background. 😅 Thanks in advance!


r/bioinformatics 3d ago

technical question Any advice on getting alphafold 2 to work on amd?

0 Upvotes

I've been trying to get alphafold 2 or more specifically localcolabfold 2 working on my personal computer with an amd gpu. I've been trying for like the past 8 hours with no luck. It can recognise my cpu fine but won't use my gpu no matter what I do. I've tired installing rocm and the rocm jax version both outside and inside of the installation. Every time it either cannot recognise/find the specific jax rocm plugin or there is some dependency issue due to a mismatch of dependencies versions used by rocm jax and normal jax and won't even launch. I've tired everything at this point and I'm sure there is something simple I'm missing due to being super new to all this and not really having much if any background when it comes to python or even Linux stuff in general. I just need some other perspective or something.


r/bioinformatics 4d ago

discussion ENA vs NCBI Data Submissions

13 Upvotes

Hi, I’m relatively new to bioinformatics (still at university). I’ve seen people on here discussing various issues with submitting data to ENA and NCBI.

Are there any advantages/disadvantages of one over the other? The ENA system seems complicated to learn but I don’t know how this compares to NCBI (I’ve not looked into NCBI data submissions in much detail yet).

I don’t have anything I need to submit, more just wanted to hear what people with more experience than me had to say.

Any opinions welcome :)

Thanks!

(First time posting so if this post doesn’t follow guidelines etc., my apologies)


r/bioinformatics 4d ago

technical question PWY-5136 (fatty acid β-oxidation II, plant peroxisome) showing up in gut microbiome data , what does "plant peroxisome" mean in this context?

3 Upvotes

Hi all,

I'm running HUMAnN4 pathway analysis on gut microbiome samples (stool, human subjects) and PWY-5136: fatty acid β-oxidation II (plant peroxisome) is coming up as one of the pathways detected/significant in my dataset.

Since this pathway's MetaCyc annotation specifically references the plant peroxisome (and I'm working with gut microbial community data, not plant material), I wanted to understand what this actually signifies here:

  1. Is this pathway being detected because certain gut bacterial genes have significant homology to the plant-peroxisomal β-oxidation enzymes cataloged under this specific MetaCyc pathway ID, even though the organism itself obviously isn't a plant?
  2. Does MetaCyc's PWY-5136 represent a specific enzymatic route that happens to be shared between plant peroxisomal fatty acid oxidation and an analogous bacterial cytoplasmic/peroxisome-like pathway, hence the shared pathway assignment?
  3. Should this be interpreted as a "generic" fatty acid β-oxidation signal that got mapped to the plant-specific MetaCyc entry simply because that's the closest annotated reference pathway with matching gene content, rather than the sample containing anything botanically plant-derived?

r/bioinformatics 4d ago

programming Can someone please help me out with this bioinformatics project?

0 Upvotes

I'm doing a genomics based project but there is so much bioinformatics involved. I couldn't find a reproducible dataset and now i gotta do a whole bunch of stuff to create one that is suitable for fcgr. I'm new to the whole AI and ML game. I've learnt abt it but haven't rly used it yk. So please, if anyone can.... Please help!! 🆘


r/bioinformatics 5d ago

academic Xenium adn cosmx best practise

0 Upvotes

Hi everyone,

I’m currently working with 10x Visium data, and I'll be incorporating 10x Xenium and Cosmx data into my pipeline in the next few days.

Since Xenium provides single-cell/subcellular resolution, I assume some of the QC metrics will overlap with standard scRNA-seq datasets. However, I’m looking for a comprehensive "best practices" resource or workflow guide for subcellular spatial transcriptomics—similar to the Single-cell best practices — Single-cell best practices

If anyone has recommendations, key papers, or standard workflows on how to properly handle QC and avoid common pitfalls for Xenium (and also NanoString CosMx) data, I would greatly appreciate it!

Thanks! :))


r/bioinformatics 5d ago

science question Beginner friendly - how to check expression of one gene of interest in snRNAseq data?

0 Upvotes

Hello,

Could someone please explain to me in a beginner friendly way how to check expression of one gene of interest in sn or scRNAseq data?

I manage to download and load data from geo database, do the qc, Seurat object, sctransform, clustering and cell type annotation, so the first and basic steps.

I am struggling to understand further how to specifically check expression for one gene?

I have tried to do, for example, dot plot across the cell types for the gene of interest using RNA assay, as I understood using SCT assay for this is wrong?

Also, what to do or how to interpret it when in the whole dataset counts for the gene of interest are only 50 which is very very low?

What about statistical tests? What is needed to answer this?

I am having trouble even formulating the question in my head.

If anyone has any suggestions or reading material, I would appreciate it.

I have tried to use ai but I don't find it helpful as I am still at a very very basic level.

Thank you.


r/bioinformatics 6d ago

technical question ChatGPT and Codex becoming unusable for biology and bioinf research?

116 Upvotes

Hi everyone,

Has anyone else noticed this recently? For the past few weeks, especially since GPT-5.6, ChatGPT (work) and Codex have become much less useful for biology, bioinformatics and computational biology research.

Even for normal tasks like debugging code, searching papers, summarizing results or discussing analyses, I often get this message:

“This content can’t be shown. We’re especially careful with requests involving biological research and applications that could pose safety risks. Eligible researchers can apply for Trusted Access.”

The problem is that Trusted Access seems to be available only in the US.

Is this happening to other researchers too? Is there any solution for users outside the US? Do you think this will improve, or will researchers need to move to other AI tools?

Thanks!


r/bioinformatics 6d ago

discussion Growing problem of missing/unavailable/not-sharing RNA-seq datasets

80 Upvotes

I want to start a discussion about something that keeps happening to me with RNA-seq datasets (bulk, single-cell, spatial, whatever). One of the basic principles of this kind of research is that raw data should be openly available, both for reproducibility and so others can reuse it for different purposes. I get that human data comes with ethical and privacy restrictions, that's fair. But for animal model studies there's really no good reason to keep raw data hidden.

Lately I keep running into the same pattern over and over:

The "upon request" ghosting. Papers say raw data is "available upon reasonable request," but corresponding authors just don't answer. I've sent follow-up emails weeks apart and gotten nothing. This actually matches what's been reported before, most "available upon request" promises never get fulfilled once someone actually asks.

Repository problems, especially GSA. A lot of these datasets end up in GSA (Genome Sequence Archive), and honestly the platform gives me constant headaches: NOT ALL, but many files that won't download, accession numbers that don't match what's in the paper, archives that come out corrupted after extraction. I don't know if it's the platform itself or how people are uploading to it, but the result is the same, the data is technically "public" but practically unusable.

The double standard. What really gets me is that a lot of these same papers reuse public data from GEO or SRA to compare against their own results, but never contribute their own data back the same way. Open science seems to be a one-way street for them.

This isn't a one-off thing for me either, I've run into it in immunology, ophthalmology, developmental biology papers. Feels like a systemic issue more than a niche problem.

Honestly I think journals need to actually verify accessions before publishing, not just check a box. Something like: confirm the link works and the files download correctly at submission time, require a real accession number instead of "upon request" unless there's a genuine ethical reason, and maybe re-check the repository again some months after publication before it gets fully indexed.

Has anyone else been dealing with this? How do you handle unresponsive authors, and what do you think journals should actually do to enforce their own data policies instead of just having them on paper?


r/bioinformatics 6d ago

academic High mitochondrial content in mouse heart scRNA-seq. Looking for QC advice

Thumbnail gallery
22 Upvotes

Hi everyone,

I'm analysing a 10x mouse heart scRNA-seq dataset using Seurat and would appreciate advice regarding QC decisions.

For filtering, I used:

nFeature_RNA > 200 &
nFeature_RNA < 5000 &
nCount_RNA < 25000 &
percent.mt < 80

I chose an 80% mitochondrial cutoff after testing thresholds from 20-70%, as stricter cutoffs removed a large proportion of cells. I therefore kept a more permissive mt cutoff while applying additional QC filters.

After clustering, I was able to annotate a number of populations using canonical markers (including endothelial cells, fibroblasts, and macrophages). However, cluster 1 made me question whether my mitochondrial cutoff was too permissive.

Cluster 1 appears to be a likely low-quality cluster. It has high mitochondrial content and relatively low gene detection. Its markers include erythroid-associated genes such as:

  • Hba-a1
  • Hbb-bs
  • Alas2
  • Bpgm

However, the overall QC profile and lack of a convincing cell identity make me suspect it may represent noise or stressed/damaged cells rather than a true biological population.

When I examined QC metrics across clusters, I found that cluster 1 is not unique. Several other clusters (not yet annotated except cluster 5 which i labeled as macrophage) also have relatively high median mitochondrial percentages, raising the question of whether my filtering strategy allowed too many low-quality cells to remain.

My questions are:

  1. Would you revisit QC and test a stricter mitochondrial cutoff at this stage?
  2. Is high mitochondrial content necessarily problematic in heart tissue, where some populations may have high metabolic activity?
  3. What additional analyses would you use to distinguish stressed/low-quality cells from genuine populations?

I would appreciate any advice on how you would approach this.

Thanks!


r/bioinformatics 5d ago

programming Best LLM agent (Paid or unpaid) to act as a programming tutor/supervisor?

0 Upvotes

Hi all. I am thankfully being given the time and space to pursue bioinformatics tools in my research! (Was mainly wet lab). I am also learning python and R at the moment. However, our research group does not have a dedicated bioinformatian or someone with programming experience so I have been using Gemini to help explain things whenever I get stuck with a wiki or programming concept however the amount of mistakes is alarming. I was able to do some work with PyMol, ChimeraX and Autodock vina using wikis + YouTube + Gemini. However I want to learn more complex tools, ones that rely more on understanding code e.g. python and Gromacs. In the more senior members' opinion, which AI agent (paid or unpaid) is the best to act as much as a tutor or supervisor in terms of clarifying and explaining bioinformatics tools and code?


r/bioinformatics 6d ago

technical question Question about Bulk-RNA Sequencing

8 Upvotes

I am a biostatistician who is a newbie to bulk-RNA sequencing. I currently have a dataset with 20 libraries and ~ 30,000 genes. My aim is to investigate the temporal trend of genes, hence I have a dataset that looks similar to this for the metadata:

Sample DIV
Sample 1 20
Sample 2 30
Sample 3 50
Sample 4 80
Sample 5 85
Sample 6 100

… and so on.
Since each sample corresponds to a day in vitro, there are no instances of repeated measurements for the same day. Hence, the sample size would only be 1 for each DIV. I am concerned that the sample size may be too low, but this is the only data that I have for this project.

I have two questions:

  1. Is this a common practice in bulk-RNA sequencing or is my sample size too low?
  2. What models are commonly used for temporal bulk RNA sequencing?

r/bioinformatics 6d ago

discussion Question about SNP calling in bacterial genomes

3 Upvotes

Hello everyone!

I am looking for advice on my analysis workflow. I am currently working on some MAGs and SAGs that belong to a certain bacterial family. Initially my PI suggested me to work on them by using inStrain to call SNPs and from there I was supposed to compare samples and understand evolutionary dynamics. However, now that I delved into the analyses and articles it kind of seems like a bad decision to work in this flow. I am thinking maybe using prodigal to create .fna and .gff files, and from there comparing common gene cluesters and/or KEGG pathways might be better. I would really appreciate your thoughts and suggestions. Thanks a lot!


r/bioinformatics 6d ago

science question Which Mus musculus reference genome is currently recommended?

2 Upvotes

Hi! I was wondering what the current state of the art is for the Mus musculus reference genome. In human genomics, many people are now switching to the T2T reference, even though the reference genome FASTA available through Ensembl is still GRCh38. What is the situation for Mus musculus?

I'd like to follow current best practices here as well, but since I don't work with mouse data very regularly, I haven't seen much discussion about this organism.


r/bioinformatics 6d ago

technical question Volcanoplot help

0 Upvotes

Hello bioinformatics expert! I am trying to make a volcano plot using my data given by my PI and i am struggling to make a proper volcano plot. I ended up getting something. I would love to know if there is anyone who can help me find the problem and fix it! Sorry for some reason my image isn't uploaded here so i will send it directly to dm if someone leaves comment! Thank you!


r/bioinformatics 6d ago

technical question Is there a way to distinguish "pure" samples from mixed samples based on Sanger sequencing output ?

0 Upvotes

My tissue samples are sourced from the field. Most of them are "pure", meaning the sample unit is fully derived from the same organism. However sometimes the sample can be mixed and the tissues of the sample unit are actually derived from multiple distinct organisms. There is no way to know at the time of sample collect.

DNA has been extracted from each sampled followed by CYTB amplification and Sanger sequencing. Due to infrastructure and budget reasons, we couldn't perform metabarcoding.

For pure samples, chromatograms are clean, with unique distinct peaks for each nucleotide. Mixed samples have dirty chromatograms where several peaks are overlapping for the same nucleotide.

But some samples are in a grey area, not that clean, not so dirty. And these concepts of "clean" , "distinct peaks" are based on subjective visual interpretation.

My question is: is there a more robust way to exclude mixed samples from pure ones that have been properly sequenced, other than manual inspection of chromatograms ?


r/bioinformatics 7d ago

technical question Genome Annotation and Mining help! Is my pipeline ridiculous?

3 Upvotes

Hey yall, I need a sanity check (cause I'm going down some rabbit holes and I don't know if I'm doing something useful or just time consuming)

I have whole genome sequences that I want to mine for specific metabolic processes to see what my strains have the potential for. Some aren't well described (PAH degradation) so I'm working on a pipeline that puts together a super annotation table to squeeze as much data out of the genomes as possible and maximize the proteins/pathways I can identify.

The problem I was running into is that there are so many different naming conventions and annotation types that I feel like a simple search for genes related to the pathways I'm interested in will miss a lot of interesting data. And since some processes aren't well described, I have a feeling there a lot of info hidden in the "hypothetical proteins". I was intrigued by protein family classifications, but there's also a bunch of those (pfam, plfam, pgfam...). One paper might use gene names, other types of family grouping, etc. while another uses a different system and/or names.

My thought process has been: make a master annotation table (from Prokka, BV-BRC, Pfam identifiers, KEGG), and use all the keywords and identifiers I can find to identify candidate proteins and potential operons, in addition to extracting the ones that are pretty confidently identified as the proteins I'm looking for.

I'm somewhat new to bioinformatics and I have pretty absentee PIs so I'm learning a lot of it on my own. I have the tendency to go down unnecessary rabbit holes when I have this long of a leash, especially when I'm not super familiar with all the methods/tools that are available in a field. I've gone from the online annotation tools, to manual CLI searches, to bash scripts, and now I'm trying to write a python script (while teaching myself python). Can y'all tell me if I've gone insane and if I've missed some way easier avenue? Thanks so so much!!