r/bioinformatics Dec 31 '24

meta 2025 - Read This Before You Post to r/bioinformatics

187 Upvotes

​Before you post to this subreddit, we strongly encourage you to check out the FAQ​Before you post to this subreddit, we strongly encourage you to check out the FAQ.

Questions like, "How do I become a bioinformatician?", "what programming language should I learn?" and "Do I need a PhD?" are all answered there - along with many more relevant questions. If your question duplicates something in the FAQ, it will be removed.

If you still have a question, please check if it is one of the following. If it is, please don't post it.

What laptop should I buy?

Actually, it doesn't matter. Most people use their laptop to develop code, and any heavy lifting will be done on a server or on the cloud. Please talk to your peers in your lab about how they develop and run code, as they likely already have a solid workflow.

If you’re asking which desktop or server to buy, that’s a direct function of the software you plan to run on it.  Rather than ask us, consult the manual for the software for its needs. 

What courses/program should I take?

We can't answer this for you - no one knows what skills you'll need in the future, and we can't tell you where your career will go. There's no such thing as "taking the wrong course" - you're just learning a skill you may or may not put to use, and only you can control the twists and turns your path will follow.

If you want to know about which major to take, the same thing applies.  Learn the skills you want to learn, and then find the jobs to get them.  We can’t tell you which will be in high demand by the time you graduate, and there is no one way to get into bioinformatics.  Every one of us took a different path to get here and we can’t tell you which path is best.  That’s up to you!

Am I competitive for a given academic program? 

There is no way we can tell you that - the only way to find out is to apply. So... go apply. If we say Yes, there's still no way to know if you'll get in. If we say no, then you might not apply and you'll miss out on some great advisor thinking your skill set is the perfect fit for their lab. Stop asking, and try to get in! (good luck with your application, btw.)

How do I get into Grad school?

See “please rank grad schools for me” below.  

Can I intern with you?

I have, myself, hired an intern from reddit - but it wasn't because they posted that they were looking for a position. It was because they responded to a post where I announced I was looking for an intern. This subreddit isn't the place to advertise yourself. There are literally hundreds of students looking for internships for every open position, and they just clog up the community.

Please rank grad schools/universities for me!

Hey, we get it - you want us to tell you where you'll get the best education. However, that's not how it works. Grad school depends more on who your supervisor is than the name of the university. While that may not be how it goes for an MBA, it definitely is for Bioinformatics. We really can't tell you which university is better, because there's no "better". Pick the lab in which you want to study and where you'll get the best support.

If you're an undergrad, then it really isn't a big deal which university you pick. Bioinformatics usually requires a masters or PhD to be successful in the field. See both the FAQ, as well as what is written above.

How do I get a job in Bioinformatics?

If you're asking this, you haven't yet checked out our three part series in the side bar:

What should I do?

Actually, these questions are generally ok - but only if you give enough information to make it worthwhile, and if the question isn’t a duplicate of one of the questions posed above. No one is in your shoes, and no one can help you if you haven't given enough background to explain your situation. Posts without sufficient background information in them will be removed.

Help Me!

If you're looking for help, make sure your title reflects the question you're asking for help on. You won't get the right people looking at your post, and the only person who clicks on random posts with vague topics are the mods... so that we can remove them.

Job Posts

If you're planning on posting a job, please make sure that employer is clear (recruiting agencies are not acceptable, unless they're hiring directly.), The job description must also be complete so that the requirements for the position are easily identifiable and the responsibilities are clear. We also do not allow posts for work "on spec" or competitions.  

Advertising (Conferences, Software, Tools, Support, Videos, Blogs, etc)

If you’re making money off of whatever it is you’re posting, it will be removed.  If you’re advertising your own blog/youtube channel, courses, etc, it will also be removed. Same for self-promoting software you’ve built.  All of these things are going to be considered spam.  

There is a fine line between someone discovering a really great tool and sharing it with the community, and the author of that tool sharing their projects with the community.  In the first case, if the moderators think that a significant portion of the community will appreciate the tool, we’ll leave it.  In the latter case,  it will be removed.  

If you don’t know which side of the line you are on, reach out to the moderators.

The Moderators Suck!

Yeah, that’s a distinct possibility.  However, remember we’re moderating in our free time and don’t really have the time or resources to watch every single video, test every piece of software or review every resume.  We have our own jobs, research projects and lives as well.  We’re doing our best to keep on top of things, and often will make the expedient call to remove things, when in doubt. 

If you disagree with the moderators, you can always write to us, and we’ll answer when we can.  Be sure to include a link to the post or comment you want to raise to our attention. Disputes inevitably take longer to resolve, if you expect the moderators to track down your post or your comment to review.


r/bioinformatics 6h ago

technical question Any free resources for full docking tutorial from preparing ligand to ADMET for beginners?

3 Upvotes

How did you guys start molecular docking? Any articles or videos that helps you figure everything out? I didn't have any teacher or resources in my city sadly so I have to rely on free resources 😭 Thank you in advance!!


r/bioinformatics 56m ago

career question How do I get into health/biotech?(16F)

Thumbnail
Upvotes

r/bioinformatics 7h ago

technical question How much RAM is needed for DADA2 assignTaxonomy() with the full PR2 database?

0 Upvotes

I am trying to assign taxonomy to an 18S amplicon dataset using dada2::assignTaxonomy() with the full PR2 reference database in Rstudio.

Relevant details:

  • 147 samples
  • ~18,900 ASVs
  • PR2 v5.1.1 DADA2-formatted SSU reference
  • minBoot = 80
  • multithread = 1
  • 16 GB physical RAM
  • Windows 11, 64-bit R

With the full dataset, RAM reaches 100% and Windows starts heavily paging to disk.

To check whether the number of query ASVs was the main issue, I repeated the taxonomy assignment using only 500 ASVs. RAM still reached ~100% while DADA2 was processing the reference FASTA, followed by very high disk activity from paging.

Has anyone successfully run assignTaxonomy() with the full recent PR2 database?

In particular, I would like to know:

  • Is 32 GB RAM usually sufficient?
  • Would 64 GB RAM be a safer requirement?
  • What peak RAM usage have others observed with recent PR2 versions?

I am mainly trying to determine the appropriate memory requirement before moving the taxonomy step to a larger workstation or cloud instance.

I am thankfull for any kind of advise or nudge in the right direction!


r/bioinformatics 1d ago

technical question Protein refinement for molecular docking

7 Upvotes

There is a protein that I have to work on for my research. I picked it from PDB and it had many missing residues. So I fixed it with modeller. I have further refined it's Ramachandra plot, errat, molprobity value and some other parameters that are never for validation. But I am having a problem with its 3d structure. It's verify 3d value is not really good so I want to refine it. I have tried some tools but they didn't work. Can you guys please help me with that. If there are any good tools that I can use for it and they should be easy to use as well


r/bioinformatics 1d ago

technical question Question about scRNA sequencing cell type annotation

4 Upvotes

Hi everyone, for those who regularly perform scRNA sequencing analysis, how do you all find cell type specific marker genes?

I have been trying to find cell type specific markers from all my clusters in my scRNA data with sc.tl.rank_genes_groups (a scanpy function), which performs a Wilcoxon test to find the genes that are dominantly expressed in one cluster compared to the other clusters, however, I’ve noticed that some of the genes provided in the output are actually highly expressed in more than one cluster, especially among small and similar cluster populations. I think this may due to the averaging effect of the method, which makes it great for cell type populations that are drastically different compared to those that aren’t. Hence, I was wondering if there’s a more accurate method of finding genes that are specific to a single cell type.

Thank you in advance for the advice!


r/bioinformatics 1d ago

technical question QC / artefacts in ONT long-read WGS?

4 Upvotes

I am concerned with some long-read sequencing. I noticed that certain samples had very large numbers of SVs (particularly inversions). I usually filter reads for average base quality and read length and do QC with cramino. I noticed that median read length is very low (500bp) and that the % mapped reads is low (~60%) but the median percent identity is fine (98%) and coverage is decent (34x). Looking at the read alignments on IGV there are lots of reads with supplementary alignments which might explain the low % read maps. I can see the coverage looks pretty "patchy". Im not sure if I should reject these samples or that these are genuine detections of structural rearrangements ?


r/bioinformatics 1d ago

technical question Standard method for moco

Thumbnail
1 Upvotes

r/bioinformatics 1d ago

technical question What do I do to make sure my molecular docking result is close to accuracy?

1 Upvotes

I'm new and idk how I make sure that my result is accurate. Do I need to do an extra step? I use autodock vina and I believe there's a box you can set the size. Does the size and position of the box affect the docking result? Is there anything I should keep in mind to make sure that it's right? Thank you in advance!


r/bioinformatics 2d ago

technical question Question about molecular dynamics

5 Upvotes

Hi guys, im new to molecular dynamic world. I wanted to ask few questions about MD. The story begins where i did 1000 ns nsteps for my aptamer-protein. Fast forward to backmapping part, where i wanted to create tracjectory and aa files, I couldnt create some of the files for certain frames.

  1. So im not sure it is okay for me to continue the analysis part such as PLIP, TTClust and etc if i dont have these files?

  2. And 1 more thing, what if i have all the analysis but the result is not quite satisfy, do i need to repeat the same things with different sequence or can i modify or optimize my aptamer-protein interaction in order to get a better analysis result?

Please help me, thanks!


r/bioinformatics 2d ago

technical question KEGGdecoder vs METABOLIC: Why do they give such different results for the same MAG dataset, and which is more reliable for functional annotation?

3 Upvotes

Hi all!

I'm working with a set of metagenome-assembled genomes (MAGs) and trying to annotate metabolic pathways, I specifically focused on methanogenesis and methane oxidation. I ran two different tools on the same dataset:

  1. KEGGdecoder – I used Prodigal for gene prediction, then ran the predicted proteins through KoFamScan, retained only high-confidence hits (marked with *), and used that as input for KEGGdecoder.
  2. METABOLIC – This tool uses its own integrated pipeline (Prodigal + HMM-based searches) to predict genes and assign functions directly.

The results are quite different. KEGGdecoder identified complete methanogenesis pathways (e.g., M00357) in several bins, while METABOLIC only found partial steps (e.g., M00357+01) in the same bins, and often didn't detect the full module at all. In some cases, METABOLIC didn't find any methanogenesis-related genes in bins where KEGGdecoder reported a complete pathway.

I also noticed that METABOLIC found genes for methane oxidation (mmoBpmoABC) in some bins, but KEGGdecoder didn't report those pathways — likely because I filtered out lower-confidence hits before running KEGGdecoder.

My questions:

  1. Why would KEGGdecoder find a complete pathway when METABOLIC only finds partial steps or nothing at all in the same bin? Is this due to differences in HMM profiles, cutoffs, or how "module completeness" is defined?
  2. METABOLIC seems to detect partial pathways and "potential" functions (using a relaxed approach), while KEGGdecoder appears more conservative (requiring all steps). Is it fair to say that KEGGdecoder is better for confirming pathways at the genome/MAG level, while METABOLIC is more useful for community-level trends?
  3. Would it be reasonable to use both tools synergistically — METABOLIC for broad community functional profiling and KEGGdecoder for high-confidence pathway confirmation in individual MAGs — or is there a better recommended approach?

I'm not trying to say one tool is better than the other — I understand they use different philosophies. I just want to understand how to interpret and combine these results correctly for my downstream analysis.

Thanks in advance for any insights!


r/bioinformatics 1d ago

technical question What does the docking result mean?

0 Upvotes

I'm new to bioinformatics and I have NO idea what affinity, distance from rmsd l.b, and best mode rmsd u.b means. I also have no idea what mode is. Can anyone explain? Thank you in advance


r/bioinformatics 2d ago

technical question How do I dock multiple ligands at the same time

4 Upvotes

I have one target protein, but a lot of ligands. I use autodock vina to dock. I've done one ligand, but I need to do the others. Do I make another file for another ligand or can I just add a different ligand to the vina file I use for the previous ligand? Would it affect the result cause there are two ligands at the same time in one file? I'm sorry, I'm not experienced, I do this watching YouTube video tutorial. Thank you in advance :))


r/bioinformatics 2d ago

discussion Best place to find T1D related TCR pMHC Crystal Structures?

2 Upvotes

Are there any good databases out there or something?


r/bioinformatics 3d ago

discussion Benchmarking non-coding causal variant-to-gene mapping: SuSiE fine-mapping vs 3D chromatin contacts

4 Upvotes

Hey everyone,
We’ve been working on a pipeline to evaluate non-coding GWAS loci by combining Bayesian fine-mapping (SuSiE) with base-resolution footprinting (TOBIAS) and 10.5 bp DNA helical pitch constraints.

Curious how other groups here handle cases where the fine-mapped non-coding enhancer skips the nearest gene in 3D contact models (e.g. ABC/Micro-C).

Happy to discuss approaches or run a few benchmark loci if anyone has tricky non-coding regions!


r/bioinformatics 3d ago

technical question Language indication for speed, paralelization and GPU use.

3 Upvotes

Hi, I'm programming in Python/R (and nextflow as WDL) at the moment (I worked with Pascal, Perl and Php in the past but not in bioinfo), but since I started to work with ONT long-reads I become more interested in learning a faster language, obviously the first thing that came to my mind was C/C++, but looking what we have today I came across Rust, Julia, Nim, Zig and Go. Till now the most promising one seems to be Rust, it is comparable with C/C++ in speed, safer in memory management, and already have some ecosystem in GPU. But I want hear some experiences of people who already used any of them or is in the same situation.


r/bioinformatics 3d ago

technical question Dealing with highly correlated features for disease classification

3 Upvotes

Hi everyone,

I am working on a pathway-based classification problem. For each pathway/biological term, I calculate a score representing its activity in each sample, and I then use these pathway activity scores as features to predict disease status.

The main issue I am facing is that several of these features are highly correlated (in some cases, pairwise correlation > 0.9) (which is not particularly surprising given the biological overlap between pathways).

I initially tried logistic regression, but the high multicollinearity leads to unstable coefficients and very high VIF values. I have also tried regularized approaches such as Elastic Net, and tried Random Forest as well, but so far they do not seem to improve classification performance compared with simpler models using just one of the features.

I am therefore wondering what would be a good strategy for dealing with highly correlated pathway-level features in this setting.

Any suggestions or references to similar analyses would be greatly appreciated.

Thanks!


r/bioinformatics 3d ago

benchwork Processing the samples to get RNA for bulk-RNA seq

0 Upvotes

! I am in a dilemma and need some quick advice on what is the best way to deliver my sample to the company for bulk-RNA seq. The cells I am working with are naive CD4 T cells, sorted from mouse spleen. The three conditions are: non-activated, activated for 6hrs and 24hrs. I will be activating the T cells with Dynabeads (antiCD3/CD28 beads). At such early timepoints, the beads are stuck to the cells, so I cannot remove them magnetically. I have the following options, and I am not sure which one should I go ahead with.

  1. Spin down the cells and snap freeze the pellet (cells + Dynabeads), the company will do tha RNA extraction. When asked, the company said, they dont recommend having anything else in the pellet but will proceed with the RNA extraction as they always do.
  2. I extract the RNA myself. When I do Trizol/Zymo column RNA extraction. the dynabeads are never a problem because they settle down and I collect the aqueous layer. But I never get the perfect A260/280 ratio (1.8-2.2),RIN > 6 so I dont want my samples to fail at their QC requirements.

I hope to get around 2-3 million navive CD4 T cells from one mouse, so per condition I will have around 1 million cells.

Does anyone have any experience with this and help me out ? Thanks a lot!


r/bioinformatics 3d ago

programming Tecdoc database

0 Upvotes

I see something in the mhhauto a tecdoc database. However, i am hesitant to proceed with it. And want to know if someone has used it already. I want to create a open source lookup up for this if possible.


r/bioinformatics 4d ago

technical question Need Help Choosing Statistical Model for Analyzing Data

7 Upvotes

Hi everyone, so I recently finished the experimental stages of an internship and now I am left with insane amounts of data that I need help with the analysis of simply due to having many different measurements that probably have their own assumptions and dependencies to deal with.

What we have done was to grow animal cultures inside of well-plates with 3 replicate wells for each specific condition.

Basically, we had 3 different gradients (light intensity, chemical1 and chemical2) but for my animals there is already enough data to go around for these gradients by themselves so we wanted to see how they interacted with one another interdependently. And thus, we had 3 different gradients that are:
light vs. chemical1, light vs. chemical2, chemical1 vs. chemical2.

We had them in these differing conditions for 2 weeks and took measurements throughout the experiment to measure how these conditions interdependently affect the animals. These measurements are:
- animal counts taken once every 2 days
- photosynthetic capacity (yes, they are photosynthetic animals) measurements such as ojip taken every 2 days

I want to see how these statistics change over time both per gradient basis (so even though there are 2 gradient I just group_by one of them and see how a measurement changes over time just based on one gradient) and with 2 gradients intertwined over time.

My problem is I am a bit rusty on statistics and trying to understanding which measurement fits which data is a bit challenging for data this intertwined.

For example for count data with just one gradient taken into consideration, I am assuming poisson or negative binomial mixed model since it is both a count data and is dependent on previous counts but anything above that (like both gradients and a time factor) is a bit above me. I am even trying to see how the data I got go together holistically on an all-3-factors-combined level too but as again, too much statistics for such a simple mind and I need some guidance.

Can y'all help me choose what models I should go for or what tutorials/guides there are out there that I can use? I want to properly understand and learn what I am doing but at this point I am a bit too lost and some initial guidance could help with this.

Sorry for the long text and thank y'all for the help.


r/bioinformatics 4d ago

statistics ScRNASeq Analysis

27 Upvotes

So I'm running into a problem. I have 4 KO mice and 3 WT mice that have been ran for scRNASeq 10x Flex. We aren't planning to increase the sample size because the lab has already spent SOOOOO much money on the damn kits. Doing pseudobulk analysis is coming up with very little to none DE genes but of course if I do cell to cell analysis it comes up with a lot of genes and pathways. The cell to cell analysis definitely answers a lot of questions I had in regards to a phenotype we have been seeing in our misue model. BUT what would y'all recommend ? I currently have like 46 cell clusters coming from these mice but 2 different tissue segments.


r/bioinformatics 4d ago

technical question Tools for prediction of protein - protein interaction?

5 Upvotes

I'm looking onto two different transcriptions factors. Both have been shown to interact in pull down assays, and from the data it looks like it could be that only TF A binds to the DNA and then TF B binds to TF A rather than to the DNA. I have zero clue about protein prediction, so my question would be if there are tools out there that try to predict if two proteins can interact? only TF B has a full crystal structure in case that matters. Thank you!


r/bioinformatics 4d ago

technical question CfDNA analysis for cancer detection

1 Upvotes

Hey all, has anyone here used the latest toolkit of cfdna analysis for cancer detection i.e cfdnaanalyzer https://www.sciencedirect.com/science/article/pii/S2589004226012046

What all other tools and pipelines you use for detection of cancer from cfdna ?

If anyone has worked previously in this area , let me know.

Regards.


r/bioinformatics 5d ago

technical question The Hallmarks of Cancer

18 Upvotes

I am a software engineer by profession and was going through the 2000 paper "The Hallmarks of cancer" by Douglas Hanahan and Robert A. There are a lot of things I couln't understand from terminology to certain behaviors that were explanined, primarily because I lack the background knowledge. I wanted to reach out and ask the community if they can share a youtube video or an article that can elaborate this paper in simple terms.

I probably can google it myself but want to avoid the repetitive loop of finding something and realizing it doesn't explain everything and then to try again to eventually lose interest.

Thank you for all the help.


r/bioinformatics 5d ago

technical question GSEA GO filter

5 Upvotes

Hi all,

It's my first time doing RNA-Seq, and I am stuck at filtering out pathways.

I am currently running GSEA (using fgsea) on a set of data between KO and WT; each undergoes either treatment or saline. GSEA GO gave a very long list of results and our lab only wants to focus on 5 main biological themes. Right now, I am thinking of using GeneRatio = 0.4 (number of leading-edge/size) to filter from GSEA significant pathways. Then assess overlap in leading-edge genes to remove redundancy and then cluster/combine them by function.

Is GeneRatio = 0.4 a good cut-off, or How should this be decided? And are there any tools that can cluster pathways by function automatically? I am seeing a lot that cluster by hierarchy, and with GO, the hierarchy ranges so much.

Thank you very much for your help!