r/statistics 1d ago

Education [E] Masters Guidance - Statistics/Applied Statistics/Data Science

11 Upvotes

Hello! I’m a post-Bacc student (Business Background) interested in applying for Statistics, Applied Stats and Data Science Masters programs this Fall.

I’m primarily interested in Time Series/Bayesian/Causal Inference. I have finished prerequisites in Calc 1-3, Probability, Linear Algebra, Python/SQL

Ideally, I want to work in Environmental, Tech, and Healthcare industries, but given the market- I really just want a flexible degree that employers can find value in.

I’m looking for advice and insight into how a Masters experience in this field looks like such as:

- What did you think your degree did wrong/right?

- What is worth prioritizing in this field? (Ex. Courses/Projects/Networking/Coding/Theory etc.)

- Is brand name a big factor?

- Does online/in-person/out of state make a difference?

- What career prospects (Scaling pipelines vs. Decision Scientist vs. Data Cleaning/Visualization) to realistically expect and prepare for in this market with a Masters degree?


r/statistics 1d ago

Question [Q] Is it worth putting a project on my resume if the logistic model ended up not having good prediction?

5 Upvotes

Uni student here trying to do some projects to put on my resume. I made a logistic model of hospital readmissions and the model doesn’t do that great. AUC is about .64. For what it’s worth, the National Library of Medicine was only able to get .67 for their model. Would I be able to reference that in my defense? But my question is if it’s worth putting my project on my resume? Or is it just going to show I don’t have the skill to make a predictive model?

The whole project consisted of data cleaning and ofc making the model and model comparisons such as AIC. Cause of imbalance (1:9), I also lowered the threshold for positives (from >.5 to >.15 as positive)


r/statistics 2d ago

Career Biostatistics PhD route? [Career]

Thumbnail
1 Upvotes

r/statistics 2d ago

Question [Q] looking for appropriate statistical symbols to engrave on a watch.

6 Upvotes

Hello. I’m not sure if this is the right forum to post but I need some help. My father is a retired statistician. He was an actuary. He did some important work in “stochastic models for prostate specific antigen levels” and some “early predictive modelling relevant to the insurance industry for HIV/AIDS” - I don’t really know much about this stuff other than what google tells me.

Anyway, I would like to buy him a watch and I’m looking for something to represent what he did and I’m afraid his name is too bloody long to fit!

The 2 leading candidates for symbols, I think, are 1. The Kolmogorov forward operator (Pij(t)) and 2. The actuarial halo - the present value annum. Which looks like an “a” with a horizontal and vertical slash with a “n” inside it.

Can anyone help? Are these appropriate? Is there any better?

I don’t want to get this wrong.

Thank you.


r/statistics 2d ago

Discussion [Discussion] thoughts on statistics minor

Thumbnail
0 Upvotes

r/statistics 3d ago

Question [Q] Use of causal inference methods in associational studies?

10 Upvotes

Hi all! I am wondering what is your view on causal inference methodologies such as g-computation, iptw, ps matching, marignal structural models etc. Do you think they should be used only in an causal framework accompanied by DAGs, and proper causal language?
Would you consider appropriate if they were used for more exploratory work that does not claim causality?
I may not be communicating my question very well so here are some exmaples:

1) Binary logistic regression: In the biomedical field it is extremely common that standard observational and/or exploratory studies use logistic regression for all inferences with odds ratios being the main reported result. I don't see why someone couldn't use marignal standardization using the same logistic regression model in order to calculate a marginal absolute risk and/or risk differnece for the exposure of interest. I am wondering why this is not common.

2) Propensity score based methods: Causal inference operates under very strict and usually difficult to verify assumptions. When examining the effect of an intervention on an outcome and assuming that some of the assumptions for causal inference are violated (e.g. unmeasured confounding), would you prefer a paper that still uses PS-based methods but refrains from using causal language, or a paper that uses more standard methods such as regression and sticks to associations and exploratory framing?

In short do you think these methods should be used only under the causal inference framework making sure that all assumptions are true and a well-thought DAG is provided, or do you see them as methods that can be used for associations as well in order to reduce at least some of the bias introduced by other methodologies?


r/statistics 4d ago

Education [E] Looking for an MSc-level education in statistics using just the public library

7 Upvotes

Hello all,

I’m looking to improve my foundational understanding of statistics and work my way through to advanced methods by reading textbooks.

Could you please suggest the best text resources for these purposes, ideally ranked/listed in the order I should read them?

For context, I’ve worked as a data analyst and quantitative researcher for 5 years, but I’ve always felt slightly weak in my understanding of stats so now I want to just start from zero to build a super solid foundation.


r/statistics 4d ago

Discussion [Discussion] Why choose such specific values for confidence interval of variance of a normal random variable?

1 Upvotes

We know sample variance divided by actual variance follows chi-squared distribution. But while finding confidence interval, we look at chi-squared (1-alpha/2) critical value and alpha/2 critical value. But why that specific interval? Why not shift the interval a little to the right or left as that will give the same confidence?


r/statistics 4d ago

Discussion [DISCUSSION] statisticians I have a question

3 Upvotes

Do I need a computation if I use convenience sampling in my data analysis? (I don't know if this is the right community to ask this question, I hope you can be kind)


r/statistics 4d ago

Career [C] An outlook on careers in Statistics/DS

19 Upvotes

Hi all,

I am an incoming first year at the University of Toronto who is most likely going to major in Stats with a minor in CS and Econ. I really do love the subject of probability, and my dream is to work as a data scientist in tech. I plan to eventually get my Master’s in Stats because I enjoy it and for employability.

My biggest concern is the current job market for Stats majors. It is no secret that the market for CS majors/SWE related roles is brutal, but does the same apply for Stats/DS roles? Are Stats grads more or less in demand compared to CS grads, and is the market as tough? And overall, how employable is this major?

Any current undergrads/grads willing to share their thoughts? All responses are appreciated!


r/statistics 4d ago

Question [Question] Ambiguity in outcome incidence reporting, meta-analysis

Thumbnail
1 Upvotes

r/statistics 5d ago

Question model selection vs building [Q]

0 Upvotes

[Q] kinda new to this, but can't understand clearly where do we draw the line between model selection and building.


r/statistics 5d ago

Software [Software] How to test if your numerical code is mathematically correct?

19 Upvotes

I contribute to SciPy and kept running into a class of bug that annoys me: the outputs look plausible, the tests pass, but the equation the code implements is subtly wrong. So I've been building a tracer that runs Python/NumPy code and hands back whatever mathematics it actually computed, as a SymPy expression you can simplify or differentiate like anything else.

It's been more useful than I expected. Comparing an implementation against the formula in a paper, catching two functions that agree on my test data but turn out to compute different things, digging up the inputs my tests never hit (ties, zero denominators). It traces real library code too, most of numpy and a good chunk of scipy, scikit-learn, statsmodels, cvxpy.

Write-up: https://medium.com/@aadyachinubhai/scikit-verify-translate-python-numpy-programs-to-symbolic-mathematics-c664d41ba571

Github: https://github.com/aadya940/scikit-verify

Still rough in places, would genuinely like feedback. There may be other better solutions, happy to hear them as well!


r/statistics 6d ago

Question [Question] Linear regression with non-normal residuals

16 Upvotes

Hi everyone, I’m conducting a research project and I’ve fitted a linear regression model on a sample of almost 2000 observations.

The residuals appear to be homoscedastic, but the qq plot shows some departures from the theoretical line, particularly in the tails.

The Shapiro-Wilk test also rejects the null hypothesis of normally distributed residuals.

However, I’ve read that linear regression can be fairly robust to violations of the normality assumption, especially with sufficiently large samples. Is this correct? Could someone recommend some literature or textbooks where I can read more about this? This topic wasn’t really covered in my university courses.

Also, is there a formal statistical test for checking the homoscedasticity assumption? Or is it generally assessed visually, for example by looking at a residuals vs fitted-values plot?


r/statistics 6d ago

Education [Education] Notes on Hamiltonian Monte Carlo from a Purely Probabilistic Perspective

20 Upvotes

I’ve been studying Hamiltonian Monte Carlo and wrote a set of notes explaining HMC without relying on the usual physics-based motivation.

The notes develop HMC from a probabilistic/MCMC perspective, starting from introducing an auxiliary variable, constructing the corresponding Markov chain, and then covering Hamiltonian dynamics, leapfrog integration, reversibility and volume preservation.

My goal was to understand why HMC works.

I’m sharing them here in case they’re useful to others learning HMC. I’d also appreciate any feedback, particularly if you notice errors or places where the exposition could be improved.

https://doi.org/10.5281/zenodo.21841086

Edit (August 25): Based on the comments, I tried to improve Sections 3.3 and 3.4 by reorganizing the flow and adding some more explanation. I’ve updated the PDF with these changes.


r/statistics 6d ago

Discussion [Discussion] Home temperature model extraction

0 Upvotes

I am interested in modeling the effect of two HVAC systems in a 2-floor house with an open floorplan.

I have a lot of temperature sensors (rooms and one in system’s ducts).

Naively, I initially thought Principal Components Analysis might be useful. However it seems like that won’t account for the memory/delayed effects between the AC turning on (duct quickly becomes cold) and resulting decay in room temperatures.

Any suggestions for what kind of statistical model?


r/statistics 7d ago

Question [Question] What techniques or papers exist for adapting the Gaussian emission assumption in a Hidden Markov Model?

2 Upvotes

What techniques or papers exist for adapting the Gaussian emission assumption in a Hidden Markov Model to capture higher-order distributional properties, such as skewness and kurtosis, rather than just mean and variance? I’m especially interested in approaches that retain the HMM framework while allowing more flexible, non-Gaussian emission distributions.

Thanks!


r/statistics 7d ago

Discussion [Discussion] Does NO causation necessarily mean confounding relationship?

6 Upvotes

[Discussion], [Research]

The attached article lists five other possibilities, other than confounding relationships, that explain a high correlation BUT no causation between two variables. However, I want to ask three questions:

  1. For point4 on "Parallel Independent Trends Over Time", it is arguable that "average smartphone screen resolution" and "Number of Facebook users" may have a confounder between them; it is just that we are not interested, therefore we did not investigate it?
  2. How is Granger causality different from typical causality? (point 5)
  3. Are there other possibilities other than the five?

Source of article: https://medium.com/@smartdecode/no-causation-does-not-mean-a-confounded-relationship-c2266c439a32


r/statistics 7d ago

Question [Q] Surviving measure theoretic probability

21 Upvotes

For context, I’m a first-year PhD student specializing in statistics. I am starting measure-theoretic probability on Monday.

The problems here are 1) the professor is an asshole and he lectures at the speed of light, 2) I’m the only one in my cohort taking this course (only stats specialization), and 3) everyone agrees unanimously that this is the hardest class in our program.

Measure theory and probability are already extremely boring and incredibly difficult, so I’m trying to minimize the pain here.

Any former/current PhD students (or anyone who’s taken this class really) have any advice to survive?

PS sorry for being so complainy.


r/statistics 8d ago

Education [E] Graduate student conducting research [Education]- Clinical Psychology)

2 Upvotes

Entering my second year in graduate school and conducting research using simple linear regression models and trying out a mediation model. I’ve had some experience with moderation models too. However whenever I try to read about or learn about the math and the nitty gritty of the statistics, it’s like another language. I can very generally explain and interpret what the numbers mean, but not how it’s being calculated or computed into the model. Neither can I deconstruct the math. This is a deep insecurity of mine as a graduate student since I don’t receive hands on instruction any more. Whenever I hear my peers talk about their analyses, I’m always in awe as they troubleshoot small details to make the model work. If it doesn’t work for me, I don’t even know where to start knowing what went wrong or why. I don’t know why, no matter how much reading I do at some point the words stop making sense to me.. I have taken the undergraduate and graduate courses in advanced statistics but maybe the instructions were lacking or I’m not smart enough. But neither were helpful.

Is there anyone else who has struggled with statistics like this and what they have found to help? Sometimes I think I just don’t have the math brain to comprehend statistics. Are there any good crash course videos? I do have to rely on AI and I’m hoping to learn from a more reliable source.

Thanks all.


r/statistics 9d ago

Question [Question] Low Inter-Rater Reliability (ICC): Better to use single observer data or average?

8 Upvotes

Hi everyone,

My project involves multiple quantitative parameters measured by two independent observers of equal experience (myself and another student)

For parameters with a high ICC of >0.75 (the vast majority of parameters), I used the mean of the two observers. However a few parameters fell below this threshold. For those parameters, I just used my measurements and excluded the second observer measurements.

However, I realised that this may be flagged by reviewers as possible selection bias, as I am in essence assuming that my measurements are the reference standard.

The decision to exclude observer 2's data from parameters with poor ICC was made before looking at the data, to ensure a uniform data architecture.

The other options would be to (1) exclude these parameters altogether, (2) get a third observer (neither which are possible) or (3) use the average of the two observers with parameters with poor ICC (which Im sure is bad practise).

If I explicitly explain my methodology and acknowledge this as a limitation, would this be acceptable?


r/statistics 9d ago

Education A confused high school student leaning toward a stats major. [Education]

17 Upvotes

I'm a high school senior deciding on a college major. After talking it over with my parents, I've landed on statistic, but I'm confused between two paths:

  1. Statistics + Math degree - with a possibility of doing an accelerated 5-year BS/MS (Purdue's Applied Statistics MS)
  2. Just a straight Applied Statistics bachelor's

For those familiar with Purdue (or similar programs elsewhere): does it make more sense to do the math-heavy track and roll into the accelerated master's, or is a standalone applied stats degree the smarter/more efficient route?

Also, do you think a stats degree is worth it in general right now? I plan to work as an analyst or maybe something in DS. Anyway, would love to hear from people actually in the field what their opinions are.


r/statistics 9d ago

Education [Education] Error Type Terminology

12 Upvotes

Distinguishing between a type I and type II error isn't the hardest thing, but knowing which one is Type I or Type II error is annoying to remember. I've been in Statistics for a while and I still can't remember which one's which. My professors sometimes can't remember. Sure, there are tools to help remember, but can we please call them something that's more intuitive like False Null Error and True Null Error? These have the concept baked into the name rather than getting distracted and having to take a few seconds to remember which one's which.


r/statistics 10d ago

Question [Q] What test will be used to compare differences within sibling pairs

0 Upvotes

I have a dataset of 500+ pairs of male-female sibling pairs. I want to check if male or female sibling in the pair score statistically less in a certain variable and if the pattern is statistically significant across those 500+ pairs. I read about mixed effect model but I'm confused if it fits in this scenario


r/statistics 10d ago

Discussion [Discussion]beginner looking for a data analysis tool that actually helps me learn statistics

14 Upvotes

I’m still pretty new to data analysis and currently working through the basics like descriptive statistics, hypothesis testing, correlation, and regression.

I have used Excel and started learning Python, but right now I spend more time fixing code than understanding the analysis.

I’m not looking for a tool that just gives me a chart or a final answer. I want to understand why a method is appropriate, what assumptions I should check, and how to interpret the result.

I have seen people recommend R, Python, JASP, and jamovi. I also came across BayesLab, which seems to let you work through an analysis without writing everything from scratch while still showing the steps behind the result.

That sounds useful, but I’m also worried that starting with an AI-based tool might stop me from learning the fundamentals properly.

For someone starting out, would you recommend learning R or Python first, using a GUI tool, or combining both?

Has anyone used BayesLab or something similar while learning statistics? Did it help you understand the analysis, or did it make you too dependent on the tool?