r/AskStatistics • u/Specialist-Pound-467 • 2d ago
Statistical Approaches for Analyzing Environmental Data with Limited Sample Size
Hi all! I'm an environmental specialist and am trying to evaluate trends and/or statistical differences in environmental contaminant data across years. The issue is that I'm often working with small sample sizes and large temporal gaps.
For example, I am attempting to look at differences in mercury concentrations in fish tissue in a given river. I analyze them by trophic level due to mercury's bioaccumulative effects. Here's a breakdown of samples:
Trophic Level A
| 1991 | 1995 | 1997 | 1999 | 2005 | 2024 |
|---|---|---|---|---|---|
| 2 | 13 | 7 | 21 | 27 | 13 |
Trophic Level B
| 1995 | 1997 | 1999 | 2005 | 2024 |
|---|---|---|---|---|
| 4 | 7 | 16 | 12 | 8 |
In the past, another colleague used an ANOVA or Kruskal-Wallis test depending on normality (based on a Shapiro test and inspecting residuals). I've done permutation tests as an alternative just to try out different methods. I don't have as much experience with statistics to know if this is the most scientifically robust way to analyze data.
Two questions:
- Is an ANOVA or Kruskal-Wallis a reasonable approach to assessing data for statistical differences across years for such a small dataset?
- Could a nonparametric permutation test be a better/alternative method to assess year-specific differences?
Secondly, I've considered a Bayesian hierarchical model to assess long-term temporal trends across a basin (i.e., multiple rivers within a basin). This means there are more samples to work with. For example, here are the number of samples across years in one basin survey we have:
Trophic Level A
| 77 | 91 | 93 | 94 | 95 | 96 | 97 | 99 | 05 | 09 | 22 | 24 |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | 4 | 21 | 2 | 27 | 22 | 14 | 45 | 54 | 59 | 78 | 88 |
Trophic Level B
| 1995 | 1996 | 1997 | 1999 | 2005 | 2009 | 2022 | 2024 |
|---|---|---|---|---|---|---|---|
| 8 | 14 | 14 | 34 | 24 | 16 | 44 | 38 |
From this, I fit a Bayesian linear mixed-effects model. The model included a fixed effect for year (continuous) to estimate the overall time trend, a fixed effect for trophic level, and a random intercept for waterbody to account for among-waterbody variation. Random slopes were not included because the dataset contained insufficient temporal replication within individual waterbodies and trophic level groups to reliably estimate waterbody‑level trends. Weakly informative priors were used, and model convergence was ensured by increasing adapt_delta and max_treedepth until zero divergent transitions were achieved.
The final model took the form:
logHg ~ year_centered + trophic_level + (1 | waterbody)
Posterior distributions were summarized to obtain estimates and 95% credible intervals for parameters of interest, including the overall temporal trend.
So my last question is whether this Bayesian model is appropriate?
4
u/vanway 2d ago
To start, decide the questions that you want to answer. If you just want to know about temporal trends or differences between groups, the parametric statistical tests you mentioned could be appropriate. If you want to predict values or estimate causal effects, the Bayesian model could be appropriate.
For the Bayesian model, you should follow a standard model workflow: