r/statistics

[Question] - Which wording is best?

I wrote in my abstract that “Medication X was not associated with significantly different subjective or objective sleep outcomes compared with medication Y.”

One of my co-supervisors suggested instead saying that “Medication X and medication Y showed similar subjective and objective sleep outcomes.”

Would you make that change? I’m hesitant because, methodologically, I’m not sure that a non-significant difference allows us to conclude that the outcomes were “similar.” Or am I overthinking this?

Thanks a lot! 🙏🏻

reddit.com
u/ElOctopusDeBadia — 20 hours ago

[C] Continue in industry or PhD in stats? Are all stats jobs this boring?

Hello everyone, I'm looking for some guidance. For context, I am n my mid 20s, I hold a master's in stats, and I have been working for a big CRO in Europe for 1 year. I'm mainly looking for feedback from people based in Europe, but any feedback is welcome!

Honestly, I'm not enjoying the work. I feel like only about 10% of what I do is actual statistics. The rest is monotonous and boring work, constantly churning out deliverables, dealing with clients (something I deeply dislike), working on timelines, regulations, administrative stuff, endless meetings, and so on. I miss programming, I miss doing more intellectually demanding work, and I miss engaging with new methodology and having the liberty to do so. I feel like my knowledge is atrophying. When there is a new interesting problem to solve from a statistical standpoint, I never have time to actually think about the methods, it's just constant pressure to deliver, deliver, deliver.

Are all stats jobs in pharma/CROs like this? For people working in data science, stats, and adjacent areas, is your job this repetitive? Do you have the chance to apply new, cool methodologies? How much time do you actually spend working with the methods rather than doing something over and over again?

I know how fortunate I am to have a job in this economy, especially one that pays well by my country's standards given my experience. Still, I've been considering applying for a PhD, and deep down I know it's something I really want to try. I'm aware of how difficult a PhD can be. I've been looking at positions in other wuropean countries that would allow for a decent standard of living and, even some savings.

My question is, would I be committing career suicide if I later wanted to return to industry? One of the main reasons I want to do a PhD in stats is that I want to move into more R&D roles, rather than doing this kind of mind-numbing work. Maybe I could even pivot into tech or some other area that is more interesting to me. Would that be a realistic prospect?

Thanks all!

reddit.com
u/Lis_7_7 — 1 day ago

[Q] Multiple statistical tests in one table?

In completing a study on left and right limbs, data came out that was parametric on left limb and non-parametric on right meaning using a rm-anova for left and Friedman's for right. Would putting these results in the same table make sense as long as it is clear what was done for each and what they mean? Or is 2 different tests in the same table a definite no?

reddit.com
u/Agent56570 — 2 days ago
▲ 3 r/statistics+1 crossposts

[Question] [Q] Trying to calculate whether marketing campaign's impact is statistically significant and financially justified

I have a data series of daily streaming counts for a song. My spreadsheet has the date in column A and the number of streams in column B. There are 960 rows/records in the spreadsheet dating from 2024/01/01 to 2026/08/17.

We hired a marketing company to promote the song for 3 weeks. Their first posts on social media started on 2026/07/27. They ran until 2026/08/17. I am trying to assess whether our money was well invested or not. Their dashboard stats are not useful to us because they show the number of views and social media engagement (e.g., TikTok), whereas we are interested in the number of times our song gets streamed on a streaming platform (e.g., Spotify).

I know how to calculate means and standard deviations, and have done so for various timeframes (e.g., yearly, during the promotion, the 22 days prior to the promotion, etc) but I do not know how to:

  1. determine if the 22-day marketing campaign had a statistically significant impact on our daily streaming numbers

  2. determine if the impact on streams, if any, justifies the money invested, call it **CampaignCost**.

Can someone advise me how to go about this? We might assume somewhere between 0.001 and 0.003 USD of revenue per stream.

I know there is surely some well known statistical test to determine whether my daily streams show a statistically significant increase or decrease, but I cannot remember what this test is called or how to calculate it.

I might add that there are potential complications:

  1. Streams have grown significantly year over year. Average daily streams in 2024 were 135k, in 2025 were 250k, and so far in 2026 are 259k.

  2. Some unexplained events in the real world have prompted surges in streaming numbers. E.g., our streams surged upward for a month around December 2024 and remained elevated for months then gradually declined. Another unexplained surge arrived around Feb 5, 2026 and persisted for months but then started gradual decline.

  3. The streams show a weekly cycle, lowest on Sundays, peaking on Thu or Fri.

Any help would be much appreciated.

reddit.com
u/sneaky_imp — 1 day ago

[Q] How to compare data from 3 periods of time

Data of event turn out, annual totals, collated into 3 groups (before organisational change, during organisational change, after organisational change). Originally planned to use chi squared goodness of fit, but one year of data is missing, so groups are now 2 years of data, 3 years of data, 3 years of data. So thinking of calculating mean average for each group to equate them. But then what stats to compare/assess any significant difference, anova? Is there a better way I am missing?

reddit.com

[E] Just a noob here — need to learn the basics of Statistics [E]

Hi everyone,

I'm currently doing a course where I have five major topics to cover, but I'm starting with these two:

  1. Comparison of means between two food-diet groups using a t-test
  2. Comparison of means among more than two food-diet groups using ANOVA

The remaining topics are:

  1. Post-hoc statistical analysis for identifying significant group differences
  2. Development of a General Linear Model (GLM) in Linear Regression
  3. Understanding Adjusted Sum of Squares and Sequential Sum of Squares in Linear Regression

My background is in biology, so I'm pretty new to statistics. I know how to run basic analyses in Minitab and SPSS, but my problem is that I don't really understand what is happening behind the numbers.

I want to properly understand the basics first — things like mean, median, variance, standard deviation, standard error, distributions, hypothesis testing, p-values, confidence intervals, etc. — and then build up to t-tests, ANOVA, post-hoc tests, and regression.

This is not really a theory-heavy course, so I'm expected to understand and interpret the statistical output rather than just calculate everything manually. But I don't want to blindly click buttons in SPSS/Minitab without understanding what the results actually mean.

Could anyone recommend a good beginner-friendly statistics textbook, notes, or learning resource that would help me build these fundamentals from scratch and eventually reach the topics above?

My professor isn't particularly helpful with recommending resources, and my seniors are currently busy with their own work, so I'm trying to find a good starting point myself.

Thanks in advance to anyone who takes the time to point me in the right direction! 🙏

reddit.com
u/Common-Tax-7833 — 2 days ago

Wilcoxon Signed-Rank Test Chart - but for a big study [Q]

Help! My sample sizes are is 40-82, and I can only find Wilcoxon test charts that go up to n=30! I've looked everywhere and can't find anything that can accommodate a bigger n. Does anyone know where to find a bigger chart?

reddit.com
u/Square_Structure5094 — 2 days ago

[R] Proof of the Riemann hypothesis just dropped!

This work is devoted to the proof of the Riemann hypothesis. The theory of compact Jacobi matrices with the off-diagonal elements given as a moment sequence of an absolutely continuous measure is constructed. Various operators and resolvent functionals associated with this class are introduced and studied.

https://zenodo.org/records/21907861

reddit.com
u/rasstrelyat — 2 days ago

[Q] New to data cleaning. I’m stuck on an unclear variable.

I’m doing a personal project right now and for the most part it’s going alright. Sometimes I delete an entire column cause too much is missing. I’ve also group together a few entries as “Unknown”. I’ve never deleted a subject yet.

Anyways, now I’m really stuck. There is a variable that is three digits (345, 078, 150…) and some of them come in as two digits (26, 88) and I’m unsure what to do with them. There are quite a few. I don’t know if they’re meant to be have a zero on the front (28 turns to 028 and 88 turns to 088) or if they are three digits (28 into 280 and 88 into 880). It’s an important variable so I can’t delete it. Should I delete the patients (probably not), group the two digits numbers into “unknown”? any input? I know each data has their own situations but what is something to generally consider in these situations?

Edit: The column represents diagnosis IDC-9 number. It’s three digits. I want to use it to build a logistic model. I was thinking of grouping the numbers to their diagnosis category for example 001–139 is infectious diseases, and so on

reddit.com
u/Unalina — 4 days ago

[Q] Are these courses feasible for first year of my MS?

**First semester!! I have only taken Mathematical Statistics I and II (Not sure if the content of those courses are universal, but it covered interval estimation, hypothesis testing, and tests involving means, variances, & proportions). I have signed up for the following:

Experimental Design

Linear Statistical Analysis I

Pattern Recognition in ML

Advanced Matrix Analysis

Wondering if these work well together/ may be too much.

Sorry in advance if this is silly, we have somewhat limited options at my school.

I have a BS in pure math for reference.

reddit.com
u/pinkfaerie0 — 4 days ago
▲ 13 r/statistics+1 crossposts

[E] Confidence Intervals — Explained

Hi there,

I've created a video here where I explain what confidence intervals are and how they differ from probabilities.

I hope some of you find it useful and as always, feedback is very welcome! :)

u/Personal-Trainer-541 — 6 days ago

[Question] Can I present both the results of log-odds and average marginal effects (logistic regression)?

Hello,

Health economics student here. I newly enter the field for my master degree.

I'm writing my master thesis and I'm running into an issue while trying to interpret my results.

Initially, I've decides to use OR (odd ratio). However, my supervisor told me the way I've interpreted it is not correct (probabilitied and chance). But he told me he didn't really know either how to interpret it correctly and had advised me to use logs odds instead!

It's ok, but I really wanted to have more "concrete results" that can be get by anyone. Log odds just show if the relation is negative or positive.

So, I've just heard about average marginal effects and that is literally What I was searching for during all this time. However, now I'm wondering if is common for scientific papers to use both log-odds and average marginal effects?

Do i need to create two tables? Or could I only keep the table with log odds and present the results of marginal effets in the text?

Thank you

reddit.com
u/Legitimate_Mud_9245 — 6 days ago

[Q] If you're in grad school for Stats (PhD or Master's) and your undergraduate was in math: 1) what do you miss about math; 2) what are you gad to have traded with stats?

Can be silly or serious, just out of curiosity for someone with a background in math contemplating stats. Like for #1, maybe you miss not having to deal with numbers. #2 refers to "trading" X in math for Y in stats (like numbers).

EDIT: "glad", not "gad".

reddit.com
u/plop_1234 — 7 days ago

[E] Randomness can be an asset or a tax depending on curvature

I wrote a short article here exploring a simple way to think about when randomness helps versus hurts.

The core idea is Jensen’s inequality: if the payoff is convex, variance can help; if it’s concave, variance can hurt.

I use examples from compounding, option-like payoffs, and a few non-finance settings.

Would be curious to hear other examples where this framing is useful?

reddit.com
u/Due_Raspberry_6269 — 5 days ago

Submitting table as image <440 pixels wide [Question] [Q]

Hello,

I am trying to submit my article for publication. Unfortunately, the journal asks for any tables to be submitted as images "provided as 72 - 300 dpi; pre-sized .BMP, .GIF, .JPG, or .PNG images only, with a maximum width of 440 pixels (no limit on length)."

I have tried exporting my table from excel to pdf, jpg, or png, and then resizing but no matter what I try, the image of the requested size ends up unreadable.

Does anyone have any ideas on how to accomplish this requirement while keeping my table-figure as readable?

reddit.com
u/idontcareenoughatm — 6 days ago

[Q] How should I do a Bayesian Update?

I'm a year 1 liberal arts undergrad, so I don't have a ton of math sense.

I just learned about Bayes' Theorem the other day for general epistemic use. I worked out a couple of example problems correctly, but the examples I found didn't include any iterative updates.

I know the theorem is:
P(H|E) = P(E|H) * P(H) / (P(E|H) * P(H) + P(¬H) * P(E|¬H))

And I know that P(H|E) becomes the new P(H) in my update, but I'm unsure whether I should be using the updated or original P(H) in the marginalization. I *think* it should be the new P(H), but I'd rather be safe than sorry.

The example question I worked out was this:

________________________________________

There's a disease that afflicts 1 / 1,000,000 people
There's a test for the disease that's right 99 / 100 times for both positive and negative results
A random person is tested as positive

P(she is afflicted | she tests positive)

= 0.99 * 0.000001 / (0.99 * 0.000001 + 0.999999 * 0.01)

= 0.000099

_________________________________________

So should the update look like this if she tests positive a second time?

P(she is afflicted | she tests positive)

= 0.99 * 0.000099 / (0.99 * 0.000099 + 0.999901 * 0.01)

= 0.009707

_________________________________________

If so, how should I approach this problem from the starting point of
P(she is afflicted | she tests positive twice)?
I can't think of how to handle that correctly, since squaring 0.99 just gets me a smaller number

Infinitely thankful <3

reddit.com
u/Xancrim — 7 days ago

[Discussion] Real Analysis (1 semester vs 2 semester sequence) for Statistics PhD Applications

Hi everyone, I’m applying to Stat PhD programs this fall. I graduated with my master's in 2021 and have been working full-time for 5 years, but I never took Real Analysis in school.

I just enrolled in Fordham's Math 3003 (Real Analysis) this semester so it’ll be on my transcript for this application cycle (though will not have a final grade since apps are due before the semester ends). My concern is that it’s only a one-semester class. Does anyone know if admissions committees strongly prefer a two-semester sequence (Real Analysis 1 & 2) over a single semester? Will this be adequate?

MATH 3003. Real Analysis. (4 Credits)

This course focuses on analysis on Euclidean spaces. Topics include limits, continuity, uniform continuity, sequences of numbers and functions, modes of convergence, differentiability, Riemann integrability, and associated theorems. Students who have not taken MATH 2004 prior to taking Real Analysis may request permission from the instructor. Note: Four-credit courses that meet for 150 minutes per week require three additional hours of class preparation per week on the part of the student in lieu of an additional hour of formal instruction.

Prerequisites: MATH 2001 and (MATH 2004 or MATH 2008).

reddit.com
u/Usual-Recipe-5415 — 7 days ago

[Q] I am looking for some VERY INTERESTING and catchy statistical concepts/paradox/theories for a presentation. Can yall suggest some?

the more unpopular, the better. But still very catchy and interesting. Thank you in advance :)

reddit.com
u/EvilSiren_03 — 10 days ago

[Career] explore/exploit problem as an undergrad

Hello!

I'm a sophomore undergrad studying applied math and stats who is broadly interested in applied stats research. Essentially, I'm trying to figure out how to balance exploration to find what exactly I want to do research in with deep-diving in order to develop skills (and resume) for grad school.

What is scary to me is how little time it feels like I have. Freshman year I didn't do anything that is particularly useful for grad school. So I now basically have two years + two summers before grad school apps. If I spend sophomore year doing something that doesn't end up being my interest, then I worry I'll only have a year to do what I actually end up doing.

I know I love solving problems with statistics and do want to continue developing my stats skills (there is even a good chance I'll go for a PhD in stats). Part of me thinks I should just focus on stats because I know I enjoy it and it'll transfer to whatever I want to do. The thing is though that as much as I love learning about stats and reading stats papers and so on, I don't want my research to be developing new statistical techniques in the abstract, I want to work with problems in the real world and develop statistical methods of attacking them. The thing I am most interested in now though is comp neuro, but I know very little about it currently. However, I worry that if I start taking neuro classes or even spend a summer doing neuro research and decide it's not what I want to do (which I basically did got econ already), then I'll have wasted the time that could have gone to learning mote stats.

My current plan is to dedicate all my coursework and formal research to stats for now (I have a position in a stats research group for this year and reading the papers they've put out it seems pretty lit!), and then do some independent reading of papers in other fields to decide if I am drawn to them​. While this is great, reading about something is very different from doing it so I fear it may not be possible to explore optimally without taking at least some risk. Thoughts on how I should navigate this?

reddit.com
u/MediumLog6435 — 6 days ago