r/datascience

Interactive tool for learning ML System design for free

Interactive tool for learning ML System design for free

I know data scientists are increasingly being asked to own models end-to-end. The problem I see (specially with juniors) is that they jump straight into Docker or other MLOps tools without building the foundation first.

I’ve been in data science for 8+ years and I think the best way to start is with ML system design.

I’ve gone through different resources over the years like Chip Huyen's "Designing Machine Learning Systems" and realized that learning system design just by reading a book or staring at diagrams is really hard. You don't really get the intuition until you can see how the components actually connect and behave.

So I built a free interactive tool based on a real system I deployed. It walks you through how the system was architected so you can build the intuition to design one yourself.

Here it is: https://futureproofds.com/tools/ml-system-map

A few things you can do with it:

  1. Play a flow and watch it run step by step (training, serving a live prediction, the nightly batch, a drift alert firing)

  2. Click any component to see how it works and why it matters in the system

  3. Follow the build order to see how the whole system comes together at different stages

Hoping some of you find it useful. Would love to hear what you think, and let me know if there's any functionality you want me to add.

u/avourakis — 15 hours ago

Another rant like interview experience

I was given a home assignment to do modeling for some adtech data. They had no explicit ask about what kind of model or how deep you have to go. Just data and they asked we want to see the modeling.

I spent lot of time in understanding the data, identifying the features, creating labels etc. When it came to modeling I picked Catboost since they handle categorical features quite well. I even mentioned how this can be further tuned and/or different models can be compared. I put it explicitly in a section for future work. Finally this was the thing that got me rejected.

Basically they expected me to compare different model families from more complex deep models to such boosting models. I have worked in this domain and actually such models (catboost) works quite well. You don't need very complex models. I remember in one of the previous jobs, they had like ensemble of 3 deep models which was super slow and was so painful to maintain. I basically replaced that with a boosting model + some probability calibration which did quite well. Also the data size I got for the task isn't big enough to justify such huge models.

In any case, I wish these tasks would be more explicit in what they are looking for. I know they also want to see how I handle ambiguity but it's really hard to assess which side of it is worth handling since I am not building a full fledged system. I explained all the decisions I made and why I did it. Also what I didn't do and why.

reddit.com
u/proof_required — 1 day ago

How does one prepare for such interviews?

https://preview.redd.it/99irhgolmzjh1.png?width=622&format=png&auto=webp&s=99a2c4b9cec59a17af5ff650cb38720f373095d0

I see posts like these on my Linkedin feed every day. At this juncture, I am not sure if this is true or just one of those AI Slops - I am assuming there's a grain of truth in them.

But now, when I am preparing for interviews and job hunting, I don't think I could have ever imagined answering it in this way, unless I have worked on specific/adjacent use cases.

How does one prepare for such questions?

reddit.com
u/JayBong2k — 2 days ago

Weekly Entering & Transitioning - Thread 17 Aug, 2026 - 24 Aug, 2026

Welcome to this week's entering & transitioning thread! This thread is for any questions about getting started, studying, or transitioning into the data science field. Topics include:

  • Learning resources (e.g. books, tutorials, videos)
  • Traditional education (e.g. schools, degrees, electives)
  • Alternative education (e.g. online courses, bootcamps)
  • Job search questions (e.g. resumes, applying, career prospects)
  • Elementary questions (e.g. where to start, what next)

While you wait for answers from the community, check out the FAQ and Resources pages on our wiki. You can also search for answers in past weekly threads.

reddit.com
u/AutoModerator — 3 days ago

How widely is R still used in industry today?

I’m a Data Science student (career changer, not in a data related role). My program is focused more on the applied statistics side, so most of my classes use R. I’m already familiar with Python since it was the main language used in my prerequisite courses, and I’ve completed projects using Python, so I’m comfortable with the syntax.

However, I’m really enjoying using and learning R in my classes and seeing what it can do. Many of the statistics textbooks I’m interested in use R as well. I’m starting to explore R more deeply on my own and plan to start using it for personal projects.

But I’m curious, is R still used in industry? I know it’s heavily used in academia. I also know that in the current AI/ML world, Python is used heavily, which is the main reason I use it for all of my personal projects at the moment.

I’d like to eventually be comfortable with both and take advantage of the strengths of each language. But, of course, there are also people who say learning R is a waste of time.

reddit.com
u/Kati1998 — 6 days ago

Typical question in the first interview?

I have a 30minute zoom meeting for a data science job and I'm just wondering what types of questions others have been asked in these interviews?

I had one a couple months ago and they did ask me a SQL question but that was the only technical one I can remember

Edit: Interview finished and doesn't look like I got it y'all! I'm a fucking idiot! There was absolutely no technical questions, just "Tell me about yourself" "What's your experience with python" "Walk me through a project" "Do you use Generative AI"

I'm not entirely sure how I messed that up but I guess my charisma stats are that low

reddit.com
u/WhatIsMyNamme — 7 days ago
▲ 284 r/datascience+1 crossposts

Laid off after 4.5 yrs at the company as Sr Data scientist. How is the job market ?

PhD computational Physics from USA and 3 yrs of Postdoc in the USA. Transitioned to DS in early 2022. Mainly worked with Text data (embedding related word2vec to Transformer based, AI solutions too but Prompt based no agent based solution), Traditional ML & NeuralNets for classification and regression. Python, SQL and PySpark tech stack, AWS & snowflake platforms. Comfortable with either Linux/Unix or windows.

  1. How is the job market ?

  2. What are the chances of finding Job by end of my 2-3 months of severance ?

  3. What should I prepare the most ? How shall I approach the job market?

Currently remote at a decent Midwest city.
Any suggestions and advice will be appreciated.

Thank you

reddit.com
u/dead_n_alive — 8 days ago

Data Science in manufacturing vs IT/consulting

I’m currently working at an IT company and will probably be leaving soon. I’m already talking with companies in banking, IT and consulting, mostly for roles close to my current experience.

But I also got an opportunity at a large factory with a small data science team. From the initial talks, their work seems to be around sensor data, predictive maintenance, anomaly detection, safety, maybe some computer vision. They manufacture some machines, so it sounds quite different from my usual IT environment.

Most of my recent work has been around LLMs, agents, GenAI, etc. I know this area pretty well, but I’m not sure I’m passionate about doing mostly that long term because of the hype. I still find things like gradient boosting, computer vision, time series and more traditional ML problems really interesting.

So I’m curious about people who have worked in manufacturing DS/ML. What is the culture and day-to-day work like? Is it generally calmer than IT/consulting, or does production bring its own kind of pressure? How is the work-life balance?

Career-wise, would moving into industrial ML be a risky switch in the current AI market, or could it actually be a good way to build a more specialized ML background? Also, what skills would you recommend learning for this kind of role?

reddit.com
u/missing-in-idleness — 8 days ago

I'm curious about people working in ranking and if you can change customer behavior

Basically I have a ranking service for b2b SaaS but basically like hotels flights etc

The models do well and I can improve accuracy pretty easily to a point

But if I want to promote options better for other metrics I'm struggling to change behavior other than people selectng the same thing lower

Just hoping for experiences for those in ranking specifically and anything they might have tried other than traditional lighting ranking etc

reddit.com
u/Xamius — 8 days ago

Just used AI for the first time. Need your advice.

I've been a data analyst since before the recent AI boom. At my previous company, AI use basically meant pasting SQL into ChatGPT and asking it to fix, join or optimize queries. It wasn't connected to our warehouse, so I still had to do everything myself.

I've now moved to a much larger company where Claude/Hex are integrated with our warehouse and semantic layer. The difference is insane. I can describe what I need and it finds the right tables/columns, figures out joins, writes and executes the SQL, explores the output, checks nulls/value distributions and helps validate the result.

It's incredibly productive, but it has me wondering:

  1. Am I deskilling myself? If AI writes my SQL every day, won't my ability to write complex queries from scratch eventually deteriorate? It sometimes feels almost like cheating

  2. What does this mean for data careers? If AI can already write SQL, explore schemas, analyze outputs and perform basic data-quality checks, how much of traditional analytics work remains?

  3. Should I automate everything with AI? Should analysts be trying to automate as much of their workflow as possible—SQL, analysis, emails, meetings, Jira, documentation, etc.—because people who don't will simply fall behind?

reddit.com
u/informatica6 — 11 days ago

Tips for Getting Information from Colleagues

I recently started working in a data scientist role for the first time, pivoting from mathematical ecology. (It's actually at an environmental organization, so the fit is great.) The job is hybrid, mostly remote. So far, it's been going really well.

Last week, they asked me to do a power analysis of a planned study. (Yay!) Of course, this requires a lot of information about measurements, expected values, outliers, what size change would be of interest, etc. I asked the necessary questions on Slack, along with some follow-ups and reminders. They were able to get me much of the information I needed and I found some in the literature, but it felt like I was bugging people (including my boss). Does anyone have communication tips on getting this kind of info from colleagues?

reddit.com
u/jaiagreen — 9 days ago

Attempted to apply creative writing skills to an explainer of Markov Chain Monte Carlo. Tell me how bad I did 😅

Lately I've been deep in a personal project by writing chapter summaries of Richard McElreath’s Statistical Rethinking textbook and applying them to wildfire models, and somehow found a way to elegantly (in my opinion) combine the two through storytelling. The tl;dr: I built a whole narrative around a wildfire forensic investigator named Prof. Markov, rolling an eight-sided die to decide which direction to search a burnt forest grid, to explain how the Metropolis-Hastings algorithm (the earliest variant of Markov Chain Monte Carlo (MCMC)) actually works.

MCMC sits at the foundation of modern Bayesian computation and probabilistic programming frameworks like PyMC and Stan so it could be genuinely useful to anyone looking to level up in these topics. Roast me, tell me what you liked and didn’t like. Regardless, it was a fun little mini-project!

https://pub.towardsai.net/explaining-markov-chain-monte-carlo-using-wildfire-forensics-a334fecaefb3

reddit.com
u/vanisle_kahuna — 8 days ago

Anyone else struggling to balance coding yourself vs. letting AI do it?

Since I got access to Claude at work, I haven’t really written much code from scratch, especially for ad hoc analyses or quick charts. I still review the code and make sure I understand everything, but it’s honestly a little scary how much better Claude’s code often is than mine. At that point, it’s hard not to wonder what the value is in writing it yourself.

On top of that, management is encouraging us to use AI to be more productive and deliver results faster, so there’s that pressure as well.

To keep my interviewing skills sharp, I still practice on LeetCode or StrataScratch from time to time. But at work, I’ve been relying on Claude pretty heavily.
Is anyone else dealing with the same dilemma?

reddit.com
u/Fig_Towel_379 — 13 days ago

How do you design a forecasting system?

Hey y'all! How do you design your forecasting system?

In my case, the company has many SKUs over a big region. We did an MVP to show our forecast improves the current process on the reported lags that are currently used by the business to monitor forecast health.

Future is looking good, but I really want to be ready with a production-grade plan. Refitting a pool of models per SKU every week, then selecting the best one, feels like overkill and very sensitive to recent flukes.

I thought of having a pool of models (i.e. config/setups) and labelling them as champion if a specific config results in the best trained model.

For the next X weeks this model will always be chosen, and after that the throne is up for grabs.

But it kind of railroads me into having a 1 SKU = 1 model setup in perpetuity.

How do you guys solve this in a responsible way? Are there books/resources you recommend?

Reasoning about a live system turns out to be a whole different cookie than the usual stats/ML etc

reddit.com
u/Berlibur — 13 days ago

Embeddings

Hi folks,

I've been thinking a lot about where embeddings and foundation models are taking data science.

I work in the geospatial/Earth Observation space, and honestly it feels like the landscape has shifted massively over the last few years. We're seeing more and more open source foundation models that are so good you can often just extract the embeddings, stick an XGBoost or regression/classification head on top (or do a light fine tune), and get really strong results. A few years ago I'd have expected to spend most of my time building models and engineering features. Now it increasingly feels like the challenge is choosing the right representation, or at least factoring that in.

It feels like quite a fundamental shift, and I'm curious whether others are seeing the same thing in their own domains.

reddit.com
u/likescroutons — 13 days ago

Weekly Entering & Transitioning - Thread 10 Aug, 2026 - 17 Aug, 2026

Welcome to this week's entering & transitioning thread! This thread is for any questions about getting started, studying, or transitioning into the data science field. Topics include:

  • Learning resources (e.g. books, tutorials, videos)
  • Traditional education (e.g. schools, degrees, electives)
  • Alternative education (e.g. online courses, bootcamps)
  • Job search questions (e.g. resumes, applying, career prospects)
  • Elementary questions (e.g. where to start, what next)

While you wait for answers from the community, check out the FAQ and Resources pages on our wiki. You can also search for answers in past weekly threads.

reddit.com
u/AutoModerator — 10 days ago