I keep hitting a wall trying to learn LLMs systematically. So I'm building an open map of the whole stack — need contributors
▲ 3 r/LargeLanguageModels+2 crossposts

I keep hitting a wall trying to learn LLMs systematically. So I'm building an open map of the whole stack — need contributors

After a year of working with LLMs, I still don't feel like I've built any real, systematic knowledge. Even when I go deep on one area — RAG, say — and track every detail, the fog around LLMs as a whole doesn't lift. It just feels equally thick.

I think most of us learn this field through news headlines and whatever project suddenly jumps into the spotlight. What's missing is a map — something that shows the whole pipeline, from raw data to the app someone actually uses, and for each layer, links both the newest tools/papers AND the older, less-famous work that the newest stuff is quietly standing on. A lot of the real foundations predate "Attention Is All You Need" and never made it into any course.

So I started building one: an open, community-maintained GitHub repo mapping the LLM stack layer by layer —

Data → Training → Model → Deployment → Inference → API → Gateway/Router → Application → User

Each layer gets:
- a plain-language definition
- current, actively maintained projects
- the foundational paper(s) that layer is built on (even if they're old and unglamorous)

Repo here: https://github.com/YKs22k/LLM-Big-Map

I'd love help from people who actually work in data curation, training infra, inference engines, or the app layer, to correct what's wrong and add what's missing. Even a single "you're missing X paper" comment helps.

If this resonates with anyone else who's felt the same fog, I'd appreciate a look.

u/FaithlessnessOdd3645 — 6 days ago
▲ 1 r/Sabermetrics+1 crossposts

I am RME. Setting on an idea and looking for a technical cofounder.

This Idea. could be separated into two components. Part A focuses on a video recognition model designed to track advanced metrics that are often overlooked, such as the precise timing and trajectory of a ball in the air. The objective of this part should be extracting “hidden” factors of a football game, such as the time football is on the air after kick-off, it accomplished by the model. Part B is a data synthesis engine that creates comprehensive player profiles. This system allows users to select specific indicators, assign them custom weights or templates, and simulate plays to predict performance results. Part B focuses on collecting data from users and community building; my plan is to send emails and try to connect with scouts or analysts working on football clubs, especially small football clubs. This part B is inspired by obsidian.md

I called it unpolished because I am not confirmed the mode of part B only the target customers (Scouting system, analysts) Are confirmed. Send me an email if you're interested in this idea. I gonna schedule a quite Q&A like meeting

u/FaithlessnessOdd3645 — 3 months ago