r/Sabermetrics

▲ 5 r/Sabermetrics+1 crossposts

High school senior building an MLB front-office portfolio on GitHub. Just finished a mock Braves/Rangers trade evaluation for Kumar Rocker and would love feedback!

github.com
u/Specialist_Fix1376 — 2 days ago
▲ 45 r/Sabermetrics+1 crossposts

Is a runner on second really more likely to score if there are 2 outs? Yes.

I tested a familiar baseball idea: with two outs, a runner on second can run immediately on contact and without worrying about getting doubled-off.

Using 47,653 MLB regular-season balls in play from 2021-2025 with second occupied, first and third empty, and home runs excluded, I measured whether the runner scored on that same play (i.e., not whether they advanced to third and scored later). Below are the probabilities the runner scores:

  • 0 outs: 14.9%
  • 1 out: 18.7%
  • 2 outs: 24.5%
  • Two outs vs. 0–1 outs: +7.4 percentage points
    • 95% CI: +6.7 to +8.2 points

A two-proportion hypothesis test evaluated the null hypothesis that the scoring probabilities were equal (2 outs vs fewer than 2 outs). The result was statistically significant (p < 0.001), so I reject the null and conclude that two outs are associated with a higher probability of scoring on the play compared to 0 or 1 outs.

The increase may reflect larger leads, running on contact, and more aggressive third-base coaching. This is a strong statistical result but not necessarily causal: runners, contact, defense, and game situations likely have an influence.

u/BadDatatude — 5 days ago

I built a percentile-normalized performance score for MLB players (wOBA/ISO/FIP/K9/WHIP/BB% weighted composite) — feedback welcome

I've been building a composite performance metric for MLB players — calling it "BAX Score". The goal is a single number that reflects "how good is this player playing right now," built only from real performance stats (no market/demand signal in the mix at all — this matters because the score also drives pricing in a fantasy stock market I'm building, and I didn't want volume to be able to move it).

Methodology, roughly:

  • Pull wOBA, ISO, FIP, K/9, WHIP, BB% (hitters and pitchers get separate axis   weightings)
  • Percentile-normalize each stat against the current league sample
  • Apply sample-size shrinkage so a guy with a hot two weeks doesn't swing the   score around
  • Weighted composite across the normalized axes → single score, updated daily   after games finalize

Pulled today's top-of-leaderboard hitters and plotted wOBA (the single most correlated input) against the full composite score to see how much the extra axes actually move things:

https://preview.redd.it/teo8si5wbqjh1.png?width=1350&format=png&auto=webp&s=227db2fe111ca9c6d4f33caba6aafafccf85d26b

A few things stand out:

  • Caminero, Harper, and Alvarez score noticeably ABOVE where wOBA alone would put them — power/ISO and BB% are dragging the composite up past what a wOBA-only ranking would suggest
  • Kurtz, Soto, and Ohtani score BELOW their wOBA-implied line — despite great wOBA, other axes (BB%/K% mix, sample shrinkage) pull them down relative to the trend

Basically: if I'd just ranked by wOBA, the order would look meaningfully different from the composite. Curious whether this sub thinks that's a feature or a red flag — is baking in ISO/BB%/K% on top of wOBA adding signal, or just re-weighting the same underlying skill twice?

Also still going back and forth on:

  • How much weight FIP vs xFIP should get for pitchers with small sample IP
  • Whether BB% deserves its own axis or should just feed into wOBA indirectly

Happy to share more of the pipeline if useful and your feedback :)

(Side project — baxeball.com — mostly posting here because I want the methodology stress-tested, not to plug it.)

reddit.com
u/Striking-Lock-4937 — 4 days ago

The Hitting Approach of Back-to-Back Champs

This era of Dodger baseball has been the best I've ever experienced (though my heart will always be with those Jim Tracy, Grady Little, Joe Torre rosters).

There are a ton of ways to slice or point to how this core roster has won back to back titles...but wanted to share one offensive angle that might be interesting if you're into hitters' approaches/mentalities, a Dodger lover, or even a Dodger hater.

A side project study, I found: by creating a scoring index of decisions per count…the Dodgers' 1 through 9 (or 10, 11, 12, etc depending on Doc's definition of everyday lineup pieces) seem to control the counts/swing decisions better than any other team in the last 2 seasons. Admitting one gap here is the pitcher, game situation, maybe even subjective scoring too. But it’s interesting to see the clusters of each team under this method…

Example: 2-0 fastball down the middle, they're consistently take a hack. 1-1 they're consistently taking a ball off the plate. 0-0, they're consistently taking a strike on the black because it's not a pitch they can groove. They make good decisions, regardless of the outcome. They hunt the same way...seemingly all bought in on an approach at the plate.

Sharing a writeup with some visuals if you're curious to dive in! (Link to article)

Need to refresh it for 2026 season-to-date if they're trending the same way. But in the meantime...cue Randy Newman!

u/youravesfriend — 5 days ago

I built a next pitch prediction bot and scored it against real, live pitches for a month (50,760 pitches). Here is what it learned about pitcher predictability, plus a "deception" metric I threw out.

I built a model that predicts the type of the next pitch (four seam, sinker, cutter, changeup, curve, slider, sweeper, splitter, knuckle) before it is thrown, wired it to live games, and have been scoring every prediction against what actually got thrown. This post is the ture, honest version of what came out, including the parts that did not work, because I think the negatives are more interesting than the headline number if i'm honest.

TLDR

  • Live accuracy over 50,760 scored pitches: 42.8% top-1, 85.7% top-3. Offline, on a frozen test set, it beats a strong pitcher and count conditioned baseline by +6.7 percentage points (38.9% to 45.6%). Live is lower than offline, which is what should happen.
  • Predictability is a repertoire trait, not a talent gap. Elite arms show up at both ends of the predictability scale.
  • The model's single biggest failure mode: when it is wrong, it guesses fastball.
  • I built a metric that looked like it measured pitcher deception. It is very reliable. Then I tested it against real batter outcomes and it predicted nothing, so I threw the interpretation out.

1. How good is it, honestly

42.8% top-1 sounds low until you compare it to the right baseline. The naive "always guess his most common pitch" is easy to beat.... the real baseline is each pitcher's own count - conditioned mix, smoothed. Beating that by ~7 points offline means the model is reading context, not just base rates. Top-3 at ~86% is the more useful number for most real uses: even when the exact call misses, the pitch is almost always in the top three.

The live number moves around by day, mostly because of *who pitched*, not model quality. I compute a composition adjusted expected accuracy (each day's pitches weighted by each pitcher's own history) and the actual line tracks it closely. A recent dip is a harder slate, not decay (Page/Hinkley drift detector - no alarm).

2. Predictable is not the same as good

This is the finding I care about most. Rank pitchers by how often the model nails the exact next pitch and you get elite arms scattered top to bottom:

Kenley Jansen and Tim Hill are basically a coin you can call in advance. Max Fried, Yoshinobu Yamamoto, and Tarik Skubal are near the bottom. All of them are good!!! Chris Sale has a low arsenal entropy and Skubal a high! Both are, quite obviously, aces. So predictability and quality are uncorrelated, and you should never fold a "predictability" number into a "how good is he" number. Predictable is a description of the arsenal, not a criticism of the pitcher.

3. When the model is wrong, it guesses fastball

Every offspeed and breaking pitch's single most common wrong prediction is a four-seam fastball:

Part of this is the model over defaulting to the majority class, and part is genuine label fuzz (Statcast's four-seam / sinker / cutter boundary is noisy). When I collapse those three into one "fastball family," a big chunk of the apparent error disappears, but not all of it, so the overcall is real, not just a labeling artifact.

4. The metric I built, liked, and killed

I wanted to measure pitcher deception: is a pitcher harder to read than his raw pitch mix diversity alone would predict? I built a residual (actual readability minus what mix entropy predicts). It looked great! :

  • It is reliable: split-half correlation of 0.92, and it survives controlling for role and count context.
  • It produces a sensible leaderboard (some guys read easier than their mix implies, some harder).

Then I did the step most people skip. I correlated it against things batters actually produce, controlling for stuff quality:

Nothing. No relationship with whiff rate, called strike rate, chase rate, or run prevention. So the residual is a reliable, stable property of how my model reads a pitcher, with no demonstrated connection to deception. Most likely it is a model blind spot, not really a hitter relevant trait. Reliability is not validity. I kept it as an internal diagnostic and built no "deception index" on top of it, which was hard to do because the story was so good.

5. What does not improve it

I ran four honest ablations trying to push accuracy up: pitch movement features, within at bat sequencing plus tunneling, an online ingame adaptation layer, and batter side features. All null or actively harmful tbh. The takeaway is that the model is near its ceiling for pitcher side features, the signal lives in tendency and rate features (how often he throws X in this count, what he threw last, pitcher and batter identity), and those are already saturated. The online adaptation experiment did surface one real thing: pitchers negatively autocorrelate (after a fastball streak they are more likely to change), which the static transition rates already capture.

6. The boring parts that make me trust it to the best of my knowledge....

Leak-safe temporal splits with a hard assertion against future leakage, probability calibration (overall ECE ~0.01; the 3-0 count is the known weak spot at ~0.05 and I am fixing it), a live trust score built from historical per-(count, pitch type) precision, drift detection on the composition adjusted residual, and empirical Bayes shrinkage on the per pitcher numbers so small samples do not lie.

Limitations

  • Live (42.8%) trails offline (45.6%), as expected from distribution shift and new pitchers.
  • The 3-0 count is genuinely miscalibrated (rare, small sample, overconfident) and is the one production issue I would flag.
  • This is pitch type, not location, and the swing/whiff model is secondary.
  • The "deception" residual failed external validation, so please do not read the readability numbers as a deception measure.

Happy to answer questions on any of it, and genuinely curious what this sub would test next. My thought is that any further accuracy is a dead end and the interesting frontier is measurement (what predictability correlates with, if anything), but I have been very, very wrong before within this project.

Also for fun - my dashboard screen shots.

https://preview.redd.it/nvhh51hmbgjh1.png?width=256&format=png&auto=webp&s=cec5847a04b8d591b1cfd9a7b67ef8d38fcdd5b3

https://preview.redd.it/vrxhomdpbgjh1.png?width=814&format=png&auto=webp&s=58a911de10405d3c2b4d78c25c2847b9ac81198e

https://preview.redd.it/jutltt3rbgjh1.png?width=1555&format=png&auto=webp&s=b5861bc1d867664e8cbf7479e9f668d6ec577adb

https://preview.redd.it/f5e5kumsbgjh1.png?width=1238&format=png&auto=webp&s=fcbeb4264f5011744bdef5b57ff4ebb504a67255

https://preview.redd.it/1k54pyavbgjh1.png?width=1535&format=png&auto=webp&s=ba0b8d44813f0e66754d35d45cf3219edc24fc36

https://preview.redd.it/0nf2mhnwbgjh1.png?width=1854&format=png&auto=webp&s=5c6716aba4276330aaa3e17ce8480075ae6627f2

https://preview.redd.it/p3fub50ybgjh1.png?width=1450&format=png&auto=webp&s=3cab37bf72f56e4ff2b748495f101e85d126c772

https://preview.redd.it/ib299ynzbgjh1.png?width=1848&format=png&auto=webp&s=5699e85fc0cef24ea0c73eecc504531f6cd5fdeb

Also included write up

 

u/Ok-Hovercraft-5701 — 5 days ago

I created a new stat RBI Opportunity Metric (ROM) to replace RBI entirely

Traditional RBIs are a flawed stat because they measure lineup luck rather than hitter efficiency. A mediocre hitter in a stacked lineup easily out-produces an elite hitter on a terrible team.

To fix this, I created ROM (RBI Opportunity Metric).

How ROM Works (The Point System)

Instead of treating all base situations equally, ROM assigns fixed point values based on how close a runner is to scoring, plus a baseline for the batter (due to the constant threat of a home run):

  • Batter Baseline: 0.25 points
  • Runner on 1st Base: +0.50 points
  • Runner on 2nd Base: +0.75 points
  • Runner on 3rd Base: +1.00 point

The Formula: You add up the points of the runners a player actually drove home (Points Earned) and divide it by the maximum points available on base (Points Possible), using strictly At-Bats (AB) as the denominator.

The "Free Points" Discipline Mechanic

Because walks (BB) and sacrifice flies (SF) do not count as At-Bats, they add 0.00 to the denominator. But if a hitter draws a bases-loaded walk or hits a sac fly, they still get the point in their numerator. It heavily rewards game-winning team baseball.

Historical Proof: The Career ROM Leaders (Post-1957)

To test the lifetime stability of the index, I calculated career data using non-overlapping game log configurations. Sustaining a high conversion efficiency across decades highlights true legendary status:

Rank Player Career ROM Batting Average (AVG) Career OPS
1 Barry Bonds .292 .298 1.051
2 Manny Ramirez .288 .312 .996
3 Albert Pujols .281 .296 .918
4 Alex Rodriguez .276 .295 .930
5 Mike Trout .272 .292 .981
6 David Ortiz .269 .286 .931
7 Miguel Cabrera .265 .306 .901
8 Ken Griffey Jr. .261 .284 .907
9 Eddie Murray .259 .287 .836
10 Vladimir Guerrero Sr. .258 .318 .931

Here is a link to a more detailed breakdown. Disclaimer: While the equations is original I did use googles AI chat bot to help format and chart inputs.

https://docs.google.com/document/d/e/2PACX-1vRUpqp9wnzWTp0elP4zQ4-z3re7m5iM7PeFEu8dc-psQrx8zSs3klvDyvVBq8ggKE4ytMfsGiypqmTY/pub

reddit.com
u/Horror_Owl_3655 — 8 days ago
▲ 6 r/Sabermetrics+1 crossposts

Solo Home Runs

I’m no sabermetric Mets fan and baseball analogics is not my game.

But how many solo home runs have the Mets hit this year?

What is the record for solo home runs when compared to multi run dingers?

And is all this my imagination when I tell my wife I’ve never seen a team hit so many solo homers in a year with so few when there’s men on base?

I understand it’s a function of how poor the offense is that they just aren’t putting people on base to be driven in, but it’s still frustrating when I watch another game when they hit 3 solo shots over the fence.

reddit.com
u/ericloz — 7 days ago

saber method

Title: Held-out calibration on 4,570 MLB starts: Brier 0.2375 vs 0.2806 baseline — looking for criticism of the validation design

I've been working on a probabilistic starting-pitcher strikeout model and wanted to post the validation methodology rather than predictions.

The entire 2025 MLB season was withheld from model fitting and used as the out-of-sample test set: 4,570 starts the model had never seen.

The main results:

Brier score: 0.2375
Naive/base-rate Brier: 0.2806
Expected Calibration Error: 0.0049
Mean PIT: 0.4977
Holdout: 4,570 starts

The piece I'm most interested in is calibration rather than raw classification accuracy.

For the reliability table, predicted probabilities were grouped into probability bands and compared against observed frequencies.

For example:

50–55% predicted → 50.9% observed
55–60% → 55.5%
60–65% → 59.3%
65–70% → 63.3%
70%+ → 70.1%

I'm deliberately trying to avoid the usual sports-model trap of saying something like "the model was 64% accurate" without establishing what was predicted, at what probability, or against what baseline.

The validation gate was designed around a few requirements:

  1. No lookahead. Inputs/fitting procedures can only use information that would have existed before the game being predicted.
  2. Full-season holdout. 2025 was excluded from training rather than randomly splitting games across seasons.
  3. Brier against a baseline. A Brier score in isolation isn't particularly informative, so I'm comparing it against a naive probability baseline on the identical events.
  4. Reliability/calibration. If the model assigns a group of events ~60%, they should occur roughly 60% of the time.
  5. PIT diagnostics. I'm looking at where actual outcomes fall within the predicted distributions, rather than only checking the mean prediction.

One thing I've tried to be careful about is separating calibration from usefulness. A model can be beautifully calibrated by staying close to the base rate and still contain very little information. Conversely, a sharp model can have useful discrimination while being overconfident.

So rather than asking whether these numbers are "good," I'd be interested in how people here would try to break this validation design.

A few questions I'm considering:

  • Is the naive/base-rate Brier benchmark the right primary baseline, or would you want additional benchmarks?
  • Would you report Brier decomposition into reliability, resolution and uncertainty?
  • Would you bootstrap confidence intervals around Brier/Brier skill rather than report point estimates?
  • For calibration, would you prefer adaptive/equal-count bins over fixed probability bands?
  • What would you use to test whether the apparent calibration survives season-to-season distribution shift?
  • Are PIT + reliability + Brier redundant here, or do you think all three earn their place?
  • What failure mode would you look for first if you were reviewing this?

I'm much more interested in finding where the methodology is weak than in defending the headline number.

Would appreciate any criticism from people who have done probabilistic baseball forecasting.

reddit.com
u/sportshackai — 6 days ago
▲ 4 r/Sabermetrics+2 crossposts

30 HR 50 or fewer strikeouts in a season

Excuse me while I reminisce about the good ole days as a Cardinals fan, watching Albert Pujols do things at the plate I’ve never seen anybody do. I can remember reading and hearing from media pundits, not only in St. Louis, but nationally, that Pujols was a throwback to the sluggers of the past like Ted Williams, Joe DiMaggio, Lou Gehrig, and Cards own legend Stan Musial, he can get you homers and not strike out a ton. And that’s not a bunch of media exaggeration- he really didn’t strike out a lot. Never once did he strike out 100 times. His first year in 2001 he struck out 93 times. He would tie that in 2017, but after that his next highest was 76. Prime Pujols, striking out 50-60 times in a season. Unreal compared to what we see now. And yes, I know, he hit into the most double plays ever, certainly a down fall of being a contact hitter with power.

Anyway, I saw in 2006 he had 50 strikeouts while crushing 49 homers. That’s his lowest total of a qualified season in his career, which got me thinking about who else has done that in recent memory and of course going back through history. List below.

Highlights, DiMaggio leads everyone with 7 seasons of 30 or more home runs with 50 or fewer strikeouts. Gehrig, Williams and Musial had 6.
Ken Williams was the first to achieve the marks in 1922.
Victor Martinez in 2014 has been the last player to achieve the feat.

=>30 Home Runs
=<50 strikeouts 

Victor Martinez 14 Tigers
Albert Pujols 06 Cardinals
Vladimir Guerrero 05 Angels
Barry Bonds 04 Giants
Barry Bonds 02 Giants 
Moises Alou 00 Astros
Gary Sheffield 92 Padres
Cal Ripken 91 Orioles
Don Mattingly 87 Yankees 
Don mattingly 86 Yankees
Don Mattingly 85 Yankees
Gary Carter 85 Mets
George Brett 85 Royals
Bob Horner 80 Braves
Hank Aaron 69 Braves
Felix Mantilla 64 Red Sox
Hank Aaron 58 Braves
Ted Williams 57 Red Sox 
Ted Kluszewski 56 Reds
Yogi Berra  56 Yankees
Ted Kluszewski 55 Reds
Stan Musial 55 Cardinals
Roy Campanella 55 Dodgers
Ted Kluszewski 54 Reds
Stan Musial 54 Cardinals 
Al Rosen 53 Indians
Ted Kluszewski 53 Reds
Stan Musial 53 Cardinals
Yogi Berra 52 Yankees
Stan Musial 51 Cardinals
Andy Pafko 51 Cubs/Dodgers
Ted Williams 51 Red Sox
Andy Pafko 50 Cubs
Joe DiMaggio 50 Yankees
Vern Stephens 50 Red Sox 
Ted Williams 49 Red Sox
Stan Musial 49 Cardinals 
Johnny Mize 48 Giants 
Stan Musial 48 Cardinals
Joe DiMaggio 48 Yankees
Sid Gordon 48 Giants
Johnny Mize 47 Giants
Willard Marshall 47 Giants
Walker Cooper 47 Giants
Ted Williams 47 Red Sox
Ted Williams 46 Red Sox 
Ted Williams 41 Red Sox
Tommy Henrich 41 Yankees 
Joe DiMaggio 41 Yankees
Johnny Mize 40 Cardinals
Joe DiMaggio 40 Yankees
Joe DiMaggio 39 Yankees 
Joe DiMaggio 38 Yankees
Mel Ott 38 Giants
Joe DiMaggio 37 Yankees
Lou Gehrig 37 Yankees
Joe Medwick 37 Cardinals
Lou Gehrig 36 Yankees
Mel Ott 36 Yankees
Lou Gehrig 35 Yankees
Lou Gehrig 34 Yankees
Hal Trotsky 34 Indians
Ripper Collins 34 Cardinals
Mel Ott 34 Giants
Earl Averill 34 Indians 
Zeke Bonura 34 White Sox
Lou Gehrig 33 Yankees
Lou Gehrig 32 Yankees
Chuck Klein 32 Phillies 
Earl Averill 32 Indians 
Mel Ott 32 Giants 
Chuck Klein 31 Phillies 
Earl Averill 31 Indians 
Chuck Klein 30 Phillies
Al Simmons 30 A’s
Mel Ott 29 Giants
Al Simmons 29 A’s
Lefty O’Doul 29 Phillies 
Don Hurst 29 Phillies
Rogers Hornsby 25 Cardinals
Ken Williams 22 Browns

reddit.com
u/UpsetBackground9093 — 7 days ago
▲ 2 r/Sabermetrics+1 crossposts

Ballista turns phones into portable baseball performance labs. Beta testers wanted.

My name is Riley. I am Co-founder and CEO of Ballista. We turn phones into portable baseball performance labs.

Using CV, custom AI, and the phones high frame rate camera, we extract EV, LA, and Distance of hit balls automatically and aggregate all the data.

It is an alternative to existing hardware systems like radar guns, Trackman, Rapsodo, Hittrax, that cost thousands of dollars.

We are getting ready for a private beta launch via TestFlight and are gathering a list of people who are interested.

If anyone is interested, head to our website and sign up for the waitlist and expect an email (coming soon) with the app testing link. Alternatively you can drop your email in the comments or PM it to me and I can add you manually.

Cheers!

u/50th-century — 9 days ago
▲ 1 r/Sabermetrics+1 crossposts

Should Blake Snell have been pulled in the 2020 World Series? I created a Degradation Index for pitcher fatigue to determine whether this was the right choice.

Ever since Kevin Cash took out Blake Snell because his velo dropped in the 2020 World Series I wanted to find a better way to determine when to pull a pitcher.

To solve this, I created a Degradation Index to quantify pitcher fatigue. This index is based on variance in arm angle and extension, and achieves a much higher specificity than standard velo drops. By tracking pitcher fatigue in mechanics, teams can save their bullpens only for when it is most important.

I used an XGBoost regressor to filter out pitcher's intentional changes in mechanics so that the isolation forest algorithm only checks the residuals for fatigue-caused anomalies.

I just published a comprehensive breakdown on Medium, and the code is open on GitHub. Any and all feedback is very appreciated.

medium.com
u/Adventurous-Ship5825 — 8 days ago
▲ 4 r/Sabermetrics+2 crossposts

I built a College Baseball Explorer for exploring D1 baseball data

I've been working on a project called College Baseball Explorer, an interactive site for exploring college baseball statistics and analytics. One thing I wanted to do differently with this project is to keep it free and open source rather than putting the data and tools behind a subscription/paywall. The goal is to make college baseball analytics accessible to anyone who wants to explore it.

It currently includes ACC, SEC, Big 12, and Big Ten data, with tools to look at performance by season, conference, team, and player. I've also added advanced metrics, percentile rankings, interactive tables/visualizations, and downloadable data.

The project started as an ACC baseball data project and gradually grew into something broader, so I recently relaunched it as the College Baseball Explorer.

I'm hoping to continue expanding it with features and perhaps conferences.

I'd love to hear what other college baseball fans, coaches, analysts, or data folks would want to see added.

College Baseball Explorer: ballclubdata.com

https://preview.redd.it/62xhkesmpxih1.png?width=2816&format=png&auto=webp&s=665726185f5d0b2ccbb3ac34e3ebf9e081f4e6e7

reddit.com
u/Other-Win5218 — 8 days ago
▲ 0 r/Sabermetrics+1 crossposts

Curt Schilling: A Sabermetric Analysis of His Career

Given how controversial Schilling is, a sabermetrics evaluation of his career is probably the one area that is least up for debate. I am not interested in his politics, post-career goings-on, or other controversies, I want to keep it strictly about his on-field performance. As I hope to be able to show, Schilling's career was perhaps the pinnacle of great regular season performance, otherworldly postseason success, and one that cannot be fully appreciated without diving into the sabermetrics. His time spent on the diamond was nothing short of legendary.

Going strictly by his on-field performance, Schilling has an extremely strong case as being the fifth-best pitcher of the last 40 years, behind the Big Four of Clemens, Maddux, Johnson, and Martinez (not necessarily in that order). Schilling's off-field views are almost certainly the reason why he is not in the Hall of Fame, but really, they never should have factored in much at all because he should have been an easy first-ballot HOF selection back in 2013, and he didn't really start opening his mouth and talking politics until 2016. In 2013, his first year on the ballot, he received only 38.8% and the next year he DROPPED by a quarter to 29.2%, which is just ridiculous. Back then voters were still enamored with pitcher wins, explaining why Glavine got elected on his first time on the ballot to the tune of 91.9%, even though Schilling, on the field, was far, FAR superior, and that's just in the regular season, let alone any postseason heroics.

You cannot talk about Schilling's pitching without mentioning K/BB ratio. Alongside home runs allowed, those three statistics form the basis of FIP. Why does this matter when talking about Schilling's place among the game's greatest? It is because he absolutely dominated this category, and when factoring in his roughly league-average HR%, the fact that he had such a high fWAR total for his career should tell you something about this guy's pinpoint control. Greg Maddux gets lauded for his accuracy, and rightfully so, but I think Schilling was even better, at least when it comes to K/BB: Maddux's 7-year peak K/BB+ (K/BB ratio adjusted for era) was a staggering 285.58, seventh-best since 1901 by my calculations. Schilling's figure was 315.07, just barely behind Pedro's 316.62 and Cy Young's 316.02. For his career, Schilling tossed to a 268.58 K/BB+, second all-time behind Young's otherworldly 291.28; Schilling's career mark is higher than Clayton Kershaw's seven-year peak figure of 261.41. While not walking hitters is crucial, I think a more holistic view of control is how efficiently you strike guys out as you pitch. Schilling did that just about better than anyone ever has.

Of course, K/BB ratio, while important, is not the end-all, be-all. Schilling's brilliance also has to be contextualized in the era he played, and the ballparks he played in, as we all know. Schilling started his career in a low-offense environment of 1988 but didn't really get significant time on the mound until 1992. While that year was tied with '88 for the lowest runs per 9 in the NL since 1968, starting in 1993, it ramped all the way up to 4.52 and stayed high for the rest of his time in the Sr. Circuit, averaging 4.70 RA and in his 4 years in the AL from 2004-2007 it was 4.89 RA. Since FIP is equated to ERA for the purposes of fWAR, and since Schilling's peak was more FIP-dominant than pure run prevention, I wanted to focus on two years in particular, in which, granted, he only made one postseason appearances, but we'll get to the playoffs later.

While 1997 was Schilling's breakout sabermetrics year, 1998 saw him go to heights seldom ever reached from an fWAR perspective. FanGraphs credits him with 8.25 fWAR, which is really good, but not otherworldly. However, my calculations give him 9.43 fWAR, which while I cannot confirm with 100% certainty, is most likely due to their use of multi-year regressed PFs and my use of single-year PFs, as the 1998 Phillies had a 101 5-year regressed PF and a 106 PF for just that year. Also, my FIP calculation is based solely on an intraleague derivation, not interleague, which is why my 2.69 FIP for '98 Schilling is slightly less than the 2.77 FIP FG credits him with. Using FG's PFs and FIP would result in a derivation of 8.55 fWAR for Schilling as opposed to 9.43, and the remaining 0.3 fWAR is probably due to a combination of immaterial PythagenPat variances between FG and BRef, which I employ the latter for my calcs. With that out of the way, let's dive into what made this season special.

Schilling had 300+ Ks for the second year in a row at a rate 57% better than the NL average. He led the majors in IPs at 268.2 and his 9.43 fWAR was .45 wins more than second-place Kevin Brown's 8.98 fWAR. He was first in pretty much all FG-related metrics this season, both volume and rate, with the exception of FIP, in which his 153.9 FIP+ was fifth. However, because of his pitching in a very hitter-friendly ballpark, his fWAA and fWAR totals get him the gold medal in those metrics. He severely underperformed his peripherals with a 3.25 ERA, and while we can play the "would coulda shoulda" game all day, had his ERA aligned to his FIP and his RA remained at 0.13 runs above his ERA, you are looking at a bWAR of 9.73 instead of what he actually had of 7.62 bWAR (based on my derivations). Again, hypotheticals all day long, but when most think of Schilling, his time with Arizona comes to mind, and while understandable, he was awesome in multiple seasons with Philadelphia and that should not be forgotten to history.

2002 was his magnum opus and one of the most dominant fWAR seasons this century, maybe the single best. I would say only Johnson's 2001 season (in which I calculated his fWAR at 11.00) eclipses it in this millennium. Over 256.2 innings, Schilling's 10.22 fWAR is the best since the aforementioned Randy, same for his 8.10 fWAA and .550 f_Yr. WL% (an approximation of an otherwise average team's winning percentage over a full season if they had this player). His K/BB was 9.58, good for an almost impossible-to-believe 494 K/BB+ figure. 2002 is also the first year with xFIP data and his was 2.20, so if you convert that to his actual FIP and ERA (and maintain an RA of 0.07 above his ERA), we are looking at a 10.63 fWAR and 11.78 bWAR. Schilling's underlying peripherals show a pitcher that, at age 35, was beyond dominant, on par with his Arizona teammate who took home his 4th straight Cy Young. His 2.31 FIP (again slightly off from FG because I use purely intraleague figures to derive it) was 80% better than average, second only to Pedro's 185 FIP+. The reason that Schilling's rate and volume fWAR totals supersede Martinez's is because of Park Factors, where Chase Field had a 108 PF while Fenway's was actually slightly pitcher-friendly at 98.

My methodology actually ranks Schilling 11th all-time, but because it only goes back to 1901, Cy Young is diminished here since I use a combination of rate and volume as well as weight peak at 43% and career at 57%. Schilling's fWAR metrics are doing a lot of the heavy lifting here, especially his rate and peak splits. For a pitcher's best seven years (not consecutive), Schilling's 153.2 FIP+ is 5th all-time, his 39.9 fWAA is fourth (behind Randy, Roger, and Pedro), his 53.9 fWAR is sixth, his f_waaWL% (the estimated win percentage an otherwise average team would have in games this pitched participated in) of .679 is fourth, and his .536 f_Yr. WL% is sixth. Yes, peak dominance is not everything, but I think it should count significantly when contextualizing all-time rankings because mere volume accumulation, while impressive in its own right, doesn't swing the needle on a season-to-season basis, and while I admit that it's a preference, dominance deserves a place.

However, I have neglected to mention bWAR-based metrics, and admittedly, Schilling does not fare as well here. In the interest of transparency, here are his bWAR placements for 7-year peaks: a 141.6 ERA+ is 41st, his .666 bwaaWL% is 16th, his .533 b_Yr. WL% is 29th, his 37.2 bWAA is 26th, and his 51.2 bWAR is 33rd. To put it frankly, those do not scream Top-15 pitcher of all-time. But we also have to include career totals as well, as I even admitted that I weight career more than peak for my ranking. On the bWAR side, he fares much better here than his peak. Below are his FG and BRef stats in bold and his 1901-present ranking is in (italicized parenthesis). The reason my WAR totals differ from FG and BRef (aside from the various slight adjustments I made as described above) is because I only included the sum of qualified seasons, not partial ones, nor did I factor in 2020 due to the short 60-game season.

FIP+: 136.5 (7th)

fWAA: 52.9 (8th)

fWAR: 77.6 (16th)

fwaaWL%: .631 (6th)

f_Yr. WL%: .526 (4th)

ERA+: 134.0 (34th)

bWAA: 55.9 (16th)

bWAR: 80.5 (24th)

bwaaWL%: .639 (9th)

b_Yr. WL%: .527 (21st)

It is completely understandable if these numbers still leave one skeptical of Schilling's status of a Top-15 pitcher ever, but if you are open to the possibility of it, I think it is closer than maybe the initial gut reaction would say, which even for me, is pessimistic on such a posit. Now would this hold if the color barrier never existed? Probably not. Willie Foster, Bullet Rogan, Satchel Paige, William Bell, and Ray Brown may have him beat. Kid Nichols has a good case as being above him, although I would draw the line at him for other pre-1901 pitchers (my apologies to Tim Keefe, John Clarkson, and Pud Galvin).

Now, to cap it off, if we are factoring in the postseason, I think it should CLEARLY make him a top-15 pitcher of all-time. His 2.23 ERA is spectacular, but it isn't the best (Koufax has a 0.98 ERA and Mathewson has him beat at 0.97). Adjusted for the era he played in, however, Schilling has a 200 ERA+ over 133.1 postseason IPs, beating even Bob Gibson who has a 182 ERA+ for his 1.89 ERA. Schilling's ERA+ is 8th all-time for starters and no pitcher above him comes within 30 innings of his total, with Mathewson's 333 ERA+ coming over the course of 101.2 IPs. His WPA is the highest ever for SPs at 4.07 WPA, with Andy Pettitte taking the silver at 3.20 WPA. The incredible part about that is Pettitte earned it over 276.2 IPs, so Schilling, in less than half the number of innings, materially supersedes him. Rivera takes the record for postseason WPA at 11.38, but considering that he was a closer and could go all-out for an inning and change at most, I am comfortable saying Schilling ranks higher than Mo among postseason hurlers. Schilling had a 3.05 WPA/100 postseason innings pitched, which is second only to Zack Wheeler all-time among pitchers with at least 10 postseason starts, and Wheeler has averaged only 17.6 outs per games, or ~5.2 innings, while Schilling averaged 21.1 outs/7 innings per outing. In the modern game with high-usage bullpens, Schilling's ability to maintain excellent rate dominance and going deep into games was another big value-add to his teams' postseason successes. While I do not generally use postseason performance in ascertaining the historical placing of a player, when a guy is that dominant over that long a stretch, I think it deserves some recognition.

reddit.com
u/UTexas2005 — 12 days ago
▲ 25 r/Sabermetrics+1 crossposts

Appreciation for Bobby Abreu

Another Hall of Fame ceremony came and went and Bobby Abreu still sits on the outside. I believe he got a 30% vote this past year. I do believe he is worthy of the hall of fame and here might be one of the best examples of the type of player he was.

Abreu is the last player in a single season (2000) to hit 10+ Doubles, triples and home runs, while also walking 100+ times. In fact, he also achieved the feat the prior year (99). To put this achievement in perspective, by my research only 16 other players have reached this mark in a season going back to the 1890’s.
Take it a step further, besides Abreu, 5 other players have done it multiple times in their career: Babe Ruth, Lou Gehrig, Dolph Camilli, Charlie Keller and Mickey Mantle.

To be on a short list with Ruth, Gehrig and Mantle is pretty special.

reddit.com
u/UpsetBackground9093 — 11 days ago

Reached-on-error run value dropped ~11% and flipped below the single from the 1980s to today — real effect or noise?

I've been comparing linear weights (event run values built by averaging RE24 over every occurrence) for two MLB eras: 1985–1990 vs 2023–2025. Most of the movement I can rationalize, but one number I can't.

Reached-on-error (ROE):

  • 1985–1990: +0.499 runs
  • 2023–2025: +0.446 runs (about −11%)

The strange part is the flip relative to a single:

  • 1985–1990: error (+0.499) > single (+0.461)
  • 2023–2025: error (+0.446) < single (+0.472)

So in the '80s a reached-on-error was worth more than a single; today it's worth a bit less. For context, most other events either held steady or moved in directions I can explain: the single barely budged (+0.461 → +0.472), walks rose (+0.304 → +0.329), and the generic out got slightly more costly (−0.265 → −0.274). ROE is the one event that both dropped meaningfully and swapped ranks with the single.

My questions:

  1. Is a ~11% shift in the ROE weight actually meaningful, or is it within the noise you'd expect given how relatively rare ROE is?
  2. If it's real, what's driving it? Fewer extra-base advances on errors today (better fields, better positioning)? A change in the type of misplays getting scored as errors? Official-scorer discretion drifting over the decades? Something about the run environment itself?
  3. Has anyone tracked ROE run value across seasons/decades? A pointer to that would be great.

Both sets are RE24-based linear weights (the "average" across all 24 base-out states) for each era. Thanks in advance.

reddit.com
u/albertop — 10 days ago

I built a self-hosted MLB analytics tool with pitch-level Statcast data — strike zone heatmaps, spray charts, leaderboards

Wanted something like Baseball Savant I could run myself and tweak however I wanted, so I built one.

It's pulling Statcast data through a dbt-duckdb pipeline, with a React/D3 frontend for the visualizations. What it does:

Pitch-level strike-zone heatmaps per player Spray charts Season rollups and leaderboards All backed by a dbt pipeline so the data modeling is transparent and versioned, not a black box

It's live if you want to poke around: https://mlb.palanbates.com

Built with FastAPI on the backend, DuckDB for storage, dbt for the transforms. Since it's dbt-based, the underlying models are all inspectable if you're curious how a given stat is actually computed — no mystery aggregations.

Would love feedback on what stats or views are missing, especially from anyone who lives in Savant regularly and has opinions about what's missing there.

If you want to join the repo feel free as well to shoot me a message!

reddit.com
u/Mysterious-Reason522 — 9 days ago
▲ 51 r/Sabermetrics+1 crossposts

wpybl v0.0.1 (WPBL stats library)

Given the recent start of the Women's Pro Baseball League (WPBL), I have just released v0.0.1 of wpybl (https://github.com/rockysnow7/wpybl), a little Python library for fetching data from the WPBL API, calculating stats, getting play-by-plays, and so on. It will hopefully be useful for sabermetrics of the WPBL. Obviously there isn't a huge amount of data yet, but it'll be interesting to see how the WPBL's stats compare to MLB/MiLB as the sample size grows.

This first version is basically the minimal version of a library like this, so I'll aim to add more features soon. Feel free to try it out if you're interested, message/email me with ideas or questions, and send PRs!

u/rockysnow7 — 11 days ago