▲ 2 r/mlops

Open-source tool for tuning inference servers: 81 → 421 tok/s on RTX 5090, 257 → 490 tok/s on H100, cost down 81% / 48%

Hello everybody,

I built Profile to make inference tuning deterministic, and save us all time. v2.2 is out today.

It reads a live vLLM server's metrics, compares them against the GPU's roofline ceiling, and names the bottleneck with the exact flag to change.

You apply, it re-measures, and prints before/after on every metric. Regressions get labeled worse, not buried. It never touches the server: no restarts, no config writes, no synthetic load.

Two runs on record, both real SWE-Bench agent traffic, no synthetic benchmarks:

RTX 5090, muse-glimmer 30B, 4 iterations:

  • 81 → 421 tok/s at 25k ctx
  • $3.41 → $0.65 per 1M output tok
  • TTFT 224ms (p95 500ms) at end of run
  • 4.72 → 1.08 J/tok

H100 80GB, Qwen3.8-27B, 3 iterations:

  • 257 → 490 tok/s at 27k ctx
  • $3.23 → $1.69 per 1M output tok
  • TTFT 1.9s → 539ms (p95 4.2s → 1.9s)
  • 2.39 → 1.00 J/tok

The honest part: on the H100 I scaled agents 10 → 285 without fixing KV first. TTFT exploded to 172s. Profile labeled it worse, named KV pressure, and the fix (fp8 KV, ctx trim, seat cut 345 → 22) recovered the run.

Both journeys on video: https://jungledesh.github.io/profile/journeys.html

Note: vLLM only today, more engines next. Single GPU, NVIDIA or AMD; multi-GPU / TP is next on the roadmap.

If you run vLLM in prod, tell me what it names on your servers, and where it's wrong.

curl --proto '=https' --tlsv1.2 -LsSf \
  https://github.com/jungledesh/profile/releases/latest/download/profile-installer.sh | sh

profile diagnose --url http://localhost:8000/metrics --duration 2m

GitHub: https://github.com/jungledesh/profile
Docs: https://jungledesh.github.io/profile/docs.html

reddit.com
u/Inevitable-Diet-1870 — 2 days ago
▲ 12 r/Qwen_AI

Qwen3.8-27B on a single H100: 257 → 490 tok/s at 27k ctx, TTFT 1.9s → 539ms. Tuned in 3 iterations with Profile

Hi all,

I built Profile to make inference tuning deterministic, and save us all time. v2.2 is out today, and one of the two runs on record is Qwen.

Profile reads your live vLLM metrics, compares them against your GPU's roofline ceiling, and names the bottleneck with the flag to change. You apply, it measures the delta. That's the loop.

The run: Qwen3.8-27B on one H100 80GB, SWE-Bench coding agents, real varying load. 3 iterations:

  • 257 → 490 tok/s
  • TTFT 1.9s → 539ms (p95 4.2s → 1.9s)
  • $3.23 → $1.69 per 1M output tok
  • 2.39 → 1.00 J/tok

The middle iteration is the honest part. I raised agents from 10 to 285 without fixing KV first: TTFT exploded to 172s. Profile printed worse next to it, named KV pressure, and the fix (fp8 KV cache, ctx 32768 → 27141, seats 345 → 22) recovered the run to 490. Full run on video.

https://preview.redd.it/tg4iz35b17kh1.png?width=2824&format=png&auto=webp&s=16dc15f110a7e5c8c4140ebfed845bdf944713ce

Final state: no issues, capped by traffic. The server wants more load than my agents give it.

# Download
curl --proto '=https' --tlsv1.2 -LsSf \
  https://github.com/jungledesh/profile/releases/latest/download/profile-installer.sh | sh

# Profile your vLLM server
profile diagnose --url http://localhost:8000/metrics --duration 2m

GitHub: https://github.com/jungledesh/profile
Docs: https://jungledesh.github.io/profile/docs.html

vLLM only today, more engines next. Multi-GPU / TP after this launch.

Running Qwen on your own hardware? Tell me what Profile names on your server!

reddit.com
u/Inevitable-Diet-1870 — 2 days ago
▲ 8 r/Vllm+1 crossposts

Profile v2.2: 421 tok/s with 25k ctx size on RTX 5090 with muse-glimmer. DFlash speculative decoding turned off.

Hi all,

I've been working on making Profile smarter, grounded in physics, and better at getting the max out of your inference server.

Profile (v2.2): a closed-loop, physics and cost aware optimizer. It finds the bottleneck, gives the fix, waits for you to apply it, and measures the delta on every change.

For this release: the core rule engine is rewritten. Eight rules on a priority DAG with mutual exclusivity, so when five alarms fire, four echoes are silenced and the one true cause survives. Less threshold hardcoding and number guessing, more reasoning from what the server is actually doing.

Deterministic: same server, same traffic, same verdict.

We now also support AMD servers (one of the most requested feature).

On my setup: 5.2x throughput (81 → 421 tok/s at 25k ctx) and 81% cost reduction ($3.41 → $0.65 per 1M output tok), with muse-glimmer agents running SWE-Bench. No DFlash.

Watch it live. One iteration regressed; Profile labeled it worse and the next fix recovered it. Regressions stay in the record.

https://preview.redd.it/xx2wsdkmm6kh1.png?width=2248&format=png&auto=webp&s=64b16e06d41e7a5dde78e3808eca45c11f40b640

The whole journey was 4 iterations, ~30 minutes end to end. The same tuning by trial and error is days of guessing.

I've also had engineers from MSFT, Google, and a few startups run it and share numbers, plus few users from this sub whose feedback shaped this release. Grateful for that.

Please give it a try and tell me how to make it help you!

# Download
curl --proto '=https' --tlsv1.2 -LsSf \
  https://github.com/jungledesh/profile/releases/latest/download/profile-installer.sh | sh

# Start profiling your vLLM server
profile diagnose --url http://localhost:8000/metrics --duration 2m

GitHub: https://github.com/jungledesh/profile
Docs: https://jungledesh.github.io/profile/docs.html

Soon: multi-GPU / TP support, cluster / k8s support, and more rules + smarter engine :)

reddit.com
▲ 6 r/Sikhpolitics+1 crossposts

What would you do if you lived in Sikh nation?

Waheguru Ji Ka Khalsa, Waheguru Ji Ki Fateh Ji,

I've always wondered, how'd we all live, behave, & navigate our lives if we lived in Sikh nation; perhaps Khalistan.

Today, we are deeply bruised by culture, propaganda, fake babas, Indian state's concentrated efforts to dilute Sikh spirit, internal attacks, and being stateless.

Qs:

- How do you see yourself living in a place where the Sikh intellect is not attacked, mocked, callled names, and is respected?

- How do you see yourself living & building your life knowing this is where you belong, you feel part of it, and you feel no pressure to assimilate, or justify yourself to dominant discourse?

- How do you see Sikh institutions in a scenario, where they serve not just Sikhs, but become the carrier of Sikh spirit of Sarbat da bhalla, by serving rightfully?

- How do you see the Sikh mass's demographic distribution across the world then?

Curios what you all have to say.

Waheguru Ji Ka Khalsa, Waheguru Ji Ki Fateh Ji.

reddit.com
u/Inevitable-Diet-1870 — 2 months ago
▲ 55 r/mlops+5 crossposts

Profile v2: A physics-grounded, cost-aware optimizer for vLLM

A simple CLI tool that helps you to fine tune your vLLM server.

Profile deeply scans your inference engine (vLLM to begin with), and GPU, calculates your HW limits using Math, & uses metrics from vLLM to give you the waste, its cause, and finally tips to fix it.

It does not stop there, it waits for you to apply the tips, and then keep on re-iterating, until you AI server is tuned to get max out of its limits, or there are no more issues.

A closed loop optimizer for vLLM.

Github: https://github.com/jungledesh/profile
Live Demo + Docs: https://jungledesh.github.io/profile/index.html

I'd love to have any feedback, and answer any q's / concerns.

u/Inevitable-Diet-1870 — 2 months ago
▲ 10 r/Vllm

Profile v2: A physics-grounded, cost-aware optimizer for vLLM.

A simple CLI tool that helps you to fine tune your vLLM server.

Profile deeply scans your inference engine (vLLM to begin with), and GPU, calculates your HW limits using Math, & uses metrics from vLLM to give you the waste, its cause, and finally tips to fix it.

It does not stop there, it waits for you to apply the tips, and then keep on re-iterating, until you AI server is tuned to get max out of its limits, or there are no more issues.

A closed loop optimizer for vLLM.

Github: https://github.com/jungledesh/profile
Live Demo + Docs: https://jungledesh.github.io/profile/index.html

I'd love to have any feedback, and answer any q's / concerns.

reddit.com
u/Inevitable-Diet-1870 — 2 months ago