r/DGX_Spark

Qwen 3.8 27b Single Spark - Recipe(s) - 32 t/s decode

Qwen 3.8 27b Single Spark - Recipe(s) - 32 t/s decode

After 2 days of testing and benchmarking every single recipe I could find on the forums and twitter, this this the current best I have put together for single spark (i tested vllm, sglang, etc). This is tuned also for use with Hermes agent at 256k context, 5 lanes.

PLEASE POST YOURS, with your results.

https://github.com/styles01/sparkrun-recipes/blob/main/recipes/qwen-38-27b.yaml

recipe_version: "4"
name: qwen-38-27b
description: "Qwen3.8 27B NVFP4 — triton_attn, MTP k=2, fp8 KV (auto from checkpoint), 256K context, vision. Mia Lab config."
model: unsloth/Qwen3.8-27B-NVFP4
runtime: vllm
container: vllm/vllm-openai:nightly-aarch64
maintainer: styles01

metadata:
  author: styles01
  created: "2026-08-14"
  updated: "2026-08-14"
  tags:
    - qwen
    - qwen3.8
    - 27b
    - nvfp4
    - mtp
    - triton-attn
    - vision
  runbook: runbooks/qwen-38-27b.md

solo_only: false
cluster_only: false

min_nodes: 1
max_nodes: 1

defaults:
  port: 8000
  host: 0.0.0.0
  tensor_parallel: 1
  pipeline_parallel: 1
  gpu_memory_utilization: 0.84
  max_model_len: 262144
  max_num_seqs: 4
  max_num_batched_tokens: 8192
  kv_cache_dtype: fp8
  attention_backend: triton_attn
  load_format: fastsafetensors
  tool_call_parser: qwen3_coder
  reasoning_parser: qwen3
  served_model_name: "qwen3.8-27b unsloth/Qwen3.8-27B-NVFP4"
  speculative_config: '{"method":"mtp","num_speculative_tokens":2}'
  async_scheduling: true
  enable_prefix_caching: true
  enable_chunked_prefill: true
  trust_remote_code: true
  enable_auto_tool_choice: true
  quantization: compressed-tensors

env:
  CUTE_DSL_ARCH: sm_121a

executor_config:
  entrypoint: ""
  auto_remove: false
  user: root

# NO Mamba patch needed — vLLM nightly-aarch64 has native MTP support for Qwen 3.8.
# NO enforce-eager — CUDA graphs work with triton_attn + NVFP4 checkpoint.
# The NVFP4 checkpoint auto-applies FP8 KV cache scheme (calibrated), doubling KV capacity.
# Using FP8 checkpoint (Qwen/Qwen3.8-27B-FP8) instead of NVFP4 causes OOM during MTP — that was our root crash cause.
# NOTE: max_num_batched_tokens=8192 is REQUIRED for GDN/Mamba cache alignment (32768 crashes under concurrency).
# See runbooks/qwen-38-27b.md for alternative configs (erdaltoprak GMU 0.50, SGLang FP8, Radix DSpark).

benchmark:
  framework: llama-benchy
  pp: [2048]
  tg: [128]
  depth: [0, 4096, 8192, 16384, 32768, 65535, 100000]
  concurrency: [1, 2, 5, 10]
  prefix_caching: true
  runs: 3

command: |
  vllm serve {model} \
    --host {host} \
    --port {port} \
    --served-model-name {served_model_name} \
    --max-model-len {max_model_len} \
    --max-num-seqs {max_num_seqs} \
    --max-num-batched-tokens {max_num_batched_tokens} \
    --trust-remote-code \
    --gpu-memory-utilization {gpu_memory_utilization} \
    --enable-auto-tool-choice \
    --tool-call-parser {tool_call_parser} \
    --reasoning-parser {reasoning_parser} \
    --quantization {quantization} \
    --load-format {load_format} \
    --attention-backend {attention_backend} \
    --async-scheduling \
    --enable-prefix-caching \
    --enable-chunked-prefill \
    --speculative-config '{speculative_config}' \
    -tp {tensor_parallel} \
    -pp {pipeline_parallel}
u/styles01 — 4 days ago
▲ 111 r/DGX_Spark+6 crossposts

Hi all,
with this post I want to talk again of AudioMuse-AI, a free and open source selfhostable software to analyze your song and automatically create playlist on your supported music server like Jellyfin, Navidrome (or open subsonic api based), Emby and Lyrion:

With this post I want to celebrate two big things, first of all AudioMuse-Ai born on May 2025, so it's stil live and fully mantained after 1 years, 217 issue closed and 182 PR closed !

We also want to celebrate the new AudioMuse-AI v1.1.0 release that introduce Lyrics Semanthics similarity throug different functionality.

I'm very proud of this release because multiple time we heard that yes the mood is similar but totally different lyrics, now you can search your song also semathically with:

  • Axis-based search: Explore songs across 5 defined semantic axes, selecting one or more values that best describe the target mood or meaning.
  • Text search: Simple natural language queries (e.g., “love”, “run”) focused on lyrical meaning, not musical groove (distinct from DCLAP search).
  • Song similarity search: Use a reference track to find similar songs, weighted by default as 75% lyrical meaning and 25% audio similarity to preserve genre consistency.

Lyrics functionality off course need lyrics, the best way is to have already them in your music server OR configure in AudioMuse-AI your favourite API in the setup wizard:

Example API formats supported in Setup Wizard:

https://api.example.com/get?artist={artist}&title={title}
https://api.example.com/v1/{artist}/{title}

Anyway as a fallback is also supported the transcription with Whisper Small and if needed can be disabled in the setup wizard by setting LYRICS_ENABLED=true

Important: after the update a new analysis will do the Lyrics analysis on the already analyzed song (if enabled, enabled by default) or a full analysis (Musicnn + Clap + Lyrics) for new song. This new analysis is mandatory to use the new functionality.

I hope you will like both of this milestone and as usual, if you want to support AudioMuse-AI, please add a start on the github repository.
Thanks to be with us for our first year!

u/Old_Rock_9457 — 7 days ago

Buyer beware Amazon MSI DGX Spark

 bought an MSI spark from Amazon about 3 months ago. It died today. SInce it was past the return period, I called MSI to get the RMA started. I gave them my serial number and they came back and said the warranty was no good because it was a Chinese product. So, this was sold my Amazon, fulfilled by Amazon and it was not a real MSI spark. After hours of yelling and speaking to 10 different people I finally was able to return them for a refund. What a pain. Here is the link to the actual product. They are still selling them: Amazon.com: msi EdgeXpert AI Mini Desktop (DGX Spark Platform), NVIDIA GB10 Grace Blackwell, 128GB LPDDR5 Unified Memory, 4TB NVMe Gen5 SSD, WiFi 7, BT 5.3, NVIDIA DGX OS (Linux): 13SUS Black : Electronics

Be careful. If you buy one verify the serial numbers. I am sure this is illegal.

reddit.com
u/Ok-Wheel128 — 7 days ago

How busy is your Spark?

Are you using it (them) 24/7?

What are you using it for?

I'm mainly experimenting with different models, fine-tuning, and running benchmarks, but my usage is ~8hrs / day.

For the rest of the time, it could be used by somebody else.

reddit.com
u/hytro36 — 8 days ago

Help me understand now vs then, and the future (Someone on the fence for purchasing)

Hey everyone,

So back in dec/jan I was seriously considering getting a spark for local LLM work. I came to the conclusion its not worth it at that time. Now I am seeing alot has changed...

Questions up front:

  1. How has your perception of the utility of the DGX spark (or multiple) compare now vs back in ~dec/jan prior to NVFP4 availability?
  2. Do you envision NVFP4 will become increasingly common from post training, as well as from QAT?
  3. Do you anticipate expansion of your # of sparks in the next 2-3 years to be possible before they become obsolete?

Overall I am just super excited now that there is so much support from the community, and really tempted to pull the trigger on getting a DGX soon! Below is just some background on my use case i

reddit.com
u/ThrwAway868686 — 10 days ago

Intrigued Amateur

Hi DGX Peeps

Feel free to tell me to bugger off. I run a business and I am interested in creating a "mini me". My brief understanding is that with a DGX I can download a model weight and have it only know me, give it eyes, ears and when confident agency to execute commands locally.

My particular interest is in things like:

Reply to customers on WhatsApp

Reply to emails

Placing orders

Maybe accounting?

Basically, any task that I do that has a reasonably determinable workflow. In my minds eye I can have it set to watch mode and then approval mode and then execute mode.

Am I dreaming, or is this something a DGX locally can feasibly do?

Any advice, pointers, or if you guys know anyone that is building mini people in dgx boxes that would be appreciated.

Cheers

reddit.com
u/NewShock2391 — 9 days ago

What do you use for telementry & monitoring?

I'm coming from devops and my first instinct is to set up Grafana & Prometheus for my box, just to see what's happening inside

Was wondering what setup you have? Are you using Ansible for provisioning? Docker containers for running inference and experiments?

Or do you just SSH and run the things you want?

reddit.com
u/hytro36 — 9 days ago

Anyone run a node of 3?

Local developer looking to make the switch to the Nvidia ecosystem and the spark seems awesome. I noticed the maximum amount that you can wire together without getting a dedicated switch is Three but I know tensor parallelism needs even numbers to operate properly. Do any of you guys run AI models via pipeline parallelism with a note of Three or is it a hard two or four type set up?

reddit.com
u/habachilles — 14 days ago

Question about choosing an OS for the DGX series (gigabytes).

I'm using the Gigabyte AI TOP GB10.

I'm trying to reinstall OS, but I have a question. Is the OS file provided by GIGABYTE better, or is the OS file provided by NVIDIA better? I’d appreciate the opinions of experienced seniors.

  1. Latest version update date

Gigabyte 2026.03.02

Endivia, 2026.07.14

There’s a big difference

  1. Purpose

Not fine-tuning, but planned to use it as a photo-coding analysis server (main)

reddit.com
u/CriticismIcy6583 — 10 days ago