▲ 0 r/ollama

Why is ollama detecting my RTX 5090 as a Vulkan device?

Hi, I am trying to see how fast Qwen 3.8 is on my RTX 5090 but I have a dual GPU setup... and somehow ollama detects the 5090 as a Vulkan instead of cuda device:

PS C:\Users\kxx> nvidia-smi -L
GPU 0: NVIDIA GeForce RTX 5060 Ti (UUID: GPU-7117d8ca-25c4-ceab-ffd2-b2714e569a49)
GPU 1: NVIDIA GeForce RTX 5090 (UUID: GPU-5e9ba80b-0394-39af-f604-df44b9c13659)
PS C:\Users\kxx> $env:CUDA_VISIBLE_DEVICES="1"
PS C:\Users\kxx> ollama serve
time=2026-08-15T14:16:25.030+02:00 level=INFO source=routes.go:1933 msg="server config" env="map[CUDA_VISIBLE_DEVICES:1 GGML_VK_VISIBLE_DEVICES: GPU_DEVICE_ORDINAL: HIP_VISIBLE_DEVICES: HSA_OVERRIDE_GFX_VERSION: HTTPS_PROXY: HTTP_PROXY: LLAMA_ARG_FIT: LLAMA_ARG_FIT_TARGET: NO_PROXY: OLLAMA_CONTEXT_LENGTH:0 OLLAMA_DEBUG:INFO OLLAMA_DEBUG_LOG_REQUESTS:false OLLAMA_EDITOR: OLLAMA_FLASH_ATTENTION:false OLLAMA_GO_TEMPLATE:true OLLAMA_GPU_OVERHEAD:0 OLLAMA_HOST:http://0.0.0.0:11434 OLLAMA_IGPU_ENABLE: OLLAMA_KEEP_ALIVE:3m0s OLLAMA_KV_CACHE_TYPE: OLLAMA_LLM_LIBRARY: OLLAMA_LOAD_TIMEOUT:5m0s OLLAMA_MAX_LOADED_MODELS:0 OLLAMA_MAX_QUEUE:512 OLLAMA_MAX_TRANSFER_STREAMS:4 OLLAMA_MODELS:C:\\Users\\kkeun\\.ollama\\models OLLAMA_NOHISTORY:false OLLAMA_NOPRUNE:false OLLAMA_NO_CLOUD:false OLLAMA_NUM_PARALLEL:1 OLLAMA_ORIGINS:[http://localhost https://localhost http://localhost:* https://localhost:* http://127.0.0.1 https://127.0.0.1 http://127.0.0.1:* https://127.0.0.1:* http://0.0.0.0 https://0.0.0.0 http://0.0.0.0:* https://0.0.0.0:* app://* file://* tauri://* vscode-webview://* vscode-file://*] OLLAMA_REMOTES:[ollama.com] OLLAMA_SCHED_SPREAD:false OLLAMA_VULKAN:true ROCR_VISIBLE_DEVICES:]"
time=2026-08-15T14:16:25.058+02:00 level=INFO source=routes.go:1935 msg="Ollama cloud disabled: false"
time=2026-08-15T14:16:25.065+02:00 level=INFO source=images.go:912 msg="total blobs: 19"
time=2026-08-15T14:16:25.067+02:00 level=INFO source=images.go:919 msg="total unused blobs removed: 0"
time=2026-08-15T14:16:25.068+02:00 level=INFO source=routes.go:1990 msg="Listening on [::]:11434 (version 0.32.13)"
time=2026-08-15T14:16:25.069+02:00 level=INFO source=runner.go:60 msg="discovering available GPUs..."
time=2026-08-15T14:16:25.080+02:00 level=WARN source=runner.go:722 msg="user overrode visible devices" CUDA_VISIBLE_DEVICES=1
time=2026-08-15T14:16:25.080+02:00 level=WARN source=runner.go:726 msg="if GPUs are not correctly discovered, unset and try again"
time=2026-08-15T14:16:25.255+02:00 level=INFO source=model_list_cache.go:112 msg="model list cache hydration complete" models=5 failures=0 elapsed=186.532ms
time=2026-08-15T14:16:25.320+02:00 level=INFO source=model_recommendations.go:177 msg="model recommendations cache sleep scheduled" wait=4h18m42.63498222s consecutive_failures=0
time=2026-08-15T14:16:30.504+02:00 level=INFO source=types.go:32 msg="inference compute" id=1 filter_id=1 library=Vulkan compute=0.0 name=Vulkan1 description="NVIDIA GeForce RTX 5090" libdirs=ollama,vulkan driver=0.0 pci_id=0000:09:00.0 type=discrete total="31.6 GiB" available="30.9 GiB"
time=2026-08-15T14:16:30.505+02:00 level=INFO source=types.go:32 msg="inference compute" id=0 filter_id=1 library=CUDA compute=12.0 name=CUDA0 description="NVIDIA GeForce RTX 5060 Ti" libdirs=ollama,cuda_v13 driver=13.3 pci_id=0000:04:00.0 type=discrete total="15.9 GiB" available="14.8 GiB"
time=2026-08-15T14:16:30.505+02:00 level=INFO source=routes.go:2040 msg="vram-based default context" total_vram="47.6 GiB" default_num_ctx=262144

The result is that I can't force ollama to use only the RTX 5090 somehow. Anyone knows how to fix this?

reddit.com
u/FrankWanders — 5 days ago

Will they ever release a 122B MoE model for 24-32GB VRAM setups?

Hi, just as a lot of you i'm quite happy with the new 3.8 27B, and first rumors of 3.8 35B A3B. But on my 5090, I get around 50 T/s with the 27B model, which is fast enough for my agentic use. So 35B A3B is nice, but the 150T/s is not essential.

It would however, be great to get something like a 122B A9B model, or something comparable, which will run at around 50T/s on an RTX 5090. The last year I found out that this is the sweet spot in terms of speed vs quality loss, so I was hoping a MoE model with bigger parameters than 35B will ever be released to be able to get more quality at that 50T/s.

Are there any companies thinking and/or working on something like this? I think this really would be a gamechanger and a local model that could be on par with the cloud models... what do you think of this?

reddit.com
u/FrankWanders — 5 days ago
▲ 137 r/Haarlem+5 crossposts

The first photo (ca 1860) of the first Dutch National Monument, 13th century castle Ruins of Brederode, and now.

For its history and a 3D reconstruction, watch the free mini documentary.

u/FrankWanders — 13 days ago

Wake up external pc on network to power my OpenClaw via telegram

Hi guys, I have a Mini PC running Linux Mint with openclaw 26.6.11 installed and working via Telegram. The bot uses an ollama server on a pc in the network, so I don´t need a cloud subscription.

To save energy, the ollama pc turns off at night and now I want to wake it up via Telegram. I have a working "wakeonlan xx:xx:xx:xx:xx:xx¨ which i can run fine in the terminal myself, but I can´t get openclaw to allow Telegram to start this command. I made a "/wake_pc2" skill, the skill.md contains:

---
name: wake-pc2
description: "Wake the second PC via Wake-on-LAN."
user-invocable: true
command-dispatch: tool
command-tool: exec
command-params:
  bin: "wakeonlan"
  args: ["xx:xx:xx:xx:xx:xx"]
---

next to that, I added this in openclaw.json, which should allow telegram / openclaw to exec wakeonlan:

"tools": {
    "web": {
      "search": {
        "enabled": false
      }
    },
    "exec": {
      "security": "allowlist",
      "ask": "off",
      "safeBins": ["wakeonlan"],
      "pathPrepend": [
        "/home/linuxbrew/.linuxbrew/bin",
        "/home/linuxbrew/.linuxbrew/sbin",
        "/home/bai/.local/bin"
      ]
    }
  },

But no matter what I try, in both telegram and the gateway chat I keep getting the same error.

❌ Tool not available: exec

All I want is to be able to execute a wakeonlan command by openclaw gateway via Telegram. After that, the LLM runs and the bot is online.

Anyone who knows how to set this up properly?

u/FrankWanders — 1 month ago

I don't get this error

Hi, after using ComfyUI quite some time I wanted to do a fresh start. So enthousiastically from https://github.com/Comfy-Org/Comfy-Desktop I downloaded the new Desktop, which looks great. However, no matter what I try... after installing the Desktop app during the installation of the comfyUI instance i keep getting this error (on a clean install, just downloaded fresh from the website, nothing else)

>> "D:\Comfy-Desktop\ComfyUI-Installs\Comfy UI\ComfyUI\.venv\Scripts\python.exe" -s ComfyUI\main.py --feature-flag show_signin_button=true --enable-manager --extra-model-paths-config "C:\Users\x\AppData\Roaming\Comfy Desktop\shared_model_paths.yaml" --input-directory D:\Comfy-Desktop\ComfyUI-Shared\input --output-directory D:\Comfy-Desktop\ComfyUI-Shared\output
[INFO] setup plugin alembic.autogenerate.schemas [INFO] setup plugin alembic.autogenerate.tables
[INFO] setup plugin alembic.autogenerate.types
[INFO] setup plugin alembic.autogenerate.constraints
[INFO] setup plugin alembic.autogenerate.defaults
[INFO] setup plugin alembic.autogenerate.comments
[INFO] Adding extra search path checkpoints D:\Comfy-Desktop\ComfyUI-Shared\models\checkpoints
[INFO] Adding extra search path classifiers D:\Comfy-Desktop\ComfyUI-Shared\models\classifiers
[INFO] Adding extra search path clip_vision D:\Comfy-Desktop\ComfyUI-Shared\models\clip_vision
[INFO] Adding extra search path configs D:\Comfy-Desktop\ComfyUI-Shared\models\configs
[INFO] Adding extra search path controlnet D:\Comfy-Desktop\ComfyUI-Shared\models\controlnet
[INFO] Adding extra search path controlnet D:\Comfy-Desktop\ComfyUI-Shared\models\t2i_adapter
[INFO] Adding extra search path diffusers D:\Comfy-Desktop\ComfyUI-Shared\models\diffusers
[INFO] Adding extra search path diffusion_models D:\Comfy-Desktop\ComfyUI-Shared\models\diffusion_models
[INFO] Adding extra search path embeddings D:\Comfy-Desktop\ComfyUI-Shared\models\embeddings
[INFO] Adding extra search path gligen D:\Comfy-Desktop\ComfyUI-Shared\models\gligen
[INFO] Adding extra search path hypernetworks D:\Comfy-Desktop\ComfyUI-Shared\models\hypernetworks
[INFO] Adding extra search path latent_upscale_models D:\Comfy-Desktop\ComfyUI-Shared\models\latent_upscale_models
[INFO] Adding extra search path loras D:\Comfy-Desktop\ComfyUI-Shared\models\loras
[INFO] Adding extra search path model_patches D:\Comfy-Desktop\ComfyUI-Shared\models\model_patches
[INFO] Adding extra search path audio_encoders D:\Comfy-Desktop\ComfyUI-Shared\models\audio_encoders
[INFO] Adding extra search path photomaker D:\Comfy-Desktop\ComfyUI-Shared\models\photomaker
[INFO] Adding extra search path style_models D:\Comfy-Desktop\ComfyUI-Shared\models\style_models
[INFO] Adding extra search path text_encoders D:\Comfy-Desktop\ComfyUI-Shared\models\text_encoders
[INFO] Adding extra search path upscale_models D:\Comfy-Desktop\ComfyUI-Shared\models\upscale_models
[INFO] Adding extra search path background_removal D:\Comfy-Desktop\ComfyUI-Shared\models\background_removal
[INFO] Adding extra search path frame_interpolation D:\Comfy-Desktop\ComfyUI-Shared\models\frame_interpolation
[INFO] Adding extra search path geometry_estimation D:\Comfy-Desktop\ComfyUI-Shared\models\geometry_estimation
[INFO] Adding extra search path optical_flow D:\Comfy-Desktop\ComfyUI-Shared\models\optical_flow
[INFO] Adding extra search path detection D:\Comfy-Desktop\ComfyUI-Shared\models\detection
[INFO] Adding extra search path vae D:\Comfy-Desktop\ComfyUI-Shared\models\vae
[INFO] Adding extra search path vae_approx D:\Comfy-Desktop\ComfyUI-Shared\models\vae_approx
[INFO] Adding extra search path clip D:\Comfy-Desktop\ComfyUI-Shared\models\clip
[INFO] Adding extra search path unet D:\Comfy-Desktop\ComfyUI-Shared\models\unet
[INFO] Setting output directory to: D:\Comfy-Desktop\ComfyUI-Shared\output
[INFO] Setting input directory to: D:\Comfy-Desktop\ComfyUI-Shared\input
[START] Security scan
[DONE] Security scan
** ComfyUI startup time: 2026-07-06 14:00:35.185
** Platform: Windows
** Python version: 3.13.12 (main, Feb 12 2026, 00:38:53) [MSC v.1944 64 bit (AMD64)]
** Python executable: D:\Comfy-Desktop\ComfyUI-Installs\Comfy UI\ComfyUI\.venv\Scripts\python.exe
** ComfyUI Path: D:\Comfy-Desktop\ComfyUI-Installs\Comfy UI\ComfyUI
** ComfyUI Base Folder Path: D:\Comfy-Desktop\ComfyUI-Installs\Comfy UI\ComfyUI
** User directory: D:\Comfy-Desktop\ComfyUI-Installs\Comfy UI\ComfyUI\user
** ComfyUI-Manager config path: D:\Comfy-Desktop\ComfyUI-Installs\Comfy UI\ComfyUI\user\__manager\config.ini
** Log path: D:\Comfy-Desktop\ComfyUI-Installs\Comfy UI\ComfyUI\user\comfyui.log
[INFO] [PRE] ComfyUI-Manager
[ERROR] Failed to import comfy_kitchen, Error: cannot import name 'TensorWiseINT8Layout' from 'comfy_kitchen.tensor' (D:\Comfy-Desktop\ComfyUI-Installs\Comfy UI\ComfyUI\.venv\Lib\site-packages\comfy_kitchen\tensor\__init__.py), fp8 and fp4 support will not be available.
[WARNING] comfy_kitchen does not support stochastic FP8 rounding, please update comfy_kitchen.
[INFO] Checkpoint files will always be loaded safely. Traceback (most recent call last):
File "D:\Comfy-Desktop\ComfyUI-Installs\Comfy UI\ComfyUI\main.py", line 227, in <module>
import execution
File "D:\Comfy-Desktop\ComfyUI-Installs\Comfy UI\ComfyUI\execution.py", line 18, in <module>
import comfy.model_management
File "D:\Comfy-Desktop\ComfyUI-Installs\Comfy UI\ComfyUI\comfy\model_management.py", line 36, in <module>
import comfy_aimdo.vram_buffer
ModuleNotFoundError: No module named 'comfy_aimdo.vram_buffer'

I tried everything, did a few complete uninstalls of the ComfyUI Dekstop app (with Revo uninstaller, so no registry left), tried to remove the instance and create another one... the only thing i do is "skip and install" (because i don't want any of the startup models.

The only thing I can think of is that I have a multi gpu setup (5060 ti and 5090) but I didn't use it at all, comfy just doesn't get through the installation steps on a clean system. AI tells me it's a memory error (for obvious reasons) and starts with all kind of "clean your system" and "tweak this and that" solutions but: this is a clean install. I did not do anything else than download the Desktop app and it doesn't get through basic installation. It seems that tweaking isn't the solution here: this should work out of the box.

Anyone who knows what's going wrong here? Is this a bug in comfyUI desktop with multiple gpu systems which causes this error? Or something else I forgot?

u/FrankWanders — 1 month ago
▲ 7 r/comfyui_elite+1 crossposts

[LTX 2.3] multi gpu workflows / advancements / experiences

Hi, I recently bought a 5060 Ti next to my 5090. The main idea is to use this to free up vram for the base model by putting the text encoder / vae on the 5060 Ti so LTX 2.3 will be reserved for the 5090.

I found this older workflow for a 3090 / 4060 ti setup: https://www.reddit.com/r/comfyui/comments/1rr67bc/ltxvideo_23_workflow_for_dualgpu_setups_3090_4060/

But since that time, no new workflows and/or experiences have been shared. I tried to run this workflow, but no matter what I try, I get an "Cannot read properties of null (reading 'replace')" error as soon as I start things up. There are no additional terminal errors so I can't find a solution.

Was wondering if anyone here has succesful LTX 2.3 experiences with multi gpu setups and/or knows possible solutions.

I know because of the setup in comfyUI multiple gpu has limitations, but it would be great to improve the things that are possible by sharing working workflows like the one shown above. An additional GPU is quite a cheap way to get extra vram, so any tips to make use of it more in ComfyUI and/or recent workflows / guides are appreciated!

reddit.com
u/FrankWanders — 2 months ago
▲ 3 r/comfyui+1 crossposts

LTX 2.3 BF16 or nvfp4 with RTX 5090.

Hi guys a short question, I thought I understand quite a bit about which things I can and can't to get an optimal performance/quality basis with my setup (RTX 5090 with 64GB of DDR4-3200).

I thought in ComfyUI for LTX, my best bet will be the nvfp4 version of the LTX 2.3 model because it's 22GB in size, leaving 10GB for context/calculations. But now, I am reading that not even fp8, but the full FP16 would be doable with my setup, which ofcourse will give full quality.

But I don't get this, the BF16 model is 46 GB in size, how is this going to fit/work?

Are the claims I read online that it should be the best choice for an RTX 5090 false, or do I not understand something? Running my ollama setup has learned i usually can run a 22-24GB model max, and the rest then is context to be able to generate text answers fast, but does this work differently in comfyUI?

And how does it work with context for the prompt and the lora's, how much vram does it use?

Thanks for answers and/or resources (video/links) to help me understand how this works!

reddit.com
u/FrankWanders — 2 months ago
▲ 351 r/ColorizedStatues+7 crossposts

Impression of Julius Caesar's face using the famous 'Tusculum bust' for reference

The reconstruction was used for the mini documentary about Atuatuca Tungrorum (Tongeren, Belgium) in the link.

u/FrankWanders — 2 months ago

Running OpenClaw locally with Caveman optimalizations?

Hi, since about 6 months I have tried to use openclaw with local models, it turned out to be useful with the free Gemini 3 Flash tier and combining that with local models on my RTX 5090 with 32GB vram. Local models just weren't good enough.

But since Gemma 4 and Qwen 3.6 27B, I finally found two models that are quite useable. With 192K context, which is acceptable, both models fit in my vram with the nvfp4 model. But ofcourse it would be the best to get 256K context and that's why I was wondering if it's already possible to implement Caveman into openclaw for local models, which will not increase context ofcourse but uses that context much more efficiently.

Anyone knows if this is possible or in the making by the openclaw devs?

reddit.com
u/FrankWanders — 2 months ago
▲ 112 r/Utrecht+4 crossposts

In 2024, the restoration of the Dom Tower in Utrecht, which stands separate from its original church, was completed.

u/FrankWanders — 2 months ago

Historical 3D reconstruction of the Roman temple in Tongeren, Belgium

The temple, likely dedicated to Jupiter, was surrounded by a colonnade which is also visualized, and was a sacred complex in Atuatuca Tungrorum around 100 A.D. In the video, the Roman history of the city is visualized (created using archeological sources & feedback from 3 (city) archeologists).

u/FrankWanders — 2 months ago
▲ 125 r/EuropeanCulture+10 crossposts

Historical 3D reconstruction of the Roman temple in Tongeren

The temple, likely dedicated to Jupiter, was surrounded by a colonnade which is also visualized, and was a sacred complex around 100 A.D. in Atuatuca Tungrorum. In the video, the Roman history of the city is visualized (created using archeological sources & feedback from 3 (city) archeologists).

u/FrankWanders — 2 months ago

"The Chzar's Palace", a color photogrom of the Kremlin dating back to 1890-1900, just before Nicholas II abdicated as last Tsar in 1917, bringing an end to the 370-year-old Russian Tsarist Empire.

u/FrankWanders — 4 months ago