Running 1T+ Models on Shared GPU Meshes?
Look, I'm probably the idiot here, but I really want to be able to use things like Kimi K3 with hardware I actually own. I suspect there may be other people that are also dreaming of this day, but ultimately have a single GPU dangling by a thread from their machine. So, I wonder if a group of us could put our GPUs into a cluster and essentially time-share our much larger cluster to actually run frontier-level models as a group? The cost is simple: you get your portion of the cluster's compute for free. If you want more, you essentially beg, borrow, or buy it.
Why hasn't this been done? Have I missed something in the market?