Do you share GPUs between workloads, and what slice size do you actually need?
by•
We just launched Cloud Acropolis Kubernetes, where the GPU option is a shared NVIDIA V100 sliced into 8, 16 or 32 GB of VRAM per workload. We'd like to hear from people running inference or fine-tuning: is a fraction of a GPU enough for your workloads, or do you always need a whole card? And which VRAM size would cover most of your models?
https://cloudacropolis.com/kubernetes
3 views

Replies