Do you share GPUs between workloads, and what slice size do you actually need?

by•

We just launched Cloud Acropolis Kubernetes, where the GPU option is a shared NVIDIA V100 sliced into 8, 16 or 32 GB of VRAM per workload. We'd like to hear from people running inference or fine-tuning: is a fraction of a GPU enough for your workloads, or do you always need a whole card? And which VRAM size would cover most of your models?

3 views

Add a comment

Replies

Be the first to comment