MIG and vGPU in corporate infrastructure: GPU partitioning models
In corporate infrastructure, a powerful GPU is far from always working on a single task. Multiple teams’ tasks, AI applications, virtual workstations, or inference tasks can run on a single accelerator. Allocating a separate GPU to each user or service is expensive, and if the load is low, part of the computing resources will remain idle.
Therefore, the task arises to divide the resources of a single GPU among several workloads while maintaining acceptable performance. With CPUs, this approach has long become the standard: the hypervisor distributes processor cores and RAM among virtual machines. With GPUs, it’s more complicated. Accelerators have their own computing units and video memory, and different tasks can simultaneously compete for the same resources.
NVIDIA offers several technologies for such separation. MIG (Multi-Instance GPU) allows you to split a physical GPU into separate, hardware‑isolated instances with allocated resources. NVIDIA vGPU is used to virtualize the GPU and provide its resources to virtual machines.
Both technologies help make more efficient use of expensive accelerators, but they solve this problem in different ways. The choice between MIG and vGPU depends on the infrastructure architecture, the type of workload, the isolation requirements, and how exactly GPU resources should be distributed among users and applications.
What is MIG?
MIG (Multi-Instance GPU) is an NVIDIA technology that allows a single physical GPU to be hardware‑divided into multiple independent instances. It appeared in the Ampere generation and is supported by a number of modern NVIDIA server accelerators.
Each MIG instance receives a dedicated portion of the GPU’s computing resources and video memory and can be used as a separate accelerator for a specific task or user. The number of instances and the available amount of resources depend on the GPU model and the MIG profiles it supports. For example, for the RTX PRO 6000 Blackwell Server Edition server with 96 GB of RAM, configurations with different memory capacities are available for each instance.
The MIG configuration can be adjusted to suit current tasks, but such resource redistribution requires stopping active loads on the GPU. Therefore, this technology is more convenient for infrastructure where the configuration is set in advance and does not require constant changes.
MIG is well‑suited for inference, multi‑user AI platforms, test environments, and services with requirements for predictable performance. Hardware isolation allows each load to be allocated its own portion of GPU resources and reduces the impact of neighboring tasks on each other.
Pic 1. Multi-Instance GPU (MIG)
What is vGPU?
If MIG shares resources at the level of the GPU itself, then NVIDIA vGPU works at the virtualization level. The technology allows multiple virtual machines to use the resources of a single physical accelerator. Each VM receives a virtual graphics adapter and works with it as with a separate device.
Resource allocation is managed by NVIDIA vGPU Manager, which is installed on the hypervisor side. Guest NVIDIA drivers are used inside the virtual machines. The administrator sets vGPU parameters using profiles: they determine the available video memory and other characteristics of the virtual GPU.
The resources of a physical accelerator can be distributed among virtual machines in various ways. One of them is time-slicing, in which several vGPUs take turns accessing the GPU’s computing resources. This approach allows for higher utilization of the accelerator and enables the placement of several virtual machines on it.
On GPUs with MIG support, another option is possible — MIG‑backed vGPU. In this case, the virtual machine is provided with the resources of a hardware‑isolated MIG instance. This approach combines the capabilities of the vGPU infrastructure with the hardware sharing of MIG resources.
The difference between time-slicing and MIG-backed vGPU is particularly important for corporate infrastructure: the chosen mode determines the density of virtual machine placement, the degree of resource isolation, and the predictability of performance. Let’s take a closer look at both options.
Pic 2. NVIDIA vGPU Architecture
Time-slicing vGPU
In time-slicing mode, several virtual machines use the computing resources of a single physical GPU. The NVIDIA vGPU Manager scheduler distributes the accelerator’s operating time among the vGPUs, sequentially granting them access to its computing blocks. A specific vGPU profile is used for each virtual machine.
The main advantage of this approach is high GPU utilization. Several virtual machines can be placed on a single accelerator, and its resources can be distributed among different users or applications without allocating a separate physical GPU for each workload.
However, computing resources remain shared. If several virtual machines simultaneously generate a high load, they begin to compete for access to the GPU. As a result, task execution times and delays may vary depending on the activity of neighboring VMs. Therefore, time‑slicing is better suited for scenarios where flexibility and density are more important than strictly predictable performance.
Typical examples include data science work environments, model development and testing, and inference without strict SLA requirements.
MIG-backed vGPU
MIG-backed vGPU combines MIG hardware partitioning with the capabilities of virtual infrastructure. First, the physical GPU is divided into several MIG instances, each of which is allocated a specific portion of computing resources and memory. Then these resources are provided to virtual machines via NVIDIA vGPU.
As a result, each VM receives its own hardware‑isolated portion of the accelerator. Unlike time‑slicing, virtual machines do not compete for the same GPU computing resources. The load on a single VM has a significantly smaller impact on the performance of neighboring machines, so the system’s behavior becomes more predictable.
At the same time, the GPU remains part of the virtual infrastructure: the administrator can allocate resources between VMs and manage them through a supported virtualization platform. This approach is suitable for corporate environments where it is necessary to combine virtualization with stricter isolation of GPU resources and stable performance.
Deploying a MIG-backed vGPU requires a compatible GPU, a supported virtualization platform, appropriate software, and configuration of MIG profiles on the nodes. Depending on the chosen configuration, NVIDIA licensing will also be required. This architecture is more complex to configure than a typical time‑slicing system, so it makes sense to use it in situations where the advantages of hardware isolation are truly important.
MIG-backed vGPU is suitable for inference with requirements for stable performance, SLA-critical AI services, and multi-user platforms. For example, several teams or clients can work on a single physical GPU, receiving hardware resources allocated for their virtual machines.
In some configurations, MIG can be combined with time-slicing. Several vGPUs use the resources of a single MIG instance, while neighboring workload groups remain separated at the MIG level. This approach allows for an increase in the density of virtual machines while maintaining hardware isolation between individual workload groups.
Comparison of GPU partitioning models
MIG and vGPU solve different problems and can be used either separately or together. MIG divides resources directly at the GPU level, while vGPU provides GPU resources to virtual machines. At the same time, vGPU can use time‑slicing or operate on top of MIG instances.
The table compares three approaches.
Criterion
MIG
Time-Sliced vGPU
MIG-Backed vGPU
Resource partitioning
Hardware partitioning of a physical GPU into independent MIG instances
GPU compute time is shared between multiple vGPUs
vGPUs are provisioned on hardware-isolated MIG instances
Resource isolation
High: each MIG instance receives a dedicated portion of GPU resources
Lower: virtual machines share the GPU’s compute resources
High: MIG hardware isolation is combined with virtualization
Performance predictability
High within the allocated resources
Depends on the activity of other vGPUs running on the same GPU
High within the resources allocated to the corresponding MIG instance
Primary use cases
Isolated AI and compute workloads, containerized environments
VDI, development, testing, data science, and workloads with variable utilization
Multi-user AI platforms and services with higher isolation requirements
Use with VMs
Possible, but MIG itself is not a VM virtualization technology
Yes, each VM receives its own vGPU
Yes, the VM receives a vGPU backed by a MIG instance
Resource allocation flexibility
Medium: changing the MIG configuration may require stopping workloads that use the GPU
High: GPU resources can be flexibly distributed across multiple VMs
Medium: allocation options are limited by the selected MIG configuration
GPU utilization efficiency
Enables a large GPU to be divided among multiple isolated workloads
Increases VM density for workloads with variable GPU utilization
Combines high virtualization density with hardware-level isolation
Infrastructure complexity
Requires MIG profile configuration and management
Requires a vGPU-compatible virtualization infrastructure
Requires both MIG and vGPU support and more careful configuration planning
The choice depends on the nature of the workload. MIG is suitable when hardware isolation of GPU resources is important. Time-sliced vGPU is convenient when flexibility and efficient use of the accelerator by multiple virtual machines are a priority. MIG-backed vGPU combines both approaches and is suitable for virtualized environments where stricter isolation of GPU resources is required.
NVIDIA vGPU Licensing
NVIDIA vGPU is a commercial technology, so when designing a virtual GPU infrastructure, you need to take into account the licensing of the software stack. It applies to the NVIDIA vGPU components that enable the operation of virtual GPUs on the hypervisor side and inside virtual machines.
The NVIDIA License System (NLS) is used to manage licenses. Depending on the infrastructure architecture, licenses can be provided via the cloud Cloud License Service (CLS) or the local Delegated License Service (DLS). The second option is suitable for isolated corporate environments where virtual machines or servers do not have direct access to the internet.
The available set of features depends on the selected software product and the type of NVIDIA license. For example, certain options are designed for virtual desktops and office applications, professional graphics and engineering tasks, as well as computational and AI workloads.
When choosing a license, it is important to take into account the intended usage scenario, the GPUs being used, the virtualization platform, and the required features. This should be done at the infrastructure design stage: licensing affects the solution architecture and its total cost.
GPU Cloud rental from ITGLOBAL.COM
To deploy a company’s own GPU infrastructure, it is necessary to select hardware, configure virtualization, organize licensing, and ensure ongoing platform support. An alternative option is to rent GPU resources in the cloud, where a significant portion of these tasks is handled by the cloud provider.
ITGLOBAL.COM provides GPU Cloud for corporate AI and computing workloads. The infrastructure supports NVIDIA vGPU and allows the use of various resource allocation models depending on the project requirements, including time‑slicing and MIG‑backed vGPU.
MIG-backed vGPU is suitable for tasks where isolation and predictable resource allocation are particularly important. In this mode, each virtual machine receives resources from a hardware‑isolated MIG instance, while still retaining the benefits of a virtual infrastructure.
At GPU Cloud ITGLOBAL.COM, MIG-backed vGPU technology is available on the NVIDIA RTX PRO 6000 Blackwell Server Edition.
text_with_btn btn="Rent a cloud server with a GPU" id="50"Request a consultation /text_with_btn
Available GPUs
ITGLOBAL.COM offers several NVIDIA GPU models for AI, machine learning, inference, virtual workstations, and other computing tasks.
Характеристика
NVIDIA H200
RTX PRO 6000 Blackwell Server Edition
NVIDIA A100/A800
NVIDIA A16
NVIDIA L40S
Архитектура
Hopper
Blackwell
Ampere
Ampere
Ada Lovelace
CUDA-ядер
16 896
24 064
6 912
1 280 × 4 GPU
18 176
Tensor-ядер
528
752
432
160 × 4 GPU
568
RT-ядер
—
188
—
40 × 4 GPU
142
Видеопамять
141 ГБ HBM3e
96 ГБ GDDR7
80 ГБ HBM2e
64 ГБ GDDR6 (16 ГБ × 4)
48 ГБ GDDR6
Поддержка MIG
✓
✓
✓
—
—
Пропускная способность памяти
4,8 ТБ/с
1,6 ТБ/с
2 ТБ/с
232 ГБ/с на GPU
864 ГБ/с
For supported configurations, ready‑made vGPU profiles with a specific amount of video memory and computing resources of the virtual machine are available. This allows you to choose a configuration tailored to a specific workload: from development and small inference tasks to resource‑intensive AI applications that require a significant portion or the entire available GPU memory.
The choice of profile depends on the accelerator model, the nature of the workload, and the requirements for performance and isolation.
vGPU profiles
For GPU Cloud ITGLOBAL.COM, several vGPU configurations are available. They differ in the amount of GPU video memory, as well as the number of vCPUs and RAM of the virtual machine. This allows you to select resources tailored to a specific workload without allocating an entire GPU in cases where its full performance is not required.
Модель GPU
VRAM
vCPU
RAM
NVIDIA A100/A800
10 ГБ
4
16 ГБ
NVIDIA A100/A800
20 ГБ
8
32 ГБ
NVIDIA A100/A800
40 ГБ
16
64 ГБ
NVIDIA A100/A800
80 ГБ
24
128 ГБ
NVIDIA RTX PRO 6000 Blackwell
12 ГБ
4
16 ГБ
NVIDIA RTX PRO 6000 Blackwell
24 ГБ
8
32 ГБ
NVIDIA RTX PRO 6000 Blackwell
48 ГБ
16
64 ГБ
NVIDIA RTX PRO 6000 Blackwell
96 ГБ
24
128 ГБ
Small profiles are suitable for development, testing, and less resource‑intensive inference. Configurations with larger VRAM can be used for more demanding AI workloads and models that require more GPU memory.
If the standard profile is not sufficient, a configuration with the full amount of available accelerator memory can be used for high‑load tasks.
Conclusion
The choice of GPU partitioning model affects performance, resource isolation, management flexibility, and the cost of enterprise AI infrastructure.
Time-slicing vGPU allows you to distribute one physical GPU among several virtual machines and increase its utilization. This option is suitable for workloads where performance changes are acceptable depending on the activity of neighboring VMs.
MIG divides the GPU into hardware‑isolated instances with dedicated resources. This approach ensures more predictable workload behavior, but the resource configuration is less flexible and requires prior planning.
MIG‑backed vGPU combines MIG’s hardware isolation with the capabilities of virtual infrastructure. It is suitable for multi‑user AI platforms, inference services, and other enterprise workloads where resource isolation and stable performance are important.
ITGLOBAL.COM provides GPU Cloud with support for various GPU resource allocation models, including MIG-backed vGPU. The client receives a ready‑made virtual infrastructure, and the provider handles the setup and maintenance of the physical GPU platform. This allows running AI workloads without purchasing their own GPU equipment and deploying the entire infrastructure independently.
FAQ
How does MIG-backed vGPU differ from MIG?
MIG divides a physical GPU into hardware‑isolated instances with dedicated resources. MIG-backed vGPU uses these MIG instances in a virtual infrastructure: a virtual machine receives a vGPU based on a dedicated MIG instance, and the administrator can manage GPU resources through a supported virtualization platform.
Thus, MIG is responsible for hardware resource partitioning, while MIG-backed vGPU adds virtualization and GPU management capabilities to it in a virtual machine environment.
Can MIG be used without a vGPU license?
Yes. MIG is a hardware feature of supported NVIDIA GPUs, so when using MIG directly, for example in a bare‑metal environment, a separate vGPU license is not required. Compatible NVIDIA drivers and a software environment with MIG support are needed for operation.
Licensing becomes relevant when MIG is used together with NVIDIA vGPU and MIG instances are provided to virtual machines via vGPU Manager. In this case, the licensing requirements depend on the NVIDIA software product being used and the configuration of the virtual infrastructure.
When is time-slicing vGPU sufficient, and when is it better to use MIG-backed vGPU?
Time-slicing is suitable for workloads where flexibility and high GPU utilization are important, and slight performance variations are acceptable. Typical scenarios include model development and testing, interactive environments, and tasks with inconsistent loads.
MIG-backed vGPU is suitable for services that require stricter resource isolation and predictable performance. It can be used for corporate inference, multi‑user AI platforms, and SLA‑critical workloads where the impact of neighboring virtual machines must be minimized.
Is MIG available on all NVIDIA GPUs?
No. MIG is supported only by certain NVIDIA GPU models. The technology appeared in the Ampere generation and is available on a number of modern accelerators, including the NVIDIA A100/A800, H200, and RTX PRO 6000 Blackwell Server Edition.
However, the availability of MIG depends on the specific GPU model: for example, the NVIDIA A16 and L40S do not support this technology. Therefore, before designing the infrastructure, it is necessary to check the support for MIG for the selected accelerator and the software stack being used.
What happens to virtual machines when the vGPU license expires?
If there is no valid NVIDIA vGPU license, the virtual machine continues to run, but the functionality and performance of the GPU may be limited. This can lead to a noticeable decrease in the performance of graphics applications, virtual workstations, and computing workloads.
Therefore, in a corporate infrastructure, it is important to monitor the status of licenses and ensure constant access to the configured NVIDIA licensing system.
What vGPU profiles are available on the RTX PRO 6000 Blackwell?
At GPU Cloud ITGLOBAL.COM, four configurations based on the RTX PRO 6000 Blackwell Server Edition are available: 12, 24, 48, and 96 GB of video memory. Each configuration corresponds to a specific amount of computing resources for the virtual machine - from 4 vCPU and 16 GB RAM to 24 vCPU and 128 GB RAM.
The profile is selected based on the workload requirements: from development, testing, and tasks with low GPU memory consumption to resource‑intensive AI applications that require the full capacity of the accelerator’s memory.
text_with_btn btn="Rent a cloud server with a GPU" id="50"Request a consultation /text_with_btn