Software - General
1866998 Members
1252 Online
110505 Solutions
New Discussion

Re: Driving GPU Efficiency and Cost Optimization with NVIDIA Multi-Instance GPU (MIG)

 
Syed_Saifulla
HPE Pro

Driving GPU Efficiency and Cost Optimization with NVIDIA Multi-Instance GPU (MIG)

NVIDIA MIG Implementation on DGX Kubernetes Cluster

Technical Knowledge-Sharing Document

Scope: 10 DGX nodes | 80 Nvidia Cards | H100/H200 GPUs | GPU Operator | Air-gapped Kubernetes

 

1. Introduction, Environment and Business Context

1.1 Purpose

This document shares the practical approach used to enable NVIDIA Multi-Instance GPU (MIG) on a DGX Kubernetes platform.

The platform uses a flexible allocation model:

  • Some GPUs remain available as complete physical cards.
  • Selected GPUs are divided into MIG instances.
  • Different MIG profiles are used based on workload memory, compute, and performance requirements.

The document is intended for platform engineers, Kubernetes administrators, GPU infrastructure teams, and application teams evaluating a similar GPU-sharing model.

1.2 Environment

The implementation covers 10 bare-metal NVIDIA DGX worker nodes running a Kubeadm-based Kubernetes cluster.

Component

Environment

DGX worker nodes

10

NVIDIA H100

4 nodes x 8 GPUs = 32 GPUs

NVIDIA H200

6 nodes x 8 GPUs = 48 GPUs

Total physical GPUs

80

NVIDIA software

GPU Operator v25.10.1

Deployment model

Offline and air-gapped

 

The 80 physical GPUs are not configured with one common MIG profile. The platform contains a combination of:

  • Full physical GPUs for large or performance-sensitive workloads
  • Smaller MIG profiles for lightweight workloads
  • Medium and larger MIG profiles for workloads requiring additional memory or compute

1.3 Business Context

H100 and H200 GPUs are high-value infrastructure resources. Assigning a complete GPU to a workload that uses only a small portion of its memory or compute capacity leaves the remaining capacity unavailable to other applications.

This can lead to:

  • Underutilized GPU capacity
  • Unnecessary full-GPU reservations
  • Longer workload scheduling queues
  • Reduced workload concurrency
  • Limited visibility into actual GPU consumption
  • Premature demand for additional hardware

The objective is to match GPU allocation with workload demand. Smaller workloads can use appropriately sized MIG instances, while workloads requiring complete-card performance can continue using full physical GPUs.

This approach can improve utilization, increase workload concurrency, and provide better capacity visibility. The expected business value is primarily capacity optimization and cost avoidance, including the possibility of deferring additional GPU procurement.

Key point: MIG does not create additional physical compute. Its value comes from making existing GPU capacity easier to allocate, share, isolate, and manage efficiently.

2. Challenge and Solution Approach

2.1 Challenge

The original allocation model assigned a complete physical GPU to each workload. This was suitable for training and high-demand applications, but smaller inference, testing, and development workloads often used only part of the available capacity.

Because the remaining capacity could not be scheduled for another application, the model resulted in:

  • Underutilized GPU resources
  • Unnecessary full-GPU reservations
  • Longer waiting times for GPU workloads
  • Limited workload concurrency
  • Difficulty distinguishing actual shortages from inefficient allocation

Workload requirements also varied significantly. Some applications required only a small portion of a GPU, while others required more memory, higher compute capacity, or direct access to a complete physical card.

2.2 Solution Approach

NVIDIA MIG was enabled on selected H100 and H200 GPUs to provide smaller, hardware-isolated GPU resources that Kubernetes can schedule independently.

A mixed allocation model was selected:

  • Full physical GPUs remain available for training, large models, and performance-sensitive workloads.
  • Smaller MIG profiles support lightweight inference, development, and testing.
  • Medium and larger MIG profiles support applications requiring additional memory or compute.
  • Profile assignments are selected according to workload demand and observed utilization.

This model provides flexibility without unnecessarily fragmenting the complete GPU fleet.

Key point: The objective is not to divide every GPU into the maximum number of MIG instances. The objective is to maintain the correct balance of full GPUs and MIG profiles based on workload demand.

2.3 Expected Value

  • Better use of existing GPU capacity
  • Increased workload concurrency
  • Reduced unnecessary full-GPU allocation
  • Shorter scheduling queues for smaller workloads
  • Improved resource-consumption visibility
  • Better capacity planning
  • Potential cost avoidance through deferred hardware expansion

3. MIG Architecture and Design

3.1 How MIG Works

NVIDIA MIG divides a supported physical GPU into multiple hardware-isolated instances. Each instance receives a defined portion of GPU memory and compute resources and is exposed to Kubernetes as an independently schedulable resource.

The platform supports:

  • Full physical GPUs for compute-intensive or performance-sensitive workloads
  • MIG instances for workloads that operate efficiently with smaller GPU allocations

3.2 Component Flow

DGX H100 or H200 Hardware
            |
            v
      NVIDIA Driver
            |
            v
       GPU Operator
            |
            v
      MIG Manager <--- MIG ConfigMap
            |
            v
 Full-GPU or MIG Configuration
            |
            v
   NVIDIA Device Plugin
            |
            v
 Kubernetes Allocatable Resource
            |
            v
   Kubernetes Scheduler
            |
            v
      Application Pod

3.3 Component Responsibilities

  • NVIDIA Driver: Manages the physical GPU and provides MIG capability.
  • GPU Operator: Manages NVIDIA software components in the Kubernetes cluster.
  • MIG ConfigMap: Defines approved MIG profiles and partitioning layouts.
  • MIG Manager: Applies the configuration requested through the node label.
  • NVIDIA Device Plugin: Advertises full-GPU and MIG resources to Kubernetes.
  • Kubernetes Scheduler: Places Pods where the requested resource is available.
  • Application Pod: Requests a full GPU or a specific MIG resource.

3.4 MIG Profiles and Geometry

A MIG profile defines the compute and memory available to an instance. A MIG geometry defines how multiple instances are arranged on a physical GPU.

Different geometries can support smaller, medium, and larger MIG profiles, while complete physical GPUs remain available for workloads requiring full-card performance. Profile availability depends on the GPU model and active configuration.

3.5 Mixed MIG Strategy

mig:
  strategy: mixed

This allows Kubernetes to advertise profile-specific resources such as:

nvidia.com/mig-1g.10gb
nvidia.com/mig-1g.18gb

It also supports different configurations across H100 and H200 nodes. A geometry change may interrupt workloads and require a GPU reset or node reboot.

3.6 Profile Selection Criteria

Profile selection should use actual workload behavior, including:

  • GPU memory consumption
  • Compute utilization
  • Application latency and throughput
  • Expected concurrency
  • Model size and workload duration
  • Requirement for full-GPU performance

Key point: The design goal is to match each workload with an appropriate resource while retaining sufficient full GPUs for complete-card workloads.

4. Implementation Approach

The implementation was completed in a controlled sequence. Each node was validated before being returned to service.

Step 1: Prepare the Node


# kubectl cordon <node>
# kubectl drain <node>  --ignore-daemonsets  --delete-emptydir-data

Before changing the GPU configuration, stop active GPU workloads and host-level GPU processes. Use a maintenance window when a reset or reboot may be required.

Step 2: Enable MIG Manager

The GPU Operator was deployed from a local Helm chart using internally mirrored container images.

# helm upgrade gpu-operator ./gpu-operator-offline  --namespace gpu-operator  --values ./gpu-operator-offline/values.yaml --set mig.strategy=mixed  --set migManager.enabled=true  --set migManager.config.name=my-custom-mig-configs

Main settings:

  • mig.strategy=mixed enables profile-specific resources.
  • migManager.enabled=true enables automated MIG configuration.
  • migManager.config.name identifies the custom ConfigMap.

Step 3: Create the MIG Configuration (Refrence)

apiVersion: v1
kind: ConfigMap
metadata:
  name: my-custom-mig-configs
  namespace: gpu-operator
data:
  config.yaml: |
    version: v1
    mig-configs:
      all-disabled:
      - devices: all
        mig-enabled: false

      h100-profile:
      - devices: all
        mig-enabled: true
        mig-devices:
          "1g.10gb": 7

      h200-profile:
      - devices: all
        mig-enabled: true
        mig-devices:
          "1g.18gb": 7

      only-gpu0-test:
      - devices: [0]
        mig-enabled: true
        mig-devices:
          "1g.10gb": 7

      - devices: [1, 2, 3, 4, 5, 6, 7]
        mig-enabled: false

Apply the configuration:

# kubectl apply -f mig-config.yaml

Additional profiles can be defined according to supported GPU geometries and measured workload requirements.

Step 4: Apply the Required Profile

For an H100 node:

kubectl label node <h100-node> \
  nvidia.com/mig.config=h100-profile \
  --overwrite

For an H200 node:

kubectl label node <h200-node> \
  nvidia.com/mig.config=h200-profile \
  --overwrite

For limited testing:

kubectl label node <test-node> \
  nvidia.com/mig.config=only-gpu0-test \
  --overwrite

Testing on a selected GPU reduces the risk of applying an incorrect configuration across the entire node.

Step 5: Monitor the Transition

kubectl get node <node> --show-labels \
  | grep nvidia.com/mig

kubectl logs -n gpu-operator \
  -l app=nvidia-mig-manager \
  --tail=200

A successful configuration should report:

nvidia.com/mig.config.state=success

If the configuration remains pending or changes to failed, review the MIG Manager logs before reapplying the profile.

Step 6: Validate the Resources

nvidia-smi
nvidia-smi -L

kubectl describe node <node> \
  | grep -i -E 'nvidia.com/gpu|nvidia.com/mig'

The resource names and counts should match the assigned profile.

Step 7: Run a Test Workload

apiVersion: v1
kind: Pod
metadata:
  name: mig-validation
  namespace: team-alpha
spec:
  restartPolicy: Never
  containers:
  - name: cuda-test
    image: nvcr.io/nvidia/cuda:12.3.0-base-ubuntu22.04
    command:
    - bash
    - -lc
    - nvidia-smi
    resources:
      limits:
        nvidia.com/mig-1g.10gb: 1

kubectl apply -f mig-validation.yaml
kubectl get pod mig-validation -n team-alpha -o wide
kubectl describe pod mig-validation -n team-alpha
kubectl exec -n team-alpha mig-validation -- nvidia-smi

Step 8: Return the Node to Service

kubectl uncordon <node>

Key point: A Running MIG Manager Pod does not confirm successful implementation. Validate the complete path from physical GPU configuration to application-level device visibility.

5. Workload Consumption, Governance and Cost Control 5.1 Resource Requests

H100 MIG request:

resources:
  limits:
    nvidia.com/mig-1g.10gb: 1

H200 MIG request:

resources:
  limits:
    nvidia.com/mig-1g.18gb: 1

Full physical GPU request:

resources:
  limits:
    nvidia.com/gpu: 1

If the requested resource is unavailable, the Pod remains Pending.

kubectl describe pod <pod-name> -n <namespace>
kubectl get events -n <namespace> --sort-by=.lastTimestamp

5.2 Namespace Resource Quota

apiVersion: v1
kind: ResourceQuota
metadata:
  name: team-mig-quota
  namespace: team-alpha
spec:
  hard:
    requests.nvidia.com/mig-1g.10gb: "7"

This prevents one namespace from consuming all available instances and establishes a clear capacity boundary.

5.3 Recommended Allocation Model

Workload type

Recommended allocation

Development and lightweight inference

Smaller MIG profile

Medium inference or memory-sensitive workloads

Medium or larger MIG profile

Training and performance-sensitive workloads

Full physical GPU

 

Profile selection should be based on measured utilization rather than assumptions.

5.4 Monitoring and Capacity Visibility

Monitoring should cover:

  • Physical GPU and MIG instance utilization
  • GPU memory usage
  • Allocation by namespace, Pod, and application
  • Pending GPU workloads
  • Full-GPU versus MIG consumption
  • Idle or consistently underused allocations

5.5 Business and Cost Value

Workload-specific allocation can reduce unnecessary full-GPU reservations, increase concurrency, preserve complete GPUs for demanding workloads, reduce scheduling delays, and improve capacity planning.

The financial benefit should be described as capacity optimization and cost avoidance. MIG does not reduce the cost of hardware already purchased, but it can reduce or delay the need for additional infrastructure.

Key point: Cost efficiency means selecting the correct resource size while maintaining a balanced pool of full GPUs and MIG instances.

6. Implementation Challenges and Lessons Learned 6.1 MIG Configuration Missing After Reboot

Observation: Expected MIG resources were unavailable after restart.

Cause: The required nvidia.com/mig.config label was missing.

kubectl label node <node> \
  nvidia.com/mig.config=<profile-name> \
  --overwrite

Confirm nvidia.com/mig.config.state=success.

6.2 Invalid Disabled Profile

A disabled profile must not contain a mig-devices section:

all-disabled:
- devices: all
  mig-enabled: false

6.3 Test Configuration Affected Every GPU

Using devices: all applies the profile to every GPU. Restrict limited testing to a specific GPU:

devices: [0]

Define the remaining GPUs separately when they must remain in full-GPU mode.

6.4 Air-Gapped Deployment Failure

Public chart or image references caused external access attempts. Use the local chart, approved values, and internally mirrored images:

helm upgrade gpu-operator ./gpu-operator-offline \
  --namespace gpu-operator \
  --values ./gpu-operator-offline/values.yaml

6.5 Custom ConfigMap Was Not Used

Configure MIG Manager to use the custom ConfigMap:

migManager:
  enabled: true
  config:
    name: my-custom-mig-configs

6.6 Configuration Required a Reboot

Cordon and drain the node, stop GPU workloads and host processes, apply the change during a maintenance window, reboot if required, and revalidate after restart.

6.7 Application Pod Remained Pending

Possible causes include unavailable or fully allocated profiles, missing advertised resources, quota restrictions, node selectors, affinity rules, or taints.

kubectl describe pod <pod-name> -n <namespace>

kubectl describe node <node> \
  | grep -i -E 'nvidia.com/gpu|nvidia.com/mig'

6.8 Key Lessons Learned

  • Select profiles using utilization, memory, latency, throughput, and concurrency data.
  • Avoid applying one profile across the entire fleet.
  • Keep full GPUs available for complete-card workloads.
  • Test new configurations on a limited GPU set.
  • Treat geometry changes as planned maintenance.
  • Use MIG Manager as the configuration authority.
  • Verify labels and resource availability after reboot.
  • Validate the complete path from hardware to application Pod.


7. Validation, Business Value and Key Takeaways

7.1 Technical Validation

# Physical GPU health
nvidia-smi

# GPU Operator components
kubectl get pods -n gpu-operator -o wide

# Applied MIG profile
kubectl get node <node> --show-labels | grep nvidia.com/mig

# Host device layout
nvidia-smi -L

# Kubernetes GPU resources
kubectl describe node <node> \
  | grep -i -E 'nvidia.com/gpu|nvidia.com/mig'

# Application validation
kubectl apply -f mig-validation.yaml
kubectl get pod mig-validation -n team-alpha -o wide
kubectl exec -n team-alpha mig-validation -- nvidia-smi

# Namespace quota
kubectl describe resourcequota -n team-alpha

7.2 Success Criteria

The implementation is successful when:

  • All expected physical GPUs are healthy.
  • The required node profile is applied.
  • MIG Manager reports a successful state.
  • The host device layout matches the selected profile.
  • Kubernetes advertises the expected resources.
  • A test Pod can request and use the assigned resource.
  • Namespace quota tracks and limits consumption.
  • Full GPUs remain available for workloads that require them.

7.3 Measuring Business Value

Use actual allocation and utilization data rather than the theoretical maximum number of MIG instances.

Useful measures include:

  • GPU and memory utilization before and after MIG
  • Number of concurrent workloads
  • Average GPU scheduling wait time
  • Full-GPU versus MIG allocation
  • Idle or underutilized allocations
  • Additional hardware demand that can be deferred

Business value statement: The implementation does not create additional physical GPU compute. It improves usable capacity by matching GPU resource size to workload demand.

7.4 Key Takeaways

  • Maintain a balanced combination of full GPUs and MIG profiles.
  • Select profiles using measured workload behavior.
  • Keep multiple resource sizes available according to demand.
  • Use namespace quotas to control and track consumption.
  • Treat geometry changes as maintenance activities.
  • Validate every change from the physical GPU to the application.
  • Measure value through utilization, concurrency, scheduling time, and cost avoidance.

7.5 Final Outcome

The implementation established a flexible GPU allocation model across the H100 and H200 platform. Selected GPUs can be partitioned into workload-specific MIG profiles, while other GPUs remain available as complete physical resources.

The resulting model provides improved GPU utilization, better workload placement, increased workload concurrency, clearer capacity visibility, and potential cost avoidance through deferred expansion.

8. References

  1. NVIDIA Multi-Instance GPU User Guide 

https://docs.nvidia.com/datacenter/tesla/mig-user-guide/latest/

  1. NVIDIA Supported MIG Profiles 

https://docs.nvidia.com/datacenter/tesla/mig-user-guide/supported-mig-profiles.html

  1. NVIDIA GPU Operator with MIG 

https://docs.nvidia.com/datacenter/cloud-native/gpu-operator/latest/gpu-operator-mig.html

  1. NVIDIA Getting Started with MIG 

https://docs.nvidia.com/datacenter/tesla/mig-user-guide/getting-started-with-mig.html

Reference note: Supported profiles and operational requirements may change across GPU models, drivers, and GPU Operator versions. Review the applicable NVIDIA documentation before using these examples in another environment.

Usage note: This document is intended for technical knowledge sharing and is not a universal production runbook. Review and adapt commands, images, node names, profiles, and configurations for the target environment. Test changes on a limited GPU or non-production node before wider implementation.



I work at HPE
HPE Support Center offers support for your HPE services and products when and how you need it. Get started with HPE Support Center today.
[Any personal opinions expressed are mine, and not official statements on behalf of Hewlett Packard Enterprise]
Accept or Kudo
1 REPLY 1
Thaufique_Mod
Community Manager

Re: Driving GPU Efficiency and Cost Optimization with NVIDIA Multi-Instance GPU (MIG)

Thanks for posting this informative topic. It is very helpful.



I work at HPE
HPE Support Center offers support for your HPE services and products when and how you need it. Get started with HPE Support Center today.
[Any personal opinions expressed are mine, and not official statements on behalf of Hewlett Packard Enterprise]
Accept or Kudo