- Community Home
- >
- Software
- >
- Software - General
- >
- Re: Driving GPU Efficiency and Cost Optimization w...
Categories
Company
Local Language
Forums
Discussions
- Integrity Servers
- Server Clustering
- HPE NonStop Compute
- HPE Apollo Systems
- High Performance Computing
Knowledge Base
Forums
- Data Protection and Retention
- Entry Storage Systems
- Legacy
- Midrange and Enterprise Storage
- Storage Networking
- HPE Nimble Storage
Discussions
Knowledge Base
Forums
Discussions
- Cloud Mentoring and Education
- Software - General
- HPE OneView
- HPE Ezmeral Software platform
- HPE OpsRamp Software
Knowledge Base
Discussions
Forums
Discussions
Discussion Boards
Discussion Boards
Discussion Boards
Discussion Boards
Discussion Boards
Discussion Boards
Discussion Boards
Discussion Boards
Discussion Boards
Discussion Boards
Discussion Boards
Discussion Boards
Discussion Boards
Discussion Boards
Discussion Boards
Discussion Boards
Discussion Boards
Discussion Boards
Discussion Boards
Discussion Boards
Discussion Boards
Discussion Boards
Discussion Boards
Community
Resources
Forums
Blogs
- Subscribe to RSS Feed
- Mark Topic as New
- Mark Topic as Read
- Float this Topic for Current User
- Bookmark
- Subscribe
- Printer Friendly Page
- Mark as New
- Bookmark
- Subscribe
- Mute
- Subscribe to RSS Feed
- Permalink
- Report Inappropriate Content
Thursday - last edited Thursday by support_s
Thursday - last edited Thursday by support_s
Driving GPU Efficiency and Cost Optimization with NVIDIA Multi-Instance GPU (MIG)
NVIDIA MIG Implementation on DGX Kubernetes Cluster
Technical Knowledge-Sharing Document
Scope: 10 DGX nodes | 80 Nvidia Cards | H100/H200 GPUs | GPU Operator | Air-gapped Kubernetes
1. Introduction, Environment and Business Context
1.1 Purpose
This document shares the practical approach used to enable NVIDIA Multi-Instance GPU (MIG) on a DGX Kubernetes platform.
The platform uses a flexible allocation model:
- Some GPUs remain available as complete physical cards.
- Selected GPUs are divided into MIG instances.
- Different MIG profiles are used based on workload memory, compute, and performance requirements.
The document is intended for platform engineers, Kubernetes administrators, GPU infrastructure teams, and application teams evaluating a similar GPU-sharing model.
1.2 Environment
The implementation covers 10 bare-metal NVIDIA DGX worker nodes running a Kubeadm-based Kubernetes cluster.
Component
Environment
DGX worker nodes
10
NVIDIA H100
4 nodes x 8 GPUs = 32 GPUs
NVIDIA H200
6 nodes x 8 GPUs = 48 GPUs
Total physical GPUs
80
NVIDIA software
GPU Operator v25.10.1
Deployment model
Offline and air-gapped
The 80 physical GPUs are not configured with one common MIG profile. The platform contains a combination of:
- Full physical GPUs for large or performance-sensitive workloads
- Smaller MIG profiles for lightweight workloads
- Medium and larger MIG profiles for workloads requiring additional memory or compute
1.3 Business Context
H100 and H200 GPUs are high-value infrastructure resources. Assigning a complete GPU to a workload that uses only a small portion of its memory or compute capacity leaves the remaining capacity unavailable to other applications.
This can lead to:
- Underutilized GPU capacity
- Unnecessary full-GPU reservations
- Longer workload scheduling queues
- Reduced workload concurrency
- Limited visibility into actual GPU consumption
- Premature demand for additional hardware
The objective is to match GPU allocation with workload demand. Smaller workloads can use appropriately sized MIG instances, while workloads requiring complete-card performance can continue using full physical GPUs.
This approach can improve utilization, increase workload concurrency, and provide better capacity visibility. The expected business value is primarily capacity optimization and cost avoidance, including the possibility of deferring additional GPU procurement.
Key point: MIG does not create additional physical compute. Its value comes from making existing GPU capacity easier to allocate, share, isolate, and manage efficiently.
2. Challenge and Solution Approach
2.1 Challenge
The original allocation model assigned a complete physical GPU to each workload. This was suitable for training and high-demand applications, but smaller inference, testing, and development workloads often used only part of the available capacity.
Because the remaining capacity could not be scheduled for another application, the model resulted in:
- Underutilized GPU resources
- Unnecessary full-GPU reservations
- Longer waiting times for GPU workloads
- Limited workload concurrency
- Difficulty distinguishing actual shortages from inefficient allocation
Workload requirements also varied significantly. Some applications required only a small portion of a GPU, while others required more memory, higher compute capacity, or direct access to a complete physical card.
2.2 Solution Approach
NVIDIA MIG was enabled on selected H100 and H200 GPUs to provide smaller, hardware-isolated GPU resources that Kubernetes can schedule independently.
A mixed allocation model was selected:
- Full physical GPUs remain available for training, large models, and performance-sensitive workloads.
- Smaller MIG profiles support lightweight inference, development, and testing.
- Medium and larger MIG profiles support applications requiring additional memory or compute.
- Profile assignments are selected according to workload demand and observed utilization.
This model provides flexibility without unnecessarily fragmenting the complete GPU fleet.
Key point: The objective is not to divide every GPU into the maximum number of MIG instances. The objective is to maintain the correct balance of full GPUs and MIG profiles based on workload demand.
2.3 Expected Value
- Better use of existing GPU capacity
- Increased workload concurrency
- Reduced unnecessary full-GPU allocation
- Shorter scheduling queues for smaller workloads
- Improved resource-consumption visibility
- Better capacity planning
- Potential cost avoidance through deferred hardware expansion
3. MIG Architecture and Design
3.1 How MIG Works
NVIDIA MIG divides a supported physical GPU into multiple hardware-isolated instances. Each instance receives a defined portion of GPU memory and compute resources and is exposed to Kubernetes as an independently schedulable resource.
The platform supports:
- Full physical GPUs for compute-intensive or performance-sensitive workloads
- MIG instances for workloads that operate efficiently with smaller GPU allocations
3.2 Component Flow
DGX H100 or H200 Hardware
|
v
NVIDIA Driver
|
v
GPU Operator
|
v
MIG Manager <--- MIG ConfigMap
|
v
Full-GPU or MIG Configuration
|
v
NVIDIA Device Plugin
|
v
Kubernetes Allocatable Resource
|
v
Kubernetes Scheduler
|
v
Application Pod
3.3 Component Responsibilities
- NVIDIA Driver: Manages the physical GPU and provides MIG capability.
- GPU Operator: Manages NVIDIA software components in the Kubernetes cluster.
- MIG ConfigMap: Defines approved MIG profiles and partitioning layouts.
- MIG Manager: Applies the configuration requested through the node label.
- NVIDIA Device Plugin: Advertises full-GPU and MIG resources to Kubernetes.
- Kubernetes Scheduler: Places Pods where the requested resource is available.
- Application Pod: Requests a full GPU or a specific MIG resource.
3.4 MIG Profiles and Geometry
A MIG profile defines the compute and memory available to an instance. A MIG geometry defines how multiple instances are arranged on a physical GPU.
Different geometries can support smaller, medium, and larger MIG profiles, while complete physical GPUs remain available for workloads requiring full-card performance. Profile availability depends on the GPU model and active configuration.
3.5 Mixed MIG Strategy
mig:
strategy: mixed
This allows Kubernetes to advertise profile-specific resources such as:
nvidia.com/mig-1g.10gb
nvidia.com/mig-1g.18gb
It also supports different configurations across H100 and H200 nodes. A geometry change may interrupt workloads and require a GPU reset or node reboot.
3.6 Profile Selection Criteria
Profile selection should use actual workload behavior, including:
- GPU memory consumption
- Compute utilization
- Application latency and throughput
- Expected concurrency
- Model size and workload duration
- Requirement for full-GPU performance
Key point: The design goal is to match each workload with an appropriate resource while retaining sufficient full GPUs for complete-card workloads.
4. Implementation Approach
The implementation was completed in a controlled sequence. Each node was validated before being returned to service.
Step 1: Prepare the Node
# kubectl cordon <node>
# kubectl drain <node> --ignore-daemonsets --delete-emptydir-data
Before changing the GPU configuration, stop active GPU workloads and host-level GPU processes. Use a maintenance window when a reset or reboot may be required.
Step 2: Enable MIG Manager
The GPU Operator was deployed from a local Helm chart using internally mirrored container images.
# helm upgrade gpu-operator ./gpu-operator-offline --namespace gpu-operator --values ./gpu-operator-offline/values.yaml --set mig.strategy=mixed --set migManager.enabled=true --set migManager.config.name=my-custom-mig-configs
Main settings:
- mig.strategy=mixed enables profile-specific resources.
- migManager.enabled=true enables automated MIG configuration.
- migManager.config.name identifies the custom ConfigMap.
Step 3: Create the MIG Configuration (Refrence)
apiVersion: v1
kind: ConfigMap
metadata:
name: my-custom-mig-configs
namespace: gpu-operator
data:
config.yaml: |
version: v1
mig-configs:
all-disabled:
- devices: all
mig-enabled: false
h100-profile:
- devices: all
mig-enabled: true
mig-devices:
"1g.10gb": 7
h200-profile:
- devices: all
mig-enabled: true
mig-devices:
"1g.18gb": 7
only-gpu0-test:
- devices: [0]
mig-enabled: true
mig-devices:
"1g.10gb": 7
- devices: [1, 2, 3, 4, 5, 6, 7]
mig-enabled: false
Apply the configuration:
# kubectl apply -f mig-config.yaml
Additional profiles can be defined according to supported GPU geometries and measured workload requirements.
Step 4: Apply the Required Profile
For an H100 node:
kubectl label node <h100-node> \
nvidia.com/mig.config=h100-profile \
--overwrite
For an H200 node:
kubectl label node <h200-node> \
nvidia.com/mig.config=h200-profile \
--overwrite
For limited testing:
kubectl label node <test-node> \
nvidia.com/mig.config=only-gpu0-test \
--overwrite
Testing on a selected GPU reduces the risk of applying an incorrect configuration across the entire node.
Step 5: Monitor the Transition
kubectl get node <node> --show-labels \
| grep nvidia.com/mig
kubectl logs -n gpu-operator \
-l app=nvidia-mig-manager \
--tail=200
A successful configuration should report:
nvidia.com/mig.config.state=success
If the configuration remains pending or changes to failed, review the MIG Manager logs before reapplying the profile.
Step 6: Validate the Resources
nvidia-smi
nvidia-smi -L
kubectl describe node <node> \
| grep -i -E 'nvidia.com/gpu|nvidia.com/mig'
The resource names and counts should match the assigned profile.
Step 7: Run a Test Workload
apiVersion: v1
kind: Pod
metadata:
name: mig-validation
namespace: team-alpha
spec:
restartPolicy: Never
containers:
- name: cuda-test
image: nvcr.io/nvidia/cuda:12.3.0-base-ubuntu22.04
command:
- bash
- -lc
- nvidia-smi
resources:
limits:
nvidia.com/mig-1g.10gb: 1
kubectl apply -f mig-validation.yaml
kubectl get pod mig-validation -n team-alpha -o wide
kubectl describe pod mig-validation -n team-alpha
kubectl exec -n team-alpha mig-validation -- nvidia-smi
Step 8: Return the Node to Service
kubectl uncordon <node>
Key point: A Running MIG Manager Pod does not confirm successful implementation. Validate the complete path from physical GPU configuration to application-level device visibility.
5. Workload Consumption, Governance and Cost Control 5.1 Resource Requests
H100 MIG request:
resources:
limits:
nvidia.com/mig-1g.10gb: 1
H200 MIG request:
resources:
limits:
nvidia.com/mig-1g.18gb: 1
Full physical GPU request:
resources:
limits:
nvidia.com/gpu: 1
If the requested resource is unavailable, the Pod remains Pending.
kubectl describe pod <pod-name> -n <namespace>
kubectl get events -n <namespace> --sort-by=.lastTimestamp
5.2 Namespace Resource Quota
apiVersion: v1
kind: ResourceQuota
metadata:
name: team-mig-quota
namespace: team-alpha
spec:
hard:
requests.nvidia.com/mig-1g.10gb: "7"
This prevents one namespace from consuming all available instances and establishes a clear capacity boundary.
5.3 Recommended Allocation Model
Workload type
Recommended allocation
Development and lightweight inference
Smaller MIG profile
Medium inference or memory-sensitive workloads
Medium or larger MIG profile
Training and performance-sensitive workloads
Full physical GPU
Profile selection should be based on measured utilization rather than assumptions.
5.4 Monitoring and Capacity Visibility
Monitoring should cover:
- Physical GPU and MIG instance utilization
- GPU memory usage
- Allocation by namespace, Pod, and application
- Pending GPU workloads
- Full-GPU versus MIG consumption
- Idle or consistently underused allocations
5.5 Business and Cost Value
Workload-specific allocation can reduce unnecessary full-GPU reservations, increase concurrency, preserve complete GPUs for demanding workloads, reduce scheduling delays, and improve capacity planning.
The financial benefit should be described as capacity optimization and cost avoidance. MIG does not reduce the cost of hardware already purchased, but it can reduce or delay the need for additional infrastructure.
Key point: Cost efficiency means selecting the correct resource size while maintaining a balanced pool of full GPUs and MIG instances.
6. Implementation Challenges and Lessons Learned 6.1 MIG Configuration Missing After Reboot
Observation: Expected MIG resources were unavailable after restart.
Cause: The required nvidia.com/mig.config label was missing.
kubectl label node <node> \
nvidia.com/mig.config=<profile-name> \
--overwrite
Confirm nvidia.com/mig.config.state=success.
6.2 Invalid Disabled Profile
A disabled profile must not contain a mig-devices section:
all-disabled:
- devices: all
mig-enabled: false
6.3 Test Configuration Affected Every GPU
Using devices: all applies the profile to every GPU. Restrict limited testing to a specific GPU:
devices: [0]
Define the remaining GPUs separately when they must remain in full-GPU mode.
6.4 Air-Gapped Deployment Failure
Public chart or image references caused external access attempts. Use the local chart, approved values, and internally mirrored images:
helm upgrade gpu-operator ./gpu-operator-offline \
--namespace gpu-operator \
--values ./gpu-operator-offline/values.yaml
6.5 Custom ConfigMap Was Not Used
Configure MIG Manager to use the custom ConfigMap:
migManager:
enabled: true
config:
name: my-custom-mig-configs
6.6 Configuration Required a Reboot
Cordon and drain the node, stop GPU workloads and host processes, apply the change during a maintenance window, reboot if required, and revalidate after restart.
6.7 Application Pod Remained Pending
Possible causes include unavailable or fully allocated profiles, missing advertised resources, quota restrictions, node selectors, affinity rules, or taints.
kubectl describe pod <pod-name> -n <namespace>
kubectl describe node <node> \
| grep -i -E 'nvidia.com/gpu|nvidia.com/mig'
6.8 Key Lessons Learned
- Select profiles using utilization, memory, latency, throughput, and concurrency data.
- Avoid applying one profile across the entire fleet.
- Keep full GPUs available for complete-card workloads.
- Test new configurations on a limited GPU set.
- Treat geometry changes as planned maintenance.
- Use MIG Manager as the configuration authority.
- Verify labels and resource availability after reboot.
- Validate the complete path from hardware to application Pod.
7. Validation, Business Value and Key Takeaways
7.1 Technical Validation
# Physical GPU health
nvidia-smi
# GPU Operator components
kubectl get pods -n gpu-operator -o wide
# Applied MIG profile
kubectl get node <node> --show-labels | grep nvidia.com/mig
# Host device layout
nvidia-smi -L
# Kubernetes GPU resources
kubectl describe node <node> \
| grep -i -E 'nvidia.com/gpu|nvidia.com/mig'
# Application validation
kubectl apply -f mig-validation.yaml
kubectl get pod mig-validation -n team-alpha -o wide
kubectl exec -n team-alpha mig-validation -- nvidia-smi
# Namespace quota
kubectl describe resourcequota -n team-alpha
7.2 Success Criteria
The implementation is successful when:
- All expected physical GPUs are healthy.
- The required node profile is applied.
- MIG Manager reports a successful state.
- The host device layout matches the selected profile.
- Kubernetes advertises the expected resources.
- A test Pod can request and use the assigned resource.
- Namespace quota tracks and limits consumption.
- Full GPUs remain available for workloads that require them.
7.3 Measuring Business Value
Use actual allocation and utilization data rather than the theoretical maximum number of MIG instances.
Useful measures include:
- GPU and memory utilization before and after MIG
- Number of concurrent workloads
- Average GPU scheduling wait time
- Full-GPU versus MIG allocation
- Idle or underutilized allocations
- Additional hardware demand that can be deferred
Business value statement: The implementation does not create additional physical GPU compute. It improves usable capacity by matching GPU resource size to workload demand.
7.4 Key Takeaways
- Maintain a balanced combination of full GPUs and MIG profiles.
- Select profiles using measured workload behavior.
- Keep multiple resource sizes available according to demand.
- Use namespace quotas to control and track consumption.
- Treat geometry changes as maintenance activities.
- Validate every change from the physical GPU to the application.
- Measure value through utilization, concurrency, scheduling time, and cost avoidance.
7.5 Final Outcome
The implementation established a flexible GPU allocation model across the H100 and H200 platform. Selected GPUs can be partitioned into workload-specific MIG profiles, while other GPUs remain available as complete physical resources.
The resulting model provides improved GPU utilization, better workload placement, increased workload concurrency, clearer capacity visibility, and potential cost avoidance through deferred expansion.
8. References
- NVIDIA Multi-Instance GPU User Guide
https://docs.nvidia.com/datacenter/tesla/mig-user-guide/latest/
- NVIDIA Supported MIG Profiles
https://docs.nvidia.com/datacenter/tesla/mig-user-guide/supported-mig-profiles.html
- NVIDIA GPU Operator with MIG
https://docs.nvidia.com/datacenter/cloud-native/gpu-operator/latest/gpu-operator-mig.html
- NVIDIA Getting Started with MIG
https://docs.nvidia.com/datacenter/tesla/mig-user-guide/getting-started-with-mig.html
Reference note: Supported profiles and operational requirements may change across GPU models, drivers, and GPU Operator versions. Review the applicable NVIDIA documentation before using these examples in another environment.
Usage note: This document is intended for technical knowledge sharing and is not a universal production runbook. Review and adapt commands, images, node names, profiles, and configurations for the target environment. Test changes on a limited GPU or non-production node before wider implementation.
I work at HPE
HPE Support Center offers support for your HPE services and products when and how you need it. Get started with HPE Support Center today.
[Any personal opinions expressed are mine, and not official statements on behalf of Hewlett Packard Enterprise]
- Tags:
- Operating System
- Mark as New
- Bookmark
- Subscribe
- Mute
- Subscribe to RSS Feed
- Permalink
- Report Inappropriate Content
Thursday
Thursday
Re: Driving GPU Efficiency and Cost Optimization with NVIDIA Multi-Instance GPU (MIG)
Thanks for posting this informative topic. It is very helpful.
I work at HPE
HPE Support Center offers support for your HPE services and products when and how you need it. Get started with HPE Support Center today.
[Any personal opinions expressed are mine, and not official statements on behalf of Hewlett Packard Enterprise]