AI-managed operations for NVIDIA DGX and HGX infrastructure.
Unified monitoring that goes beyond Linux to see the GPU, ECC, XID, NVLink/NVSwitch and BMC. GIIP AI analyses the blast radius and the cause of an anomaly, then proposes, records and executes the response you need.
DGX OS · Ubuntu · HGX · DCGM · Prometheus · Redfish/IPMI supported
DCGM, Prometheus and Redfish collect server state, and GIIP AI analyses logs, metrics and change history to propose the cause of an incident and how to respond. Safe, routine work runs as a Runbook; high-risk work such as GPU Reset, reboot or firmware update runs only after customer approval.
You cannot operate a GPU server with ordinary Linux monitoring alone.
CPU and memory graphs will not tell you a GPU is failing. DGX and HGX bring failure modes that live below the OS.
OS metrics miss GPU faults
Monitoring only CPU, memory and disk at the OS level leaves ECC errors, XID events and thermal throttling invisible.
No one to read XID / ECC / NVLink
Few teams have an engineer who can judge an XID code, an ECC pattern or an NVLink fault and decide what it means.
Driver / CUDA / toolkit drift
Driver, CUDA, Container Toolkit and Fabric Manager can fall out of a supported combination and quietly break workloads.
Unclear lines of responsibility
Between the data center, the server vendor, NVIDIA and your own team, it is often unclear who owns what.
Integrated monitoring, AI root-cause analysis, and approval-gated execution.
DCGM, Prometheus and Redfish collect the state; GIIP AI turns it into a cause and a recommended response; safe work runs automatically and risky work waits for your approval.
One view across every layer
Linux, GPU, GPU faults, NVLink/NVSwitch fabric and BMC in a single unified view instead of four disconnected tools.
AI correlates, not just alerts
GIIP combines logs, metrics and change history to infer a likely cause — not a single-threshold alert.
Approval-gated execution
Safe Runbooks run automatically; GPU Reset, reboot and firmware updates run only after you approve them.
Recorded and reported
Every action is written to an issue and a monthly report, so it can be traced afterwards.
GIIP does not replace DCGM or Prometheus
GIIP does not replace your monitoring stack. It sits on top of it and adds a unified view, AI root-cause analysis, issue tracking, and an approval / execution / audit trail. Each component keeps its own role:
| DCGM Exporter | Collects GPU metrics and fault signals (utilisation, temperature, ECC, XID, and so on). |
| node_exporter | Collects Linux-side metrics — CPU, memory, disk, network. |
| Prometheus | Stores time-series metrics and evaluates alert rules. |
| Alertmanager | Routes, deduplicates and silences the alerts that fire. |
| Redfish | First choice for reading and controlling the BMC (a standard API). |
| IPMI 2.0 | Fallback for hardware that does not support Redfish. |
| Grafana | Dashboards for operators to dig deep. |
| GIIP | A unified view across all of the above, AI root-cause analysis, issue creation, proposal, approval, execution and audit trail. |
GIIP is not a replacement for DCGM or Prometheus. It keeps your existing monitoring assets and adds a judgement-and-operations layer on top.
What we monitor
We monitor the GPU, fabric and BMC layers that plain Linux monitoring cannot see, unified into one view.
Linux
- CPU
- Memory
- Filesystem
- NVMe
- Network
- Process
- Service
GPU (per GPU)
- Utilization
- Memory
- Temperature
- Power
- Clock
- Throttle
GPU faults
- ECC
- XID
- PCIe Replay
- Retired Pages
- Row Remapping
Fabric
- NVLink
- NVSwitch
- Fabric Manager
BMC
- PSU
- Fan
- Temperature
- Voltage
- System Event Log
Monitoring base
- Exporter down
- Metric gaps
- Prometheus scrape failure
Software compatibility
- DGX OS
- Ubuntu
- Driver
- CUDA
- Container Toolkit
AI operation flow
GIIP AI combines multiple signals — logs, metrics and change history — rather than firing on a single threshold. It is not full autonomy, and it does not perform unconditional auto-recovery.
- 1
Collect state from DCGM, Prometheus and Redfish.
- 2
GIIP analyses logs, metrics and change history together.
- 3
Present the blast radius, the likely cause and the recommended response.
- 4
Run safe Runbooks; request approval for high-risk work.
- 5
Record the outcome and the prevention measures in an issue and the monthly report.
AI correlates several signals to infer a cause, but work that is hard to reverse — GPU Reset, reboot — runs only after customer approval.
Recommended architecture
We respect your existing hardware layout and lay down the monitoring and control paths safely.
DGX/HGX OS
node_exporter, DCGM Exporter, NVSM and Fabric Manager run here.
BMC
Redfish is preferred; hardware without it falls back to IPMI 2.0.
Central
Prometheus, Alertmanager, Grafana and GIIP are aggregated here.
Connection
Management VLAN → VPN or GIIP Connector → Central.
We never expose the BMC directly to the public internet. It is reachable only over the management network via VPN or a dedicated connector.
Pricing
The base unit is one DGX/HGX (8-GPU physical node). The ja and en pages also show a KRW reference. Actual contracts are quoted in JPY in Japan and KRW in Korea; there is no automatic FX conversion.
| Plan | Monthly (tax excl.) | Scope |
|---|---|---|
| DGX Monitoring | 400,000 KRW | 24×365 automatic monitoring, notification, monthly report. |
| DGX Standard | 800,000 KRW | Monitoring, OS / GPU software management, first-line analysis, NVIDIA / OEM inquiry. |
| DGX Premium | 1,500,000 KRW | Emergency response, CUDA / container management, performance & fault analysis, recovery support. |
| AI Platform Managed | 2,500,000 KRW~ | Kubernetes / Slurm / tenant / queue / model-runtime operations. |
DGX Monitoring
400,000 KRW24×365 automatic monitoring, notification, monthly report.
DGX Standard
800,000 KRWMonitoring, OS / GPU software management, first-line analysis, NVIDIA / OEM inquiry.
DGX Premium
1,500,000 KRWEmergency response, CUDA / container management, performance & fault analysis, recovery support.
AI Platform Managed
2,500,000 KRW~Kubernetes / Slurm / tenant / queue / model-runtime operations.
"24×365" means automatic monitoring. Without a separate SLA it does not guarantee always-on human response or a recovery time.
Recommended standard package
The setup we recommend to most customers to start with.
- Initial integration
- 1,500,000 KRW
- Monthly operations
- 1,000,000 KRW / node
- Contract term
- 12 months
- Tax
- Excluded
- Config changes & tech support
- Up to 5h / month
Prices are in KRW and exclude tax. We quote individually by node count and workload.
Included
- Inventory
- Health Check
- Exporter / Prometheus / Alertmanager / Redfish / IPMI / GIIP integration
- AI fault analysis
- Approved Remote Recovery
- NVIDIA / OEM diagnostic materials
- Monthly report
Quoted separately
- Rack / power / cooling / network
- Hardware purchase
- NVIDIA License / Support
- On-site work
- Parts
- Backup Storage
- Kubernetes / Slurm / Run:ai
- Customer Model / RAG / Inference apps
- Large Upgrade / Migration
- Forensics
- Warranty / SLA
- NVL72 / Multi-node / Multi-site / Liquid Cooling
Safety, approval & responsibility
Work that is hard to reverse always goes through customer approval.
Automatic (GIIP)
- State collection
- Anomaly detection
- Analysis
- Notification
- Issue creation
- Approved Runbook execution
- Reporting
Customer approval required
- GPU Reset
- Power Cycle
- Reboot
- Driver / CUDA / Firmware update
- Job stop
- Permission / Firewall change
- Delete or irreversible operations
DC / NVIDIA / OEM scope
- Power
- Cooling
- Cabling
- On-site inspection
- Part replacement
- Warranty
- Paid support
GIIP does not claim unverified titles such as "NVIDIA official partner" or "NVIDIA certified service".
What this gives you
GPU · NVLink · BMC
Integrated view beyond Linux
24×365
Automatic monitoring (SLA separate)
Approval-gated
High-risk work after your approval
Audit trail
Every action recorded and traceable
Frequently asked questions
How is this different from general Ubuntu server management?
General server management watches the OS — CPU, memory, disk. This adds the GPU layer: DCGM metrics, ECC/XID faults, NVLink/NVSwitch fabric and the BMC, plus AI analysis of what a GPU signal actually means.
Do you support HGX and OEM GPU servers, not only DGX?
Yes. DGX is the reference, but HGX baseboards and OEM GPU servers running Ubuntu are in scope as long as DCGM, node_exporter and Redfish or IPMI are available.
Can we reuse our existing Prometheus and Grafana?
Yes. GIIP does not replace them — it sits on top and adds a unified view, AI root-cause analysis, approval and an audit trail. Your existing Prometheus and Grafana keep their roles.
How do you choose between Redfish and IPMI?
Redfish is the first choice because it is the standard API. IPMI 2.0 is a fallback for hardware that does not support Redfish. The BMC is never exposed directly to the internet.
Do you auto-reboot on a GPU fault?
No. Detection, analysis and notification are automatic, but a reboot, GPU Reset or power cycle runs only after customer approval. There is no unconditional auto-recovery.
What if we have no NVIDIA Support contract?
We still monitor, analyse and support recovery, and we prepare the diagnostic materials NVIDIA or the OEM would ask for. Warranty, paid support and NVIDIA licensing remain the customer's contracts.
What is the scope for Kubernetes and Slurm?
Base plans cover the node, GPU and BMC. Kubernetes, Slurm, tenant/queue and model-runtime operations are the AI Platform Managed plan, quoted separately.
Who is responsible for data-center power and cooling?
Power, cooling, cabling, on-site inspection and part replacement belong to the data center, NVIDIA or the OEM. GIIP handles monitoring, analysis and approved remote recovery.
Is there 24-hour human on-call?
24×365 refers to automatic monitoring. Without a separate SLA it does not guarantee an always-on human responder or a fixed recovery time; those are agreed as an SLA add-on.
How is pricing calculated for multiple nodes?
The base unit is one DGX/HGX 8-GPU node. Multiple nodes are quoted per node; the recommended standard package is 1,500,000 KRW initial integration plus 1,000,000 KRW per node per month on a 12-month term, tax excluded.
Let AI monitor and operate your GPU servers.
Tell us what DGX or HGX you run and how it is wired. We will show you the monitoring scope, the responsibility split and a per-node quote.