giip
SES 商機登記
AI GPU SERVER MANAGED OPERATIONS

AI-managed operations for NVIDIA DGX and HGX infrastructure.

Unified monitoring that goes beyond Linux to see the GPU, ECC, XID, NVLink/NVSwitch and BMC. GIIP AI analyses the blast radius and the cause of an anomaly, then proposes, records and executes the response you need.

DGX OS · Ubuntu · HGX · DCGM · Prometheus · Redfish/IPMI supported

DCGM, Prometheus and Redfish collect server state, and GIIP AI analyses logs, metrics and change history to propose the cause of an incident and how to respond. Safe, routine work runs as a Runbook; high-risk work such as GPU Reset, reboot or firmware update runs only after customer approval.
The problem

You cannot operate a GPU server with ordinary Linux monitoring alone.

CPU and memory graphs will not tell you a GPU is failing. DGX and HGX bring failure modes that live below the OS.

!

OS metrics miss GPU faults

Monitoring only CPU, memory and disk at the OS level leaves ECC errors, XID events and thermal throttling invisible.

!

No one to read XID / ECC / NVLink

Few teams have an engineer who can judge an XID code, an ECC pattern or an NVLink fault and decide what it means.

!

Driver / CUDA / toolkit drift

Driver, CUDA, Container Toolkit and Fabric Manager can fall out of a supported combination and quietly break workloads.

!

Unclear lines of responsibility

Between the data center, the server vendor, NVIDIA and your own team, it is often unclear who owns what.

How GIIP operates it

Integrated monitoring, AI root-cause analysis, and approval-gated execution.

DCGM, Prometheus and Redfish collect the state; GIIP AI turns it into a cause and a recommended response; safe work runs automatically and risky work waits for your approval.

One view across every layer

Linux, GPU, GPU faults, NVLink/NVSwitch fabric and BMC in a single unified view instead of four disconnected tools.

AI correlates, not just alerts

GIIP combines logs, metrics and change history to infer a likely cause — not a single-threshold alert.

Approval-gated execution

Safe Runbooks run automatically; GPU Reset, reboot and firmware updates run only after you approve them.

Recorded and reported

Every action is written to an issue and a monthly report, so it can be traced afterwards.

COMPONENT ROLES

GIIP does not replace DCGM or Prometheus

GIIP does not replace your monitoring stack. It sits on top of it and adds a unified view, AI root-cause analysis, issue tracking, and an approval / execution / audit trail. Each component keeps its own role:

DCGM ExporterCollects GPU metrics and fault signals (utilisation, temperature, ECC, XID, and so on).
node_exporterCollects Linux-side metrics — CPU, memory, disk, network.
PrometheusStores time-series metrics and evaluates alert rules.
AlertmanagerRoutes, deduplicates and silences the alerts that fire.
RedfishFirst choice for reading and controlling the BMC (a standard API).
IPMI 2.0Fallback for hardware that does not support Redfish.
GrafanaDashboards for operators to dig deep.
GIIPA unified view across all of the above, AI root-cause analysis, issue creation, proposal, approval, execution and audit trail.

GIIP is not a replacement for DCGM or Prometheus. It keeps your existing monitoring assets and adds a judgement-and-operations layer on top.

MONITORING SCOPE

What we monitor

We monitor the GPU, fabric and BMC layers that plain Linux monitoring cannot see, unified into one view.

Linux

  • CPU
  • Memory
  • Filesystem
  • NVMe
  • Network
  • Process
  • Service

GPU (per GPU)

  • Utilization
  • Memory
  • Temperature
  • Power
  • Clock
  • Throttle

GPU faults

  • ECC
  • XID
  • PCIe Replay
  • Retired Pages
  • Row Remapping

Fabric

  • NVLink
  • NVSwitch
  • Fabric Manager

BMC

  • PSU
  • Fan
  • Temperature
  • Voltage
  • System Event Log

Monitoring base

  • Exporter down
  • Metric gaps
  • Prometheus scrape failure

Software compatibility

  • DGX OS
  • Ubuntu
  • Driver
  • CUDA
  • Container Toolkit
AI OPERATION FLOW

AI operation flow

GIIP AI combines multiple signals — logs, metrics and change history — rather than firing on a single threshold. It is not full autonomy, and it does not perform unconditional auto-recovery.

  1. 1

    Collect state from DCGM, Prometheus and Redfish.

  2. 2

    GIIP analyses logs, metrics and change history together.

  3. 3

    Present the blast radius, the likely cause and the recommended response.

  4. 4

    Run safe Runbooks; request approval for high-risk work.

  5. 5

    Record the outcome and the prevention measures in an issue and the monthly report.

AI correlates several signals to infer a cause, but work that is hard to reverse — GPU Reset, reboot — runs only after customer approval.

ARCHITECTURE

Recommended architecture

We respect your existing hardware layout and lay down the monitoring and control paths safely.

DGX/HGX OS

node_exporter, DCGM Exporter, NVSM and Fabric Manager run here.

BMC

Redfish is preferred; hardware without it falls back to IPMI 2.0.

Central

Prometheus, Alertmanager, Grafana and GIIP are aggregated here.

Connection

Management VLAN → VPN or GIIP Connector → Central.

We never expose the BMC directly to the public internet. It is reachable only over the management network via VPN or a dedicated connector.

PRICING

Pricing

The base unit is one DGX/HGX (8-GPU physical node). The ja and en pages also show a KRW reference. Actual contracts are quoted in JPY in Japan and KRW in Korea; there is no automatic FX conversion.

DGX Monitoring

400,000 KRW

24×365 automatic monitoring, notification, monthly report.

DGX Standard

800,000 KRW

Monitoring, OS / GPU software management, first-line analysis, NVIDIA / OEM inquiry.

DGX Premium

1,500,000 KRW

Emergency response, CUDA / container management, performance & fault analysis, recovery support.

AI Platform Managed

2,500,000 KRW~

Kubernetes / Slurm / tenant / queue / model-runtime operations.

"24×365" means automatic monitoring. Without a separate SLA it does not guarantee always-on human response or a recovery time.

Recommended standard package

The setup we recommend to most customers to start with.

Initial integration
1,500,000 KRW
Monthly operations
1,000,000 KRW / node
Contract term
12 months
Tax
Excluded
Config changes & tech support
Up to 5h / month

Prices are in KRW and exclude tax. We quote individually by node count and workload.

Included

  • Inventory
  • Health Check
  • Exporter / Prometheus / Alertmanager / Redfish / IPMI / GIIP integration
  • AI fault analysis
  • Approved Remote Recovery
  • NVIDIA / OEM diagnostic materials
  • Monthly report

Quoted separately

  • Rack / power / cooling / network
  • Hardware purchase
  • NVIDIA License / Support
  • On-site work
  • Parts
  • Backup Storage
  • Kubernetes / Slurm / Run:ai
  • Customer Model / RAG / Inference apps
  • Large Upgrade / Migration
  • Forensics
  • Warranty / SLA
  • NVL72 / Multi-node / Multi-site / Liquid Cooling
SAFETY & RESPONSIBILITY

Safety, approval & responsibility

Work that is hard to reverse always goes through customer approval.

Automatic (GIIP)

  • State collection
  • Anomaly detection
  • Analysis
  • Notification
  • Issue creation
  • Approved Runbook execution
  • Reporting

Customer approval required

  • GPU Reset
  • Power Cycle
  • Reboot
  • Driver / CUDA / Firmware update
  • Job stop
  • Permission / Firewall change
  • Delete or irreversible operations

DC / NVIDIA / OEM scope

  • Power
  • Cooling
  • Cabling
  • On-site inspection
  • Part replacement
  • Warranty
  • Paid support

GIIP does not claim unverified titles such as "NVIDIA official partner" or "NVIDIA certified service".

What this gives you

GPU · NVLink · BMC

Integrated view beyond Linux

24×365

Automatic monitoring (SLA separate)

Approval-gated

High-risk work after your approval

Audit trail

Every action recorded and traceable

Frequently asked questions

How is this different from general Ubuntu server management?

General server management watches the OS — CPU, memory, disk. This adds the GPU layer: DCGM metrics, ECC/XID faults, NVLink/NVSwitch fabric and the BMC, plus AI analysis of what a GPU signal actually means.

Do you support HGX and OEM GPU servers, not only DGX?

Yes. DGX is the reference, but HGX baseboards and OEM GPU servers running Ubuntu are in scope as long as DCGM, node_exporter and Redfish or IPMI are available.

Can we reuse our existing Prometheus and Grafana?

Yes. GIIP does not replace them — it sits on top and adds a unified view, AI root-cause analysis, approval and an audit trail. Your existing Prometheus and Grafana keep their roles.

How do you choose between Redfish and IPMI?

Redfish is the first choice because it is the standard API. IPMI 2.0 is a fallback for hardware that does not support Redfish. The BMC is never exposed directly to the internet.

Do you auto-reboot on a GPU fault?

No. Detection, analysis and notification are automatic, but a reboot, GPU Reset or power cycle runs only after customer approval. There is no unconditional auto-recovery.

What if we have no NVIDIA Support contract?

We still monitor, analyse and support recovery, and we prepare the diagnostic materials NVIDIA or the OEM would ask for. Warranty, paid support and NVIDIA licensing remain the customer's contracts.

What is the scope for Kubernetes and Slurm?

Base plans cover the node, GPU and BMC. Kubernetes, Slurm, tenant/queue and model-runtime operations are the AI Platform Managed plan, quoted separately.

Who is responsible for data-center power and cooling?

Power, cooling, cabling, on-site inspection and part replacement belong to the data center, NVIDIA or the OEM. GIIP handles monitoring, analysis and approved remote recovery.

Is there 24-hour human on-call?

24×365 refers to automatic monitoring. Without a separate SLA it does not guarantee an always-on human responder or a fixed recovery time; those are agreed as an SLA add-on.

How is pricing calculated for multiple nodes?

The base unit is one DGX/HGX 8-GPU node. Multiple nodes are quoted per node; the recommended standard package is 1,500,000 KRW initial integration plus 1,000,000 KRW per node per month on a 12-month term, tax excluded.

Explore related capabilities

Let AI monitor and operate your GPU servers.

Tell us what DGX or HGX you run and how it is wired. We will show you the monitoring scope, the responsibility split and a per-node quote.

contact@littleworld.net