What You'll Do
Define the architecture for NVIDIA HGX platforms and establish compatibility requirements. Lead the qualification and burn-in process for new GPU clusters and analyze fleet-wide distributions to identify failures. Diagnose health and performance issues using NVIDIA diagnostics and telemetry, and validate workloads to set baselines for performance metrics.
What We're Looking For
Strong hands-on experience with NVIDIA GPU infrastructure, including HGX or comparable platforms. Demonstrated experience in designing and executing burn-in and qualification across multi-node GPU clusters. Proficiency in Linux systems administration and troubleshooting, along with experience in diagnosing GPU and firmware problems.
Additional Information
Experience Level
Lead / Principal
Employment Type
permanent
Work Mode
On-site