Breakmark

AI Compute Engineer

Breakmark

What You'll Do

Define the architecture for NVIDIA HGX platforms and establish compatibility requirements. Lead the qualification and burn-in process for new GPU clusters and analyze fleet-wide distributions to identify failures. Diagnose health and performance issues using NVIDIA diagnostics and telemetry, and validate workloads to set baselines for performance metrics.

What We're Looking For

Strong hands-on experience with NVIDIA GPU infrastructure, including HGX or comparable platforms. Demonstrated experience in designing and executing burn-in and qualification across multi-node GPU clusters. Proficiency in Linux systems administration and troubleshooting, along with experience in diagnosing GPU and firmware problems.

Additional Information

Experience Level

Lead / Principal

Employment Type

permanent

Work Mode

On-site