Technical Expertise: Experience building, operating, and optimizing distributed infrastructure in production environments. Deep hands-on expertise in one or more of the following:
- DCGM
- BMC / Redfish
- Firmware Lifecycle Management
- Driver Lifecycle Management
Networking
- InfiniBand
- High-Speed Ethernet
- NCCL
- UFM
High-Performance Storage
- Lustre
- IBM Storage Scale (GPFS)
- WEKA
- VAST Data
- Similar Enterprise Storage Platforms
Platform Experience
Hands-on experience with:
- Kubernetes
- Slurm
- GPU Scheduling
- Multi-Tenancy Architecture
Observability
- Prometheus
- Grafana
- OpenTelemetry
Automation & Infrastructure as Code
- Terraform
- Ansible
- Argo CD
- Similar Automation Frameworks
Operating Systems & Scripting
- Linux Administration
- Python
- Bash
- Comparable scripting languages
Soft Skills
- Strong root-cause analysis and troubleshooting skills.
- Ability to communicate complex technical findings clearly.
- Experience leading technical initiatives without direct authority.
- Strong customer-facing communication skills.
- Ability to manage multiple partner engagements simultaneously.
—
—