IF
Job Overview
Cluster Site Reliability Engineer Cluster & SRE Own the physical reality of the platform in one of our seven regions. You bring up new GPU racks, validate InfiniBand fabric end-to-end, and keep the cluster running at the SLA. This is a hands-on role; you will see the hardware. Team Cluster & SRE Location: On-site Region of choice: Cologix region of choice (Toronto, Montreal, Columbus, Vancouver, Ashburn) Job Type and Tier Full-time IC4-IC6 Stack Linux, InfiniBand (NDR / XDR), NCCL / RCCL, Kuber…
What you'll do
- Manage physical infrastructure across assigned region
- Bring up new GPU racks
- Validate InfiniBand fabric end-to-end
- Maintain cluster uptime and SLA compliance
- Hands-on hardware management and troubleshooting
- Own platform reliability and performance in region
What you'll need
- On-site work required
- Location in one of seven Cologix regions (Toronto, Montreal, Columbus, Vancouver, or Ashburn)
- Full-time position
- IC4-IC6 level experience
- Hands-on hardware experience
- Knowledge of Linux systems
- Understanding of InfiniBand fabric (NDR/XDR)
- Experience with NCCL/RCCL
- Kubernetes knowledge
About the Company
IF