IF

Cluster Site Reliability Engineer

Toronto, Ontario
On-site
Full-time
No salary posted2 weeks ago

Job Overview

Cluster Site Reliability Engineer Cluster & SRE Own the physical reality of the platform in one of our seven regions. You bring up new GPU racks, validate InfiniBand fabric end-to-end, and keep the cluster running at the SLA. This is a hands-on role; you will see the hardware. Team Cluster & SRE Location: On-site Region of choice: Cologix region of choice (Toronto, Montreal, Columbus, Vancouver, Ashburn) Job Type and Tier Full-time IC4-IC6 Stack Linux, InfiniBand (NDR / XDR), NCCL / RCCL, Kuber…

What you'll do

  • Manage physical infrastructure across assigned region
  • Bring up new GPU racks
  • Validate InfiniBand fabric end-to-end
  • Maintain cluster uptime and SLA compliance
  • Hands-on hardware management and troubleshooting
  • Own platform reliability and performance in region

What you'll need

  • On-site work required
  • Location in one of seven Cologix regions (Toronto, Montreal, Columbus, Vancouver, or Ashburn)
  • Full-time position
  • IC4-IC6 level experience
  • Hands-on hardware experience
  • Knowledge of Linux systems
  • Understanding of InfiniBand fabric (NDR/XDR)
  • Experience with NCCL/RCCL
  • Kubernetes knowledge

About the Company

IF

iFrame

Cluster Site Reliability Engineer at iFrame — Toronto, Ontario | Jobily