Get all C2C Jobs / hotlists 🔥 Alerts

Urgent hiring :: Site Reliability Engineering/ Core Platform Engineer (SRE) :: Sunnyvale, CA or San Jose, CA.

Contract

Centraprise

Job Title : Site Reliability Engineering / Core Platform Engineer (SRE)
Location : Sunnyvale, CA or San Jose, CA (5 days On-site) (Onsite 5 days onsite)

Type : Contract
VISA : H1B & H4EAD 

Experience : 11+ Years 

 

Must have :

Candidate can clearly explain a significant infrastructure project and their direct contributions.

Demonstrates expertise in GPU, Compute, SDN, Storage, or Virtualization.

 

Job Description:

As a Core Platform Engineer, we will serve as the first responder for production incidents, orchestrate incident management, drive reliability improvements, and establish SRE best practices across Compute, Networking, Storage, and GPU infrastructure teams. Candidates should possess hands-on infrastructure experience and sufficient technical depth to identify affected systems, engage the right subject matter experts, and drive incident resolution processes using data and observability signals.

Responsibilities

Act as first responder during infrastructure incidents.
Lead incident bridges and coordinate cross-functional response efforts.
Perform incident triage and identify impacted infrastructure domains.
Gather evidence and telemetry to route incidents to the correct SME team.
Drive incident communications and stakeholder updates.
Improve reliability processes across platform engineering teams.
Define and promote SRE best practices and operational standards. (SLO,SLI)
Identify observability gaps and implement improvements.
Build automation for incident response workflows.
Manage and optimize incident management tooling (e.g., incident.io).
Support change management and operational readiness processes.
Assist foundation engineering teams in identifying reliability risks and trends.
Participate in on-call activities and operational reviews.
Technical Environment

The Platform Engineering organization supports:
Kubernetes (GKE)
Cloud Hypervisor
Linux KVM/QEMU
Open vSwitch (OVS)
OVN
Lightbits Storage
Pure Storage
Large-scale GPU infrastructure
Mixed bare-metal and virtualized environments
Required Qualifications

5 to 10+ years of experience in Site Reliability Engineering, Platform Engineering,
Infrastructure Operations, or Systems Engineering.
Strong infrastructure troubleshooting experience.
Deep expertise in at least one of the following:
GPU infrastructure
KVM/virtualization
SDN (OVN/OVS)
Storage (Lightbits/Pure Storage)
Proven incident management and operational leadership experience.
Experience running high-severity production incidents.
Strong understanding of observability, monitoring, SLIs, and SLOs.
Experience building operational automation.
 

 

Thanks & Regards.
Upesh ,
Please send an email If i miss your phone call..
E: upesh@centraprise.com

 

To apply for this job email your details to upesh@centraprise.com

×

Post your C2C job instantly

Quick & easy posting in 10 seconds

Keep it concise - you can add details later
Please use your company/professional email address
Simple math question to prevent spam