Role: SRE Engineer with ML Ops and Arize product
Location: Malvern, PA (Onsite)
Duration: Long Term Contract
ROLE_DESCRIPTION –
AWS Cloud Platform Expertise – EC2, EKS, ECS, Lambda, CloudWatch, SNS, SQS, EventBridge for highly available and scalable services.
Arize AI Observability & Monitoring – Model performance monitoring, drift detection, evaluation analytics, and AI/LLM observability.
Site Reliability Engineering (SRE) – SLI/SLO definition, Error Budgets, reliability improvement, service resilience, and uptime management.
Incident Management & RCA – Major Incident Management (MIM), outage tracking, root cause analysis, problem management, and MTTR reduction.
Failure Analysis & Risk Assessment – FMEA, risk quantification, reliability assessments, and proactive mitigation of platform failures.
Testing & Validation Engineering – Scenario testing, regression testing, impact analysis, release validation, and upstream change testing.
Monitoring, Alerting & Automation – CloudWatch, Grafana, Prometheus, PagerDuty, automated notifications, dashboards, and operational metrics.
DevOps & MLOps Practices – Kubernetes, Terraform, CI/CD pipelines, Python scripting, AI/ML platform operations, and LLM reliability optimization.
Preferred Technologies: AWS, Arize AI, Kubernetes (EKS), Cloudformation, Python, CloudWatch, Grafana, Prometheus, PagerDuty, GitHub Actions/Jenkins.
Skills: Digital : Amazon Web Service(AWS) Cloud Computing~Digital : Machine Learning
Contact Information
Email: sonu.chauhan@1rpo.net
Click the email address to contact the job poster directly.