Site Reliability Engineer (Senior SRE)
Atlanta GA Need Local
Long Term
Contact
• Design and operate highly available AWS infrastructure for large-scale, distributed applications.
• Improve system reliability and scalability by identifying single points of failure, capacity constraints, and performance bottlenecks.
• Build and maintain infrastructure using services such as EC2, EKS, ECS, Lambda, S3, RDS, DynamoDB, VPC, IAM, CloudWatch, and Route 53.
• Red Hat OpenShift Service on AWS (ROSA): Managed and supported OpenShift/Kubernetes clusters on AWS, including cluster provisioning, upgrades, scaling, networking, IAM/STS integration, workload deployments, and integration with AWS services. observability
• Develop Infrastructure as Code (IaC) using tools such as Terraform, AWS
• Create and maintain CI/CD pipelines for safe, repeatable application and infrastructure deployments.
• Develop automation and operational tooling using Python,, Bash, or similar languages.
• Establish Dynatrace monitoring, logging, alerting, and observability using CloudWatch and other monitoring platforms.
• Define and track SLIs, SLOs, and SLAs and use reliability data to prioritize engineering work.
• Participate in architecture/design reviews and influence engineering standards across teams.
• Participate in on-call rotations and respond to production incidents.
• Lead or participate in incident response, troubleshooting, root-cause analysis, and post-incident reviews.
• Automate repetitive operational tasks and eliminate toil.
• Perform capacity planning, performance testing, load testing, and disaster-recovery planning.
• Improve security, access controls, networking, and compliance across AWS environments.
• Work with development teams to make applications more observable, fault-tolerant, and production-ready.
• Design and test high-availability and disaster-recovery architectures, including multi-AZ and, where appropriate, multi-region deployments.
• Conduct operational readiness reviews before major services or releases go into production.
• Troubleshoot complex issues involving Linux, networking, DNS, databases, containers, Kubernetes, distributed systems, and cloud infrastructure.
• Mentor junior and mid-level engineers and provide technical leadership on reliability initiatives.
Munesh
770-838-3829,
CYBER SPHERE LLC