Job Description — Site Reliability Engineer (Observability)
Job Location: New York City, NY (Hybrid — Onsite Required)
Job Type: Long-Term Contract
Client – Gspann / End client not disclosed at this moment.
Key Skills: Retail domain with skills like Splunk, LogicMonitor/Dynatrace/Datadog tool capabilities.
Role Summary
· We are seeking an experienced Site Reliability Engineer (SRE) with strong Observability expertise to support large-scale, customer-facing retail platforms. The ideal candidate has hands-on experience with enterprise monitoring toolchains and a track record of ensuring system reliability and uptime in high-traffic retail/eCommerce environments. This is a hybrid, long-term contract role based in NYC with required onsite attendance.
Key Responsibilities
- Design and maintain observability pipelines covering logs, metrics, traces, and synthetic monitoring across production and non-production environments.
- Build dashboards, alerts, and SLO/SLI frameworks using Splunk and one or more of LogicMonitor, Dynatrace, or Datadog.
- Partner with application, infrastructure, and platform teams to define monitoring coverage for retail-critical systems (order management, checkout, inventory, POS, fulfillment).
- Lead incident response and root cause analysis, using observability data to reduce MTTD and MTTR.
- Establish proactive alerting strategies that reduce noise while maintaining signal fidelity for critical services.
- Support high-traffic seasonal events (holiday peak, promotions, flash sales) with readiness reviews and real-time monitoring support.
- Automate observability configuration (monitoring-as-code) and integrate observability into CI/CD pipelines.
- Document runbooks, escalation paths, and post-incident reviews.
Required Skills & Experience
- 7+ years in SRE, DevOps, or Observability/Monitoring engineering roles.
- Hands-on expertise with Splunk (search, dashboards, alerting, log pipeline management).
- Practical experience with at least one of: LogicMonitor, Dynatrace, or Datadog.
- Prior experience supporting retail or eCommerce platforms, ideally including peak-load event support.
- Strong understanding of distributed systems, microservices, and cloud infrastructure (AWS/Azure/GCP).
- Experience with incident management, on-call rotations, and postmortem/RCA processes.
- Scripting proficiency (Python, Shell, or similar) for monitoring/alerting automation.
- Familiarity with containerized environments (Kubernetes, Docker) and CI/CD tooling.
- Solid understanding of SLIs, SLOs, error budgets, and reliability engineering principles.
Preferred
- Additional observability tools (New Relic, Grafana, Prometheus, ELK).
- APM instrumentation experience (OpenTelemetry, tracing frameworks).
- Infrastructure-as-code experience (Terraform, Ansible).
- Familiarity with retail systems: POS, OMS, WMS, inventory/fulfillment platforms.
- Relevant cloud or observability tool certifications.
Thanks,
_______________________________________
Pratiksha Hatkar | New York Technology Partners
120 Wood Avenue S | Suite 504 | Iselin NJ 08830
Direct: 201.604.3826 EXT: 441