Position: Senior Observability Operations Engineer
Location: Phoenix, AZ 85054
Duration: 6 months
Interview: Video
We are looking for an experienced Senior Observability Operations Engineer to manage, optimize, and enhance enterprise observability platforms that support mission-critical applications. The ideal candidate will have strong expertise in Dynatrace, Splunk, OpenSearch/Elasticsearch, Kubernetes, Linux, and cloud-native observability, with exposure to AIOps, Generative AI, and automation being highly desirable.
Key Responsibilities
Administer and optimize Dynatrace, Splunk, OpenSearch/Elasticsearch, and related observability platforms.
Design and maintain enterprise monitoring, logging, tracing, alerting, and dashboard solutions.
Manage large-scale OpenSearch/Elasticsearch environments, including performance tuning, indexing, capacity planning, and backup strategies.
Configure and support Dynatrace capabilities such as OneAgent, ActiveGate, APM, RUM, DEM, Synthetic Monitoring, and Davis AI.
Administer Splunk components including Indexers, Search Heads, Universal Forwarders, Cluster Manager, Deployment Server, and ITSI.
Support Linux, Kubernetes, Docker, OpenShift/Rancher, and cloud-based environments.
Implement observability best practices using OpenTelemetry, distributed tracing, metrics, logs, and events.
Troubleshoot production issues, perform root cause analysis, and drive operational excellence.
Automate operational processes using Python, Shell scripting, REST APIs, Terraform, and Ansible.
Lead platform upgrades, patching, security compliance, and reliability initiatives.
Leverage AI-driven observability, automation, and self-healing capabilities to improve incident detection and resolution.
Required Skills
Observability: Dynatrace, Splunk Enterprise, OpenSearch, Elasticsearch, Grafana, Prometheus, Kibana, Jaeger, OpenTelemetry
Infrastructure: Linux, Kubernetes, Docker, OpenShift/Rancher, Networking, System Administration
Cloud & DevOps: AWS/Azure/GCP, CI/CD, Git, Terraform, Ansible, REST APIs
Scripting: Python, Bash/Shell (PowerShell preferred)
Preferred AI & Automation Experience
Generative AI tools such as ChatGPT, GitHub Copilot, Amazon Q, or Microsoft Copilot
AIOps and AI-driven observability platforms
Dynatrace Davis AI
AI-assisted incident analysis, log analytics, and operational runbooks
Knowledge of RAG, vector databases, embeddings, LLMs, and prompt engineering
Python-based AI integrations (e.g., LangChain, LangGraph)
Qualifications
Bachelor’s degree in Computer Science, IT, Engineering, or equivalent experience
7-10+ years of IT infrastructure/observability operations experience
4+ years administering Dynatrace, Splunk, OpenSearch, or Elasticsearch
Strong Linux administration and enterprise production support experience
Excellent troubleshooting, communication, and stakeholder management skills
Preferred Certifications
Dynatrace Associate/Professional
Splunk Enterprise Certified Administrator
Elastic Certified Engineer
CKA/CKAD
AWS/Azure/GCP Certification
ITIL Foundation
AI/ML or Generative AI Certification
Key Attributes
Strong ownership and accountability
Analytical and problem-solving mindset
Ability to work independently in fast-paced environments
Strong collaboration and continuous learning focus
Contact Information
Email: tanu.kumari@anviktek.com
Click the email address to contact the job poster directly.