Vbeyond
VBeyond Corporation || PARTNERING FOR GROWTH
Job Title: SRE AI Engineer:
Austin , TX / Charlotte, NC
Job Description:
We are currently seeking a highly skilled SRE hands-on AI Engineer with solid experience in AI Observability and instrumentation approaches for AI systems, AI Agents development to perform detection, diagnosis and autonomous self-healing operates on AI Control Plane.
Observability data collection and automation to help lead transformational initiatives within IT operations, encompassing development as well. As a crucial figure in this role, you will participate/help with various technology domain groups and cross functional teams on unified observability gap analysis and solutioning (automation and manual fixes)
Responsibilities:
Incorporate GenAI tooling and agentic capabilities to strengthen reliability outcomes across monitoring/alerting, rapid incident response, change management/testing, and DevOps/deployment processes.
Experience building agentic workflows using LLMs, tool-calling, function-calling, multi-agent orchestration, and event-driven automation.
Experience with Agent-to-Agent communication, AI agent federation, and enterprise AI control-plane concepts.
Experience implementing AI control-plane governance, including policy-based execution, approval workflows, audit trails, guardrails, and risk-based remediation controls.
Expertise in Observability as a service, Dashboard as a services, monitoring as a services and alert as a service in all technology domains (application, infrastructure, database, security, middleware, network etc.,) Telemetry data collection using Dynatrace APM, SolarWinds, CISCO Switches, F5, Databases, Open-Source tools (Prometheus and Grafana), Log Aggregations (Kibana or Splunk) and AIOPS Tools.
Practical experience implementing Golden Signals (latency, traffic, errors, saturation) using related telemetry sources.
Configure application performance monitoring (APM), infrastructure monitoring, synthetic monitoring, RUM, and log monitoring.
Integrate Dynatrace with CI/CD pipelines, alerting tools, ITSM systems, and incident automation frameworks.
Tune alert thresholds, baselines, and AI-driven anomaly detection to reduce noise and improve actionable insights.
Deeper understanding of Login authentication mechanisms using Ping, ForgeRock and SiteMinder technologies (session management and cookie management)
Define best practices and principles for SRE, including monitoring, alerting, and automation.
Collaborate with development teams on resiliency to ensure that services and applications are designed with operational reliability in mind.
Implement monitoring systems to assess the performance of applications and infrastructure and proactively identifying areas for optimization.
Ability to develop close relationship with other operational teams to integrate SRE practices and drive overall operational improvements across enterprise.
Stay up to date on industry trends, new technologies, and best practices in SRE and applying relevant advancements to the organization.
Ability to build strong working relationships across different levels, client focus mindset.
Â
Qualifications:
Around 11+ years of SRE hands on experience with AI OPS, cloud technologies, development, SRE toolsets and automation
Hands-on experience implementing Retrieval-Augmented Generation using approved enterprise knowledge sources such as runbooks, SOPs, RCA documents, incident history, architecture documents, and knowledge articles.
Expertise SPEC driven and Prompting using Ai IDE tools – Cursor, Kiro and Good Antigravity
Hands-on experience AI LLM’s – GPT – Open AI, Claude, Gemini etc.,
Experience with LangChain, LangGraph, Bedrock Agents, Azure AI Foundry, or Vertex AI Agent Builder
Experience with vector databases such as OpenSearch, Pinecone, Chroma, Redis Vector, or pgvector.
Experience performing Observability current-state assessments, gap analysis and solutioning (automation and manual fixes) in all technology domains (application, infrastructure, database, security, middleware, network etc.,),
Strong hands-on automation experience in Observability as a code, dashboard as a code, monitoring as a code, alert as a code (Instrumentation, templates, automatic deployment, visualization and alerting)
Strong hands-on experience with any Cloud Technology (AWS): Control Tower, Project Setup, Creating Accounts, RDS, SSO
Monitoring & alerting setup experience with Splunk, Prometheus, Grafana, Kibana, ELK, with pref. for APM (Dynatrace).
Strong skills in APM, distributed tracing, synthetic & real user monitoring, log monitoring, and Davis AI configuration
Own the design, configuration, CICD deployment, and optimization for enterprise-wide observability tools.
Experience integrating, automation, and cloud platforms (AWS, Azure, GCP).
Extended experience instrumenting OTEL Framework.
Hands on experience with Dynatrace Plug-and-play observability modules (OKit) development for Observability Developers Java and .Net applications.
Define monitoring standards, best practices, and governance to ensure consistency and scalability.
Experience to deploy and tune OneAgent, build end-to-end PurePath tracing, and leverage Smartscape topology for proactive performance monitoring and root-cause analysis.
Collaborate with application and infrastructure teams to troubleshoot performance issues and implement permanent fixes.
Good to have:
Any of the relevant professional certifications – AIOPS related certifications, Certified Site Reliability Engineer (CSRE), Certified Kubernetes Administrator (CKA), AWS Certified DevOps Engineer Professional, , Google Cloud Professional; DevOps Engineer
Â
AbihaaranL@VBeyond.com
To apply for this job email your details to AbihaaranL@VBeyond.com