Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Site Reliability Engineer

Hard Rock Digital

What are we building?

Hard Rock Digital is a team focused on becoming the best online sportsbook, casino, and social gaming company in the world. We're building a team that resonates passion for learning, operating, and building new products and technologies for millions of consumers. We care about each customer interaction, experience, behavior, and insight and strive to ensure we're always acting authentically.

Rooted in the kindred spirits of Hard Rock and the Seminole Tribe of Florida, Hard Rock Digital taps a brand known the world over as the leader in gaming, entertainment, and hospitality. We're taking that foundation of success and bringing it to the digital space - ready to join us?

What's the position?

We are looking for a Senior Site Reliability Engineer who combines deep infrastructure expertise with a forward-thinking approach to AI-driven operations. In this role you will maintain and improve the reliability, scalability, and performance of our Java-based applications while pioneering the use of large language models (LLMs), agentic workflows, and intelligent automation to transform how we monitor, respond to, and prevent incidents.

You will design and build autonomous and semi-autonomous AI agents that consume observability data, triage alerts, generate runbooks, automate incident response steps, and surface actionable insights-reducing toil and accelerating mean time to resolution. This is a hands-on engineering role for someone who is equally comfortable tuning a JVM, writing PromQL, and prototyping an agentic pipeline with tool-calling LLMs.

Key Responsibilities

Application Reliability & Performance
  • Ensure the availability, reliability, and performance of high-traffic Java-based applications in a distributed environment.
  • Troubleshoot and resolve complex issues across production and non-production environments.
  • Participate in pre- and post-deployment performance testing and monitoring to continuously improve application performance.
  • Optimize Java application performance with a focus on JVM tuning, efficient resource utilization, and horizontal scaling.
Monitoring, Observability & AIOps
  • Deploy and manage the Grafana stack (Grafana, Prometheus, Loki, Mimir, Alloy) to deliver real-time monitoring, logging, and alerting.
  • Implement and refine observability strategies that enhance visibility into application and infrastructure health.
  • Create and maintain dashboards, alerts, and log queries for comprehensive system health monitoring.
  • Integrate AI/ML models into the observability pipeline for anomaly detection, predictive alerting, and intelligent alert correlation and noise reduction.
AI & Agentic Workflow Engineering
  • Design, build, and operate agentic AI workflows that automate operational tasks such as alert triage, root cause analysis, runbook execution, and incident summarization.
  • Develop tool-calling LLM agents that interact with infrastructure APIs (Kubernetes, Grafana, Jira, Slack, PagerDuty) to execute diagnostic and remediation actions autonomously or with human-in-the-loop approval.
  • Build and maintain MCP (Model Context Protocol) servers and integrations that expose internal systems as tool surfaces for AI agents.
  • Evaluate, select, and operationalize LLM frameworks and orchestration platforms (e.g., LangChain, LangGraph, CrewAI, n8n, or custom solutions) for production-grade agentic systems.
  • Implement guardrails, evaluation harnesses, and feedback loops to ensure AI agent outputs are accurate, safe, and continuously improving.
  • Champion the adoption of AI-assisted development and operations practices across the SRE and broader engineering organization.
Incident Management & Root Cause Analysis
  • Support the operations team's incident response efforts, conduct post-mortems, and identify root causes to prevent recurrence.
  • Leverage AI tools to accelerate incident timelines, auto-generate post-mortem drafts, and surface patterns across historical incidents.
  • Document and share lessons learned, contributing to a culture of continuous improvement.
Automation & Toil Reduction
  • Identify repetitive operational workflows and engineer AI-augmented or fully automated replacements.
  • Build self-service tools and chatbot interfaces that allow engineering teams to query system status, retrieve logs, and execute standard operating procedures through natural language.
  • Measure and report on toil reduction metrics to quantify the impact of automation initiatives.
Collaboration & Cross-functional Support
  • Work closely with developers, architects, and data/ML engineers to design solutions that improve reliability and leverage AI capabilities.
  • Collaborate with DevOps and NOC teams to support the application platform.
  • Communicate SRE practices, AI/automation capabilities, and operational insights to technical and non-technical stakeholders.
  • Provide feedback on application performance, potential improvements, and observability metrics.
Why This Role Is Different

This is not a traditional SRE position with AI bolted on as an afterthought. We are building a team that treats AI and agentic automation as core competencies-on par with Kubernetes expertise or observability design. You will have the autonomy to experiment with cutting-edge AI tools, the backing of leadership to deploy them in production, and a mandate to measurably reduce operational toil through intelligent systems.

What are we looking for?

Core SRE & Infrastructure (Required)
  • Degree in Computer Science or a related field, or equivalent professional experience.
  • 5+ years in SRE, DevOps, or similar infrastructure roles with experience managing large-scale, high-availability production systems.
  • 3+ years hands-on experience managing production Kubernetes clusters, including deep understanding of architecture, networking, storage, and security.
  • Experience with cluster autoscaling (Karpenter), upgrades, and multi-cluster management.
  • Proficiency with kubectl, Helm, Kubernetes operators, and container orchestration troubleshooting.
  • Advanced expertise with the Grafana observability stack: dashboards, alerting, visualization, and Grafana Alloy for telemetry collection.
  • Proficiency in PromQL and experience with Loki for log aggregation and analysis.
  • Hands-on experience managing Java-based applications in distributed environments, including JVM tuning and optimization.
  • Cloud platform expertise (AWS preferred; GCP or Azure also valued).
  • Familiarity with Infrastructure as Code tools such as Terraform/Terragrunt or Ansible.
  • ArgoCD proficiency for GitOps workflows and continuous deployment.
  • Strong scripting abilities in Python, Bash, or Go, with experience building CI/CD pipelines and deployment automation.
  • Proven track record with on-call rotations, incident response, and root cause analysis.
AI, Automation & Agentic Systems (Required)
  • 1+ years of practical experience building or operating AI/LLM-powered tools, agents, or workflows in a production or production-adjacent context.
  • Demonstrated ability to design agentic systems that use tool calling, retrieval-augmented generation (RAG), or multi-step reasoning to accomplish operational tasks.
  • Experience integrating LLM APIs (e.g., Anthropic Claude, OpenAI, or open-source models) into backend services or automation pipelines.
  • Familiarity with at least one agentic orchestration framework or workflow engine (LangChain, LangGraph, CrewAI, n8n, Temporal, or equivalent).
  • Understanding of prompt engineering best practices, including structured outputs, system prompts, and few-shot examples.
  • Familiarity with AI-assisted coding tools (Claude Code, Codex, Cursor) and their integration into engineering workflows.
  • Experience building or consuming MCP (Model Context Protocol) servers to expose internal tools to AI agents.
  • Awareness of AI safety, hallucination mitigation, and human-in-the-loop design patterns for autonomous systems.
Preferred / Bonus
  • Hands-on experience with vector databases (Pinecone, Weaviate, pgvector) for RAG-based knowledge retrieval.
  • Experience with LLM evaluation frameworks (e.g., Galileo, LangSmith, Braintrust) for monitoring agent quality in production.
  • Contributions to open-source AI/ML or SRE tooling projects.
  • Background in data engineering or ML pipelines that complements SRE responsibilities.
Soft Skills
  • Strong communication skills (written and verbal) with the ability to translate complex AI and infrastructure concepts for diverse audiences.
  • Proactive problem-solver with a bias toward automation and continuous improvement.
  • Ability to mentor junior team members on both traditional SRE practices and emerging AI-driven approaches.
  • Positive attitude and openness to constructive feedback.
What's in it for you?

We offer our employees more than just competitive compensation. Our team benefits include:
  • Competitive pay and benefits
  • Flexible vacation allowance
  • A hybrid / remote working environment
  • Startup culture backed by a secure, global brand

Roster of Uniques

We care deeply about every interaction our customers have with us, and trust and empower our staff to own and drive their experience. Our vision for our business and customers is built on fostering a diverse and inclusive work environment where regardless of background or beliefs you feel able to be authentic and bring all your talent into play. We want to celebrate you being you (we are an equal opportunity employer).
Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Senior Site Reliability Engineer in United States vacancy
  •  ...Evaluate applications, platforms, and vendors to assess resiliency, reliability, and operational risk.Design and implement processes that...  ...and reliability tooling.Actively participate in reliability engineering and resilience communities of practice, contributing to... 
    Senior
    Full time

    Vanguard

    Dallas, TX
    9 hours ago
  • Inspire Brands is hiring two Senior Site Reliability Engineers to help build and scale reliable, resilient, and observable systems supporting high-traffic, customer-facing digital platforms. These role blends software engineering, systems thinking, and operational excellence... 
    Senior
    Worldwide

    Inspire Brands

    Atlanta, GA
    1 day ago
  • $104.9k - $174.7k

    About the role:A FinOps Site Reliability Engineer (SRE) bridges the gap between engineering, operations, and financial governance by embedding cost optimization into infrastructure design, automation, monitoring, and operational processes. A FinOps SRE proactively identifies... 
    Senior
    Full time
    Local area

    LexisNexis Risk Solutions Group

    Boca Raton, FL
    1 day ago
  • $170k - $220k

    Who We're Looking ForWe’re looking for a hands-on, high-agency Site Reliability Engineer to help shape and scale the reliability layer of our stack. You'll own the release pipeline end-to-end — managing daily releases, weekly deploys, and hotfixes — while also automating... 
    Senior

    Supio

    Seattle, WA
    2 days ago
  • $174k - $252k

     ...systems by pushing for changes that improve reliability and velocity.Practice sustainable...  ...:Bachelor’s degree in Computer Science, Engineering, a related field, or equivalent practical...  ...degree in Computer Science or Engineering.Site Reliability Engineering (SRE) is what you... 
    Senior

    Google

    Sunnyvale, TX
    2 days ago
  • $130k - $200k

    IXL Learning, developer of personalized learning products used by millions of people globally, is seeking a Senior Site Reliability Engineer to join our team, and help maintain the reliability and optimal performance of our products. We are seeking engineers with a passion... 
    Senior
    Full time
    Work at office
    Immediate start

    IXL Learning

    San Mateo, CA
    9 hours ago
  • $104.9k - $174.7k

     ...Data Management. You can learn more about LexisNexis Risk at the link below, About the Role:We are hiring a hands-on Senior Site Reliability Engineer (SRE) to actively build, operate, and improve the reliability of our production systems. This is not a purely advisory... 
    Senior
    Full time
    Work at office
    Local area
    Remote work
    Work from home

    LexisNexis Risk Solutions Group

    Atlanta, GA
    1 day ago
  • About the RoleWe’re looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and operational excellence at the core. You’ll partner with engineers and data scientists to build, automate, and maintain the infrastructure... 
    Senior

    Alembic

    San Francisco, CA
    2 days ago
  •  ...TechMContact: Meghana GorusuCompany: SRI Tech SolutionsJob Title: Senior Site Reliability EngineerLocation: Plano , TX (remote)Years of Experience: 8...  ...are seeking a highly skilled Senior Site Reliability Engineer (SRE) to join our dynamic team. The ideal candidate will... 
    Senior
    Remote work

    SRI Tech

    Plano, TX
    2 days ago
  • LeanData helps the world’s fastest-growing companies automate, simplify, and accelerate revenue.We are looking for a Senior Site Reliability Engineer to lead the strategic evolution of our cloud infrastructure. Reporting directly to the SVP of Engineering, this role is... 
    Senior
    Full time
    Work at office
    2 days per week

    LeanData

    Santa Clara, CA
    2 days ago
  •  ...Lambda’s designated work from home day is currently Tuesday.Engineering at Lambda is responsible for building and scaling our cloud offering...  ...and SLIs for Kubernetes services, workloads, and platform reliability.You6+ years of experience in a SRE, operations engineer, or... 
    Senior
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    9 hours ago
  •  ...work from home day is currently Tuesday.Engineering at Lambda is responsible for building and...  ...and networking teams to improve service reliability and deployment workflowsDeploy and...  ...rotationYouHave 5+ years of experience in Site Reliability Engineering, Production Engineering... 
    Senior
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda Labs

    San Jose, CA
    4 days ago
  • $152.6k - $191.5k

     ...responsible for partnering with leaders across engineering and technology to define objective reliability goals for services. Key responsibilities include...  ...and continuous improvement.Position Summary:The Senior GCP Site Reliability Engineer acts as an advanced senior... 
    Senior
    Full time
    Work at office
    Day shift

    Bank of America

    Charlotte, MI
    9 hours ago
  • $158.5k - $172k

     ...the exceptional value they deserve.About The OpportunityAs a Senior Engineer on the Runtime Automation team, you will design, automate,...  ...environment. This is a high-impact position driving continuous reliability, deep system optimization, and automation across our entire... 
    Senior
    Full time
    Temporary work
    Work at office
    Flexible hours
    3 days per week

    GrubHub

    Chicago, IL
    2 days ago
  • Reliability Engineering Design, implement, and operate scalable, resilient, and highly available systems on Google Cloud Platform. Improve service...  ...Skills, and Abilities Three or more years of experience in Site Reliability Engineering, platform engineering, DevOps, cloud... 
    Senior
    Remote work

    Patterson-UTI

    Houston, TX
    9 hours ago
  •  ...candidates that are particularly strong in a few areas, and have some interest and capabilities in others.About the Role:As a Site Reliability Engineer, you’ll join the global Platform SRE team responsible for building, operating, and scaling Kong’s multi-region SaaS... 
    Senior
    Temporary work

    Kong

    Washington DC
    2 days ago
  •  ...Georgia, and serves customers in more than 35 countries worldwide.Position OverviewWe are seeking a highly experienced Senior Site Reliability Engineer (Unified Observability) to lead the design, implementation, and operational maturity of the F1 Next Generation... 
    Senior
    Full time
    Worldwide
    Flexible hours

    NCR

    Atlanta, GA
    1 day ago
  • Senior Site Reliability Engineer ILocationSan Jose, Costa Rica - RemoteSummary of roleOwn availability, the most important product feature, by continually striving for sustained operational excellence of Sumo’s planet-scale observability and security products. Work with... 
    Senior
    Flexible hours

    Sumo Logic

    San Jose, CA
    9 hours ago
  • $119.8k - $234.7k

     ...yearEmployment type: Full-TimeWork site: 3 days / week in-officeRole type: Individual...  ...EngineeringDiscipline: Site Reliability EngineeringCompany: MicrosoftOverviewMicrosoft...  ...’s most demanding workloads. As a Senior Site Reliability Engineer, you will lead reliability... 
    Senior
    Ongoing contract
    Local area
    3 days per week

    Microsoft

    Redmond, WA
    3 days ago
  • $148k - $235.75k

     ...see how you can make a lasting impact on the world.Join our team of innovative engineers who are building an AI Data Center AIOps platform that turns raw, high-volume telemetry into reliable, job-centric insights and automation for GPU fleets. We’re hiring a DevOps Engineer... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    4 days ago
  • $166k - $258k

     ...Seattle office a minimum of 4 days/week in order to be considered for this position.Nordstrom is looking for a Senior Engineer 2 to join our Site Reliability Engineering (SRE) team — and we think that person could be you.You'll help build the scalable, reliable, and resilient... 
    Senior
    Full time
    Work at office

    Nordstrom

    Seattle, WA
    1 day ago
  • $168k - $270.25k

    NVIDIA is looking for a Senior Site Reliability Engineer (SRE) to join its GeForce Now (GFN) team. SRE at NVIDIA ensures that our internal and external-facing GPU cloud gaming services have reliability and uptime as promised to the users and at the same time enables developers... 
    Senior
    Full time

    Nvidia

    Santa Clara, CA
    3 days ago
  • $165k - $225.6k

     ...From core infrastructure to enterprise platforms, we partner across functions to drive scale, reliability, and innovation through technology.The Senior Site Reliability Engineer OpportunityReporting to the Manager, Site Reliability Engineering, this role will help build,... 
    Senior
    Permanent employment
    Local area
    Worldwide
    Flexible hours

    Okta

    San Francisco, CA
    2 days ago
  • $80k - $140k

    Job DescriptionRBC Wealth Management Technology is seeking a Senior Site Reliability Engineer to join its Wealth Management SRE Team. This team is responsible for ensuring the performance, availability, resilience, and operational excellence of critical applications and... 
    Senior
    Full time
    Flexible hours
    Shift work

    Royal Bank of Canada

    Minneapolis, MN
    3 days ago
  • Job Description:Note: Fidelity will not provide immigration sponsorship for this positionThe RoleOur Site Reliability Engineering group within Enterprise Infrastructure combines Operations Excellence with the Development Experience to deliver services at high scale, high... 
    Senior
    Full time

    Fidelity Investments

    Durham, NC
    1 day ago
  • $210k - $230k

    GovCIO is currently hiring for a Senior Site Reliability Engineer (SRE) to design, implement, and maintain highly available, scalable, and resilient infrastructure systems. The ideal candidate will bridge the gap between development and operations, focusing on automation... 
    Senior
    Currently hiring
    Remote work

    Govcio

    Arlington, VA
    3 days ago
  • $139k - $257.55k

    The ChallengeThe Adobe Creative Community CCM organization is seeking an outstanding Senior Site Reliability Engineer (SRE) to support innovation through machine learning, autonomous AI workflows, and cloud-native infrastructure. Adobe Stock gives designers and businesses... 
    Senior
    Full time
    Temporary work
    Local area
    Remote work
    Worldwide

    Adobe Systems

    New York, NY
    4 days ago
  • $90k - $180k

     ...nutritionals and branded generic medicines. Our 115,000 colleagues serve people in more than 160 countries.About the RoleThis Senior Site Reliability Engineer position works on-site out of our Sylmar, CA or Sunnyvale, CA location in the Cardiac Rhythm Management Division.We... 
    Senior
    Remote work

    Abbott

    Sunnyvale, CA
    4 days ago
  • $98k - $176k

     ...joy of everyday life. We bring that vision to life through our values and culture. Learn more about Target here. As a Senior Site Reliability Engineer within Digital Enablement, you specialize in building and supporting the platforms and tools that enable teams to deliver... 
    Senior
    Full time
    Temporary work
    Work experience placement
    Flexible hours

    Target

    Brooklyn Park, MN
    9 hours ago
  • $104.9k - $174.7k

    Are you passionate about improving reliability, scalability, and resilience in complex database...  ....Own prioritization of reliability engineering tasks within team backlogs.Lead incident...  ...a Service (IaaS).Background in DevOps, site reliability engineering practices, or related... 
    Senior
    Full time
    Local area

    RELX Group

    Illinois
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!