Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer

GrabJobs

Job description What are we building? Hard Rock Digital is a team focused on becoming the best online sportsbook, casino, and social gaming company in the world. We’re building a team that resonates passion for learning, operating, and building new products and technologies for millions of consumers. We care about each customer interaction, experience, behavior, and insight and strive to ensure we’re always acting authentically. Rooted in the kindred spirits of Hard Rock and the Seminole Tribe of Florida, Hard Rock Digital taps a brand known the world over as the leader in gaming, entertainment, and hospitality. We’re taking that foundation of success and bringing it to the digital space - ready to join us? What’s the position? We are looking for a Senior Site Reliability Engineer who combines deep infrastructure expertise with a forward-thinking approach to AI-driven operations. In this role you will maintain and improve the reliability, scalability, and performance of our Java-based applications while pioneering the use of large language models (LLMs), agentic workflows, and intelligent automation to transform how we monitor, respond to, and prevent incidents. You will design and build autonomous and semi-autonomous AI agents that consume observability data, triage alerts, generate runbooks, automate incident response steps, and surface actionable insights—reducing toil and accelerating mean time to resolution. This is a hands-on engineering role for someone who is equally comfortable tuning a JVM, writing PromQL, and prototyping an agentic pipeline with tool-calling LLMs. Key Responsibilities Application Reliability & Performance Ensure the availability, reliability, and performance of high-traffic Java-based applications in a distributed environment. Troubleshoot and resolve complex issues across production and non-production environments. Participate in pre- and post-deployment performance testing and monitoring to continuously improve application performance. Optimize Java application performance with a focus on JVM tuning, efficient resource utilization, and horizontal scaling. Monitoring, Observability & AIOps Deploy and manage the Grafana stack (Grafana, Prometheus, Loki, Mimir, Alloy) to deliver real-time monitoring, logging, and alerting. Implement and refine observability strategies that enhance visibility into application and infrastructure health. Create and maintain dashboards, alerts, and log queries for comprehensive system health monitoring. Integrate AI/ML models into the observability pipeline for anomaly detection, predictive alerting, and intelligent alert correlation and noise reduction. AI & Agentic Workflow Engineering Design, build, and operate agentic AI workflows that automate operational tasks such as alert triage, root cause analysis, runbook execution, and incident summarization. Develop tool-calling LLM agents that interact with infrastructure APIs (Kubernetes, Grafana, Jira, Slack, PagerDuty) to execute diagnostic and remediation actions autonomously or with human-in-the-loop approval. Build and maintain MCP (Model Context Protocol) servers and integrations that expose internal systems as tool surfaces for AI agents. Evaluate, select, and operationalize LLM frameworks and orchestration platforms (e.g., LangChain, LangGraph, CrewAI, n8n, or custom solutions) for production-grade agentic systems. Implement guardrails, evaluation harnesses, and feedback loops to ensure AI agent outputs are accurate, safe, and continuously improving. Champion the adoption of AI-assisted development and operations practices across the SRE and broader engineering organization. Incident Management & Root Cause Analysis Support the operations team’s incident response efforts, conduct post-mortems, and identify root causes to prevent recurrence. Leverage AI tools to accelerate incident timelines, auto-generate post-mortem drafts, and surface patterns across historical incidents. Document and share lessons learned, contributing to a culture of continuous improvement. Automation & Toil Reduction Identify repetitive operational workflows and engineer AI-augmented or fully automated replacements. Build self-service tools and chatbot interfaces that allow engineering teams to query system status, retrieve logs, and execute standard operating procedures through natural language. Measure and report on toil reduction metrics to quantify the impact of automation initiatives. Collaboration & Cross-functional Support Work closely with developers, architects, and data/ML engineers to design solutions that improve reliability and leverage AI capabilities. Collaborate with DevOps and NOC teams to support the application platform. Communicate SRE practices, AI/automation capabilities, and operational insights to technical and non-technical stakeholders. Provide feedback on application performance, potential improvements, and observability metrics. Why This Role Is Different This is not a traditional SRE position with AI bolted on as an afterthought. We are building a team that treats AI and agentic automation as core competencies—on par with Kubernetes expertise or observability design. You will have the autonomy to experiment with cutting-edge AI tools, the backing of leadership to deploy them in production, and a mandate to measurably reduce operational toil through intelligent systems. Job requirements What are we looking for? Core SRE & Infrastructure (Required) Degree in Computer Science or a related field, or equivalent professional experience. 5+ years in SRE, DevOps, or similar infrastructure roles with experience managing large-scale, high-availability production systems. 3+ years hands-on experience managing production Kubernetes clusters, including deep understanding of architecture, networking, storage, and security. Experience with cluster autoscaling (Karpenter), upgrades, and multi-cluster management. Proficiency with kubectl, Helm, Kubernetes operators, and container orchestration troubleshooting. Advanced expertise with the Grafana observability stack: dashboards, alerting, visualization, and Grafana Alloy for telemetry collection. Proficiency in PromQL and experience with Loki for log aggregation and analysis. Hands-on experience managing Java-based applications in distributed environments, including JVM tuning and optimization. Cloud platform expertise (AWS preferred; GCP or Azure also valued). Familiarity with Infrastructure as Code tools such as Terraform/Terragrunt or Ansible. ArgoCD proficiency for GitOps workflows and continuous deployment. Strong scripting abilities in Python, Bash, or Go, with experience building CI/CD pipelines and deployment automation. Proven track record with on-call rotations, incident response, and root cause analysis. AI, Automation & Agentic Systems (Required) 1+ years of practical experience building or operating AI/LLM-powered tools, agents, or workflows in a production or production-adjacent context. Demonstrated ability to design agentic systems that use tool calling, retrieval-augmented generation (RAG), or multi-step reasoning to accomplish operational tasks. Experience integrating LLM APIs (e.g., Anthropic Claude, OpenAI, or open-source models) into backend services or automation pipelines. Familiarity with at least one agentic orchestration framework or workflow engine (LangChain, LangGraph, CrewAI, n8n, Temporal, or equivalent). Understanding of prompt engineering best practices, including structured outputs, system prompts, and few-shot examples. Familiarity with AI-assisted coding tools (Claude Code, Codex, Cursor) and their integration into engineering workflows. Experience building or consuming MCP (Model Context Protocol) servers to expose internal tools to AI agents. Awareness of AI safety, hallucination mitigation, and human-in-the-loop design patterns for autonomous systems. Preferred / Bonus Hands-on experience with vector databases (Pinecone, Weaviate, pgvector) for RAG-based knowledge retrieval. Experience with LLM evaluation frameworks (e.g., Galileo, LangSmith, Braintrust) for monitoring agent quality in production. Contributions to open-source AI/ML or SRE tooling projects. Background in data engineering or ML pipelines that complements SRE responsibilities. Soft Skills Strong communication skills (written and verbal) with the ability to translate complex AI and infrastructure concepts for diverse audiences. Proactive problem-solver with a bias toward automation and continuous improvement. Ability to mentor junior team members on both traditional SRE practices and emerging AI-driven approaches. Positive attitude and openness to constructive feedback. What’s in it for you? We offer our employees more than just competitive compensation. Our team benefits include: Competitive pay and benefits Flexible vacation allowance A hybrid / remote working environment Startup culture backed by a secure, global brand Roster of Uniques We care deeply about every interaction our customers have with us, and trust and empower our staff to own and drive their experience. Our vision for our business and customers is built on fostering a diverse and inclusive work environment where regardless of background or beliefs you feel able to be authentic and bring all your talent into play. We want to celebrate you being you (we are an equal opportunity employer). All done! Your application has been successfully submitted! Other jobs You've already applied for this job We appreciate your interest in this position. Unfortunately, you have already applied for this job.

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in Denver, CO vacancy
  • $105.6k - $145.2k

     ...Future of Enterprise Cloud: Become Our Next Enterprise Cloud Engineer! Ready to make a tangible impact on global industries using...  ...performance through seamless Azure integrations and proactive site reliability. About Us Trimble is a global technology company that connects... 
    Suggested
    Ongoing contract
    Full time
    Work at office
    Local area
    Worldwide

    Trimble Inc.

    Westminster, CO
    3 days ago
  • $62k - $141k

    Site Reliability Engineer The Opportunity: Engineering to make a system more resilient and efficient frees up time and money to build more capabilities. Whether you come from a background in network engineering, systems administration, or software development, if you... 
    Suggested
    Full time
    Contract work
    Part time
    Work at office
    Local area
    Remote work

    Booz Allen Hamilton

    Aurora, CO
    2 days ago
  • $62k - $141k

    Site Reliability EngineerThe Opportunity: Engineering to make a system more resilient and efficient frees up time and money to build more capabilities. Whether you come from a background in network engineering, systems administration, or software development, if you have... 
    Suggested
    Full time
    Contract work
    Part time
    Work at office
    Local area
    Remote work

    Booz Allen Hamilton

    Aurora, CO
    2 days ago
  • $86.9k - $198k

    Site Reliability Engineer, SeniorThe Opportunity: Engineering to make a system more resilient and efficient frees up time and money to build more capabilities. Whether you come from a background in network engineering, systems administration, or software development, if... 
    Suggested
    Full time
    Contract work
    Part time
    Work at office
    Local area
    Remote work

    Booz Allen Hamilton

    Aurora, CO
    5 days ago
  • $104.43k - $156.65k

     ...Comcast. (In most cases, Comcast prefers to have employees on-site collaborating unless the team has been designated as virtual...  ..., Fox, Disney, NBC, Paramount+, and many others.Our Site Reliability Engineering (SRE) team is at the heart of our mission to deliver seamless... 
    Suggested
    Permanent employment
    Full time
    Work at office
    Remote work
    Worldwide
    Flexible hours

    Comcast

    Centennial, CO
    12 hours ago
  • $100k - $115k

     ...Internal Developer Platform (IDP) as a product, treating engineering teams as customers and optimizing for reliability, usability, and delivery velocity.Define and...  ....4+ years of experience in Platform Engineering, Site Reliability Engineering, DevOps, or Systems Engineering... 
    Temporary work

    Analytic Partners

    Denver, CO
    4 days ago
  • $127k - $249k

    The TeamPlatform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions...  ...fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper).... 
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    Denver, CO
    12 hours ago
  •  ...foundation of success and bringing it to the digital space - ready to join us? What’s the position? We are looking for a Senior Site Reliability Engineer who combines deep infrastructure expertise with a forward-thinking approach to AI-driven operations. In this role you will... 
    Remote work
    Flexible hours
    Night shift

    GrabJobs

    Denver, CO
    1 day ago
  • $87.4k - $123.4k

     ...the U.S. We are unable to sponsor or take over sponsorship of an employment visa at this time, including CPT/OPT.*** The Site Reliability Engineer will help ensure the reliability, scalability, and performance of Empower’s financial services platform. This person will... 
    16 hours
    Contract work
    Temporary work
    Work experience placement
    Casual work
    Work at office
    Local area
    Remote work
    Work from home
    Work visa
    Flexible hours

    Empower Retirement

    Greenwood Village, CO
    12 hours ago
  • $98.58k - $138.02k

     ...Site Reliability Engineer II Restaurant365 is a SaaS company disrupting the restaurant industry! Our cloud-based platform provides a unique, centralized solution for accounting and back-office operations for restaurants. Restaurant365's culture is focused on empowering... 
    Work at office

    Restaurant365

    Denver, CO
    4 hours ago
  • $131k - $227.13k

     ...Description: The 1LMX MES COE is seeking an engineer who will own infrastructure‑as‑code, cloud platform, and reliability for the Apriso environment on AWS. This role blends full‑stack development, DevOps, and Site Reliability Engineering (SRE) practices to deliver... 
    Full time
    Temporary work
    Work experience placement
    Work at office
    Remote work
    Relocation
    Flexible hours
    Shift work
    3 days per week

    Lockheed Martin Corporation

    Littleton, CO
    3 days ago
  • $110k - $145k

     ...content reflecting our world. NBCU's Distribution engineering is responsible for the automation and reliability of NBCU's Live sources. Reasonable for the...  ...Distribution Engineering is looking to add a talented Site Reliability Engineer to be part of our Video Streaming... 
    Work experience placement
    Work at office
    Local area

    NBCUniversal

    Greenwood Village, CO
    4 hours ago
  •  ...focusing on private cloud systems supporting 5G wireless systems. This position will focus on platform monitoring, logging, and reliability aspects supporting the Mobile Core team. A critical goal is to gather metrics of the platform during stress and load events to ensure... 

    Software Technology Inc

    Denver, CO
    2 days ago
  • $160k - $190k

     ...Site Reliability Engineer (Classified Deployments) Location: Southern California or Washington, D.C. Clearance: Active Secret required; TS/SCI strongly preferred Work Mode: Hybrid/On-site with government customers Citizenship: U.S. Citizen Compensation:... 

    Zachary Piper Solutions

    Arvada, CO
    4 days ago
  • $112.5k - $187.5k

     ...NoticePersonal Information We CollectYour Privacy ChoicesTeam OverviewAt TransUnion, this role will report to a DevOps Director. The Site Reliability Engineering team drives reliability strategy, elevates engineering standards, and owns some of the most complex and consequential... 
    Full time
    Temporary work
    Work experience placement
    Work at office
    Flexible hours
    2 days per week

    TransUnion

    Greenwood Village, CO
    1 day ago
  • $130k - $180k

     ...alongside some of the most experienced and innovative leaders and engineers in the field. Where we work Headquartered in Amsterdam and...  ...an in-house AI R&D team. The role Nebius is looking for a Site Reliability Engineer in Hardware Infrastructure team. You’re welcome to... 
    Temporary work
    Work at office
    Immediate start
    Remote work
    Flexible hours

    GrabJobs

    Denver, CO
    2 days ago
  •  ...scale, we invite you to bring your talents to Zscaler to help shape the future of cybersecurity. Role We are looking for a Site Reliability Engineer-SkillBridge Intern (San JosA Ca or Bellevue WA) to join our Zero Trust Exchange team. This is a remote role based in San... 
    Internship
    Work at office
    Local area
    Remote work
    Worldwide

    GrabJobs

    Aurora, CO
    6 hours ago
  • $114k - $165.3k

     .... We are unable to sponsor or take over sponsorship of an employment visa at this time, including CPT/OPT.*** The Lead Site Reliability Engineer will combine deep technical expertise with team leadership to drive reliability across Empower's financial services platform... 
    16 hours
    Contract work
    Temporary work
    Work experience placement
    Casual work
    Work at office
    Local area
    Remote work
    Work from home
    Work visa
    Flexible hours

    Empower Retirement

    Greenwood Village, CO
    1 day ago
  • $160k - $180k

     ...with headquarters in Denver, Colorado, and offices across the U.S., Canada, and India. We are seeking a Principal Site Reliability Engineer to define the strategic vision and own the enterprise-wide reliability, scalability, and performance of our critical... 
    Contract work
    Temporary work
    Work at office
    Work from home
    Flexible hours

    Vertafore

    Denver, CO
    3 days ago
  • $148k - $193k

     ...Staff Site Reliability Engineer Denver, CO, USA DAT is an award-winning employer of choice and a next-generation SaaS technology company that has been at the leading edge of innovation in transportation supply chain logistics for 45 years. We continue to transform... 
    Temporary work
    Work experience placement
    Work at office
    Local area
    Immediate start
    Flexible hours

    DAT Freight & Analytics

    Denver, CO
    1 day ago
  • $175k - $220k

     ...is global, with headquarters in Denver, Colorado, and offices across the U.S., Canada, and India. The Director, Site Reliability Engineering (SRE) will lead reliability, performance, and observability initiatives for a portfolio of Vertafore products. This role... 
    Contract work
    Temporary work
    Work at office
    Work from home
    Flexible hours

    Vertafore

    Denver, CO
    2 days ago
  •  ...PNC is seeking a Software Engineering Manager—Site Reliability Engineering to lead a 24x7 production support team and drive reliability across mission-critical platforms powering PNC's digital experiences. You will manage incident response, RCA programs, and cross-functional... 

    Fairygodboss

    Lakewood, CO
    2 days ago
  • $110k - $155k

     ...global, with headquarters in Denver, Colorado, and offices across the U.S., Canada, and India. We are seeking a Senior Site Reliability Engineer to own the reliability, scalability, performance, and operational integrity of critical production services. This role is... 
    Contract work
    Work at office
    Work from home
    Flexible hours

    Vertafore

    Denver, CO
    17 days ago
  • $130k - $200k

    Summary Position Title: Site Reliability Engineer Position ID: TA247 Location(s): On-site; Aurora, CO; Herndon, VA Application Deadline: August 31, 2026 Security Clearance Requirement: TS/SCI Security Clearance with Polygraph Job Description Trusted Space... 
    Full time
    Temporary work
    Local area

    Trusted Space, LLC

    Aurora, CO
    12 hours ago
  •  ...Position - Senior Site Reliability Engineer Location - 100% Remote Experience - 10+ Years Full Time Hiring Job Description - Must Have Technical/Functional Skills: • 10+ years of experience in SRE, DevOps, or infrastructure engineering... 
    Full time
    Remote work

    vaaridatech

    Denver, CO
    8 days ago
  • $160k - $200k

     ...Job Description Job Description Description TL;DR Kharon is seeking a full-time Senior Site Reliability Engineer based in Denver. This role requires in-office attendance at least 3 days a week.  RESPONSIBILITIES: Spearhead the full lifecycle management of our... 
    Full time
    Work at office
    Immediate start
    Flexible hours
    3 days per week

    Kharon

    Denver, CO
    24 days ago
  • $90k - $105k

     ...Job Description Job Description Site Reliability EngineerAbout Acquire Learning Acquire Learning is a learning management platform built...  ...Role Acquire is hiring its first dedicated Site Reliability Engineer, a mid-level role with a clear path to Lead SRE as we grow.... 
    Work at office
    Local area
    Work from home

    BehaviorSpan

    Denver, CO
    a month ago
  • $107.5k - $204.5k

     ...demands of a rapidly evolving global market. Collins is seeking Engineers to be part of an engineering team for the AF DCGS High Band...  ...active TS/SCI security clearance. Travel to customer CONUS & OCONUS sites as required, but no more than 25%. This position is an onsite... 
    Contract work
    Temporary work
    Work experience placement
    Work at office
    Remote work
    Relocation package
    Flexible hours

    Raytheon

    Aurora, CO
    1 day ago
  • $160.8k - $214.1k

     ...observability needs of modern infrastructure. The Customer Reliability Engineering team is the deep technical escalation tier for Cisco Hypershield...  .../fix and reliability cases escalated by Cisco TAC, applying Site Reliability Engineering practices across the full stack: the... 
    Full time
    Temporary work
    Local area
    Remote work
    Flexible hours

    CISCO Systems

    Denver, CO
    5 days ago
  • $86.8k - $198k

     ...specifications makes you an integral part of delivering a customer-focused engineering solution. If this sounds like you, come join Booz Allen’s...  ...total benefits by visiting the Resource page on our Careers site and reviewing Our Employee Benefits page.Salary at Booz Allen... 
    Full time
    Contract work
    Part time
    For contractors
    Work at office
    Local area
    Remote work

    Booz Allen Hamilton

    Aurora, CO
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!