Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer

GrabJobs

Job description What are we building? Hard Rock Digital is a team focused on becoming the best online sportsbook, casino, and social gaming company in the world. We’re building a team that resonates passion for learning, operating, and building new products and technologies for millions of consumers. We care about each customer interaction, experience, behavior, and insight and strive to ensure we’re always acting authentically. Rooted in the kindred spirits of Hard Rock and the Seminole Tribe of Florida, Hard Rock Digital taps a brand known the world over as the leader in gaming, entertainment, and hospitality. We’re taking that foundation of success and bringing it to the digital space - ready to join us? What’s the position? We are looking for a Senior Site Reliability Engineer who combines deep infrastructure expertise with a forward-thinking approach to AI-driven operations. In this role you will maintain and improve the reliability, scalability, and performance of our Java-based applications while pioneering the use of large language models (LLMs), agentic workflows, and intelligent automation to transform how we monitor, respond to, and prevent incidents. You will design and build autonomous and semi-autonomous AI agents that consume observability data, triage alerts, generate runbooks, automate incident response steps, and surface actionable insights—reducing toil and accelerating mean time to resolution. This is a hands-on engineering role for someone who is equally comfortable tuning a JVM, writing PromQL, and prototyping an agentic pipeline with tool-calling LLMs. Key Responsibilities Application Reliability & Performance Ensure the availability, reliability, and performance of high-traffic Java-based applications in a distributed environment. Troubleshoot and resolve complex issues across production and non-production environments. Participate in pre- and post-deployment performance testing and monitoring to continuously improve application performance. Optimize Java application performance with a focus on JVM tuning, efficient resource utilization, and horizontal scaling. Monitoring, Observability & AIOps Deploy and manage the Grafana stack (Grafana, Prometheus, Loki, Mimir, Alloy) to deliver real-time monitoring, logging, and alerting. Implement and refine observability strategies that enhance visibility into application and infrastructure health. Create and maintain dashboards, alerts, and log queries for comprehensive system health monitoring. Integrate AI/ML models into the observability pipeline for anomaly detection, predictive alerting, and intelligent alert correlation and noise reduction. AI & Agentic Workflow Engineering Design, build, and operate agentic AI workflows that automate operational tasks such as alert triage, root cause analysis, runbook execution, and incident summarization. Develop tool-calling LLM agents that interact with infrastructure APIs (Kubernetes, Grafana, Jira, Slack, PagerDuty) to execute diagnostic and remediation actions autonomously or with human-in-the-loop approval. Build and maintain MCP (Model Context Protocol) servers and integrations that expose internal systems as tool surfaces for AI agents. Evaluate, select, and operationalize LLM frameworks and orchestration platforms (e.g., LangChain, LangGraph, CrewAI, n8n, or custom solutions) for production-grade agentic systems. Implement guardrails, evaluation harnesses, and feedback loops to ensure AI agent outputs are accurate, safe, and continuously improving. Champion the adoption of AI-assisted development and operations practices across the SRE and broader engineering organization. Incident Management & Root Cause Analysis Support the operations team’s incident response efforts, conduct post-mortems, and identify root causes to prevent recurrence. Leverage AI tools to accelerate incident timelines, auto-generate post-mortem drafts, and surface patterns across historical incidents. Document and share lessons learned, contributing to a culture of continuous improvement. Automation & Toil Reduction Identify repetitive operational workflows and engineer AI-augmented or fully automated replacements. Build self-service tools and chatbot interfaces that allow engineering teams to query system status, retrieve logs, and execute standard operating procedures through natural language. Measure and report on toil reduction metrics to quantify the impact of automation initiatives. Collaboration & Cross-functional Support Work closely with developers, architects, and data/ML engineers to design solutions that improve reliability and leverage AI capabilities. Collaborate with DevOps and NOC teams to support the application platform. Communicate SRE practices, AI/automation capabilities, and operational insights to technical and non-technical stakeholders. Provide feedback on application performance, potential improvements, and observability metrics. Why This Role Is Different This is not a traditional SRE position with AI bolted on as an afterthought. We are building a team that treats AI and agentic automation as core competencies—on par with Kubernetes expertise or observability design. You will have the autonomy to experiment with cutting-edge AI tools, the backing of leadership to deploy them in production, and a mandate to measurably reduce operational toil through intelligent systems. Job requirements What are we looking for? Core SRE & Infrastructure (Required) Degree in Computer Science or a related field, or equivalent professional experience. 5+ years in SRE, DevOps, or similar infrastructure roles with experience managing large-scale, high-availability production systems. 3+ years hands-on experience managing production Kubernetes clusters, including deep understanding of architecture, networking, storage, and security. Experience with cluster autoscaling (Karpenter), upgrades, and multi-cluster management. Proficiency with kubectl, Helm, Kubernetes operators, and container orchestration troubleshooting. Advanced expertise with the Grafana observability stack: dashboards, alerting, visualization, and Grafana Alloy for telemetry collection. Proficiency in PromQL and experience with Loki for log aggregation and analysis. Hands-on experience managing Java-based applications in distributed environments, including JVM tuning and optimization. Cloud platform expertise (AWS preferred; GCP or Azure also valued). Familiarity with Infrastructure as Code tools such as Terraform/Terragrunt or Ansible. ArgoCD proficiency for GitOps workflows and continuous deployment. Strong scripting abilities in Python, Bash, or Go, with experience building CI/CD pipelines and deployment automation. Proven track record with on-call rotations, incident response, and root cause analysis. AI, Automation & Agentic Systems (Required) 1+ years of practical experience building or operating AI/LLM-powered tools, agents, or workflows in a production or production-adjacent context. Demonstrated ability to design agentic systems that use tool calling, retrieval-augmented generation (RAG), or multi-step reasoning to accomplish operational tasks. Experience integrating LLM APIs (e.g., Anthropic Claude, OpenAI, or open-source models) into backend services or automation pipelines. Familiarity with at least one agentic orchestration framework or workflow engine (LangChain, LangGraph, CrewAI, n8n, Temporal, or equivalent). Understanding of prompt engineering best practices, including structured outputs, system prompts, and few-shot examples. Familiarity with AI-assisted coding tools (Claude Code, Codex, Cursor) and their integration into engineering workflows. Experience building or consuming MCP (Model Context Protocol) servers to expose internal tools to AI agents. Awareness of AI safety, hallucination mitigation, and human-in-the-loop design patterns for autonomous systems. Preferred / Bonus Hands-on experience with vector databases (Pinecone, Weaviate, pgvector) for RAG-based knowledge retrieval. Experience with LLM evaluation frameworks (e.g., Galileo, LangSmith, Braintrust) for monitoring agent quality in production. Contributions to open-source AI/ML or SRE tooling projects. Background in data engineering or ML pipelines that complements SRE responsibilities. Soft Skills Strong communication skills (written and verbal) with the ability to translate complex AI and infrastructure concepts for diverse audiences. Proactive problem-solver with a bias toward automation and continuous improvement. Ability to mentor junior team members on both traditional SRE practices and emerging AI-driven approaches. Positive attitude and openness to constructive feedback. What’s in it for you? We offer our employees more than just competitive compensation. Our team benefits include: Competitive pay and benefits Flexible vacation allowance A hybrid / remote working environment Startup culture backed by a secure, global brand Roster of Uniques We care deeply about every interaction our customers have with us, and trust and empower our staff to own and drive their experience. Our vision for our business and customers is built on fostering a diverse and inclusive work environment where regardless of background or beliefs you feel able to be authentic and bring all your talent into play. We want to celebrate you being you (we are an equal opportunity employer). All done! Your application has been successfully submitted! Other jobs You've already applied for this job We appreciate your interest in this position. Unfortunately, you have already applied for this job.

Vacancy posted 19 hours ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in Reno, NV vacancy
  •  ...and AI agent a cryptographically secured identity, improving engineering velocity while maintaining security. We make trusted computing...  ...problems that allow our customers to trust us for secure and reliable access to their infrastructure. Excellent security is table stakes... 
    Suggested
    Work at office
    Local area
    Remote work
    Sleeping nights

    GrabJobs

    Reno, NV
    2 days ago
  • $114k - $148k

     ...Site Reliability Engineer Location: Remote, United States Employment Type: Full-Time Benefits Offered: Vision, Medical, Life, Dental, 401K Gross Annual Base Salary: USD 114,000-148,000 Additional variable compensation and benefits may apply. Total compensation is based... 
    Suggested
    Full time
    Temporary work
    Work experience placement
    Remote work

    GrabJobs

    Reno, NV
    3 days ago
  • $153k - $210k

     ...building resilient, highly available cloud platforms that enable engineering teams to move quickly and confidently? Do you enjoy...  ...so, we invite you to be a part of our innovative team. As a Site Reliability Engineer, you’ll help ensure the reliability, scalability, and... 
    Suggested

    Ridgeline

    Reno, NV
    1 day ago
  • $160.8k - $214.1k

     ...observability needs of modern infrastructure. The Customer Reliability Engineering team is the deep technical escalation tier for Cisco Hypershield...  .../fix and reliability cases escalated by Cisco TAC, applying Site Reliability Engineering practices across the full stack: the... 
    Suggested
    Full time
    Temporary work
    Local area
    Remote work
    Flexible hours

    CISCO Systems

    Reno, NV
    3 days ago
  • $115k - $160k

     ...Hardware Reliability EngineerReno, Nevada, United States; San Francisco, California, United...  ...X-ray, etc.) and coordinate with senior engineers on complex investigations.Assist in building...  ...note: This role requires working on-site 5 days a week. We do not offer hybrid or... 
    Suggested
    Temporary work
    Local area
    Shift work

    Ampersand

    Reno, NV
    2 days ago
  • $192.3k - $248.1k

     ...unrelenting focus on our customers' success, we are Cisco's growth engine and shape the company’s future. Our values of Customer-Driven...  ...coverage, and basic life insurance. Please see the Cisco careers site to discover more benefits and perks. Employees may be eligible... 
    Full time
    Temporary work
    Local area
    Relocation
    Flexible hours

    CISCO Systems

    Reno, NV
    1 day ago
  • $108.5k - $149.18k

    Do you enjoy developing new products and services? Join us! Our Software Engineers work in an agile, collective environment. As a Software Engineer III, you will lead the design, development, and optimization of complex software systems for aerospace applications. You will... 
    Permanent employment
    Full time
    Work experience placement

    Sierra Nevada Corporation

    Sparks, NV
    1 day ago
  • $293.9k - $406.8k

     ...the TeamYou will join Cisco’s Identity Engineering Group, a foundational organization responsible...  ...domain, with a strong emphasis on reliability, interoperability, and long-term scalability...  ...insurance. Please see the Cisco careers site to discover more benefits and perks.... 
    Full time
    Temporary work
    Local area
    Remote work
    Flexible hours

    CISCO Systems

    Reno, NV
    2 days ago
  • $183.8k - $263.6k

     ...orchestration, and secure service integration. You will work closely with engineers across control plane, data plane, and platform teams to deliver...  ..., and basic life insurance. Please see the Cisco careers site to discover more benefits and perks. Employees may be eligible... 
    Full time
    Temporary work
    Local area
    Remote work
    Flexible hours

    CISCO Systems

    Reno, NV
    3 days ago
  • $101.38k - $168.96k

     ...is the world's leading company delivering sustainable design, engineering, and consultancy solutions for natural and built assets.We are...  ...Additionally, this position may involve travel of up to 30% to client sites and customer meetings, both within North America and... 
    Full time
    Part time
    For contractors

    Arcadis

    Reno, NV
    14 hours ago
  • $250.6k - $362.6k

     ...comprehensive security outcomes, as a Principal Engineer. The team delivers secure, scalable...  ...networking, with a strong emphasis on reliability, interoperability, and long-term...  ...insurance. Please see the Cisco careers site to discover more benefits and perks. Employees... 
    Full time
    Temporary work
    Local area
    Remote work
    Flexible hours

    CISCO Systems

    Reno, NV
    2 days ago
  • $111.07k - $148.1k

     ...demonstrated knowledge and experience in system architecture and engineering disciplines. •Recommends optimized solutions to support current...  ...capabilities. •Supports due diligence activities including site surveys, design, design review, bill of materials creation, statement... 
    Full time
    Temporary work
    Remote work
    1 day per week

    Lumen

    Sparks, NV
    2 days ago
  • $250.6k - $362.6k

     ...visibility and intelligence across diverse environments. As a Principal Engineer, you will be at the forefront of advancing Tetragon’s...  ...coverage, and basic life insurance. Please see the Cisco careers site to discover more benefits and perks. Employees may be eligible to... 
    Full time
    Temporary work
    Local area
    Flexible hours

    CISCO Systems

    Reno, NV
    3 days ago
  •  ...that drive operational excellence? Join Spectrum as a Systems Engineer II and help shape enterprise-wide Operations Support Systems applications...  ...scalable solutions that strengthen Spectrum’s performance and reliability. Your work will directly influence our technological... 
    Work at office
    Local area
    Visa sponsorship
    Shift work

    Spectrum Charter

    Sparks, NV
    2 days ago
  • $110k - $165k

     ...future of our communities. This is a Lead Cloud & Infrastructure Engineering position at the Vice President level, which is part of the job...  ...end user functions are delivered on a scalable, secure, and reliable infrastructure composed of seamlessly integrated datacenter,... 
    Temporary work
    Local area
    Remote work
    Flexible hours

    Morgan Stanley

    Reno, NV
    14 hours ago
  •  ...relocation to the area. We recruit nationally and provide financial relocation assistance. Responsibilities As a Technical Solutions Engineer at Epic, you’ll work on software that impacts 305 million patients around the world. Together with customer counterparts, you’ll... 
    Relocation
    Visa sponsorship
    Relocation package

    Epic

    Reno, NV
    19 hours ago
  • $110k - $125k

     ...comfortable and excited to come into the office once or twice a week. You have 2+ years of experience in a similar role (Solutions Engineer / Technical Sales Engineer / etc.) –ideally in the SaaS world. You are technically strong. You have at least a few technical... 
    Summer work
    Work at office
    Local area
    Remote work
    Home office

    GrabJobs

    Reno, NV
    19 hours ago
  • $105k - $140k

     ...Inclusivity. Today, nearly 200 people around the globe work on Speechify in a 100% distributed setting. These include frontend and backend engineers, AI research scientists, and others from Amazon, Microsoft, and Google, leading PhD programs like Stanford, high growth startups... 

    GrabJobs

    Reno, NV
    19 hours ago
  •  ...Akamai Technologies, Inc. is seeking a Solutions Engineer II in the NA Pre-Sales Team. You will be the customer’s trusted technical advisor, collaborating with sales to drive adoption and account growth while turning goals into scalable, secure architectures. You will... 

    Akamai

    Reno, NV
    1 day ago
  •  ...OpportunityWSP is seeking a Mine Waste Lead Civil/Geotechnical Engineer to join our team in Reno, NV. The following locations will also...  ...and heap leach facilities including initial layout and design, site selection and facility layout, geotechnical and hydrologic analyses... 
    For subcontractor
    Work at office
    Local area
    Flexible hours

    WSP Group

    Reno, NV
    2 days ago
  •  ...to provide the exceptional service our customers expect and contribute positively to our community. Job Summary The Software Engineer I is an essential member of the Software Development team, responsible for assisting in the design, development, and maintenance of... 
    Internship

    1 CLICK LOGISTICS

    Sparks, NV
    7 days ago
  •  ...to provide the exceptional service our customers expect and contribute positively to our community. Job Summary The Software Engineer II is a high impact role within the newly structured Software Development team, responsible for the architectural integrity and performance... 

    1 CLICK LOGISTICS

    Sparks, NV
    3 days ago
  • $93.09k - $124.12k

     ...deliver meaningful impact, and help shape the future of AI‑ready connectivity, join us today. The Role The Senior IT Systems Engineer provides advanced Tier II support by troubleshooting and repairing network devices, tools, and services for a nationwide fiber... 
    Full time
    Temporary work
    Work at office
    Remote work
    Shift work
    Night shift

    Lumen

    Sparks, NV
    19 hours ago
  • $77.2k - $96.5k

     ...Software Engineer I The software engineer I participates in the design, programming, testing, documentation, and implementation of computer applications and systems. Evaluates software packages, provides recommendations to management and business clients, and identifies... 
    Permanent employment
    Full time
    Work experience placement
    Internship
    H1b
    Local area

    BHE Renewables

    Reno, NV
    19 hours ago
  • $100k - $175k

     ...Software Engineer Reno, Nevada, United States; San Francisco, California, United States Company Overview Amperesand is reinventing...  ...team success. Please note: This role requires working on-site 5 days a week. We do not offer hybrid or remote options. SF... 
    Temporary work
    Local area
    Worldwide
    Shift work

    Ampersand

    Reno, NV
    3 days ago
  •  ...environment. Do you have what it takes? Find out today - About the Position ITS has an immediate opening for a Solutions Engineer. The Solutions Engineer will be responsible for collaborating with the sales team to understand client needs, designing and... 
    Work at office
    Immediate start
    Flexible hours
    Shift work

    ITS Logistics, LLC

    Reno, NV
    28 days ago
  •  ...Are you a backend engineer with a passion for clean, scalable systems and a deep appreciation for well-modeled financial data? Do you enjoy solving complex, data-rich problems and collaborating closely with product, design, and engineering peers to build industry-defining... 
    Full time

    Ridgeline

    Reno, NV
    19 hours ago
  •  ...Mobile Building Engineer Keep facilities running like clockwork. As a CBRE Mobile Engineer, you'll handle hands-on maintenance and...  ...the move and on point. You'll be the frontline expert ensuring reliability, safety, and performance wherever you're needed. What You'... 
    Work at office
    Shift work

    CBRE Group

    Reno, NV
    3 days ago
  • $150k - $190k

    About Reveal Technology Founded in 2019, Reveal is a dynamic startup revolutionizing field operations by delivering software tools and intelligence to operators in remote, disconnected, and extreme environments.Reveal is the Digital Arms Room for the Modern Warrior ...
    Remote work
    Home office

    GrabJobs

    Reno, NV
    2 days ago
  • Our Benefits - Designed with You in Mind Comprehensive Health & Well-being Coverage From your very first day, you’ll have access to medical, dental, vision, and prescription drug coverage - ensuring you and your family stay healthy and protected. Generous Paid Time...
    Full time
    Immediate start

    Stellantis

    Sparks, NV
    19 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!