Remote Senior Site Reliability Engineer
GrabJobs
Job description What are we building? Hard Rock Digital is a team focused on becoming the best online sportsbook, casino, and social gaming company in the world. We’re building a team that resonates passion for learning, operating, and building new products and technologies for millions of consumers. We care about each customer interaction, experience, behavior, and insight and strive to ensure we’re always acting authentically. Rooted in the kindred spirits of Hard Rock and the Seminole Tribe of Florida, Hard Rock Digital taps a brand known the world over as the leader in gaming, entertainment, and hospitality. We’re taking that foundation of success and bringing it to the digital space - ready to join us? What’s the position? We are looking for a Senior Site Reliability Engineer who combines deep infrastructure expertise with a forward-thinking approach to AI-driven operations. In this role you will maintain and improve the reliability, scalability, and performance of our Java-based applications while pioneering the use of large language models (LLMs), agentic workflows, and intelligent automation to transform how we monitor, respond to, and prevent incidents. You will design and build autonomous and semi-autonomous AI agents that consume observability data, triage alerts, generate runbooks, automate incident response steps, and surface actionable insights—reducing toil and accelerating mean time to resolution. This is a hands-on engineering role for someone who is equally comfortable tuning a JVM, writing PromQL, and prototyping an agentic pipeline with tool-calling LLMs. Key Responsibilities Application Reliability & Performance Ensure the availability, reliability, and performance of high-traffic Java-based applications in a distributed environment. Troubleshoot and resolve complex issues across production and non-production environments. Participate in pre- and post-deployment performance testing and monitoring to continuously improve application performance. Optimize Java application performance with a focus on JVM tuning, efficient resource utilization, and horizontal scaling. Monitoring, Observability & AIOps Deploy and manage the Grafana stack (Grafana, Prometheus, Loki, Mimir, Alloy) to deliver real-time monitoring, logging, and alerting. Implement and refine observability strategies that enhance visibility into application and infrastructure health. Create and maintain dashboards, alerts, and log queries for comprehensive system health monitoring. Integrate AI/ML models into the observability pipeline for anomaly detection, predictive alerting, and intelligent alert correlation and noise reduction. AI & Agentic Workflow Engineering Design, build, and operate agentic AI workflows that automate operational tasks such as alert triage, root cause analysis, runbook execution, and incident summarization. Develop tool-calling LLM agents that interact with infrastructure APIs (Kubernetes, Grafana, Jira, Slack, PagerDuty) to execute diagnostic and remediation actions autonomously or with human-in-the-loop approval. Build and maintain MCP (Model Context Protocol) servers and integrations that expose internal systems as tool surfaces for AI agents. Evaluate, select, and operationalize LLM frameworks and orchestration platforms (e.g., LangChain, LangGraph, CrewAI, n8n, or custom solutions) for production-grade agentic systems. Implement guardrails, evaluation harnesses, and feedback loops to ensure AI agent outputs are accurate, safe, and continuously improving. Champion the adoption of AI-assisted development and operations practices across the SRE and broader engineering organization. Incident Management & Root Cause Analysis Support the operations team’s incident response efforts, conduct post-mortems, and identify root causes to prevent recurrence. Leverage AI tools to accelerate incident timelines, auto-generate post-mortem drafts, and surface patterns across historical incidents. Document and share lessons learned, contributing to a culture of continuous improvement. Automation & Toil Reduction Identify repetitive operational workflows and engineer AI-augmented or fully automated replacements. Build self-service tools and chatbot interfaces that allow engineering teams to query system status, retrieve logs, and execute standard operating procedures through natural language. Measure and report on toil reduction metrics to quantify the impact of automation initiatives. Collaboration & Cross-functional Support Work closely with developers, architects, and data/ML engineers to design solutions that improve reliability and leverage AI capabilities. Collaborate with DevOps and NOC teams to support the application platform. Communicate SRE practices, AI/automation capabilities, and operational insights to technical and non-technical stakeholders. Provide feedback on application performance, potential improvements, and observability metrics. Why This Role Is Different This is not a traditional SRE position with AI bolted on as an afterthought. We are building a team that treats AI and agentic automation as core competencies—on par with Kubernetes expertise or observability design. You will have the autonomy to experiment with cutting-edge AI tools, the backing of leadership to deploy them in production, and a mandate to measurably reduce operational toil through intelligent systems. Job requirements What are we looking for? Core SRE & Infrastructure (Required) Degree in Computer Science or a related field, or equivalent professional experience. 5+ years in SRE, DevOps, or similar infrastructure roles with experience managing large-scale, high-availability production systems. 3+ years hands-on experience managing production Kubernetes clusters, including deep understanding of architecture, networking, storage, and security. Experience with cluster autoscaling (Karpenter), upgrades, and multi-cluster management. Proficiency with kubectl, Helm, Kubernetes operators, and container orchestration troubleshooting. Advanced expertise with the Grafana observability stack: dashboards, alerting, visualization, and Grafana Alloy for telemetry collection. Proficiency in PromQL and experience with Loki for log aggregation and analysis. Hands-on experience managing Java-based applications in distributed environments, including JVM tuning and optimization. Cloud platform expertise (AWS preferred; GCP or Azure also valued). Familiarity with Infrastructure as Code tools such as Terraform/Terragrunt or Ansible. ArgoCD proficiency for GitOps workflows and continuous deployment. Strong scripting abilities in Python, Bash, or Go, with experience building CI/CD pipelines and deployment automation. Proven track record with on-call rotations, incident response, and root cause analysis. AI, Automation & Agentic Systems (Required) 1+ years of practical experience building or operating AI/LLM-powered tools, agents, or workflows in a production or production-adjacent context. Demonstrated ability to design agentic systems that use tool calling, retrieval-augmented generation (RAG), or multi-step reasoning to accomplish operational tasks. Experience integrating LLM APIs (e.g., Anthropic Claude, OpenAI, or open-source models) into backend services or automation pipelines. Familiarity with at least one agentic orchestration framework or workflow engine (LangChain, LangGraph, CrewAI, n8n, Temporal, or equivalent). Understanding of prompt engineering best practices, including structured outputs, system prompts, and few-shot examples. Familiarity with AI-assisted coding tools (Claude Code, Codex, Cursor) and their integration into engineering workflows. Experience building or consuming MCP (Model Context Protocol) servers to expose internal tools to AI agents. Awareness of AI safety, hallucination mitigation, and human-in-the-loop design patterns for autonomous systems. Preferred / Bonus Hands-on experience with vector databases (Pinecone, Weaviate, pgvector) for RAG-based knowledge retrieval. Experience with LLM evaluation frameworks (e.g., Galileo, LangSmith, Braintrust) for monitoring agent quality in production. Contributions to open-source AI/ML or SRE tooling projects. Background in data engineering or ML pipelines that complements SRE responsibilities. Soft Skills Strong communication skills (written and verbal) with the ability to translate complex AI and infrastructure concepts for diverse audiences. Proactive problem-solver with a bias toward automation and continuous improvement. Ability to mentor junior team members on both traditional SRE practices and emerging AI-driven approaches. Positive attitude and openness to constructive feedback. What’s in it for you? We offer our employees more than just competitive compensation. Our team benefits include: Competitive pay and benefits Flexible vacation allowance A hybrid / remote working environment Startup culture backed by a secure, global brand Roster of Uniques We care deeply about every interaction our customers have with us, and trust and empower our staff to own and drive their experience. Our vision for our business and customers is built on fostering a diverse and inclusive work environment where regardless of background or beliefs you feel able to be authentic and bring all your talent into play. We want to celebrate you being you (we are an equal opportunity employer). All done! Your application has been successfully submitted! Other jobs You've already applied for this job We appreciate your interest in this position. Unfortunately, you have already applied for this job.
$120k - $175k
...together? We are seeking a highly skilled and experienced Senior Site Reliability Engineer to join our team. We are passionate about delivering... ...applicants from anywhere in the U.S. and are willing to consider remote candidates. #LI-Remote Working at PrizePicks: The...Remote workSeniorFull timeWork visaFlexible hours$189k - $283.6k
...proactively and reactively improve the reliability of Block's platform and critical infrastructure... ...strong desire to perform and grow as an engineer * 5+ years of software development... ...career while building the life you want. Remote work, medical insurance, flexible time...Remote workSeniorFull timeLocal areaRelocation packageFlexible hoursShift work$118.6k - $195.68k
...the Job The Red Hat IT OpenShift team is looking for a Senior Site Reliability Engineer (SRE) to design, develop, scale, and operate our Red Hat... ...for bonus, commission, and/or equity. For positions with Remote-US locations, the actual salary range for the position may...Remote workSeniorPermanent employmentFull timeContract workWork experience placementWork at officeFlexible hours- ...Meghana GorusuCompany: SRI Tech SolutionsJob Title: Senior Site Reliability EngineerLocation: Plano , TX (remote)Years of Experience: 8 to 15 yearsSkillsKubernetes... ...seeking a highly skilled Senior Site Reliability Engineer (SRE) to join our dynamic team. The ideal...Remote workSenior
$65 - $75 per hour
DescriptionKforce has a client seeking a remote Senior Site Reliability Engineer to be a l be a leading member of the team working with a diverse range of technologies. You will enjoy working in a friendly environment and benefit from our investment in staff. The role also...Remote workSenior$104.9k - $174.7k
...the link below, About the Role:We are hiring a hands-on Senior Site Reliability Engineer (SRE) to actively build, operate, and improve the reliability... ...may work a hybrid schedule. If not, this role is fully remote. We do not restrict applicants based on job site or...Remote workSeniorFull timeWork at officeLocal areaWork from home$90k - $180k
...serve people in more than 160 countries.About the RoleThis Senior Site Reliability Engineer position works on-site out of our Sylmar, CA or Sunnyvale... ..., and operational excellence of Merlin.net — a remote monitoring platform designed to help doctors, cardiologists...Remote workSenior$15k
...beautiful modern office, daily catered lunches, and more.As a Senior Cluster Site Reliability Engineer (SRE), you will help scale our research compute cluster... ...Compensation Range: $205K - $235KLocationBerkeley, CA; Remote, United StatesEmployment TypeFull timeLocation...Remote workSeniorWork at officeLocal area$86.9k - $198k
Site Reliability Engineer, SeniorThe Opportunity: Engineering to make a system more resilient and efficient frees up time and money to build more... ...expected to have their cameras on during meetings.Remote: If this position is listed as remote, there may still be occasions...Remote workSeniorFull timeContract workPart timeWork at officeLocal area$150k - $180k
...operates through three business units: Remote Sensing (the data), Space Systems (the... ...are seeking an experienced SeniorSite Reliability Engineer to help design, build, operate, and scale... ...organization.This position is based on-site in either our Arlington, VA office, Reston...Remote workSeniorPermanent employmentFull timeWork at officeLocal areaWorldwide- ...A leading livestream shopping platform is seeking a Senior Software Engineer for the Logistics Platform team. This role focuses on improving logistical... ...operational debugging. The position offers flexibility for remote work and benefits including health insurance and generous...Remote workSenior
$117k - $209.33k
...6Position OverviewWant to help make a better world? As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable,... ...internally (not on this external site).SummaryLocation: Idaho, USA - Remote; AMER - United States - Texas - PlanoType: Full timeRemote workSeniorFull timeFor contractors- Site Reliability Engineers are responsible for ensuring the availability, reliability, scalability, and performance of the firm’s most critical customer... ....This is an on-site position located in Springfield, MO. Remote work is not an option for this position.Primary...Remote workSeniorLocal areaFlexible hoursShift work
$140k - $150k
WORK OPTION: Remote_________________The NBA is hiring a Senior Site Reliability Engineer (SRE) - Messaging & Collaboration to ensure the availability, performance, and reliability of enterprise messaging and collaboration platforms, including Microsoft Exchange Online (...Remote workSeniorFull timeTemporary workLocal areaWeekend work$127k - $249k
...on a hybrid basis, or it can be fully remote while working from a location based in... ...zones. We are looking for an experienced Senior Engineer for our SRE, Atlas team to support,... ...Role OverviewWe are seeking a talented Site Reliability Engineer (SRE) with a strong...Remote workSeniorLocal areaWorldwideFlexible hours$149.4k - $202k
...Noctua Technology is seeking a Senior Software Engineer specializing in Site Reliability Engineering to join their team. This role focuses on the reliability and... ...in software engineering. The position is primarily remote but requires candidates to be in the CA or DC Metro...Remote workSenior$130k - $180k
...belonging at iManage. Mondays and Fridays are reserved for (remote-friendly) focus time to get things done. Have the best of... ...belonging, collaboration, and accomplishment.Being a Senior Site Reliability Engineer at iManage Means… You are an engineer, a builder, and a systems...Remote workSeniorWork at officeLocal areaWorldwideMonday to FridayFlexible hours$127k - $249k
We are looking for an experienced Senior or Staff Engineer for our SRE, InfraSec team, to guide the security of our cloud-based infrastructure.... ...San Francisco offices on a hybrid basis, or it can be fully remote while working from a location based in either Eastern or...Remote workSeniorLocal areaWorldwideFlexible hours- ...software developers, platform engineers, and IT staff to improve... ...requirements, service quality, reliability, security, and compliance needs... ...Required: 8+ years of experience in Site Reliability Engineering,... ...or equivalent experience REMOTE WORK NOTICE: This position may...Remote workSeniorWork at office
$175k - $250k
...Senior Cloud Infrastructure Engineer Location: San Francisco, CA. Remote unavailable. Modality: On‑Site only. Must live within commuting distance of San Francisco or be willing to... ...ensuring scalability, performance, and reliability across environments. What You’ll Do Design...Remote workSeniorFull timeRelocationRelocation package- ...security, user fund transparency, trading engine speed, deep liquidity, and an... ...around the world. We’re looking for a Senior Site Reliability Engineer Engineer to take ownership of... ...and performance. This is a full-time remote role , with a preference for candidates...Remote workSeniorFull timeWork from home
- ...Cassandra, SQL Server, My SQL and Mongo DB Seniority level Seniority level Mid-Senior... ...in to set job alerts for “Senior Site Reliability Engineer” roles. Bellevue, WA $204,000.00-$259... ...ago Site Reliability Engineer (SRE, Remote US) Seattle, WA $120,000.00-$160,000....Remote workSeniorContract work
$165k - $195k
...employees a variety of ways to work, ranging from a fully remote experience to working full-time in one of our physical... ...or all of the time. About your role We're looking for a Senior Site Reliability Engineer II to help us scale our infrastructure and reliability practices...Remote workSeniorFull timeWork at officeLocal areaWork from homeFlexible hours- ...Senior Site Reliability Engineer (Enterprise Platform) Location: Remote - US - Open to Europe if happy to overlap with EST Compensation: Competitive We are a high-growth software company supporting the development of a premier open-source, EVM-compatible public ledger...Remote workSeniorContract workCurrently hiring
$149.4k - $202k
...Senior Software Engineer- Site Reliability Engineering (SRE) DC, MD, VA, CA The Site Reliability Engineering discipline at Noctua Technology, LLC is... ...applications and infrastructure. Location : Primarily Remote. Candidates must be based in CA or DC Metro Area for...Remote workSenior$141.8k - $195k
...massive, fast‑moving market. With a global workforce, we’re remote‑first and grounded in a simple idea: software is a people... ...herd. Why You’ll Love This Role Cribl Inc is seeking a Senior Site Reliability Engineer to join our mission where you will unlock the value of...Remote workSeniorTemporary work- ...Senior Site Reliability Engineer Company Overview: Arctiq is a global, intelligence-driven technology services company delivering professional... ...value to clients across diverse industries. This is a remote, contract opportunity for a project Arctiq is delivering...Remote workSeniorContract work
- ...education and literacy. About the Role We're looking for a Senior Site Reliability Engineer to drive the stability, observability, and reliability of... ...and workflows running reliably. This is a fully remote, US-based role working closely with a global engineering...Remote workSenior
$125.04k - $187.56k
...Digital and E-commerce, Technology and more. Overview The Site Reliability Engineer (SRE) III is responsible for ensuring the scalability, reliability... ...includes 3 in-person days at our Chicago office and 2 remote days. Responsibilities Design and implement infrastructure...Remote workSeniorFull timeWork at officeFlexible hours$160k - $240k
...way IT organizations work. We are currently looking for a Senior Site Reliability Engineer to join our SRE team in the Platform Engineering... ...availability of our services. Location - We are flexible on remote working from home, if you are located in the USA and reside...Remote workSeniorPermanent employmentFull timeWork from homeRelocationFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Remote Senior Site Reliability Engineer. Be the first to apply!
- remote clinical Bakersfield, CA
- remote auto claims adjuster Bakersfield, CA
- entry level project manager remote Bakersfield, CA
- remote epic analyst Bakersfield, CA
- remote legal Bakersfield, CA
- remote work from home Bakersfield, CA
- remote customer service agent Bakersfield, CA
- remote virtual Bakersfield, CA
- junior python remote Bakersfield, CA
- remote data entry part time Bakersfield, CA

