Senior Site Reliability Engineer
GrabJobs
Job description What are we building? Hard Rock Digital is a team focused on becoming the best online sportsbook, casino, and social gaming company in the world. We’re building a team that resonates passion for learning, operating, and building new products and technologies for millions of consumers. We care about each customer interaction, experience, behavior, and insight and strive to ensure we’re always acting authentically. Rooted in the kindred spirits of Hard Rock and the Seminole Tribe of Florida, Hard Rock Digital taps a brand known the world over as the leader in gaming, entertainment, and hospitality. We’re taking that foundation of success and bringing it to the digital space - ready to join us? What’s the position? We are looking for a Senior Site Reliability Engineer who combines deep infrastructure expertise with a forward-thinking approach to AI-driven operations. In this role you will maintain and improve the reliability, scalability, and performance of our Java-based applications while pioneering the use of large language models (LLMs), agentic workflows, and intelligent automation to transform how we monitor, respond to, and prevent incidents. You will design and build autonomous and semi-autonomous AI agents that consume observability data, triage alerts, generate runbooks, automate incident response steps, and surface actionable insights—reducing toil and accelerating mean time to resolution. This is a hands-on engineering role for someone who is equally comfortable tuning a JVM, writing PromQL, and prototyping an agentic pipeline with tool-calling LLMs. Key Responsibilities Application Reliability & Performance Ensure the availability, reliability, and performance of high-traffic Java-based applications in a distributed environment. Troubleshoot and resolve complex issues across production and non-production environments. Participate in pre- and post-deployment performance testing and monitoring to continuously improve application performance. Optimize Java application performance with a focus on JVM tuning, efficient resource utilization, and horizontal scaling. Monitoring, Observability & AIOps Deploy and manage the Grafana stack (Grafana, Prometheus, Loki, Mimir, Alloy) to deliver real-time monitoring, logging, and alerting. Implement and refine observability strategies that enhance visibility into application and infrastructure health. Create and maintain dashboards, alerts, and log queries for comprehensive system health monitoring. Integrate AI/ML models into the observability pipeline for anomaly detection, predictive alerting, and intelligent alert correlation and noise reduction. AI & Agentic Workflow Engineering Design, build, and operate agentic AI workflows that automate operational tasks such as alert triage, root cause analysis, runbook execution, and incident summarization. Develop tool-calling LLM agents that interact with infrastructure APIs (Kubernetes, Grafana, Jira, Slack, PagerDuty) to execute diagnostic and remediation actions autonomously or with human-in-the-loop approval. Build and maintain MCP (Model Context Protocol) servers and integrations that expose internal systems as tool surfaces for AI agents. Evaluate, select, and operationalize LLM frameworks and orchestration platforms (e.g., LangChain, LangGraph, CrewAI, n8n, or custom solutions) for production-grade agentic systems. Implement guardrails, evaluation harnesses, and feedback loops to ensure AI agent outputs are accurate, safe, and continuously improving. Champion the adoption of AI-assisted development and operations practices across the SRE and broader engineering organization. Incident Management & Root Cause Analysis Support the operations team’s incident response efforts, conduct post-mortems, and identify root causes to prevent recurrence. Leverage AI tools to accelerate incident timelines, auto-generate post-mortem drafts, and surface patterns across historical incidents. Document and share lessons learned, contributing to a culture of continuous improvement. Automation & Toil Reduction Identify repetitive operational workflows and engineer AI-augmented or fully automated replacements. Build self-service tools and chatbot interfaces that allow engineering teams to query system status, retrieve logs, and execute standard operating procedures through natural language. Measure and report on toil reduction metrics to quantify the impact of automation initiatives. Collaboration & Cross-functional Support Work closely with developers, architects, and data/ML engineers to design solutions that improve reliability and leverage AI capabilities. Collaborate with DevOps and NOC teams to support the application platform. Communicate SRE practices, AI/automation capabilities, and operational insights to technical and non-technical stakeholders. Provide feedback on application performance, potential improvements, and observability metrics. Why This Role Is Different This is not a traditional SRE position with AI bolted on as an afterthought. We are building a team that treats AI and agentic automation as core competencies—on par with Kubernetes expertise or observability design. You will have the autonomy to experiment with cutting-edge AI tools, the backing of leadership to deploy them in production, and a mandate to measurably reduce operational toil through intelligent systems. Job requirements What are we looking for? Core SRE & Infrastructure (Required) Degree in Computer Science or a related field, or equivalent professional experience. 5+ years in SRE, DevOps, or similar infrastructure roles with experience managing large-scale, high-availability production systems. 3+ years hands-on experience managing production Kubernetes clusters, including deep understanding of architecture, networking, storage, and security. Experience with cluster autoscaling (Karpenter), upgrades, and multi-cluster management. Proficiency with kubectl, Helm, Kubernetes operators, and container orchestration troubleshooting. Advanced expertise with the Grafana observability stack: dashboards, alerting, visualization, and Grafana Alloy for telemetry collection. Proficiency in PromQL and experience with Loki for log aggregation and analysis. Hands-on experience managing Java-based applications in distributed environments, including JVM tuning and optimization. Cloud platform expertise (AWS preferred; GCP or Azure also valued). Familiarity with Infrastructure as Code tools such as Terraform/Terragrunt or Ansible. ArgoCD proficiency for GitOps workflows and continuous deployment. Strong scripting abilities in Python, Bash, or Go, with experience building CI/CD pipelines and deployment automation. Proven track record with on-call rotations, incident response, and root cause analysis. AI, Automation & Agentic Systems (Required) 1+ years of practical experience building or operating AI/LLM-powered tools, agents, or workflows in a production or production-adjacent context. Demonstrated ability to design agentic systems that use tool calling, retrieval-augmented generation (RAG), or multi-step reasoning to accomplish operational tasks. Experience integrating LLM APIs (e.g., Anthropic Claude, OpenAI, or open-source models) into backend services or automation pipelines. Familiarity with at least one agentic orchestration framework or workflow engine (LangChain, LangGraph, CrewAI, n8n, Temporal, or equivalent). Understanding of prompt engineering best practices, including structured outputs, system prompts, and few-shot examples. Familiarity with AI-assisted coding tools (Claude Code, Codex, Cursor) and their integration into engineering workflows. Experience building or consuming MCP (Model Context Protocol) servers to expose internal tools to AI agents. Awareness of AI safety, hallucination mitigation, and human-in-the-loop design patterns for autonomous systems. Preferred / Bonus Hands-on experience with vector databases (Pinecone, Weaviate, pgvector) for RAG-based knowledge retrieval. Experience with LLM evaluation frameworks (e.g., Galileo, LangSmith, Braintrust) for monitoring agent quality in production. Contributions to open-source AI/ML or SRE tooling projects. Background in data engineering or ML pipelines that complements SRE responsibilities. Soft Skills Strong communication skills (written and verbal) with the ability to translate complex AI and infrastructure concepts for diverse audiences. Proactive problem-solver with a bias toward automation and continuous improvement. Ability to mentor junior team members on both traditional SRE practices and emerging AI-driven approaches. Positive attitude and openness to constructive feedback. What’s in it for you? We offer our employees more than just competitive compensation. Our team benefits include: Competitive pay and benefits Flexible vacation allowance A hybrid / remote working environment Startup culture backed by a secure, global brand Roster of Uniques We care deeply about every interaction our customers have with us, and trust and empower our staff to own and drive their experience. Our vision for our business and customers is built on fostering a diverse and inclusive work environment where regardless of background or beliefs you feel able to be authentic and bring all your talent into play. We want to celebrate you being you (we are an equal opportunity employer). All done! Your application has been successfully submitted! Other jobs You've already applied for this job We appreciate your interest in this position. Unfortunately, you have already applied for this job.
$87.12k - $151.25k
...thinking organization, apply now.We are currently seeking a Digital Site Reliability Sr Engineer - Remote to join our team in Memphis, Tennessee (US-TN), United States (US).Digital Site Reliability Senior EngineerWe are seeking a Senior Site Reliability Engineer to drive...SeniorTemporary workWork at officeRemote workFlexible hours- ...ensure a bright future for NTT DATA Services and for the people who work here.NTT DATA Services currently seeks a DevOps Release Senior Engineer to join our team in Memphis, Tennessee (US-TN), United States (US).Roles and ResponsibilitiesForward Plan the release windows...SeniorFor contractors
- ...industry-leading security, user fund transparency, trading engine speed, deep liquidity, and an unmatched portfolio of digital... ...access for people around the world. We’re looking for a Senior Site Reliability Engineer Engineer to take ownership of building and evolving...SuggestedFull timeRemote workWork from home
$114k - $148k
...Site Reliability Engineer Location: Remote, United States Employment Type: Full-Time Benefits Offered: Vision, Medical, Life, Dental, 401K Gross Annual Base Salary: USD 114,000-148,000 Additional variable compensation and benefits may apply. Total compensation is based...SuggestedFull timeTemporary workWork experience placementRemote work$138k - $181.13k
...redefine the future of how work gets done.We are looking for a Senior Solution Engineer who is accustomed to solving customer’s most complex... ...States, please visit the job posting on the Snowflake Careers Site for salary and benefits information: careers.snowflake.comCompensation...SeniorRemote work$40 per hour
...revolutionizing the hospitality industry around the world! As a Senior Software Engineer, you will bring your technical skills to a hospitality... ...satisfy established performance, security, and reliability standards.What It Takes to Make the StayYou have these minimum...SeniorWork experience placementWork at officeWorldwideNight shift$125.84k - $238.16k
The Senior Staff Software Engineer manages assigned software systems through strategic direction, technical design, implementation, and operational health.The Principal Software Engineer anticipates, identifies, and provides solutions for complex problems through deep...SeniorFull timeLocal area- Senior Software EngineerLocation: Memphis, TNPosition OverviewWe are seeking a highly skilled Senior Software Engineer with a strong background in Full Stack Development and expertise in React.JS. The ideal candidate will be responsible for designing, developing, and maintaining...Senior
$115k - $145k
...work in a collaborative environment? As an experienced SAP PP-ME Senior Consultant, you will have the ability to share new ideas and... ...real-time manufacturing visibility.Partner with production, engineering, quality, and supply chain teams to define and maintain execution...SeniorWork at officeLocal areaVisa sponsorship- ...• Bachelor's degree (B.A.) in Computer Science, MIS or related degree and a minimum of five (5) years of relevant development or engineering experience or a combination of education, training and experience. • Financial Services experience preferred. • Experience in the...SeniorShift workNight shift
$116.2k - $229.1k
...of change and join us on this transformative journey.Recruiting for this role ends on 9/11/2026.Work you'll doAs a ServiceNow HRSD Senior Consultant you will be responsible for:Supporting ServiceNow HR Service Delivery implementations and optimization efforts across design...SeniorLocal area- ...deployments. The ideal candidate possesses excellent problem-solving skills, strong communication abilities, and a passion for building reliable, scalable, and high-performing enterprise applications.Requirements:8+ years of Java development experienceStrong hands-on...Senior
- ...Senior Software EngineerThe Senior Software Engineer is an experienced full stack developer with a proven track record of developing both front and back-end... ...of code metrics, system risk analysis, software reliability analysis, scalability analysis, performance analysis...SeniorTemporary workWork experience placementLocal area
$105.4k - $207.8k
Position Summary Our Deloitte AI & Engineering team works to transform technology... ...through innovation.Work you'll doAs a Senior Engineering Management Specialist - DevOps... ...security tools, and address issues affecting reliability and availabilityManage cloud services...SeniorLocal area- ...Senior DevOps Engineer We have an urgent requirement for a Senior DevOps Engineer for our client. This is a hands-on role. Strong hands-on experience in DevOps practices, CI/CD pipelines, cloud platforms, infrastructure automation, and monitoring tools is a MUST. The...SeniorLocal area
- ...environment? As an experiencedSAP Commerce Cloud Senior Developer, you will have the ability to... ...Manage and update client's B2B Commerce site, built on SAP Commerce Cloud with... ...scientists, operators, creatives, designers, engineers, and architects. Our team balances...SeniorVisa sponsorship
- ...Senior ServiceNow Developer Duration : Long-Term W2 Contract Openings: 4 Positions Location: 100% Remote (Must Work EST Hours) Work Authorization: USC or GCH Only We are seeking experienced Senior ServiceNow Developers to join a...SeniorLong term contractLocal areaRemote workFlexible hours
- ...Foley is seeking a Senior DevOps Engineer to help modernize and scale the infrastructure that powers... ...developer productivity, platform reliability, and security. What you'll do CI/CD Platform... ...are 5+ years of experience in DevOps, Site Reliability Engineering (SRE),...SeniorContract workRemote work
- Location: On site at location listed in job posting. The IT Developer Senior (Full Stack .Net Developer - Wealth) will develop program logic for new applications... ...requires a Bachelor's degree in Electronics Engineering a related field, or a foreign equivalent. Must have...SeniorFull time
- ...interview process will be initiated as soon as possible. We are excited to hear back from you. Job Description: Role: Senior Azure Cloud Engineer Salary: $$ plus Bonus potential (6%-10%) Location: 775 Ridge Lake Boulevard; Memphis, TN - 100% Onsite -...SeniorPermanent employmentFull timeLocal areaImmediate startRelocation
$105.4k - $207.8k
...Recruiting for this role ends on 12/31/2026. Work you'll do As a Senior Consultant, Strategy, Growth, and Transformation on the Cloud... ...in Computer Science, Cyber Security, Information Security, Engineering, or Information TechnologyExperience in a consulting roleMicrosoft...SeniorLocal areaVisa sponsorship- ...AutoZone is seeking an experienced Systems Engineer to join our expanding eCommerce B2B team, supporting AutoZonePro.com and our Mobile... ...practices. Collaborate with IT and business partners to deliver scalable, reliable solutions for commercial customers. #J-18808-Ljbffr...Senior
$110.7k - $218.3k
...shareholder value, and optimize operational efficiency. As an Oracle Senior Consultant at Deloitte, you will help clients define their... ...higher in Computer Science, Information Technology, Software Engineering, or a related field.Ability to travel 50%, on average, based...SeniorVisa sponsorship$115k - $150k
...an experienced Onshore Oracle Retail Cloud Functional Lead / Senior Consultant, you will have the ability to share new ideas and collaborate... ...in Computer Science, Information Technology, Computer Engineering, or related IT discipline; or equivalent experienceLimited...SeniorLocal areaRemote workVisa sponsorship- ...SUMMARY: The Armstrong Company is looking for a Senior NetSuite Developer who loves building on the platform. You've written... ...solutions. Help raise the bar through code review and sharing good engineering practices across the NetSuite environment. MINIMUM...SeniorWork at office
- ...Senior Web Developer Hybrid - Memphis, TN (Local candidates only) 6+-Month Contract with Long-Term Extension Must be GC/USC - Unfortunately no sponsorship available We're looking for a collaborative, forward-thinking Senior Web Developer to join a Memphis-based...SeniorContract workLocal area
$86.32k - $154.96k
...Research Computing (HPRC) and the Center for Bioimage Informatics (CBI) at St. Jude Children's Research Hospital are seeking a Senior AI Engineer to lead our efforts in advanced AI models, including large language models (LLMs), agentic AI systems, and multi-modal...SeniorFull time- noRole Responsibilities: Analyzes business requirements/processes and system integration considerations to determine appropriate technology solutions for internal and external customers. Designs, evaluates, codes, configures, tests and documents applications based on system...SeniorWork experience placement
$105.4k - $207.8k
...could be the place for you.We are looking for a hands-on Data Engineer to build and operate the governed data foundations powering cyber... ...Recruiting for this role ends on 12/31/2026.Work you'll doAs a Senior Consultant, Strategy, Growth and Transformation on the Cyber...SeniorLocal areaVisa sponsorship- ...department heads to align operational goals with company objectives. Monitor key performance indicators and prepare regular reports for senior management. Lead project management initiatives to drive continuous improvement. Foster a positive team environment and...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!
- senior lead project manager Memphis, TN
- senior robotics software engineer Memphis, TN
- senior devops engineer remote Memphis, TN
- senior sas administrator Memphis, TN
- sr project manager Memphis, TN
- senior windows systems engineer Memphis, TN
- consultant senior consultant Memphis, TN
- senior network engineer remote Memphis, TN
- senior account manager Memphis, TN
- senior accountant controller Memphis, TN


