Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Site Reliability Engineer

GrabJobs

Job description What are we building? Hard Rock Digital is a team focused on becoming the best online sportsbook, casino, and social gaming company in the world. We’re building a team that resonates passion for learning, operating, and building new products and technologies for millions of consumers. We care about each customer interaction, experience, behavior, and insight and strive to ensure we’re always acting authentically. Rooted in the kindred spirits of Hard Rock and the Seminole Tribe of Florida, Hard Rock Digital taps a brand known the world over as the leader in gaming, entertainment, and hospitality. We’re taking that foundation of success and bringing it to the digital space - ready to join us? What’s the position? We are looking for a Senior Site Reliability Engineer who combines deep infrastructure expertise with a forward-thinking approach to AI-driven operations. In this role you will maintain and improve the reliability, scalability, and performance of our Java-based applications while pioneering the use of large language models (LLMs), agentic workflows, and intelligent automation to transform how we monitor, respond to, and prevent incidents. You will design and build autonomous and semi-autonomous AI agents that consume observability data, triage alerts, generate runbooks, automate incident response steps, and surface actionable insights—reducing toil and accelerating mean time to resolution. This is a hands-on engineering role for someone who is equally comfortable tuning a JVM, writing PromQL, and prototyping an agentic pipeline with tool-calling LLMs. Key Responsibilities Application Reliability & Performance Ensure the availability, reliability, and performance of high-traffic Java-based applications in a distributed environment. Troubleshoot and resolve complex issues across production and non-production environments. Participate in pre- and post-deployment performance testing and monitoring to continuously improve application performance. Optimize Java application performance with a focus on JVM tuning, efficient resource utilization, and horizontal scaling. Monitoring, Observability & AIOps Deploy and manage the Grafana stack (Grafana, Prometheus, Loki, Mimir, Alloy) to deliver real-time monitoring, logging, and alerting. Implement and refine observability strategies that enhance visibility into application and infrastructure health. Create and maintain dashboards, alerts, and log queries for comprehensive system health monitoring. Integrate AI/ML models into the observability pipeline for anomaly detection, predictive alerting, and intelligent alert correlation and noise reduction. AI & Agentic Workflow Engineering Design, build, and operate agentic AI workflows that automate operational tasks such as alert triage, root cause analysis, runbook execution, and incident summarization. Develop tool-calling LLM agents that interact with infrastructure APIs (Kubernetes, Grafana, Jira, Slack, PagerDuty) to execute diagnostic and remediation actions autonomously or with human-in-the-loop approval. Build and maintain MCP (Model Context Protocol) servers and integrations that expose internal systems as tool surfaces for AI agents. Evaluate, select, and operationalize LLM frameworks and orchestration platforms (e.g., LangChain, LangGraph, CrewAI, n8n, or custom solutions) for production-grade agentic systems. Implement guardrails, evaluation harnesses, and feedback loops to ensure AI agent outputs are accurate, safe, and continuously improving. Champion the adoption of AI-assisted development and operations practices across the SRE and broader engineering organization. Incident Management & Root Cause Analysis Support the operations team’s incident response efforts, conduct post-mortems, and identify root causes to prevent recurrence. Leverage AI tools to accelerate incident timelines, auto-generate post-mortem drafts, and surface patterns across historical incidents. Document and share lessons learned, contributing to a culture of continuous improvement. Automation & Toil Reduction Identify repetitive operational workflows and engineer AI-augmented or fully automated replacements. Build self-service tools and chatbot interfaces that allow engineering teams to query system status, retrieve logs, and execute standard operating procedures through natural language. Measure and report on toil reduction metrics to quantify the impact of automation initiatives. Collaboration & Cross-functional Support Work closely with developers, architects, and data/ML engineers to design solutions that improve reliability and leverage AI capabilities. Collaborate with DevOps and NOC teams to support the application platform. Communicate SRE practices, AI/automation capabilities, and operational insights to technical and non-technical stakeholders. Provide feedback on application performance, potential improvements, and observability metrics. Why This Role Is Different This is not a traditional SRE position with AI bolted on as an afterthought. We are building a team that treats AI and agentic automation as core competencies—on par with Kubernetes expertise or observability design. You will have the autonomy to experiment with cutting-edge AI tools, the backing of leadership to deploy them in production, and a mandate to measurably reduce operational toil through intelligent systems. Job requirements What are we looking for? Core SRE & Infrastructure (Required) Degree in Computer Science or a related field, or equivalent professional experience. 5+ years in SRE, DevOps, or similar infrastructure roles with experience managing large-scale, high-availability production systems. 3+ years hands-on experience managing production Kubernetes clusters, including deep understanding of architecture, networking, storage, and security. Experience with cluster autoscaling (Karpenter), upgrades, and multi-cluster management. Proficiency with kubectl, Helm, Kubernetes operators, and container orchestration troubleshooting. Advanced expertise with the Grafana observability stack: dashboards, alerting, visualization, and Grafana Alloy for telemetry collection. Proficiency in PromQL and experience with Loki for log aggregation and analysis. Hands-on experience managing Java-based applications in distributed environments, including JVM tuning and optimization. Cloud platform expertise (AWS preferred; GCP or Azure also valued). Familiarity with Infrastructure as Code tools such as Terraform/Terragrunt or Ansible. ArgoCD proficiency for GitOps workflows and continuous deployment. Strong scripting abilities in Python, Bash, or Go, with experience building CI/CD pipelines and deployment automation. Proven track record with on-call rotations, incident response, and root cause analysis. AI, Automation & Agentic Systems (Required) 1+ years of practical experience building or operating AI/LLM-powered tools, agents, or workflows in a production or production-adjacent context. Demonstrated ability to design agentic systems that use tool calling, retrieval-augmented generation (RAG), or multi-step reasoning to accomplish operational tasks. Experience integrating LLM APIs (e.g., Anthropic Claude, OpenAI, or open-source models) into backend services or automation pipelines. Familiarity with at least one agentic orchestration framework or workflow engine (LangChain, LangGraph, CrewAI, n8n, Temporal, or equivalent). Understanding of prompt engineering best practices, including structured outputs, system prompts, and few-shot examples. Familiarity with AI-assisted coding tools (Claude Code, Codex, Cursor) and their integration into engineering workflows. Experience building or consuming MCP (Model Context Protocol) servers to expose internal tools to AI agents. Awareness of AI safety, hallucination mitigation, and human-in-the-loop design patterns for autonomous systems. Preferred / Bonus Hands-on experience with vector databases (Pinecone, Weaviate, pgvector) for RAG-based knowledge retrieval. Experience with LLM evaluation frameworks (e.g., Galileo, LangSmith, Braintrust) for monitoring agent quality in production. Contributions to open-source AI/ML or SRE tooling projects. Background in data engineering or ML pipelines that complements SRE responsibilities. Soft Skills Strong communication skills (written and verbal) with the ability to translate complex AI and infrastructure concepts for diverse audiences. Proactive problem-solver with a bias toward automation and continuous improvement. Ability to mentor junior team members on both traditional SRE practices and emerging AI-driven approaches. Positive attitude and openness to constructive feedback. What’s in it for you? We offer our employees more than just competitive compensation. Our team benefits include: Competitive pay and benefits Flexible vacation allowance A hybrid / remote working environment Startup culture backed by a secure, global brand Roster of Uniques We care deeply about every interaction our customers have with us, and trust and empower our staff to own and drive their experience. Our vision for our business and customers is built on fostering a diverse and inclusive work environment where regardless of background or beliefs you feel able to be authentic and bring all your talent into play. We want to celebrate you being you (we are an equal opportunity employer). All done! Your application has been successfully submitted! Other jobs You've already applied for this job We appreciate your interest in this position. Unfortunately, you have already applied for this job.

Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Senior Site Reliability Engineer in Orlando, FL vacancy
  • $175k - $215k

     ...ways to enhance these exciting experiences. Sr. Manager, Site Reliability Engineer provides strategic leadership across multiple SRE teams and...  ...functions, driving resilience and innovation. Influences senior internal and external stakeholders to secure funding, shape... 
    Senior

    Disney Experiences Careers

    Orlando, FL
    2 days ago
  • We’re seeking a future team member for the role of Sr. Site Reliability Automation Engineer to join our Technology team. This role is located in Lake Mary, FL and Pittsburgh, PA.In this role, you’ll make an impact in the following ways:•Design and implement end-to-end... 
    Senior
    Flexible hours

    The Bank of New York Mellon

    Lake Mary, FL
    5 days ago
  • $175.5k - $235.4k

     ...lines of business as well as other initiatives including MyDisneyExperience and Hey, Disney!This role sits in the Commerce Site Reliability Engineering (SRE) specifically supporting Ecommerce , Consumer Products and Licensing and Publishing organization within Technology... 
    Suggested
    Worldwide

    Disney Interactive

    Orlando, FL
    3 days ago
  • Company Overview By Light Professional IT Services LLC readies warfighters and federal agencies with technology and systems engineered to connect, protect, and prepare individuals and teams for whatever comes next. Headquartered in McLean, VA, By Light supports defense... 
    Suggested
    Contract work
    Temporary work
    Work experience placement
    Worldwide
    Flexible hours
    Shift work

    By Light Professional IT Services

    Orlando, FL
    4 days ago
  • $134.6k - $210.1k

    Position Summary The Senior Principal Engineer, DevOps plays an integral role in implementing and executing cloud practices for build management, product release and operation processes. The role is responsible for managing and automating the build and deployment process... 
    Senior
    Full time
    Temporary work
    Work experience placement
    Work at office
    Immediate start
    Flexible hours
    Night shift

    JetBlue Airways Corporation

    Orlando, FL
    2 days ago
  • $114k - $148k

     ...Site Reliability Engineer Location: Remote, United States Employment Type: Full-Time Benefits Offered: Vision, Medical, Life, Dental, 401K Gross Annual Base Salary: USD 114,000-148,000 Additional variable compensation and benefits may apply. Total compensation is based... 
    Full time
    Temporary work
    Work experience placement
    Remote work

    GrabJobs

    Orlando, FL
    4 days ago
  •  ...scale, we invite you to bring your talents to Zscaler to help shape the future of cybersecurity. Role We are looking for a Site Reliability Engineer-SkillBridge Intern (Virginia) to join our Zero Trust Exchange team. This is an onsite role based in Crystal City, Virginia... 
    Internship
    Work at office
    Local area
    Remote work
    Night shift

    GrabJobs

    Orlando, FL
    1 day ago
  • $109.4k - $146.7k

     ...Job Posting Title: Site Reliability Engineer Req ID: 10154082 Job Description: "We Power the Magic!" That's our motto at Disney Experiences (DX). Our team creates world-class immersive digital experiences for the Company's premier vacation brands including... 
    Full time
    Work experience placement
    Worldwide

    The Walt Disney Company

    Orlando, FL
    12 hours ago
  • $1,500 per month

     ...game worlds they inhabit. Our approach is centered around World Engine, our state-of-the-art onchain game server framework. World...  ...architecture to keep our platform secure. Own delivery, scalability, and reliability of our backend infrastructure. Advise and collaborate with the... 
    Full time
    Flexible hours

    GrabJobs

    Orlando, FL
    1 day ago
  • Job DescriptionHi,We have a new long term contract open for an h1 candidate near Orlando, FL! It is for a mid - senior level Groovy / Grails Developer. Please send me candidates with the following experience along with their references. Thank you!Only USC, GC and GC EAD... 
    Senior
    Long term contract

    USM Systems

    Orlando, FL
    5 days ago
  • $148.3k - $198.8k

     ...including MyDisneyExperience and Hey, Disney! The US Parks Site Reliability Organization is accountable for the reliability and...  ...critical applications and services. We partner with product, engineering, and SRE teams across the organization to define and evolve... 
    Work experience placement
    Worldwide
    Shift work

    The Walt Disney Company

    Orlando, FL
    5 days ago
  • $70 - $80 per hour

     ...Job Summary Our client, a leading entertainment and technology provider, is seeking a Lead Site Reliability Engineer to join their team! This position is a 96-week contract and can be based in Orlando, Florida; Burbank, California; or Seattle, Washington. Candidates... 
    Hourly pay
    Contract work
    Local area

    KellyMitchell Group

    Orlando, FL
    2 days ago
  • $96.8k - $161.5k

    OverviewA Senior Software Engineer provides advanced technical expertise in the design, development, integration, testing, deployment, and sustainment of software solutions supporting a complex systems and mission-focused program. This position is responsible for developing... 
    Senior
    Work at office
    Remote work

    AMERICAN SYSTEMS

    Orlando, FL
    4 days ago
  • $130k - $165k

     ...Job Description Job Description We are seeking a Mid-Senior Transmission Line Engineer to support the design and execution of high-voltage transmission line projects. This position will be responsible for engineering analysis, design development, and project delivery... 
    Senior
    Remote work

    FindTalent

    Orlando, FL
    6 days ago
  •  ...approach of blocking the exploits of application vulnerabilities. POSITION OVERVIEW We are seeking a  Windows  Kernel Driver Engineer with extensive experience in  filter driver development and  Windows system internals to join our cybersecurity product team. In... 
    Senior
    Full time
    Work at office

    Threatlocker

    Orlando, FL
    1 day ago
  •  ...innovation, ERP and CRM counselling, Product Engineering, Business Intelligence, Data Management,...  ...011). We are a project-driven firm that reliably meets the IT needs of our State and...  ...candidate near Orlando, FL! It is for a mid - senior level Groovy / Grails Developer. Please... 
    Senior
    Long term contract
    Worldwide

    USM Systems

    Orlando, FL
    6 hours ago
  •  ...mission readiness.The WorkLockheed Martin is looking for software developers who want to utilize the industry’s most powerful game engine to develop TLS’ Prepar3D Fuse product to save lives and prepare our customers for their toughest missions. Developers will work on... 
    Senior
    Part time
    Work at office
    Remote work
    Flexible hours

    Lockheed Martin

    Orlando, FL
    2 days ago
  • $86.8k - $198k

     ...inclusive of health benefits. We encourage you to learn more about our total benefits by visiting the Resource page on our Careers site and reviewing Our Employee Benefits page.Salary at Booz Allen is determined by various factors, including but not limited to location... 
    Senior
    Full time
    Contract work
    Part time
    Work at office
    Local area
    Remote work

    Booz Allen Hamilton

    Orlando, FL
    2 days ago
  •  ...Senior Software Engineer We are hiring a Senior Software Engineer to join the Workflow Engineering Team to develop applications and services that orchestrate and automate business process logic workflows for enterprise applications under the Cloud and Data Transformation... 
    Senior
    Local area
    Remote work

    RIT Solutions

    Orlando, FL
    4 days ago
  •  ...with technology and systems engineered to connect, protect, and prepare...  ...OverviewWe are seeking a Senior DevOps Engineer to lead the design...  ...for improving the speed, reliability, repeatability, and quality of...  ...engineers.This position is on-site.ResponsibilitiesLead the... 
    Senior
    Contract work
    Temporary work
    Worldwide

    By Light Professional IT Services

    Orlando, FL
    1 day ago
  •  ...software solutions. You will work internally with other software engineers, program managers, and product owners to deliver quality...  ...Attributes: ~ Innovative and creative thinking Travel: Occasional customer site visits Optional attendance at conferences... 
    Senior

    OneArc

    Orlando, FL
    1 day ago
  • $77k - $202k

     ...SectorNot ApplicableSpecialismSAPManagement LevelSenior AssociateJob Description & SummaryThe OpportunityAs a SAP BRIM Consultant, Senior Associate, you will play a pivotal role in helping clients optimize their operational efficiency by analyzing client needs, implementing... 
    Senior
    Full time
    H1b

    PwC

    Orlando, FL
    1 day ago
  • Company OverviewBy Light Professional IT Services LLC readies warfighters and federal agencies with technology and systems engineered to connect, protect, and prepare individuals and teams for whatever comes next. Headquartered in McLean, VA, By Light supports defense,... 
    Senior
    Contract work
    Temporary work
    Work experience placement
    Local area
    Worldwide

    By Light Professional IT Services

    Orlando, FL
    1 day ago
  • $120.38k - $204.66k

     .... Combining a diversity of talents, we master the decisive moments that matter to passengers and airlines. Whatever it takes.Senior AI Engineer Orlando, FL (Hybrid)Position SummaryThales is looking for a Senior AI Engineer to serve as a senior technical contributor within... 
    Senior
    Full time
    Work at office
    Local area
    Worldwide
    Monday to Friday

    Thales Group

    Orlando, FL
    4 days ago
  •  ...then consider a career in Advisory.KPMG is currently seeking a Senior Specialist to join our Federal Advisory practice.Responsibilities...  ...benefits can be found towards the bottom of our KPMG US Careers site at Benefits & How We Work.Follow this link to obtain salary ranges... 
    Senior
    Local area

    KPMG

    Orlando, FL
    5 days ago
  •  .... Job Description The Opportunity: Versant's Sports & Entertainment Digital Products division is seeking a Senior Site Reliability Engineer to help drive the reliability, scalability, and usability of internal developer platforms, tooling, and engineering workflows... 
    Senior
    Local area
    Remote work
    Worldwide

    Versant Media

    Orlando, FL
    2 days ago
  • $55k - $151.47k

     ...SummaryAt PwC, our people in data and analytics engineering focus on leveraging advanced technologies and...  ...AI agents and data pipelines, promoting reliability, scalability, and alignment with PwC standards. As a Senior Associate you will analyze complex problems, mentor... 
    Senior
    Full time
    Work experience placement
    H1b
    Remote work

    PwC

    Orlando, FL
    1 day ago
  •  ...their trusted partner. If you're ready to make an impact, you're in the right place. Job Details: Job Title: Senior Software Engineer - Front End Location: Orlando, FL (Hybrid) Duration: Fulltime Skills - React, NextJs, Python, Django Job... 
    Senior
    Full time

    Staffworxs Inc

    Orlando, FL
    1 day ago
  •  ...Senior Software Engineer At Disney, we're storytellers. We make the impossible, possible. The Walt Disney Company is a world-class entertainment and technological leader. Walt's passion was to continuously envision new ways to move audiences around the world—a passion... 
    Senior
    Work experience placement

    Walt Disney Company

    Orlando, FL
    4 days ago
  •  ...KPMG is currently seeking an AI Engineer to join our Audit Technology...  ...requirementsCollaborate with senior engineers and product...  ...patternsEnsure code quality and reliability by writing unit tests, participating...  ...of our KPMG US Careers site at Benefits & How We Work.Follow... 
    Senior
    H1b
    Local area

    KPMG

    Orlando, FL
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!