Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer

GrabJobs

Job description What are we building? Hard Rock Digital is a team focused on becoming the best online sportsbook, casino, and social gaming company in the world. We’re building a team that resonates passion for learning, operating, and building new products and technologies for millions of consumers. We care about each customer interaction, experience, behavior, and insight and strive to ensure we’re always acting authentically. Rooted in the kindred spirits of Hard Rock and the Seminole Tribe of Florida, Hard Rock Digital taps a brand known the world over as the leader in gaming, entertainment, and hospitality. We’re taking that foundation of success and bringing it to the digital space - ready to join us? What’s the position? We are looking for a Senior Site Reliability Engineer who combines deep infrastructure expertise with a forward-thinking approach to AI-driven operations. In this role you will maintain and improve the reliability, scalability, and performance of our Java-based applications while pioneering the use of large language models (LLMs), agentic workflows, and intelligent automation to transform how we monitor, respond to, and prevent incidents. You will design and build autonomous and semi-autonomous AI agents that consume observability data, triage alerts, generate runbooks, automate incident response steps, and surface actionable insights—reducing toil and accelerating mean time to resolution. This is a hands-on engineering role for someone who is equally comfortable tuning a JVM, writing PromQL, and prototyping an agentic pipeline with tool-calling LLMs. Key Responsibilities Application Reliability & Performance Ensure the availability, reliability, and performance of high-traffic Java-based applications in a distributed environment. Troubleshoot and resolve complex issues across production and non-production environments. Participate in pre- and post-deployment performance testing and monitoring to continuously improve application performance. Optimize Java application performance with a focus on JVM tuning, efficient resource utilization, and horizontal scaling. Monitoring, Observability & AIOps Deploy and manage the Grafana stack (Grafana, Prometheus, Loki, Mimir, Alloy) to deliver real-time monitoring, logging, and alerting. Implement and refine observability strategies that enhance visibility into application and infrastructure health. Create and maintain dashboards, alerts, and log queries for comprehensive system health monitoring. Integrate AI/ML models into the observability pipeline for anomaly detection, predictive alerting, and intelligent alert correlation and noise reduction. AI & Agentic Workflow Engineering Design, build, and operate agentic AI workflows that automate operational tasks such as alert triage, root cause analysis, runbook execution, and incident summarization. Develop tool-calling LLM agents that interact with infrastructure APIs (Kubernetes, Grafana, Jira, Slack, PagerDuty) to execute diagnostic and remediation actions autonomously or with human-in-the-loop approval. Build and maintain MCP (Model Context Protocol) servers and integrations that expose internal systems as tool surfaces for AI agents. Evaluate, select, and operationalize LLM frameworks and orchestration platforms (e.g., LangChain, LangGraph, CrewAI, n8n, or custom solutions) for production-grade agentic systems. Implement guardrails, evaluation harnesses, and feedback loops to ensure AI agent outputs are accurate, safe, and continuously improving. Champion the adoption of AI-assisted development and operations practices across the SRE and broader engineering organization. Incident Management & Root Cause Analysis Support the operations team’s incident response efforts, conduct post-mortems, and identify root causes to prevent recurrence. Leverage AI tools to accelerate incident timelines, auto-generate post-mortem drafts, and surface patterns across historical incidents. Document and share lessons learned, contributing to a culture of continuous improvement. Automation & Toil Reduction Identify repetitive operational workflows and engineer AI-augmented or fully automated replacements. Build self-service tools and chatbot interfaces that allow engineering teams to query system status, retrieve logs, and execute standard operating procedures through natural language. Measure and report on toil reduction metrics to quantify the impact of automation initiatives. Collaboration & Cross-functional Support Work closely with developers, architects, and data/ML engineers to design solutions that improve reliability and leverage AI capabilities. Collaborate with DevOps and NOC teams to support the application platform. Communicate SRE practices, AI/automation capabilities, and operational insights to technical and non-technical stakeholders. Provide feedback on application performance, potential improvements, and observability metrics. Why This Role Is Different This is not a traditional SRE position with AI bolted on as an afterthought. We are building a team that treats AI and agentic automation as core competencies—on par with Kubernetes expertise or observability design. You will have the autonomy to experiment with cutting-edge AI tools, the backing of leadership to deploy them in production, and a mandate to measurably reduce operational toil through intelligent systems. Job requirements What are we looking for? Core SRE & Infrastructure (Required) Degree in Computer Science or a related field, or equivalent professional experience. 5+ years in SRE, DevOps, or similar infrastructure roles with experience managing large-scale, high-availability production systems. 3+ years hands-on experience managing production Kubernetes clusters, including deep understanding of architecture, networking, storage, and security. Experience with cluster autoscaling (Karpenter), upgrades, and multi-cluster management. Proficiency with kubectl, Helm, Kubernetes operators, and container orchestration troubleshooting. Advanced expertise with the Grafana observability stack: dashboards, alerting, visualization, and Grafana Alloy for telemetry collection. Proficiency in PromQL and experience with Loki for log aggregation and analysis. Hands-on experience managing Java-based applications in distributed environments, including JVM tuning and optimization. Cloud platform expertise (AWS preferred; GCP or Azure also valued). Familiarity with Infrastructure as Code tools such as Terraform/Terragrunt or Ansible. ArgoCD proficiency for GitOps workflows and continuous deployment. Strong scripting abilities in Python, Bash, or Go, with experience building CI/CD pipelines and deployment automation. Proven track record with on-call rotations, incident response, and root cause analysis. AI, Automation & Agentic Systems (Required) 1+ years of practical experience building or operating AI/LLM-powered tools, agents, or workflows in a production or production-adjacent context. Demonstrated ability to design agentic systems that use tool calling, retrieval-augmented generation (RAG), or multi-step reasoning to accomplish operational tasks. Experience integrating LLM APIs (e.g., Anthropic Claude, OpenAI, or open-source models) into backend services or automation pipelines. Familiarity with at least one agentic orchestration framework or workflow engine (LangChain, LangGraph, CrewAI, n8n, Temporal, or equivalent). Understanding of prompt engineering best practices, including structured outputs, system prompts, and few-shot examples. Familiarity with AI-assisted coding tools (Claude Code, Codex, Cursor) and their integration into engineering workflows. Experience building or consuming MCP (Model Context Protocol) servers to expose internal tools to AI agents. Awareness of AI safety, hallucination mitigation, and human-in-the-loop design patterns for autonomous systems. Preferred / Bonus Hands-on experience with vector databases (Pinecone, Weaviate, pgvector) for RAG-based knowledge retrieval. Experience with LLM evaluation frameworks (e.g., Galileo, LangSmith, Braintrust) for monitoring agent quality in production. Contributions to open-source AI/ML or SRE tooling projects. Background in data engineering or ML pipelines that complements SRE responsibilities. Soft Skills Strong communication skills (written and verbal) with the ability to translate complex AI and infrastructure concepts for diverse audiences. Proactive problem-solver with a bias toward automation and continuous improvement. Ability to mentor junior team members on both traditional SRE practices and emerging AI-driven approaches. Positive attitude and openness to constructive feedback. What’s in it for you? We offer our employees more than just competitive compensation. Our team benefits include: Competitive pay and benefits Flexible vacation allowance A hybrid / remote working environment Startup culture backed by a secure, global brand Roster of Uniques We care deeply about every interaction our customers have with us, and trust and empower our staff to own and drive their experience. Our vision for our business and customers is built on fostering a diverse and inclusive work environment where regardless of background or beliefs you feel able to be authentic and bring all your talent into play. We want to celebrate you being you (we are an equal opportunity employer). All done! Your application has been successfully submitted! Other jobs You've already applied for this job We appreciate your interest in this position. Unfortunately, you have already applied for this job.

Vacancy posted 11 hours ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in Los Angeles, CA vacancy
  • $125k - $150k

     ...SpaceX is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SITE RELIABILITY ENGINEER (RAPTOR)SpaceX is looking for a Site Reliability Engineer with a strong drive to solve challenging problems in the Raptor... 
    Suggested
    Permanent employment
    Temporary work

    SpaceX

    Hawthorne, CA
    2 days ago
  • $125k - $145k

     ...SpaceX is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SITE RELIABILITY ENGINEER, GNCSpaceX’s mission is to make humanity multiplanetary by developing fully and rapidly reusable launch systems capable of... 
    Suggested
    Permanent employment
    Temporary work
    Flexible hours
    Weekend work

    SpaceX

    Hawthorne, CA
    4 days ago
  • $165k - $265k

     ...is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SR. SITE RELIABILITY ENGINEER (STARLINK)At SpaceX we’re leveraging our experience in building rockets and spacecraft to deploy Starlink, the world’s most... 
    Suggested
    Permanent employment
    Temporary work
    Worldwide
    Weekend work

    SpaceX

    Hawthorne, CA
    2 days ago
  • $165k - $230k

     ...is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SR. SITE RELIABILITY ENGINEER (STARSHIELD)Starshield leverages SpaceX’s Starlink technology and launch capability to support national security efforts.... 
    Suggested
    Permanent employment
    Temporary work
    Immediate start
    Weekend work

    SpaceX

    Hawthorne, CA
    4 days ago
  •  ...your big ideas, and your desire to team up with some of the best and brightest in technology and entertainment. The RoleThe Site Reliability Engineer (SRE) II is responsible for designing, implementing, and maintaining scalable and reliable systems and applications. Focus... 
    Suggested
    Full time
    Local area
    Worldwide
    Flexible hours

    AXS Group

    Los Angeles, CA
    2 days ago
  • $125k - $145k

     ...SpaceX is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SITE RELIABILITY ENGINEER - TOP SECRET CLEARANCEAs a Site Reliability Engineer, you will design, develop, and test key aspects of an in-house... 
    Permanent employment
    Temporary work
    Weekend work

    SpaceX

    Hawthorne, CA
    2 days ago
  • $165k - $265k

     ...is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SR. SITE RELIABILITY ENGINEER - TOP SECRET CLEARANCE (STARLINK)At SpaceX we’re leveraging our experience in building rockets and spacecraft to deploy Starlink... 
    Permanent employment
    Temporary work
    Worldwide
    Weekend work

    SpaceX

    Hawthorne, CA
    1 day ago
  • $145k - $175k

     ...SpaceX is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SITE RELIABILITY ENGINEER (TOP SECRET CLEARANCE)As a member of the Classified IT Systems Engineering team, the Site Reliability Engineer is involved... 
    Permanent employment
    Temporary work
    Weekend work

    SpaceX

    Hawthorne, CA
    2 days ago
  •  ...SRE Support Engineer While this position is not currently open, we are interviewing strong candidates for upcoming opportunities on this team. Location: Remote | Time Zone: (iNDIA)(8AM–5PM IST) Domain: Compute(Linux Fundamentals, Linux Networking, Kubernetes, Docker)... 
    Remote work

    GrabJobs

    Glendale, CA
    2 days ago
  •  ...Senior Site Reliability Engineer (SRE) Our client is a global technology consulting and digital solutions company that enables enterprises across industries to reimagine business models, accelerate innovation, and maximize growth by harnessing digital technologies.... 
    Local area

    E-Solutions

    Los Angeles, CA
    1 day ago
  •  ...A leading livestream shopping platform is seeking a Senior Software Engineer for the Logistics Platform team. This role focuses on improving logistical data systems, enhancing buyer and seller experience, and fostering collaboration across departments. Ideal candidates... 
    Remote work

    Whatnot

    Los Angeles, CA
    2 days ago
  • $180.5k - $236.91k

     ...Hi, we're Oscar. We're hiring a Senior Software Engineer, Cloud Infrastructure / SRE to join our Engineering team. Oscar is the...  ...on your team's business and technical domains such as DevOps, site reliability, and cloud best practices Lead the planning, execution and... 
    Full time
    Work at office
    Remote work

    GrabJobs

    Glendale, CA
    11 hours ago
  •  ...What you will do: Partner with a team of high-performing engineers and developers who are focused on delivering best in class software...  ...our shift to a SecDevOps culture, solving for security, reliability, cost-effectiveness, and observability Building Zero trust... 
    Full time
    Contract work
    Local area
    Flexible hours
    Shift work

    DISQO

    Los Angeles, CA
    a month ago
  • $140k - $180k

     ...fundamentally different class of spacecraft. Engineered to survive the harshest radiation...  ...create highly available, deployable, and reliable products Reduce operational toil through...  ...experience in Software Engineering, Site Reliability Engineering or DevOps ~ Deep... 
    Permanent employment
    Shift work

    K2 Space

    Los Angeles, CA
    22 days ago
  • $32 - $35 per hour

     ...and assignment.) Key Responsibilities: In this role, you will help ensure the reliability, performance, and stability of key restaurant-facing platforms by working closely with engineering and infrastructure teams. You will use observability tools such as DataDog, Grafana... 
    Contract work
    Local area
    Immediate start

    Pyramid Consulting

    Los Angeles, CA
    11 hours ago
  • $30.53 - $56.48 per hour

    Job Title:Associate Site Reliability EngineerRequisition ID:R027696Job Description:Job Title: Associate Site Reliability EngineerReporting...  ...TechnologyLocation: Santa Monica, CaOverviewThe Associate Site Reliability Engineer helps keep Marketing Technology services reliable, observable... 
    Hourly pay
    Full time
    Temporary work
    Part time
    Internship
    Local area
    Worldwide
    Relocation package

    Activision

    Santa Monica, CA
    4 days ago
  • $164k - $270k

     ...for the 21st century and beyond.The Role What You’ll DoOwn the reliability of our robotics systems, from PLCs through ROS2/middleware to...  ...remediation.Partner with controls, robotics, and platform engineering teams to bake reliability in early. Review designs, develop SLOs... 
    Permanent employment
    Full time
    Local area
    Flexible hours

    Hadrian

    Los Angeles, CA
    1 day ago
  • $164k - $270k

    Hadrian - Manufacturing the FutureHadrian is building autonomous factories that help aerospace and defense companies manufacture rockets, satellites, jets, and ships up to 10x faster and up to 2x cheaper. By combining advanced software, robotics, and full-stack manufacturing...
    Permanent employment
    Full time
    Local area
    Remote work
    Flexible hours

    Hadrian

    Los Angeles, CA
    2 days ago
  •  ...Role: Site Reliability Engineering (SRE) Location: Los Angeles, CA Remote position Fulltime position JD Site Reliability Engineer Experience in Cloud platforms (AWS, Azure, Google Cloud) and hybrid environments. Proficiency... 
    Full time
    Remote work

    SARIAN Co

    Los Angeles, CA
    11 hours ago
  •  ...scale, we invite you to bring your talents to Zscaler to help shape the future of cybersecurity. Role We are looking for a Site Reliability Engineer-SkillBridge Intern (San JosA Ca or Bellevue WA) to join our Zero Trust Exchange team. This is a remote role based in San... 
    Internship
    Work at office
    Local area
    Remote work
    Worldwide

    GrabJobs

    Los Angeles, CA
    11 hours ago
  • $114k - $148k

     ...Site Reliability Engineer Location: Remote, United States Employment Type: Full-Time Benefits Offered: Vision, Medical, Life, Dental, 401K Gross Annual Base Salary: USD 114,000-148,000 Additional variable compensation and benefits may apply. Total compensation is based... 
    Full time
    Temporary work
    Work experience placement
    Remote work

    GrabJobs

    Los Angeles, CA
    4 days ago
  • $141k - $208k

     ...be a part of our journey! About the role We are committed to providing our customers with reliable and secure services so we are expanding our central Site Reliability Engineering team. You will be responsible for building and leading processes to ensure the reliability... 
    Local area
    Remote work
    Home office
    Flexible hours

    GrabJobs

    Glendale, CA
    1 day ago
  •  ..., and thrive! KēSTA I.T. is actively seeking a Principal Engineer for an immediate full-time opportunity with our industry creating...  ...An innovative technology company is seeking experienced Site Reliability Engineers to take ownership of building reliable, scalable platforms... 
    Permanent employment
    Full time
    Temporary work
    Immediate start

    KēSTA I.T.

    Beverly Hills, CA
    17 days ago
  • $100k - $200k

     ...DevOps / Site Reliability Engineer Los Angeles, CA General Matter is enriching uranium in America. Our mission is to restore our country's ability to make nuclear fuel. Our fuel will help power AI, manufacturing, and other critical industries. It will power our next... 
    Full time
    Weekend work

    General Matter

    Los Angeles, CA
    1 day ago
  •  ...Principal Site Reliability Engineer Join the Commerce Site Reliability Engineering team as a Principal SRE, supporting Ecommerce, Consumer Products, Licensing and Publishing, and Consumer Production platforms at Disney Experiences. You will drive best-in-class infrastructure... 
    Work experience placement

    Disney Experiences

    Glendale, CA
    3 days ago
  • $197k - $291k

     ...troubleshooting distributed systems. Preferred qualifications Master's degree in Computer Science or Engineering. 1 year of people management experience. About The Job Site Reliability Engineering (SRE) combines software and systems engineering to build and run large-scale,... 
    Full time

    Google

    Los Angeles, CA
    2 days ago
  • $100k - $117.5k

     ...with the ultimate goal of enabling human life on Mars.BUILD RELIABILITY ENGINEER (FALCON/STARSHIP VALVES AND COMPONENTS) The Build Reliability...  ...in quarantining and rework activities across multiple sites BASIC QUALIFICATIONS:Bachelor's degree in mechanical engineering... 
    Permanent employment
    Temporary work
    Internship
    Work at office
    Immediate start
    Weekend work

    SpaceX

    Hawthorne, CA
    4 days ago
  •  ...Job Description Job Description Forhyre is looking for engineers who can bring unique perspectives and innovative ideas to all areas...  ...evangelize cloud best practices while building a culture of reliability and observability Engage in and improve the end to end lifecycle... 

    Forhyre

    Los Angeles, CA
    26 days ago
  • $120k - $180k

     ...organizations that test and validate complex systems—think drones, rocket engines, satellites, and nuclear reactors. Supported by leading...  ...to roll : Frequently traveling to spend time with end-users on-site (e.g. rocket test stands, spacecraft clean rooms, automated... 
    Full time
    Temporary work
    Work experience placement

    Nominal

    Los Angeles, CA
    11 hours ago
  • $120k - $145k

     ...the ultimate goal of enabling human life on Mars. SOFTWARE ENGINEER, SATELLITE SYSTEMS (STARSHIELD) Starshield leverages SpaceX’s...  ...hosted payloads. The Starshield software team is building highly reliable in-space mesh networks, designing secure systems to guarantee... 
    Permanent employment
    Full time
    Temporary work
    Weekend work

    Spacex

    Hawthorne, CA
    11 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!