Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Site Reliability Engineer

GrabJobs

Job description What are we building? Hard Rock Digital is a team focused on becoming the best online sportsbook, casino, and social gaming company in the world. We’re building a team that resonates passion for learning, operating, and building new products and technologies for millions of consumers. We care about each customer interaction, experience, behavior, and insight and strive to ensure we’re always acting authentically. Rooted in the kindred spirits of Hard Rock and the Seminole Tribe of Florida, Hard Rock Digital taps a brand known the world over as the leader in gaming, entertainment, and hospitality. We’re taking that foundation of success and bringing it to the digital space - ready to join us? What’s the position? We are looking for a Senior Site Reliability Engineer who combines deep infrastructure expertise with a forward-thinking approach to AI-driven operations. In this role you will maintain and improve the reliability, scalability, and performance of our Java-based applications while pioneering the use of large language models (LLMs), agentic workflows, and intelligent automation to transform how we monitor, respond to, and prevent incidents. You will design and build autonomous and semi-autonomous AI agents that consume observability data, triage alerts, generate runbooks, automate incident response steps, and surface actionable insights—reducing toil and accelerating mean time to resolution. This is a hands-on engineering role for someone who is equally comfortable tuning a JVM, writing PromQL, and prototyping an agentic pipeline with tool-calling LLMs. Key Responsibilities Application Reliability & Performance Ensure the availability, reliability, and performance of high-traffic Java-based applications in a distributed environment. Troubleshoot and resolve complex issues across production and non-production environments. Participate in pre- and post-deployment performance testing and monitoring to continuously improve application performance. Optimize Java application performance with a focus on JVM tuning, efficient resource utilization, and horizontal scaling. Monitoring, Observability & AIOps Deploy and manage the Grafana stack (Grafana, Prometheus, Loki, Mimir, Alloy) to deliver real-time monitoring, logging, and alerting. Implement and refine observability strategies that enhance visibility into application and infrastructure health. Create and maintain dashboards, alerts, and log queries for comprehensive system health monitoring. Integrate AI/ML models into the observability pipeline for anomaly detection, predictive alerting, and intelligent alert correlation and noise reduction. AI & Agentic Workflow Engineering Design, build, and operate agentic AI workflows that automate operational tasks such as alert triage, root cause analysis, runbook execution, and incident summarization. Develop tool-calling LLM agents that interact with infrastructure APIs (Kubernetes, Grafana, Jira, Slack, PagerDuty) to execute diagnostic and remediation actions autonomously or with human-in-the-loop approval. Build and maintain MCP (Model Context Protocol) servers and integrations that expose internal systems as tool surfaces for AI agents. Evaluate, select, and operationalize LLM frameworks and orchestration platforms (e.g., LangChain, LangGraph, CrewAI, n8n, or custom solutions) for production-grade agentic systems. Implement guardrails, evaluation harnesses, and feedback loops to ensure AI agent outputs are accurate, safe, and continuously improving. Champion the adoption of AI-assisted development and operations practices across the SRE and broader engineering organization. Incident Management & Root Cause Analysis Support the operations team’s incident response efforts, conduct post-mortems, and identify root causes to prevent recurrence. Leverage AI tools to accelerate incident timelines, auto-generate post-mortem drafts, and surface patterns across historical incidents. Document and share lessons learned, contributing to a culture of continuous improvement. Automation & Toil Reduction Identify repetitive operational workflows and engineer AI-augmented or fully automated replacements. Build self-service tools and chatbot interfaces that allow engineering teams to query system status, retrieve logs, and execute standard operating procedures through natural language. Measure and report on toil reduction metrics to quantify the impact of automation initiatives. Collaboration & Cross-functional Support Work closely with developers, architects, and data/ML engineers to design solutions that improve reliability and leverage AI capabilities. Collaborate with DevOps and NOC teams to support the application platform. Communicate SRE practices, AI/automation capabilities, and operational insights to technical and non-technical stakeholders. Provide feedback on application performance, potential improvements, and observability metrics. Why This Role Is Different This is not a traditional SRE position with AI bolted on as an afterthought. We are building a team that treats AI and agentic automation as core competencies—on par with Kubernetes expertise or observability design. You will have the autonomy to experiment with cutting-edge AI tools, the backing of leadership to deploy them in production, and a mandate to measurably reduce operational toil through intelligent systems. Job requirements What are we looking for? Core SRE & Infrastructure (Required) Degree in Computer Science or a related field, or equivalent professional experience. 5+ years in SRE, DevOps, or similar infrastructure roles with experience managing large-scale, high-availability production systems. 3+ years hands-on experience managing production Kubernetes clusters, including deep understanding of architecture, networking, storage, and security. Experience with cluster autoscaling (Karpenter), upgrades, and multi-cluster management. Proficiency with kubectl, Helm, Kubernetes operators, and container orchestration troubleshooting. Advanced expertise with the Grafana observability stack: dashboards, alerting, visualization, and Grafana Alloy for telemetry collection. Proficiency in PromQL and experience with Loki for log aggregation and analysis. Hands-on experience managing Java-based applications in distributed environments, including JVM tuning and optimization. Cloud platform expertise (AWS preferred; GCP or Azure also valued). Familiarity with Infrastructure as Code tools such as Terraform/Terragrunt or Ansible. ArgoCD proficiency for GitOps workflows and continuous deployment. Strong scripting abilities in Python, Bash, or Go, with experience building CI/CD pipelines and deployment automation. Proven track record with on-call rotations, incident response, and root cause analysis. AI, Automation & Agentic Systems (Required) 1+ years of practical experience building or operating AI/LLM-powered tools, agents, or workflows in a production or production-adjacent context. Demonstrated ability to design agentic systems that use tool calling, retrieval-augmented generation (RAG), or multi-step reasoning to accomplish operational tasks. Experience integrating LLM APIs (e.g., Anthropic Claude, OpenAI, or open-source models) into backend services or automation pipelines. Familiarity with at least one agentic orchestration framework or workflow engine (LangChain, LangGraph, CrewAI, n8n, Temporal, or equivalent). Understanding of prompt engineering best practices, including structured outputs, system prompts, and few-shot examples. Familiarity with AI-assisted coding tools (Claude Code, Codex, Cursor) and their integration into engineering workflows. Experience building or consuming MCP (Model Context Protocol) servers to expose internal tools to AI agents. Awareness of AI safety, hallucination mitigation, and human-in-the-loop design patterns for autonomous systems. Preferred / Bonus Hands-on experience with vector databases (Pinecone, Weaviate, pgvector) for RAG-based knowledge retrieval. Experience with LLM evaluation frameworks (e.g., Galileo, LangSmith, Braintrust) for monitoring agent quality in production. Contributions to open-source AI/ML or SRE tooling projects. Background in data engineering or ML pipelines that complements SRE responsibilities. Soft Skills Strong communication skills (written and verbal) with the ability to translate complex AI and infrastructure concepts for diverse audiences. Proactive problem-solver with a bias toward automation and continuous improvement. Ability to mentor junior team members on both traditional SRE practices and emerging AI-driven approaches. Positive attitude and openness to constructive feedback. What’s in it for you? We offer our employees more than just competitive compensation. Our team benefits include: Competitive pay and benefits Flexible vacation allowance A hybrid / remote working environment Startup culture backed by a secure, global brand Roster of Uniques We care deeply about every interaction our customers have with us, and trust and empower our staff to own and drive their experience. Our vision for our business and customers is built on fostering a diverse and inclusive work environment where regardless of background or beliefs you feel able to be authentic and bring all your talent into play. We want to celebrate you being you (we are an equal opportunity employer). All done! Your application has been successfully submitted! Other jobs You've already applied for this job We appreciate your interest in this position. Unfortunately, you have already applied for this job.

Vacancy posted 19 hours ago
Similar jobs that could be interesting for youBased on the Senior Site Reliability Engineer in Newark, NJ vacancy
  • $175k - $185k

     ...together. Come join our team as we develop new ways to improve the lives of working Americans. About the role: This Senior Database Reliability Engineer is responsible for supporting all production database systems so that they run smoothly. We run MySQL 8.4 on Google... 
    Senior
    Daily paid
    Remote work
    Home office
    Flexible hours

    GrabJobs

    Newark, NJ
    3 days ago
  •  ...foundation of success and bringing it to the digital space - ready to join us? What’s the position? We are looking for a Senior Site Reliability Engineer who combines deep infrastructure expertise with a forward-thinking approach to AI-driven operations. In this role you... 
    Senior
    Remote work
    Flexible hours
    Night shift

    GrabJobs

    Newark, NJ
    3 days ago
  • $210k - $220k

     ...secure and private by design, it’s popular with security, IT, engineering, finance, and other security-focused teams. At Tines, we'...  ..., and we’re looking for others to join us on our journey. Senior Site Reliability Engineer - Government Cloud You'll join the team... 
    Senior
    Work at office
    Remote work

    GrabJobs

    Newark, NJ
    2 days ago
  • $140k - $150k

    WORK OPTION: Remote_________________The NBA is hiring a Senior Site Reliability Engineer (SRE) - Messaging & Collaboration to ensure the availability, performance, and reliability of enterprise messaging and collaboration platforms, including Microsoft Exchange Online... 
    Senior
    Full time
    Temporary work
    Local area
    Remote work
    Weekend work

    National Basketball Association

    Secaucus, NJ
    3 days ago
  • Elevate your engineering prowess to unprecedented levels by joining a team of exceptionally gifted professionals and position yourself among the top echelon in site reliability.As a Sr Lead Site Reliability Engineer at JPMorgan Chase within the Consumer & Community Banking... 
    Senior

    JP Morgan Chase

    Jersey City, NJ
    4 days ago
  •  ...PVH (Tommy Hilfiger/Calvin Klein) seeks a Senior Software Engineer to own the reliability and performance of our Kubernetes-based data platform across multi-region deployments. You will design scalable infrastructure, optimize deployment pipelines, and strengthen security... 
    Senior

    PVH (Tommy Hilfiger/Calvin Klein)

    Livingston, NJ
    2 days ago
  •  ...We are seeking an experienced Site Reliability Engineer (SRE) – Microsoft Hyper-V & Private Cloud to operate highly available private cloud and Virtual Desktop Infrastructure (VDI) platforms based on Microsoft Hyper-V. This role combines deep Hyper-V expertise with modern... 
    Senior
    Local area

    2T Consulting

    Fairview, NJ
    3 days ago
  •  ...are looking for people just like you. Join our team and help us develop game-changing, high-quality solutions.As a Senior Lead Site Reliability Engineer at JPMorganChase within the Core Engineering Solutions team of Consumer and Community Banking, you are an integral part... 
    Senior

    JP Morgan Chase

    Jersey City, NJ
    4 days ago
  • $120k - $175k

     ...level of sports fandom. Ready to reimagine the DFS industry together? We are seeking a highly skilled and experienced Senior Site Reliability Engineer to join our team. We are passionate about delivering cutting-edge solutions and pushing the boundaries of what's... 
    Senior
    Full time
    Remote work
    Work visa
    Flexible hours

    GrabJobs

    Jersey City, NJ
    2 days ago
  • Elevate your engineering prowess to unprecedented levels by joining a team of exceptionally gifted professionals and position yourself among the top echelon in site reliability.As a Senior Lead Site Reliability Engineer at JPMorgan Chase within the enterprise technology... 
    Senior

    JP Morgan Chase

    Jersey City, NJ
    4 days ago
  • $113.3k - $205.52k

     ...is important to maintain our strong culture, achieve our goals, and thrive as #OneJamf. What you'll do at Jamf: As a Senior Site Reliability Engineer, you'll help us balance development velocity with the reliability our customers depend on. You'll partner with engineering... 
    Senior
    Work at office
    Remote work
    Worldwide
    Flexible hours

    GrabJobs

    Jersey City, NJ
    4 days ago
  • $130k - $180k

     ...alongside some of the most experienced and innovative leaders and engineers in the field. Where we work Headquartered in Amsterdam and...  ...an in-house AI R&D team. The role Nebius is looking for a Site Reliability Engineer in Hardware Infrastructure team. You’re welcome to... 
    Temporary work
    Work at office
    Immediate start
    Remote work
    Flexible hours

    GrabJobs

    Newark, NJ
    3 days ago
  •  ...want to work is how we deliver award-winning services to our customers and ultimately build customer value. We’re seeking a Senior Software Engineer to join our stellar team! You will have the opportunity to part of Development of market leading products in the Capital... 
    Senior
    Full time
    Work at office
    Local area
    Remote work

    Broadridge

    Newark, NJ
    2 days ago
  •  ...team.Broadridge is growing! We are seeking an enthusiastic Senior Software Engineer (Java) to join our team. In this role, you will work with internal...  ...will be organized, collaborative, and motivated to deliver reliable, high‑quality applications. ResponsibilitiesDesign, develop... 
    Senior
    Full time
    Local area

    Broadridge

    Newark, NJ
    19 hours ago
  • $158.1k - $213.8k

     ...future with us.ABOUT THE TEAMOur AI/ML Engineering team is at the forefront of transformative...  ...a highly skilled and innovative Senior AI/ML Engineer to design, architect, and...  ...design or architecture (design patterns, reliability and scaling) of new and existing systems... 
    Senior
    Internship
    Flexible hours

    Audible

    Newark, NJ
    3 days ago
  • $158.1k - $213.8k

     ...THIS ROLEThis opportunity is for a Senior Software Development Engineer for Audible’s Consumer Domains group...  ...leading business and build the sites and services (APIs) across desktop...  ...or architecture (design patterns, reliability and scaling) of new and existing systems... 
    Senior
    Internship
    Flexible hours

    Audible

    Newark, NJ
    19 hours ago
  •  ...Company Description Maania consultancy services Job Description Our Fortune client is looking for Senior iOS Software Developer in Newark, NJ. If you are interested please Apply.   Required Skill: • 3+ years of iOS development experience • Excellent... 
    Senior
    Full time

    Maania Consultancy Services

    Newark, NJ
    19 hours ago
  •  ...practicesPerform code reviews troubleshooting root cause analysis and production support activitiesMentor junior developers and promote engineering excellence across the teamWork within Agile ScrumKanban teams contributing to iterative and incremental deliveryParticipate... 
    Senior

    LTM

    Newark, NJ
    4 days ago
  • Job Description: The Senior Manager in Statistics responsible for perform statistical activities in clinical trials from protocol development to final study report and statistical activities in drug in-licensing, regulatory filings and marketing. Responsibilities: Review... 
    Senior

    Katalyst Healthcares & Life Sciences

    Newark, NJ
    4 days ago
  •  ...About Kiddie Kredit Our startup has raised a pre-seed round and is looking to hire engineers for exciting, unique projects. As an early member of the engineering team you will have input on decisions and a lot of potential impact. This position is fully remote and... 
    Senior
    Hourly pay
    Full time
    Remote work

    GrabJobs

    Newark, NJ
    4 days ago
  •  ...responsible for partnering with leaders across engineering and technology to define objective reliability goals for services. Key responsibilities include...  ...continuous improvement. Position Summary: The Senior GCP Site Reliability Engineer acts as an advanced senior... 
    Senior
    Work at office
    Shift work
    Day shift

    Bank of America Corporation

    Jersey City, NJ
    3 days ago
  • $140k - $200k

     ...- Speechify has no office. These include frontend and backend engineers, AI research scientists, and others from Amazon, Microsoft, and...  ...→ testing → release → maintenance. Ensure quality, reliability, and consistency across releases. Identify, diagnose, and resolve... 
    Senior
    Work at office

    Speechify

    Newark, NJ
    19 hours ago
  • $150k - $170k

     ...Intralinks' portfolio of applications in a highly available and secure SaaS environment. Overview : We are looking for a senior software engineer who has significant design and implementation experience with web applications development. You embrace the principles of... 
    Senior
    Ongoing contract

    SS&C Technologies

    Newark, NJ
    2 days ago
  • $118k - $178k

     ...Mission As the world's number 1 job site*, our mission is to help people get jobs...  ...March 2025) Day to Day As a Software Engineer III on the AI Gateway & Guardrails team...  ...architectural decisions, drive service reliability through SLOs and operational readiness,... 
    Senior
    Work experience placement
    Local area

    Indeed

    Newark, NJ
    3 days ago
  •  ...Brand Foundry & American Family Ventures. We are seeking a meticulous Quality Assurance Engineer to join our dynamic team and take ownership of ensuring the quality and reliability of our software products. As a Quality Assurance Engineer, you will play a crucial role... 
    Senior
    Remote work

    GrabJobs

    Newark, NJ
    20 hours ago
  • $111.3k - $139k

     ...Job Title: Senior Software Engineer Team: Systems Location: Hybrid in Chicago (IL) or Newark (NJ) Employment Type: Full-time FLSA Classification Exempt Start Date: ASAP About Braven Braven is a national nonprofit that prepares promising college students to secure a strong... 
    Senior
    Hourly pay
    Full time
    Work at office
    Immediate start
    Remote work
    Visa sponsorship
    Work visa
    Monday to Friday
    2 days per week
    3 days per week

    BRAVEN

    Newark, NJ
    3 days ago
  • $90k - $120k

    As a Performance II-Epic, your role is to provide reliability engineering services through observability and performance engineering techniques....  ...passion for optimizing operational efficiency. You will use Site Reliability Engineering practices to deliver a seamless user... 
    Full time
    Part time
    Work experience placement
    Remote work
    Flexible hours

    Quest Diagnostics

    Secaucus, NJ
    2 days ago
  •  ...and applying your skillsets to drive innovation and modernize the world's most complex and mission-critical systems.As a Site Reliability Engineer III at JPMorgan Chase within the within the Consumer & Community Banking Data and Analytics team, you will solve complex and... 

    JP Morgan Chase

    Jersey City, NJ
    4 days ago
  •  ...while helping others along the way, come join the Broadridge team.Broadridge is growing! We are seeking an enthusiastic Senior Director, Software Engineering (Java & MuleSoft) to join our team. The person In this role, designs, develops, modifies, adapts and implements... 
    Senior
    Full time
    Temporary work
    Local area

    Broadridge

    Newark, NJ
    2 days ago
  • $100k - $144k

     ...Resource Innovations is seeking Senior Software Engineer to join our growing Software as a Service (SaaS) team.As a hands-on technical lead at Resource Innovations, you will be instrumental in the design, development and deployment of innovative cloud-based enterprise... 
    Senior
    Work at office
    Local area
    Remote work

    GrabJobs

    Newark, NJ
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!