Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Staff Site Reliability Engineer

Engg

About Crunchyroll

Founded by fans, Crunchyroll delivers the art and culture of anime to a passionate community. We super-serve over 100 million anime and manga fans across 200+ countries and territories, and help them connect with the stories and characters they crave. Whether that experience is online or in-person, streaming video, theatrical, games, merchandise, events and more, it’s powered by the anime content we all love. Join our team, and help us shape the future of anime!

About the role

We are hiring a Staff Site Reliability Engineer (SRE) to join the Center for Data & Insights (CDI) in the US and play a critical role in advancing the reliability, scalability, performance, and security of Crunchyroll's consumer-facing data platforms. As a senior technical leader, you will partner closely with Engineering, Data, Infrastructure, Product, and Security teams to design and operate resilient cloud-native systems that power critical business and customer experiences. You will drive initiatives across observability, incident management, automation, capacity planning, disaster recovery, and operational excellence while helping teams adopt modern SRE practices such as SLIs, SLOs, and error budgets. The ideal candidate combines deep expertise in large-scale distributed systems with a strong sense of ownership, collaboration, and service leadership. You are passionate about building highly reliable platforms, eliminating operational toil through automation, and enabling engineering teams to move quickly and safely. In addition, you will champion SecOps best practices by driving vulnerability management, supporting penetration testing initiatives, improving security observability, strengthening cloud and Kubernetes security controls, and ensuring operational readiness for emerging threats. This is a unique opportunity to shape reliability and security engineering practices across CDI while helping build a world-class data and insights ecosystem that enables informed decision‑making throughout Crunchyroll.

Core Areas of Responsibility
  • Reliability Engineering : Define, measure, and continuously improve the reliability, availability, and performance of CDI platforms through SLIs, SLOs, and error budgets.
  • Operational Excellence : Establish and drive best practices for incident management, root cause analysis, postmortems, and service ownership across engineering teams.
  • Observability & Monitoring : Build and evolve comprehensive monitoring, logging, tracing, and alerting capabilities to enable proactive issue detection and rapid resolution.
  • Automation : Identify operational inefficiencies and develop automation, self-service capabilities, and self-healing mechanisms to improve engineering productivity.
  • Platform Scalability : Design and optimize cloud-native infrastructure and services to support growing business demands while maintaining performance and cost efficiency.
  • Infrastructure Engineering : Drive Infrastructure as Code (IaC), platform standardization, and deployment automation to improve consistency, reliability, and operational agility.
  • Capacity Planning & Performance : Lead capacity planning and performance optimization initiatives to ensure platforms can scale predictably and efficiently.
  • Disaster Recovery & Resilience : Develop and regularly validate disaster recovery, backup, and business continuity strategies to ensure platform resiliency.
  • Security Operations (SecOps) : Partner with Crunchyroll's security team to integrate security controls, operational risk management, and security best practices into platform operations and engineering workflows.
  • Vulnerability Management : Own the triage and remediation of identified vulnerabilities across infrastructure, platform, container, and application security vulnerabilities through established Crunchyroll vulnerability management processes.
  • Penetration Testing & Security Remediation : Support penetration test scoping activities by providing technical context on CDI platforms. Own the triage, prioritization, and remediation of resulting findings to drive timely resolution and strengthen platform security posture.
  • Cloud & Kubernetes Security : Implement and maintain secure cloud, container, and Kubernetes environments following least-privilege, defense-in-depth, and Zero Trust principles.
  • Cross-Functional Leadership : Collaborate with Engineering, Data, Product, Infrastructure, and Security teams to drive reliability, scalability, and security initiatives across CDI.
  • Mentorship & Engineering Excellence : Mentor engineers and champion a culture of operational excellence, reliability, ownership, continuous improvement, and security awareness.
About You
  • 12+ years of experience in Site Reliability Engineering (SRE), Platform Engineering, Infrastructure Engineering, or related disciplines, with a proven track record of operating and scaling production-critical systems.
  • Deep expertise in Kubernetes and GCP , including the design, deployment, and operation of highly available, cloud-native platforms at scale.
  • Strong Infrastructure as Code (IaC) experience , preferably with Terraform, and a commitment to automation, standardization, and operational efficiency.
  • Solid foundation in Linux systems administration, networking, and distributed systems , with the ability to troubleshoot complex production issues across multiple layers of the technology stack.
  • Proficiency in one or more programming and scripting languages , such as Go, Python, Java, or Shell, with a focus on automation and platform engineering.
  • Hands‑on experience with modern observability platforms and practices , including Prometheus, Grafana, OpenTelemetry, Datadog, or equivalent monitoring and telemetry solutions.
  • Demonstrated expertise in incident management, service reliability, capacity planning, performance optimization, and operational excellence , including the implementation of SLIs, SLOs, and error budgets.
  • Strong understanding o
#J-18808-Ljbffr
Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Staff Site Reliability Engineer in Los Angeles, CA vacancy
  • $125k - $145k

     ...SpaceX is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SITE RELIABILITY ENGINEER, GNCSpaceX’s mission is to make humanity multiplanetary by developing fully and rapidly reusable launch systems capable of... 
    Suggested
    Permanent employment
    Temporary work
    Flexible hours
    Weekend work

    SpaceX

    Hawthorne, CA
    4 days ago
  • $165k - $270k

     ...is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SR. SITE RELIABILITY ENGINEER (STARSHIELD)Starshield leverages SpaceX’s Starlink technology and launch capability to support national security efforts.... 
    Suggested
    Permanent employment
    Temporary work
    Immediate start
    Weekend work

    SpaceX

    Hawthorne, CA
    4 days ago
  • $165k - $265k

     ...is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SR. SITE RELIABILITY ENGINEER - TOP SECRET CLEARANCE (STARLINK)At SpaceX we’re leveraging our experience in building rockets and spacecraft to deploy Starlink... 
    Suggested
    Permanent employment
    Temporary work
    Worldwide
    Weekend work

    SpaceX

    Hawthorne, CA
    1 day ago
  • $155k - $195k

     ...you to join us on our mission of providing humankind access to the galaxy beyond our planet. About the RoleWe are seeking a Site Reliability Engineer to join our Ground Software team. As a Site Reliability Engineer, you will design, build, and operate the ground and site... 
    Suggested
    Full time
    Work at office

    Apex Technology

    Los Angeles, CA
    4 days ago
  • $125k - $150k

     ...possible, with the ultimate goal of enabling human life on Mars.SITE RELIABILITY ENGINEER (RAPTOR)SpaceX is looking for a Site Reliability Engineer...  ...infrastructure systems.Work with propulsion engineering staff to solve critical bottlenecks.Coordinate and communicate with... 
    Suggested
    Permanent employment
    Temporary work

    SpaceX

    Hawthorne, CA
    2 days ago
  • $145k - $195k

     ...SpaceX is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SITE RELIABILITY ENGINEER (TOP SECRET CLEARANCE)As a member of the Classified IT Systems Engineering team, the Site Reliability Engineer is involved... 
    Permanent employment
    Temporary work
    Weekend work

    SpaceX

    Hawthorne, CA
    2 days ago
  •  ...your big ideas, and your desire to team up with some of the best and brightest in technology and entertainment. The RoleThe Site Reliability Engineer (SRE) II is responsible for designing, implementing, and maintaining scalable and reliable systems and applications. Focus... 
    Full time
    Local area
    Worldwide
    Flexible hours

    AXS Group

    Los Angeles, CA
    2 days ago
  • $165k - $265k

     ...is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SR. SITE RELIABILITY ENGINEER (STARLINK)At SpaceX we’re leveraging our experience in building rockets and spacecraft to deploy Starlink, the world’s most... 
    Permanent employment
    Temporary work
    Worldwide
    Weekend work

    SpaceX

    Hawthorne, CA
    2 days ago
  • $180k - $200k

     ...Senior Site Reliability Engineer If you have ever watched television or enjoyed a movie on your phone or tablet, this experience was likely brought to you through an Ateme solution created by our award-winning engineering teams. Ateme is the global leader in video... 

    ATEME

    Culver City, CA
    4 days ago
  • $113.3k - $205.52k

     ...important to maintain our strong culture, achieve our goals, and thrive as #OneJamf. What you'll do at Jamf: As a Senior Site Reliability Engineer, you'll help us balance development velocity with the reliability our customers depend on. You'll partner with engineering... 
    Work at office
    Remote work
    Worldwide
    Flexible hours

    GrabJobs

    Glendale, CA
    2 days ago
  • $164k - $270k

     ...for the 21st century and beyond.The Role What You’ll DoOwn the reliability of our robotics systems, from PLCs through ROS2/middleware to...  ...remediation.Partner with controls, robotics, and platform engineering teams to bake reliability in early. Review designs, develop SLOs... 
    Permanent employment
    Full time
    Local area
    Flexible hours

    Hadrian

    Los Angeles, CA
    1 day ago
  • $140k - $180k

     ...Senior Site Reliability Engineer Los Angeles, CA K2 is building the largest and highest-power satellites ever flown, unlocking performance levels previously out of reach across every orbit. Backed by over $1 billion in total funding from leading investors including... 
    Permanent employment
    Shift work

    K2 Space

    Los Angeles, CA
    3 days ago
  • $181k - $225k

     ...Senior Site Reliability Engineer Los Angeles, CA Altruist is transforming the multi-trillion dollar wealth management industry by building an AI platform for wealth professionals. We partner with financial advisors nationwide, empowering them to grow, optimize time... 
    Work at office
    Immediate start
    3 days per week

    Altruist

    Los Angeles, CA
    4 days ago
  • $175k - $285k

    Hadrian - Manufacturing the FutureHadrian is building autonomous factories to reindustrialize America. By combining AI, advanced software, robotics, and full-stack manufacturing, we help aerospace and defense companies build rockets, satellites, aircraft, ships, and other...
    Permanent employment
    Full time
    Local area
    Remote work
    Flexible hours

    Hadrian

    Los Angeles, CA
    2 days ago
  • $139.9k - $199.3k

     ...Lead Site Reliability Engineer We're looking for talented professionals to join us in bringing smart money management and payment solutions to everyone's fingertips. This position is classified as structured hybrid, with an expectation of a minimum of three (3) days... 
    Work experience placement
    Work at office
    Remote work
    3 days per week

    Green Dot

    Los Angeles, CA
    1 day ago
  • $145k - $160k

     ...We are seeking a specialized Observability & Infrastructure Engineer to lead the monitoring, telemetry, and platform-as-code initiatives critical to our multi-region disaster recovery roadmap. You will architect and implement robust observability pipelines, ensure deep... 
    Temporary work
    Remote work
    Flexible hours

    EPAM Systems Inc

    Los Angeles, CA
    2 days ago
  •  ...to operate with clarity, control, and confidence across the reimbursement journey. About the Role We’re hiring a Staff Site Reliability Engineer to define and strengthen how reliability, scalability, and operational excellence are built into Pivotal’s platform.... 
    Remote work
    Flexible hours

    Pivotal Health

    Santa Monica, CA
    3 days ago
  • $30.53 - $56.48 per hour

     ...Job Title:Associate Site Reliability EngineerRequisition ID:R Job Description:Job Title: Associate Site Reliability EngineerReporting To:Manager...  ...: Santa Monica, CaOverviewThe Associate Site Reliability Engineer helps keep Marketing Technology services reliable, observable... 
    Hourly pay
    Full time
    Temporary work
    Part time
    Internship
    Local area
    Worldwide
    Relocation package

    Activision

    Santa Monica, CA
    6 hours ago
  • $125k - $195k

     ...SpaceX is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars. SITE RELIABILITY ENGINEER (HIGH PERFORMANCE COMPUTING) SpaceX HPC is a shared compute platform used across the company — vehicle and structures... 
    Permanent employment
    Full time
    Temporary work
    Weekend work

    SpaceX

    Hawthorne, CA
    2 days ago
  •  ...What you will do: Partner with a team of high-performing engineers and developers who are focused on delivering best in class software...  ...our shift to a SecDevOps culture, solving for security, reliability, cost-effectiveness, and observability Building Zero trust... 
    Full time
    Contract work
    Local area
    Flexible hours
    Shift work

    DISQO

    Los Angeles, CA
    a month ago
  •  ..., and thrive! KēSTA I.T. is actively seeking a Principal Engineer for an immediate full-time opportunity with our industry creating...  ...An innovative technology company is seeking experienced Site Reliability Engineers to take ownership of building reliable, scalable platforms... 
    Full time
    Contract work
    Immediate start
    Work from home
    Flexible hours

    KēSTA I.T.

    Beverly Hills, CA
    9 days ago
  •  ...automations, system integrations, and data pipelines across our ERP, engineering, and business platforms. This is a hands-on, execution-focused...  ...closely with the IT Business Systems team and Enterprise IT staff, travel occasionally, and provide some afterhours support. This... 
    Full time
    Work experience placement
    Remote work

    Arete Associates

    Los Angeles, CA
    2 days ago
  •  ...Job Description Job Description Forhyre is looking for engineers who can bring unique perspectives and innovative ideas to all areas...  ...evangelize cloud best practices while building a culture of reliability and observability Engage in and improve the end to end lifecycle... 

    Forhyre

    Los Angeles, CA
    4 days ago
  •  ...infrastructure, DevOps, SRE, and platform engineering. You will test AI-generated commands,...  ...and deployment workflows for accuracy and reliability. Work with AWS, Azure, GCP,...  ...Azure DevOps Cloud Infrastructure Site Reliability Engineering (SRE) Platform... 
    Remote job
    For contractors

    YO AI Labs

    Los Angeles, CA
    a month ago
  • $270k - $410k

     ...when you merge creativity, intuition and cutting-edge technology. Come be a part of what’s next.About Strategic Technical Solutions Engineering (STS)For Netflix to continue to scale and produce quality content, the Studio and Corporate business teams work closely with... 
    Hourly pay
    Full time
    Temporary work
    Work at office
    Immediate start
    Flexible hours

    Netflix

    Los Angeles, CA
    3 days ago
  • $175k - $285k

    Job Description Job Description Hadrian - Manufacturing the Future Hadrian is building autonomous factories to reindustrialize America. By combining AI, advanced software, robotics, and full-stack manufacturing, we help aerospace and defense companies build rockets...
    Permanent employment
    Full time
    Remote work
    Relocation package
    Flexible hours

    Hadrian Automation

    Los Angeles, CA
    9 days ago
  • $110k - $164k

     ...lookout for exceptional builders, fast learners, and ambitious engineers. Whether your passion lies in spacecraft systems, avionics, ML/...  ...Salary range: $110,000 - $164,000 / per year. This role is on-site in Hawthorne, CA Benefits Equity Unlimited PTO... 
    Full time
    Internship
    Worldwide
    Weekend work

    Oligo Space

    Hawthorne, CA
    more than 2 months ago
  • $209k - $313k

     ...Bitmoji, Saturn, and other digital services.We are looking for a Staff, Revenue Growth to join Snap Inc.’s global Small and Medium...  ...array of cross-functional groups (Marketing, Operations, Product, Engineering, Advertiser Support, and Data Analytics).What you’ll do:Be the... 
    Full time
    Live in
    Work at office
    Local area
    Shift work

    Snap

    Los Angeles, CA
    4 days ago
  •  ...while also being an effective team player  About The Role We’re looking for a Senior Systems Software Engineer who can take ownership of the reliability, automation, and evolution of our core systems. You’ll work across the stack, with a strong emphasis on backend... 
    Full time
    Local area
    Remote work

    Zoo

    Los Angeles, CA
    more than 2 months ago
  • $50k - $75k

     ...LA Area Campus Ministry Staff Could This Be You? Do you have a burden to see lost students hear and respond to the gospel? Do you get fired up seeing teenagers empowered to share the gospel and reach their whole schools for Christ? Can you cast a compelling... 
    Full time
    Local area

    Decision Point

    Los Angeles, CA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Staff Site Reliability Engineer. Be the first to apply!