Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer, Apple Data Platform / Multi-Cloud Infrastructure

Jobleads-US

Site Reliability Engineer, Apple Data Platform / Multi-Cloud Infrastructure

Austin, Texas, United States Software and Services

The Apple Services Engineering team (ASE) is one of the most exciting examples of Apple's long-held passion for combining art and technology. These are the people who power the App Store, Apple TV, Apple Music, Apple Podcasts, and Apple Books — at extensive scale, meeting high expectations to deliver a huge variety of entertainment in over 35 languages to more than 150 countries.Within ASE, the Apple Data Platform SRE team keeps a massive, multi-cloud platform running for thousands of internal engineers building the next generation of data and AI products at Apple. We sit at the intersection of infrastructure, automation, and customer success — running incident response, providing hands‑on support to internal teams, and partnering with developers to make cutting‑edge services like Spark, Flink, Airflow, Ray, Notebooks, and LLM‑based agent platforms reliable at scale across AWS, GCP, and on‑premise Kubernetes.

Description

This is a rare opportunity to build deep expertise across one of the most technically diverse platforms at Apple — while specialising in the cloud infrastructure that underpins all of it. As an SRE on Apple Data Platform, you'll operate and support the team's full portfolio, from big data pipelines to ML/AI platform services, and grow into the team's go‑to expert for multi‑cloud infrastructure — including AWS core services (IAM, EKS, RDS, S3, VPC networking, autoscaling, EBS), Kubernetes administration at scale, and the Infrastructure‑as‑Code and GitOps tooling that keeps it all reconciled and reliable. You'll be the person other engineers turn to when an IAM policy misfires, a cluster hits a scheduling wall, or a GitOps reconciliation drifts out of sync — and the driving force behind making those failure modes rarer over time.We're looking for a self‑motivated engineer who thrives on ownership — someone who wants a set of services to call their own, the autonomy to drive their reliability roadmap, and the collaborative instinct to keep that work aligned with the team's broader direction. If you love going deep on cloud and Kubernetes internals, enjoy being the trusted expert customers turn to, and want a front‑row seat to Apple Data Platform's multi‑cloud evolution, this role offers real room to grow your scope and impact over time.

Responsibilities

  • Operate, monitor, and triage production and non-production environments across the ADP portfolio — data processing, ML/AI, and multi-cloud infrastructure.
  • Participate in a rotating on‑call schedule across supported services, including occasional weekday and weekend coverage.
  • Own the operational health of multi‑cloud infrastructure as SME — driving reliability, support, and customer guidance for AWS services, EKS clusters, and cross‑cloud networking.
  • Provide Slack‑based support to internal customers; screen, triage, and resolve service related issues.
  • Debug production incidents involving IAM permission errors, storage quota limits, control‑plane/data‑plane namespace separation, and cluster‑wide disruptions.
  • Partner with dev teams across time zones to onboard new services — understanding architecture, then designing monitoring, alerting, and dashboards (Prometheus, Grafana, Splunk).
  • Maintain and evolve Infrastructure‑as‑Code (Crossplane, Terraform) and GitOps (Flux) workflows, troubleshooting state drift and reconciliation issues.
  • Build automation and self‑healing tooling that reduces manual toil and scales the team's operational capacity.
  • Identify, escape, and resolve production issues to protect platform reliability and customer experience.
  • Collaborate with SRE and dev partner teams, engineering, and program management to align execution with team and org goals.

Minimum Qualifications

  • Minimum Qualifications
  • * Bachelor's Degree in Computer Science, an engineering‑related field, or equivalent related experience.
  • 1‑4 years in a Site Reliability Engineering, DevOps, or Infrastructure‑focused role.
  • Proficient in Python; working knowledge of Golang a plus.
  • Kubernetes administration experience — RBAC, node/pod scheduling, autoscalers, PriorityClasses/PDBs, and troubleshooting cluster‑wide disruptions.
  • Strong communication skills and composure under pressure during incidents.
  • Solid grounding in SRE principles, with prior on‑call or production‑support experience.

Preferred Qualifications

  • Experience with Infrastructure‑as‑Code (Crossplane and/or Terraform), including debugging state drift and composition/controller issues.
  • Experience with GitOps workflows (Flux or similar) — HelmRepository/reconciliation troubleshooting and Helm chart deployment.
  • Multi‑cloud exposure (GCP) — parity and migration scenarios are emerging areas of focus.
  • Experience with Splunk for log pipeline debugging (e.g., fluent‑bit).
  • Familiarity with Spark/Flink running on Kubernetes (executor scheduling, node affinity).
  • Comfort with GitHub PR review workflows in an infrastructure‑as‑code / GitOps context.
  • A track record of automating manual operations through scripting or tooling.
  • Intellectual curiosity and a drive to keep learning — for yourself, your team, and the org.

Apple is an equal opportunity employer that is committed to inclusion and diversity. We seek to promote equal opportunity for all applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, Veteran status, or other legally protected characteristics. Learn more about your EEO rights as an applicant

At Apple, we believe accessibility is a fundamental human right. You’ll find that idea reflected in everything here — in our culture, our benefits and our digital tools. By welcoming as many perspectives as possible, we help you build a career where you feel like you belong.

Learn about accessibility in Apple’s workplace

Learn about reasonable accommodations for job applicants

Apple accepts applications to this posting on an ongoing basis.

#J-18808-Ljbffr Jobleads-US
Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer, Apple Data Platform / Multi-Cloud Infrastructure in Austin, TX vacancy
  • $127k - $249k

    Who We AreAt MongoDB, we are building the future of data platforms with MongoDB Atlas; our fully managed, multi-cloud database service. The Infrastructure Security (InfraSec) team sits within Site Reliability Engineering (SRE) in the CTO organization. We are an... 
    Platform
    Cloud
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours
    3 days per week

    MongoDB

    Austin, TX
    2 days ago
  •  ...industry leader in data-driven, client-to-cloud networking for...  ..., such as Best Engineering Team, Best...  ...experienced Senior Infrastructure Cloud Systems...  ...complex hybrid, multi-cloud...  ...networking, and platform interconnectivity...  ...enterprise AI.Reliability, Performance &... 
    Platform
    Cloud

    Arista Networks

    Austin, TX
    3 days ago
  • $218k - $323.95k

     ...of payments on our platform on behalf of our...  ...authority for Venmo’s data platform infrastructure and performance engineering. The Principal...  ...the development of multi-year business strategiesExercises...  ...in production cloud environments (AWS...  ...requiring extreme reliability and millisecond-... 
    Platform
    Cloud
    Full time
    Work at office
    Local area
    Immediate start
    Flexible hours

    PayPal

    Austin, TX
    1 day ago
  • $195k - $290k

     ...advanced AI-native platform. We work on...  ...CrowdStrike's Data Platform is...  ...spans multiple clouds and regions, including...  ...platform. Reliability, cost...  ...compliance are engineering requirements here...  ...grow it into a multi-team organization...  ..., and test infrastructure. It partners hands... 
    Platform
    Cloud
    Full time
    Work experience placement
    Work at office
    Local area
    Shift work

    CrowdStrike

    Austin, TX
    1 day ago
  • $132.4k - $238.4k

     ...comprehensive engineering, supply chain,...  ...network of over 100 sites worldwide,...  ...integrity of both the data center products...  ...-compliant, platform-based, and...  ...EngineerOrganization: Data Center Infrastructure (DCI)...  ...Participate in multi-discipline...  ...with various cloud providers across... 
    Platform
    Cloud
    Temporary work
    Work at office
    Local area
    Remote work
    Worldwide

    Jabil Circuit

    Austin, TX
    1 day ago
  • $168k - $200k

     ...Datavant is the data collaboration platform trusted for healthcare. Guided...  ...looking for a Senior Site Reliability Engineer to join our Data & ML...  ...SRE mindset with deep cloud infrastructure and data platform experience...  .... Familiarity with multi-cloud or hybrid cloud... 
    Platform
    Cloud

    Datavant

    Austin, TX
    1 day ago
  • Req ID:375454NTT DATA strives to hire exceptional...  ...currently seeking a Site Reliability Engineer to join our team in...  ...in Terraform, Cloud Infrastructure, DevOps, Automation...  ...applications across hybrid and multi-cloud environments....  ...-Code (IaC), cloud platforms, CI/CD pipelines,... 
    Platform
    Cloud
    Temporary work
    Work at office
    Remote work
    Flexible hours

    NTT DATA

    Austin, TX
    4 days ago
  • $127k - $249k

     ...TeamPlatform Engineering sits within SRE...  ...builds the core infrastructure powering...  ...everything from our multi-cloud Kubernetes...  ...observability platforms.Within Platform...  ...ensuring customer data remains safe...  ...the reliable, globally connected...  ...talented Senior Site Reliability Engineer... 
    Platform
    Cloud
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    Austin, TX
    3 days ago
  • $127k - $249k

    Platform Engineering is the department within SRE that is responsible...  ...a range of critical infrastructure and operational...  ...Among these are our multi-cloud-provider Kubernetes...  ...that ensure cluster reliability and security (e.g.,...  ...have redefined the data platform for the AI... 
    Platform
    Cloud
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    Austin, TX
    4 days ago
  • $152k - $241.5k

     ...global services platform. At NVIDIA, you’ll...  ...fabrics.Use IaC(Infrastructure‑as‑Code) and config...  ...distributed, multi‑cloud hybrid environment...  ...management, fleet reliability/auto-healing, E2E...  ...observability or data-driven operations...  ...Ruby.Mentored other engineers and influenced... 
    Platform
    Cloud
    Full time

    Nvidia

    Austin, TX
    4 days ago
  • $127k - $249k

     ...experienced Senior Engineer for our SRE,...  ...the Atlas platform. As a senior SRE...  ...a talented Site Reliability Engineer (SRE)...  ...with a strong infrastructure background. This...  ...with a major cloud provider (AWS,...  ...systems in a multi-cloud environmentA...  ...redefined the data platform for... 
    Platform
    Cloud
    Local area
    Remote work
    Worldwide
    Flexible hours

    MongoDB

    Austin, TX
    5 days ago
  • ## Site Reliability EngineerApplyremote type:...  ...management and data protection...  ...Site Reliability Engineer to ensure the...  ...in the public cloud. This product...  ...*Automation & Infrastructure as Code:** Design...  ...of the platform.* **Incident Management...  ...versatile and multi-tasking... 
    Platform
    Cloud
    Full time
    Local area
    3 days per week

    Thales Group

    Austin, TX
    1 day ago
  •  ...within the US  Staff Site Reliability Engineer / Cloud SME Location: 100% remote...  ...across multiple cloud platforms (Azure and AWS) and container...  ...services to enhance our infrastructure. Reliability...  ...computing platforms. Strong multi-cloud expertise required... 
    Platform
    Cloud
    Long term contract
    Remote work

    ASCENDING

    Austin, TX
    a month ago
  • $143.5k - $212.85k

     ...completion of payments on our platform on behalf of our...  ...Summary:The Venmo Data Infrastructure team supports Venmo's...  ...secure, scalable, reliable and performant platform...  ...Executive, Product, Engineering and Analytics teams to...  ...from AWS (or similar cloud provider), including:... 
    Platform
    Cloud
    Full time
    Work at office
    Local area
    Immediate start
    Flexible hours

    PayPal

    Austin, TX
    4 days ago
  •  ...SRE to join our Platform Engineering team where you’ll own the reliability, scalability, and...  ...users Partner with data engineering and...  ...Develop and maintain infrastructure-as-code (...  ...best practices, and multi-environment deployments...  ...working in cloud environments (AWS... 
    Platform
    Cloud
    Full time

    Dimensional Fund Advisors

    Austin, TX
    1 day ago
  •  ...Recognized as the No. 1 site trusted by real...  ...a Senior Site Reliability Engineer to join our...  ...excellence of our platform infrastructure serving millions...  ...(ECS), and multi-region architectures...  ...Skills ~ Cloud & Infrastructure...  ...and policies ~ Data‑driven decision... 
    Platform
    Cloud
    Work at office
    Local area
    Immediate start
    Flexible hours

    realtor.com

    Austin, TX
    1 day ago
  • $224k - $356.5k

    NVIDIA is hiring engineers to scale up the introduction...  ...into its EDA Infrastructure. We expect you to...  ...between cloud and on-premProviding...  ..., clear and reliable architecture specificationTranslate...  ...practices and/or Platform Engineering....  .... Developing multi-cloud... 
    Platform
    Cloud
    Full time

    Nvidia

    Austin, TX
    1 day ago
  • $150k - $170k

     ...through an API-based platform, empowering our...  ...the Manager of Site Reliability Engineering, you'll lead a team...  ...priorities. Infrastructure as Code : Set architectural...  ..., Helm, and multi-cluster patterns. Cloud Native Expertise...  ...workflows. Data & Middleware &... 
    Platform
    Cloud
    Full time
    Work at office
    Remote work
    Worldwide
    Flexible hours

    DriveWealth

    Austin, TX
    4 days ago
  • $184k - $287.5k

     ...looking for a Systems & Software Engineer interested in building and running reliable large scale infrastructure platform services. In this role you...  ...scale private or public cloud systems in productionIn depth...  ...working with or developing multi-cloud infrastructure services... 
    Platform
    Cloud
    Full time
    Remote work

    Nvidia

    Austin, TX
    3 days ago
  • $140k - $215k

     ...advanced AI-native platform. We work on...  ...standard in cloud‑native cybersecurity...  ...and Embedded Reliability charters:...  ...with product engineering teams and their...  ...and building infrastructure‑as‑code tooling...  ...define and drive multi‑year...  ...experience in data structures, algorithms... 
    Platform
    Cloud
    Full time
    Work experience placement
    Work at office
    Local area
    2 days per week
    3 days per week

    CrowdStrike

    Austin, TX
    1 day ago
  • $73.4k - $136.3k

     ...IT, optimizing data architectures,...  ..., and hybrid cloud environments.The...  ...and Systems Engineer to design, implement...  ...on-premises infrastructure environments....  ...that improve reliability, scalability,...  ...tools, platforms, and best practices...  ...access or use this site as a result of... 
    Platform
    Cloud
    Minimum wage
    Full time
    Contract work
    Flexible hours

    DXC Technology

    Austin, TX
    17 hours ago
  • $135.2k - $306.4k

    Oracle Cloud Infrastructure (OCI) Compute delivers bare metal...  ...systems that are multi-tenant, highly available...  ...to deliver new platform features focusing...  ...improve overall services reliability.Mentor and guide engineers in distributed...  ...design, high-scale data processing, and operational... 
    Platform
    Cloud
    Temporary work
    Flexible hours

    Oracle Corporation

    Austin, TX
    1 day ago
  •  ...Are:The Global AI Infrastructure team is at the center...  ...expertise across cloud, on-premises, and...  ...that powers AI platforms, GPU-accelerated workloads...  ...IT systems, data pipelines, security...  ...LLM inference engines (TensorRT-LLM), production...  ...layers including multi-node training and... 
    Platform
    Cloud
    Full time
    Work experience placement
    Live in
    Work at office
    Local area

    Accenture

    Austin, TX
    4 days ago
  • $160k - $250k

     ...’s most advanced AI-native platform. We work on large scale distributed...  ...is seeking a Sr Engineering Manager (L8) to lead a host...  ...coverage in MLOps & broader data infrastructure, as well as data analytics...  ...Trillion events/day) both in our cloud environments as well as... 
    Platform
    Cloud
    Full time
    Work experience placement
    Work at office
    Local area
    Remote work
    2 days per week

    CrowdStrike

    Austin, TX
    3 days ago
  • $184k - $287.5k

     ...a Senior Software Engineer to lead the bring-...  ...across NVIDIA GPU platforms at the largest scales...  ...efficiently and reliably at scale. You will...  ...investigations on multi-GPU and multi-node...  ...AI clusters, infrastructure, and end-to-end workloads...  ...workloads using data, tensor, pipeline,... 
    Platform
    Cloud
    Full time
    Remote work

    Nvidia

    Austin, TX
    2 days ago
  •  ...capabilities in digital, cloud and security....  .... The Global AI Infrastructure team enables...  ...automation workflows, and platform capabilities that...  ...systems, data platforms, security...  ...issues across multi-node AI training,...  ...capacity management, reliability engineering, change... 
    Platform
    Cloud
    Full time
    Work experience placement
    Live in
    Work at office
    Local area

    Accenture

    Austin, TX
    3 days ago
  • $147k - $237.5k

     ...observability platform built for control...  ...focus on the data and insights...  ...Impact Areas:Our Infrastructure organization...  ...Principal Engineers to drive the future...  ...Engineering (Reliability & Scale)The...  ...massive distributed cloud resources to...  ...AI agents, multi-agent frameworks... 
    Platform
    Cloud
    Full time
    Remote work

    Palo Alto Networks

    Austin, TX
    2 days ago
  • $132.4k - $238.4k

     ...offering comprehensive engineering, supply chain,...  ...of over 100 sites worldwide, Jabil...  ...across critical infrastructure products and automation platforms on the Data Center Infrastructure...  ...the performance, reliability, and scalability...  ...SCADA, CMMS, and cloud-based monitoring... 
    Platform
    Cloud
    Temporary work
    Work at office
    Local area
    Remote work
    Worldwide

    Jabil Circuit

    Austin, TX
    17 hours ago
  • $79.2k - $209.5k

     ...performance tests; leverages data plane platforms and distributed...  ...are met.Oracle Cloud Infrastructure Workflow is a Tier 0...  ...execution of distributed, multi-step work in a fault...  ...tolerant manner. An engineer on this team is...  ...Architecture - System Reliability Design:-Collaborates... 
    Platform
    Cloud
    Temporary work
    Flexible hours

    Oracle Corporation

    Austin, TX
    4 days ago
  •  ...Own reliability for client platforms: SLOs, observability and incident practice. You will join a...  ...team working directly with client engineers. We do not run a bench — everyone...  ...operations ~ Deep experience with cloud infrastructure A human reads every one. We... 
    Platform
    Cloud

    RTS RTech Solutions

    Austin, TX
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer, Apple Data Platform / Multi-Cloud Infrastructure. Be the first to apply!