Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Lead Cloud Platform Engineer (Data & Execution Platform) (Austin)

Part-time

Fidelity Investments

Job Description:Note: Fidelity is not providing immigration sponsorship for this position. The Role We are seeking a hands-on Lead Cloud Platform Engineer to implement, scale, and operate cloud-native infrastructure and services that power large-scale data processing systems. This role focuses on translating defined architectures into production-grade platforms that are reliable, observable, secure, and performant. You will lead the implementation and operation of a modern execution platform built on Apache Spark for distributed compute and an Airflow orchestration layer and DAG execution environment. The ideal candidate brings deep production experience in Spark and Airflow, and excels at troubleshooting, tuning, and operationalizing distributed systems in AWS environments, while leveraging modern developer productivity tools such as AI-assisted coding and LLM-based workflows.The Expertise and Skills You BringImplement and operate cloud-native platform services for distributed data systemsScale fault-tolerant, high-throughput systems aligned with architectural patternsOwn Spark data pipelines and Airflow orchestration layer and DAG executionTune Spark workloads (partitioning, memory, execution plans, shuffle optimization)Troubleshoot Spark jobs and Airflow DAGs across performance and failuresOperate and optimize Kubernetes-based execution environments, including node group scaling, workload placement, and resource utilizationTroubleshoot Kubernetes infrastructure and workload issues, including scheduling, networking, and runtime performanceLeverage developer productivity tools (e.g., GitHub Copilot, LLMs) to accelerate development, debugging, and operational workflows.Drive operational excellence including monitoring, incident response, and RCAImplement observability (metrics, logging, tracing, dashboards, alerting)Define and manage SLIs/SLOs for platform reliabilityDeploy solutions using AWS services (EKS, EC2, S3, Lambda, RDS, etc.) (Implement secure networking (VPCs, IAM, subnets, load balancing)Maintain CI/CD pipelines and deployment automationLead execution across planning, delivery, and cross-team coordinationMentor engineers and promote reliability and scalability best practicesStrong understanding of distributed systems (fault tolerance, scalability, consistencyExpertise in Apache Spark (tuning, debugging, optimization)Expertise in Apache Airflow (DAG execution, orchestration, troubleshooting)Strong experience operating Kubernetes (EKS preferred) including cluster scaling and lifecycle managementHands-on management of node groups, autoscaling, and capacity planningDeep understanding of Kubernetes networking and security (security groups, network policies, ingress/egress)Experience with Kubernetes resources (Deployments, StatefulSets, Jobs, CronJobs)Familiarity with Custom Resources (CRDs) and advanced configuration via annotations and labelsExperience monitoring Kubernetes clusters (metrics, logs, events) and integrating with observability toolsTroubleshooting Kubernetes workloads (scheduling failures, resource contention, networking issues)Experience with AWS services and cloud-native design patternsProficiency in Python, Java, or GoExperience with Docker and KubernetesHands-on observability (metrics, logging, tracing)Experience with SLI/SLO-based reliability modelsPractical experience using AI-assisted development tools (e.g., GitHub Copilot, LLMs) to improve code quality, debugging, and productivityNetworking fundamentals (DNS, TCP/IP, TLS, VPC design)Strong troubleshooting and performance tuning skillsStrong communication and leadership skillsBachelor’s or Master’s degree in Computer Science or related field (or equivalent experience)8 plus years in software, platform, or cloud engineering rolesExperience operating large-scale distributed systems in productionStrong experience with AWS cloud platformsMandatory hands-on experience with Apache Spark and Apache Airflow in productionExperience supporting ETL, data platforms, or workflow execution systems at scaleFidelity’s Onsite Working ModelFidelity is transitioning to a full-time onsite working model through a phased rollout across regions and roles. Currently, some roles and locations require 100% onsite presence, while others require less. Onsite expectations are likely to evolve as the rollout continues. This transition does not apply to fully remote roles.Certifications:Category:Information TechnologyPlease be advised that Fidelity’s business is governed by the provisions of the Securities Exchange Act of 1934, the Investment Advisers Act of 1940, the Investment Company Act of 1940, ERISA, numerous state laws governing securities, investment and retirement-related financial activities and the rules and regulations of numerous self-regulatory organizations, including FINRA, among others. Those laws and regulations may restrict Fidelity from hiring and/or associating with individuals with certain Criminal Histories.SummaryLocation: Merrimack, NH; Westlake, TXType: Full time

Vacancy posted more than 2 months ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Lead Cloud Platform Engineer (Data & Execution Platform) (Austin). Be the first to apply!