Sr. Staff Engineer Software, Infrastructure Reliability (Chronosphere)
$126k - $204.5kPalo Alto Networks
Our MissionAt Palo Alto Networks, we’re united by a shared mission—to protect our digital way of life. We thrive at the intersection of innovation and impact, solving real-world problems with cutting-edge technology and bold thinking. Here, everyone has a voice, and every idea counts. If you’re ready to do the most meaningful work of your career alongside people who are just as passionate as you are, you’re in the right place.Who We AreIn order to be the cybersecurity partner of choice, we must trailblaze the path and shape the future of our industry. This is something our employees work at each day and is defined by our values: Disruption, Collaboration, Execution, Integrity, and Inclusion. We weave AI into the fabric of everything we do and use it to augment the impact every individual can have. If you are passionate about solving real-world problems and ideating beside the best and the brightest, we invite you to join us!This role is remote, but distance is no barrier to impact. Our hybrid teams collaborate across geographies to solve big problems, stay close to our customers, and grow together. You will be part of a culture that values trust, accountability, and shared success where your work truly matters.Job SummaryJob SummaryWe're looking for an Infrastructure Engineer to build developer tooling that enables developer velocity and reliability for the entire engineering organization. Our team owns the full end-to-end development lifecycle, from local development tools, CI/CD, dynamic testing environments, deployments, and releases. You’ll be working with the latest cloud-native technologies as you build an infrastructure platform that enables the engineering organization to deliver observability for some of the largest technology companies.Key ResponsibilitiesArchitect & Build: Design and maintain high-scale developer tooling and backend services that improve productivity and reliability across a distributed cloud environment. We operate in a 100% modern, cloud-native ecosystem and you will work exclusively with ephemeral infrastructure and containerized microservices.Infrastructure as Code (IaC): Treat infrastructure as a first-class citizen. You will define, deploy, and manage entire environments using declarative IaC (Terraform), ensuring our platform is reproducible and version-controlled.Drive Systemic Quality: Identify and eliminate systemic bottlenecks in the software development lifecycle (SDLC) through architectural changes or advanced tooling.Scale & Reliability: Ensure our infrastructure remains resilient under massive traffic loads, optimizing for performance, cost-efficiency, and near real-time telemetry processing.Strategic Leadership & Mentorship: Define platform standards and reference architectures that span a 1–3 year horizon, balancing feature velocity with long-term technical debt. Act as the "glue" across teams, consulting on infrastructure best practices and up-leveling the organization through mentorship.Qualifications Required Qualifications8+ years of relevant experience with the following:Foundational Technical Proficiency: Strong experience in at least one backend language (e.g., Go, Java, Python, or Rust). We value fluency and the ability to write modular, testable code over knowing a specific syntax.Strong proficiency in at least one backend language (e.g., Go, Java, Python, or Rust). We prioritize the ability to write modular, testable code over knowledge of a specific language syntax.Deep Systems Expertise: You go beyond "using" the cloud; you understand how it works. This includes:Cloud-Native: A solid understanding of cloud-native concepts and experience working with cloud providers like AWS or GCP. You should be comfortable navigating Kubernetes and container-level logic.Operating Systems & Compute: Deep knowledge of Linux internals, process management, and resource isolation.Networking & Security: Understanding of the OSI model, service meshes, load balancing, and "zero-trust" security architectures.Distributed Systems: Experience building and debugging systems that deal with CAP theorem trade-offs, eventual consistency, and distributed tracing.Reliable Execution: A track record of completing assigned tasks/tickets reliably and estimating work effectively within a sprint. You take ownership of features from local development through to basic testing and delivery.Analytical Debugging & Quality: The ability to debug your own code efficiently using logs and tests, while proactively identifying edge cases (like nulls or limits) during the design phase.Collaborative Spirit: Strong communication skills to keep teammates informed, raise blockers early, and contribute meaningfully to code reviews and design discussions.Continuous Learning: A proactive approach to learning new tools, processes, and libraries. You are open to feedback and use incidents or code reviews as opportunities to up-level your skills.AI-Native Development: You embrace the future of engineering. Experience or interest in using AI coding assistants (like Cursor or Claude) to improve productivity and automate boilerplate tasks.#LI-NP1#LI-REMOTECompensation DisclosureThe compensation offered for this position will depend on qualifications, experience, and work location. For candidates who receive an offer at the posted level, the starting base salary (for non-sales roles) or base salary + commission target (for sales/com-missioned roles) is expected to be the annual range listed below. The offered compensation may also include restricted stock units and a bonus. A description of our employee benefits may be found here.$126,000.00 - $204,500.00/yrOur Commitment We’re trailblazers that dream big, take risks, and challenge cybersecurity’s status quo. It’s simple: we can’t accomplish our mission without diverse teams innovating, together.We are committed to providing reasonable accommodations for all qualified individuals with a disability. If you require assistance or accommodation due to a disability or special need, please contact us at View email address on click.appcast.io Alto Networks is an equal opportunity employer. We celebrate diversity in our workplace, and all qualified applicants will receive consideration for employment without regard to age, ancestry, color, family or medical care leave, gender identity or expression, genetic information, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran status, race, religion, sex (including pregnancy), sexual orientation, or other legally protected characteristics.All your information will be kept confidential according to EEO guidelines.Is role eligible for Immigration Sponsorship?: YesSummaryLocation: San Francisco, United States of America; Denver, United States of America; Austin, United States of America; Jacksonville, United States of America; Bridgeport, United States of America; Seattle, United States of America; Boston, United States of America; New York City, United States of AmericaType: Full time
$147k - $237.5k
...containerized world. Chronosphere empowers customers to... ...Two Impact Areas:Our Infrastructure organization is scaling... ...-class Principal Engineers to drive the future of... ...Production Engineering (Reliability & Scale)The Core... ...developer velocity and software reliability for the entire...SuggestedFull timeRemote work- ...containerized world. Chronosphere empowers customers to... ...focuses on providing a reliable, fast and intuitive... ...to alerts.Our users—software developers—need to troubleshoot... ...seeking experienced engineers who can craft... ...tooling, and deployment infrastructure that improve...SeniorFull timeRemote workFlexible hours
$350k
...Join a rapidly growing AI infrastructure provider delivering large... ...providers to build reliable, high-performance platforms... ...This opportunity is for a Staff Site Reliability Engineer to lead the reliability of... ..., and Megatron Strong software engineering skills in Go,...SuggestedFull time$127k - $249k
...are looking for an experienced Senior or Staff Engineer for our SRE, InfraSec team, to guide the security of our cloud-based infrastructure. As a Staff SRE, you will be very hands-... ...transform, and disrupt industries with software. MongoDB’s unified data platform, the most...SeniorLocal areaRemote workWorldwideFlexible hours- DescriptionWe are looking for a Senior or Staff level Site Reliability Engineer to strengthen the reliability, scalability, and operational maturity... ..., debugging, and capacity planning.• Contribute to infrastructure and delivery workflows across AWS, Terraform, Ansible,...Senior
- Epoch Biodesign in San Francisco is seeking a Senior Staff Cloud Support Engineer to lead technical escalations and improve cloud infrastructure. You will mentor engineers and influence architectural decisions while ensuring high availability for AI workloads. The ideal...Senior
- ...world's best data and AI infrastructure platform so our customers... ...their business. Founded by engineers — and customer obsessed —... ...Observability and Reliability systems.As a Sr. Staff Production Engineer, you... ...production-level experience as a Software Engineer or SRE in highly...SeniorWorldwide
- United States Digital Space LLC is seeking a Senior Software Engineer, Infrastructure (Release Engineering) based in San Francisco. This role involves... ...processes. With a strong emphasis on performance and reliability, successful candidates will have at least 7 years of...SeniorWork at officeRemote work
- Pivotal Health in San Francisco is seeking a Senior Platform Engineer to design, scale, and harden the foundation powering our platform... ...with engineering teams to evolve cloud architecture, improve reliability and security, and ensure scalable, event-driven systems. You’...Senior
$184.7k - $324.8k
Staff / Sr. Machine Learning Engineer, AI, Search & Knowledge Platforms San Francisco... ...large-scale data and ML infrastructure. You will build and optimize... ...cost, throughput, and reliability including model serving... ...years of experience in software engineering or ML engineering...SeniorRelocation$127k - $249k
We are hiring an experienced Security Software Engineer (Staff or Senior) for our Infrastructure Security team to design and build scalable security controls... ...cloud infrastructure.The team sits within the Site Reliability Engineering organization and works with other...SeniorWork at officeLocal areaRemote workWorldwideFlexible hours$232k - $319k
...building the trusted, neutral infrastructure that enables organizations... ...service with great people and reliable, cost-effective, and... ...processes, and tooling. As the Sr. Manager of Infrastructure Platform... ...velocity of SRE and product engineering by developing robust...SeniorPermanent employmentLocal areaWorldwideFlexible hours$129.5k - $186.1k
...primary consultant to multiple product engineering teams and enterprise groups across UKG... ...of failure and designing self-healing infrastructure on modern cloud platforms, this role... ...,MA,United StatesTravel:Up to 25%Role:Sr. Staff Cloud Resilience Engineer - SecurityDepartment...Senior$223k - $278k
...the Role:We’re hiring seasoned engineers to join our teams that work... ...Gusto's Tax platform. As a Gusto Software Engineer at this level, you’... ...processes and build the infrastructure that keeps taxes accurate and on time. By combining reliable systems with thoughtful design...SeniorFull timeWork at officeLocal area2 days per week3 days per week- ...About the Team We’re hiring software engineers to join our broader Infrastructure organization, which supports multiple high-impact teams. Depending on your... ...with a shared mandate to raise the bar on safety, reliability, and velocity across OpenAI. About the Role You...Full time
$210k - $230k
...the Role:We're looking for a Senior Staff Security Engineer to lead Gusto's edge and network security... ...the security org, partnering with infrastructure and product teams to make high-impact... ...other than a Gusto office, a secure, reliable, and consistent internet connection is...SeniorFull timeWork at officeLocal areaRemote work2 days per week3 days per week$295k - $405.5k
...Wisconsin.About this roleAs the Senior Staff Machine Learning Platform Engineer, you will own the technical... ...and MLflowOptimize performance, reliability, and cost of the ML platformEvaluate... ...in distributed systems, ML infrastructure, and cloud architecture.Demonstrated...SeniorWork experience placementWork at officeLocal areaRemote workMonday to FridayFlexible hours3 days per week- Slope in San Francisco is looking for a reliability engineer focused on managing call completion for its Voice AI platform. You will be key in establishing incident management processes and improving system stability through effective monitoring and capacity planning....
$240k - $360k
About the RoleWe're looking for a Senior Staff Data Engineer to be the technical backbone of our... ...pipelines, serving layers, and the infrastructure that makes ML models production-ready... ...teams and engineers build on — balancing reliability, performance, cost, and long-term...SeniorWork at officeLocal areaImmediate startRemote workWorldwide3 days per week- ...NVIDIA DGX SuperPOD built on Grace Blackwell infrastructure — one of the fastest private supercomputers in... ...infrastructure that has to perform under real-world scale, reliability, and security demands — and we're looking for an engineer who wants to own the foundation it runs on....Senior
$183k - $200k
...We're looking for an experienced infrastructure engineer to help evolve how Gusto stores and accesses... ...through performance tuning and reliability at scale, and help stand up and operate... ...what we're looking for: ~6+ years of software engineering experience building and...SeniorFull timeWork at officeLocal areaRemote work2 days per week3 days per week$195k - $257.5k
...assets, payment applications, and programmable blockchain infrastructure. Circle’s platform includes the world’s largest... ...everyone is a stakeholder.What you’ll be responsible for:As a Staff Site Reliability Engineer on Circle’s Platform team, you’ll design, build, and...Flexible hours$194k - $267k
...potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era.... ...on new concepts and tools.Position Overview:The Site Reliability Engineer (SRE) will play a key role in building and managing...Permanent employmentWork at officeLocal areaWorldwideFlexible hours$194k - $267k
...potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era.... ...:We are seeking a highly technical StaffObservabilitySite Reliability Engineer with a specialty in Splunk to own and evolve our Splunk...Permanent employmentWork at officeLocal areaWorldwideFlexible hours$174k - $239k
...building the trusted, neutral infrastructure that enables organizations... ...functions to drive scale, reliability, and innovation through technology.The Staff Site Reliability Engineer OpportunityOkta Federal,... ...Python, while leveraging secure software development practices....Work experience placementLocal areaWorldwideFlexible hours$220k - $235k
....We are seeking a strategic, high-output Staff/Senior Staff SRE to define the future of our cloud platform and champion engineering excellence across Ironclad. In this role,... ...leadership and strategic direction for the Site Reliability Engineering team and our broader Cloud...Full timeContract workWork at office$150k - $250k
...As a founding member of our engineering team, you will have a direct... ...isn’t just “models.” It’s the software layer that turns: Orders →... ...those systems observable, reliable, and scalable. We are also... ...build the backend systems and infrastructure that power the factory of the...SeniorFull timeContract work- ...firms and enterprises. We’re seeking a Production Engineer to design, operate, and scale Harvey’s global compute and networking infrastructure, Kubernetes platform, and production foundations. You will improve reliability, capacity planning, automation, and observability...Senior
- DoorDash is hiring a Senior Software Engineer to lead the Spark Platform, setting the long-term direction for our in-house Spark deployment... ...spanning runtime, shuffle service, and scheduler, ensuring reliability at scale across data, analytics, and ML workloads. You will...Senior
- A leading IoT infrastructure provider in San Francisco is seeking a Senior Platform Engineer. In this role, you will own the systems that drive engineering productivity and reliability. You will focus on building TypeScript tooling, optimizing PostgreSQL, and ensuring...Senior
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Sr. Staff Engineer Software, Infrastructure Reliability (Chronosphere). Be the first to apply!
- assistant engineering manager San Francisco, CA
- assistant mechanical engineer San Francisco, CA
- assistant engineer San Francisco, CA
- staff engineer San Francisco, CA
- staff data engineer San Francisco, CA
- software engineer staff San Francisco, CA
- assistant electrical engineer San Francisco, CA
- assistant chief engineer San Francisco, CA
- staff design engineer San Francisco, CA
- senior staff engineer San Francisco, CA


