Principal Distributed Systems Engineer - Observability
$222.9k - $334.3kWorkday
Your work days are brighter here.We’re obsessed with making hard work pay off, for our people, our customers, and the world around us. As a Fortune 500 company and a leading AI platform for managing people, money, and agents, we’re shaping the future of work so teams can reach their potential and focus on what matters most. The minute you join, you’ll feel it. Not just in the products we build, but in how we show up for each other. Our culture is rooted in integrity, empathy, and shared enthusiasm. We’re in this together, tackling big challenges with bold ideas and genuine care. We look for curious minds and courageous collaborators who bring sun-drenched optimism and drive. Whether you're building smarter solutions, supporting customers, or creating a space where everyone belongs, you’ll do meaningful work with Workmates who’ve got your back. In return, we’ll give you the trust to take risks, the tools to grow, the skills to develop and the support of a company invested in you for the long haul. So, if you want to inspire a brighter work day for everyone, including yourself, you’ve found a match in Workday, and we hope to be a match for you too.About the TeamThe Data Platform and Observability Engineering (DPOE) team is building Workday's next-generation, multi-petabyte scale Observability Platform. We own the libraries, distributed services, and infrastructure that power ingestion, storage, and query across the observability stack — Iceberg, ClickHouse, Tempo, Grafana, S3, Kafka, and Elasticsearch — including LangSmith for LLM/agentic tracing and evaluation serving traces, metrics, and logs for every workload at Workday. Our roadmap directly shapes how the company detects, diagnoses, and eventually predicts operational issues at scale.About the RoleTo own the technical vision and architecture for distributed tracing as a first-class pillar of Workday's Observability Platform, built on ClickHouse and/or Grafana Tempo, backed by a big-data pipeline (Kafka, Spark/Flink, Iceberg, Clickhouse, Tempo,S3) running on AWS. This is a hands-on, high-autonomy role for an engineer who can design and build multi-petabyte, low-latency tracing infrastructure end-to-end — and who is equally excited to help define where Observability AI goes next: using traces, logs, and metrics as the substrate for automated root-cause analysis, anomaly detection, and AI-driven incident triage.You'll set technical direction across multiple teams, mentor senior and staff engineers, and act as the primary architect and escalation point for the tracing subsystem — from ingestion and storage design through query performance and platform reliability.Architect and build Workday's distributed tracing platform on ClickHouse/Tempo, designed for multi-petabyte scale ingestion and sub-second interactive query performance.Own the big-data pipeline feeding tracing data — Kafka-based ingestion, Spark/Flink stream and batch processing, and Iceberg-on-S3 storage — including schema design, partitioning, compaction, and lifecycle management.Drive performance and scaling across ingestion and query paths: storage format optimization (Parquet/Iceberg), compression strategy, partitioning/indexing, and query engine tuning under real production load.Lead HA/DR design for tracing services — multi-region/multi-AZ resilience, failover, backup/restore, and recovery time/point objectives appropriate to a tier-1 platform.Design security architecture for the platform, including authentication/authorization (authn/authz) for multi-tenant data access across ingestion and query layers.Own operational excellence for distributed tracing: monitoring, logging, alerting, capacity planning, and participation in an on-call rotation for the platform.Evaluate and introduce new technologies — open source and cloud-native — that materially improve the platform's scalability, cost efficiency, or capability.Shape the future of Observability AI: Extend distributed tracing to LLM and agentic workflows using LangSmith and LangChain, enabling observability into multi-step agent execution, tool calls, and prompt/response chains.Evangelize the platform: publish best practices, mentor engineers across DPOE and partner teams, and act as a technical thought leader for the modern observability/data stack internally.Operate with high autonomy in a fast-moving, ambiguous environment — setting technical direction with minimal oversight while aligning with broader platform strategy.About YouBasic Qualification 14+ years experience in software development engineering.6+ years experience specifically focused on designing, building, and operating complex distributed system architectures, evidenced by successful deployment of systems with high availability (e.g., 99.9% uptime) and fault tolerance.8+ years experience with at least two of the following programming languages (e.g., Java, Python, Go), including experience in writing production-level code for distributed systems.Bachelor’s degree in a relevant field such as Computer Science, Engineering, or a related discipline; a Master's degree (e.g., MS in Computer Science, Distributed Systems, or related field) is strongly preferred or equivalent practical experience.Other QualificationExpert-level ability in Algorithmic Thinking to architect highly efficient and scalable solutions for complexDeep expertise in API Development, including understanding of advanced API protocols or architectural patternsDeep understanding of Distributed Systems Software principles, like distributed consensus or fault tolerance mechanismsProven ability to design and implement High Availability strategies for critical distributed systemsExperience with LLM observability and tracing tools such as LangSmith, and familiarity with LangChain or similar agent orchestration frameworksExtensive experience with Large Scale Data Processing technologies and frameworksDeep understanding of Large Scale Systems design principles like distributed data management or scalability strategiesStrong understanding of System Security principles and best practices relevant to securing complex distributed environmentsProven ability to lead Team Collaboration within and across distributed software development teams and drive architectural directionStrong skills in creating Technical Writing Documentation and PresentationWorkday Pay Transparency StatementThe annualized base salary ranges for the primary location and any additional locations are listed below. Workday pay ranges vary based on work location. As a part of the total compensation package, this role may be eligible for the Workday Bonus Plan or a role-specific commission/bonus, as well as annual refresh stock grants. Recruiters can share more detail during the hiring process. Each candidate’s compensation offer will be based on multiple factors including, but not limited to, geography, experience, skills, job duties, and business need, among other things. For more information regarding Workday’s comprehensive benefits, please click here.Primary Location: USA.CA.PleasantonPrimary Location Base Pay Range: $222,900 USD - $334,300 USDAdditional US Location(s) Base Pay Range: $187,100 USD - $334,300 USDOur Approach to Flexible WorkWith Flex Work, we’re combining the best of both worlds: in-person time and remote. Our approach enables our teams to deepen connections, maintain a strong community, and do their best work. We know that flexibility can take shape in many ways, so rather than a number of required days in-office each week, we simply spend at least half (50%) of our time each quarter in the office or in the field with our customers, prospects, and partners (depending on role). This means you'll have the freedom to create a flexible schedule that caters to your business, team, and personal needs, while being intentional to make the most of time spent together. Those in our remote "home office" roles also have the opportunity to come together in our offices for important moments that matter.Pursuant to applicable Fair Chance law, Workday will consider for employment qualified applicants with arrest and conviction records.Workday is an Equal Opportunity Employer including individuals with disabilities and protected veterans.Workday is committed to providing reasonable accommodations for qualified individuals during our application process, in order to perform one or more essential functions of their job, as well as regarding the use of AI tools for employment decision-making to any degree. Please see below for more details including how to request an accommodation as a qualified veteran, due to a disability or for religious reasons, or as otherwise provided under applicable law. Workday prohibits taking adverse action against any candidate or employee for reporting a possible violation of this policy, requesting one or more work accommodations, exercising a privacy right, or cooperating in an investigation in accordance with applicable law. Any employee who retaliates against a candidate or employee for doing so may be subject to disciplinary action, up to and including termination of employment, to the fullest extent allowable under applicable law. If you require a reasonable accommodation, you may email View email address on us.fitly.work, as far in advance as possible.Are you being referred to one of our roles? If so, ask your connection at Workday about our Employee Referral process!At Workday, we value our candidates’ privacy and data security. Workday will never ask candidates to apply to jobs through websites that are not Workday Careers. Please be aware of sites that may ask for you to input your data in connection with a job posting that appears to be from Workday but is not.In addition, Workday will never ask candidates to pay a recruiting fee, or pay for consulting or coaching services, in order to apply for a job at Workday.SummaryLocation: USA, CA, PleasantonType: Full Time
$222.9k - $334.3k
....About the TeamThe Data Platform and Observability Engineering (DPOE) team is building Workday's next... ...Platform. We own the libraries, distributed services, and infrastructure that power... ...building, and operating complex distributed system architectures, evidenced by...PrincipalFull timeWork at officeRemote workHome officeFlexible hours$148k - $222k
...TeamThe Data Platform and Observability Engineering (DPOE) team is building Workday... .... We own the libraries, distributed services, and... ...expertise in distributed systems and big data while contributing... ...work closely with Senior and Principal engineers to level up your...SuggestedFull timeWork at officeRemote workHome officeFlexible hours$196k - $294k
...on Workday's trusted systems of record, deep finance... ..., and enterprise distribution to create autonomous... ...They work directly with engineering and go deep into customer... ...hiring Senior and Principal Product Managers to... ...evaluation harnesses, observability, human-in-the-loop controls...PrincipalFull timeWork at officeRemote workHome officeFlexible hours$119.04k - $192.42k
...is a water resource management and engineering firm focused exclusively on water.... ...OR office locations. ( Senior - Principal Engineer - Water Systems & Water Supply Planning The Senior... ...as a technical leader in water distribution system hydraulic evaluations utilizing...PrincipalTemporary workWork experience placementWork at officeLocal areaRemote workFlexible hoursNight shiftAfternoon shift$130.7k - $261.3k
...scientists.The OpportunityThe Principal Architect is responsible... ...across multiple engineering teams and products, ensuring... ...architecture for a highly distributed ecosystem consisting of microservices... ..., resilient, secure, and observable distributed systems supporting millions of...PrincipalWorldwide$168k - $252k
....About the TeamThe Data Platform and Observability Engineering (DPOE) team is building Workday's next... ...Platform. We own the libraries, distributed services, and infrastructure that power... ...Workday depends on us to see inside their systems — from a single service's latency...Full timeWork at officeRemote workHome officeFlexible hours$246k - $370k
...product leaders, AI engineers, and full-stack builders... ....About the RoleAs a Principal/Senior AI UX Lead in... ...: scalable, reliable systems that simplify complex... ...architecture of modern distributed systemsKnowledge of... ..., automated testing, observability)Hands-on experience with...PrincipalFull timeWork at officeRemote workHome officeFlexible hours- ...Reference24-148901TypeContractRemote100% RemoteJob - Scanning Systems Engineer Location - Remote Must Haves:Oversee planning, design,... ...processing centers from an IT perspective.· Experience with distributed capture and ad hoc scanning solutions· Experience working with...Remote work
$196k - $294k
...evaluation frameworks for agentic systems, collaborating directly with... ...-by-side with central AI engineering teams to power the future of... ...design patterns that enable distributed engineering and domain teams... ...platform products.For Principal Level: Minimum of 12 years of...PrincipalFull timeWork at officeRemote workHome officeFlexible hours- ...scalable frameworks, govern model development, and mentor teams while partnering with Business and Engineering leaders to deliver measurable AI value. The role emphasizes distributed training, large-model deployments, and enterprise-grade inference across cloud platforms,...Principal
- ...Job Description We are seeking an experienced Sr. Electrical Engineer to support the development of innovative drug-device... ...engineering activities for advanced drug delivery and medical device systems, with particular emphasis on connected and intelligent devices...PrincipalFull timeContract workLocal areaRemote work
$166k - $343k
Presales Systems Engineer - HPE Networking (Northern California)This role has been designated as ‘Remote/Teleworker’, which means you will... ...technologies with an emphasis on datacenter, campus and distributed branch networks. The Systems Engineer will consult with their...Full timeWork experience placementLocal areaImmediate startRemote workWork from home$175k - $195k
Windows Systems Engineer (MTS) Location: Hybrid - Pleasanton, CA Reporting to: Sr Manager – Service Desk, NOC/SOC, Systems & Network... ...available ~401(k) Plan ~ Unlimited Paid Time Off (PTO) ~ Observed Holidays Paid ~ Cell Phone Allowance ~ Collaborative,...Permanent employmentFull timeWork at officeNight shift$214k - $290k
...manageability of hyperscale AI servers. This Principal-level role owns BMC firmware... ...across EVT/DVT/PVT, debug at board and system level, and drive root-cause analysis and... ...development methodology, review code, and mentor engineers; represent firmware architecture to...PrincipalPermanent employmentFull time$142.5k - $253.1k
...Much as data and privacy engineers translate legal and policy requirements... ...into working technical systems, Responsible AI Systems... ...work closely with the Lead Principal Responsible AI Systems Engineer... ...increase efficiency, consistency, observability, and scalability. Support...Full timeWork at officeRemote workHome officeFlexible hours- ...execution-oriented product leader to serve as Principal Product Manager - Digital Product... ...with Product, Design, Analytics, Engineering, and business teams to quickly generate... ...1,732 pharmacies, 405 fuel centers, 22 distribution facilities, and 19 manufacturing plants...PrincipalMinimum wageLocal areaFlexible hours
$114k - $228k
...Databricks, AWS, Azure, modern data architectures, and data engineering practices with strong business acumen and leadership... ...solutions.Strong understanding of data architecture patterns, distributed systems, data modeling, and enterprise integration concepts.Experience...PrincipalWorldwide$60 - $78 per hour
...17Location: Pleasanton, CACompany: Planet Pharma GroupContact: ApplicationsEmail: ****@*****.*** Senior Systems Support Engineer provides service on instruments and networked systems used for Research, Development, and Clinical Operations activities located...Work experience placementRemote workRelocationFlexible hours- FormFactor, Inc. seeks a Sr Principal Technical Support Engineer in Livermore, CA, with strong expertise in probe card technology and wafer testing. You will serve as a key interface among customers, product engineering, and field service to ensure rapid issue resolution...PrincipalRemote job
$170k - $221k
...Senior Systems Engineer Hybrid- Fremont, CA Agility's commercially deployed humanoids operate alongside teams in warehouses, manufacturing facilities, and distribution centers—tackling physically demanding and repetitive tasks while enabling workers to focus on...Full timeTemporary workWork at officeRelocation packageFlexible hours$104.6k - $264.1k
...Customer Success Services (CSS) is seeking a Senior Principal Forward Deployed Engineer to lead the delivery of strategic AI and data transformation... ...architecture, configuration, optimization, tuning, distributed systems, and large-scale performance engineering.Demonstrated...PrincipalTemporary workFlexible hours- ...maintaining efficient clinical and administrative information systems; driving the liaising between clinical areas, ensuring alignment... ...business priorities and deadlines; coordinates, obtains and distributes resources. Removes obstacles that impact performance; guides performance...PrincipalTemporary work
$175.53k - $222.56k
...Reference #: REF8825KJob Code: SES.3 Science & Engineering MTS 3Organization: ComputingPosition... ...a Cloud and Application Infrastructure Systems Engineer. You will architect, implement... ...We DesireExperience with observability platforms such as Datadog, Dynatrace, AppDynamics...Minimum wageFull timeFor contractorsLocal areaWork from homeRelocationFlexible hours1 day per week$76k - $146.9k
...Laboratories is the nation's premier science and engineering lab for national security and... ...HVAC, Plumbing, and our Facility Control System (FCS). This team currently oversees,... ...maintenance for electrical high and low voltage distribution, fire protection systems, lighting,...PrincipalPart timeFor contractorsApprenticeshipWork at officeRemote workWork from homeWorldwideRelocation packageFlexible hours$137k - $287k
...ofYou will join the Reliability Engineering team within Infrastructure... ...to the availability of the systems Lam's engineering,... ...demand forecasting for globally distributed production environments across... ...production. Working depth in observability tooling (Prometheus, Grafana...Full timeLocal areaRemote workWorldwideFlexible hours2 days per week3 days per week1 day per week- ...method development and validation for PK and/or biomarker analysis is preferred . Functional experience utilizing LIMS and QMS systems for GLP/GCLP bioanalysis is preferred Familiarity with additional bioanalytical platforms (e.g. LBA, PCR, Flow Cytometry) is a plus...Principal
- The System Engineer II supports the reliability, security, availability, and continuous improvement of enterprise cloud infrastructure and... ...recommend optimization opportunities.Manage monitoring and observability tools such as Azure Monitor, Application Insights, and...
- Pleasanton, CaliforniaOnsiteContractTitle: Systems Engineer IIDescription: Description: Interview process: 1. phone 2 in-person interview Position Summary We are seeking a highly motivated Senior System Engineer to join our End-to-End Solution Integration Chapter supporting...Full timeTemporary workFlexible hours
$157.03k - $212k
...payment solutions include the issuance and distribution of gift cards, egifts, corporate payouts... ...need to make a real impact.Overview:The Principal Compensation Analyst will demonstrate an... ...requirements, project timelines, system requirements and configuration, data validation...PrincipalMinimum wageFull timeTemporary workWork experience placementWork at officeLocal areaRemote workFlexible hours- ...forGLP/GCLPsample testing and data reporting. Present and interpret data internally and/or externally as needed. Serve as the Principal Investigator responsible for interaction with the client from the studydesign to scheduling, conducting, reporting, and transferringdata...Principal
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Principal Distributed Systems Engineer - Observability. Be the first to apply!
- general engineer Pleasanton, CA
- chief engineer Pleasanton, CA
- principal developer Pleasanton, CA
- engineering director Pleasanton, CA
- senior civil engineer project manager Pleasanton, CA
- data center chief engineer Pleasanton, CA
- hotel chief engineer Pleasanton, CA
- principal engineer Pleasanton, CA
- director software engineering Pleasanton, CA
- healthcare systems engineer Pleasanton, CA



