Principal Site Reliability Engineer
Eli Lilly
Lilly Technology Organization
At Lilly, the work is demanding because patients are waiting. We unite caring with discovery to help make life better for people around the world, knowing that every decision, every detail, and every day matters. Headquartered in Indianapolis, Indiana, our over 50,000 employees around the globe take on complex challenges to discover and deliver life-changing medicines, strengthen how health is understood and managed, and support the communities we serve. This is hard, urgent, selfless work—but it's work worth doing. If you're driven by purpose and ready to bring your best to work that truly matters for patients, we invite you to join us.
About the Technology Organization
Technology at Lilly builds and operates mission-critical digital products and platforms that support the discovery, development, and delivery of medicines that make life better for people around the world. Our teams operate in highly regulated, high-availability environments, where operational excellence, reliability, and quality are non-negotiable.
About the Team
Technology at Lilly builds and maintains capabilities using pioneering technologies like the most prominent tech companies. What differentiates Lilly IT is that we redefine what's possible through tech to advance our purpose, creating medicines that make life better for people around the world, including data-driven drug discovery, connected clinical trials, resilient enterprise platforms, and intelligent digital operations. We hire the best technology professionals from a variety of backgrounds, so they can bring an assortment of knowledge, skills, and diverse thinking to deliver creative solutions in every area of our business.
The Digital Core team leads Lilly's transformation into the Digital and AI era. They inspire digitally empowered teams to new ways of working and accelerate innovation and agility. This team powers and advances the entire company by building and maintaining world-class technology capabilities and platforms.
The Reliability Engineering team is the engineering-first function that owns the stability, observability, and operational quality of a multi-application production estate. It operates in close partnership with the engineering team that builds the agentic automation platform, and is in active transition from human-executed operations to engineering-led, agent-assisted reliability.
Role Summary
You build the self-healing automation, author the runbooks, and turn root-cause analysis into durable engineering fixes that let the production estate heal itself instead of paging a human. You work within the standards the Senior Principal SRE Engineer sets — SLOs, error budgets, observability — and you're the one who encodes them into working automation and documented procedure.
This is a hands-on individual-contributor role focused on execution and codification rather than cross-estate reliability strategy. You decide, in partnership with the Senior Principal SRE Engineer, which recurring patterns warrant a self-healing investment versus a documented manual runbook, and you build whichever is right. You are an individual contributor. You do not manage people. You partner daily with the Senior Principal SRE Engineer, the agentic automation engineering team, and Operations on validating outcomes. Success is measured by self-healing coverage, runbooks authored and adopted, reduced recurrence of known failure modes, and the safety record of every automation you sign off.
What You'll Be Doing
1) Self-healing automation & resilience patterns
- Design and build self-healing automation — circuit breakers, graceful degradation, automated remediation — for the failure modes that recur most across the estate.
- Run resilience or chaos testing to validate that self-healing patterns behave correctly before they're trusted in production.
- Continuously expand self-healing coverage as new failure modes are identified and proven safe to automate.
- Partner with the Senior Principal SRE Engineer on which failure modes justify self-healing investment versus a documented manual runbook.
2) Runbook authorship & validation
- Author and validate the remediation runbooks for the production estate: safe execution order, rollback steps, and exception handling for every documented fix.
- Keep the runbook library current as systems, dependencies, and failure modes evolve, retiring runbooks that no longer apply.
- Define and apply the graduation criteria that let a runbook move from human-executed to agent-assisted to autonomous.
3) RCA to durable fix
- Lead or contribute to root-cause analysis for significant incidents, and drive the blameless postmortem process to a durable engineering fix — not just a narrative.
- Convert recurring incident patterns into codified runbooks and, where appropriate, self-healing automation.
- Track fix effectiveness against recurrence, and escalate to the Senior Principal SRE Engineer when a fix needs broader engineering investment.
- Participate in high-severity incident response, including acting as incident commander for escalations within your area.
4) Partnership with agentic automation & operations
- Partner with the Agentic Automation Engineering team on which fixes are safe to hand off as agent-assisted remediations, and on the confidence thresholds and human-in-the-loop boundaries that keep them safe.
- Sign off on agent graduation criteria (accuracy over volume, zero P1/P2 caused) before an automation moves to a higher autonomy tier.
- Partner with Operations on outcome validation, feeding what's learned back into the runbook library and self-healing patterns.
5) Incident response & regulated-environment practice
- Ensure runbooks and self-healing automation meet Lilly's change-control, audit, and validated-environment standards.
- Document procedures so that audit evidence falls out of normal operation, not a special exercise.
- Mentor other reliability and automation engineers on runbook quality and self-healing design.
- Contribute proven patterns back to the broader reliability practice, in partnership with the Senior Principal SRE Engineer and Senior Architect.
How You Will Succeed
At the principal engineering level for reliability, success is defined by the durability and safety of what you build:
- Be recognized as the engineer who turns incidents into durable fixes, not repeat pages.
- Demonstrate measurable growth in self-healing coverage and runbook adoption, with falling recurrence of known failure modes.
- Maintain a clean safety record: automations you sign off don't cause P1/P2 incidents.
- Build runbooks and automation that make good practice the default, not a personal habit.
Your Basic Qualifications
- Bachelor's degree in Computer Science, Information Technology, or a related technical engineering discipline, including Software Engineering, Computer Engineering, Information Systems, Cybersecurity, Information Science, Network Engineering, Systems Engineering, Computer Information Systems (CIS), Management Information Systems (MIS), Cloud Computing, Data Science
- 5+ years of progressive engineering experience, with meaningful time as a Site Reliability Engineer, Production Engineer, or equivalent, including hands-on ownership of self-healing automation or runbook-driven remediation for a multi-application production estate.
- Production reliability experience in a regulated or audited environment (GxP, SOX, HIPAA, PCI, or equivalent), including familiarity with change-control discipline, audit evidence, and validated-system constraints.
- Hands-on experience authoring and validating runbooks: safe execution order, rollback steps, and exception handling for real remediation procedures.
- Demonstrated hands-on experience designing, implementing, and operating enterprise-scale SRE platforms, including observability solutions (Splunk, Datadog, New Relic, or Grafana/Prometheus), infrastructure-as-code with Terraform, CI/CD pipeline hardening, Kubernetes-based container platforms, and production workloads hosted on AWS, Azure, or GCP.
- Experience designing self-healing patterns (circuit breakers, graceful degradation, automated remediation) and validating them before they're trusted in production.
- Qualified applicants must be authorized to work in the United States on a full-time basis. Lilly will not provide support for or sponsor work authorization or visas for this role, including but not limited to F-1 CPT, F-1 OPT, F-1 STEM OPT,J-1, H-1B, TN, O-1, E-3, H-1B1, or L-1.
What You Should Bring:
- Hands-on experience designing self-healing automation and running chaos engineering or resilience-testing programs (AWS Fault Injection Service, Gremlin, LitmusChaos, or equivalent) tied to measurable reliability gains.
- Deep AWS fluency across reliability-relevant services (EKS, ECS, Lambda, CloudWatch, X-Ray, Systems Manager, Route 53),
$169.3k - $304.7k
...building and maintaining fast, efficient, scalable, and reliable routing software and infrastructure that is responsible... ...the growth and stability of our global platform. As a Principal Site Reliability Engineer - Network, you will be responsible for: Architecting...PrincipalWork experience placementWork at office$113.3k - $205.52k
...important to maintain our strong culture, achieve our goals, and thrive as #OneJamf. What you'll do at Jamf: As a Senior Site Reliability Engineer, you'll help us balance development velocity with the reliability our customers depend on. You'll partner with engineering...SuggestedWork at officeRemote workWorldwideFlexible hours$81.1k - $187k
...Job Description We are looking for a Site Reliability Engineer 3 to support mission-critical cloud services and production operations. The role focuses on improving service reliability, reducing operational risk, automating repetitive tasks, and driving faster detection...SuggestedTemporary workImmediate startFlexible hoursShift work$129k - $231k
...where operational excellence, reliability, and quality are non-... ...platforms. The Reliability Engineering team is the engineering-first... ...Role Summary: As the Senior Principal SRE Engineering Lead, you are... ...judgment of reliability at this site are yours. You are a senior...PrincipalContract workFlexible hoursShift work$152.07k - $202.76k
...AI‑ready connectivity, join us today. The Role The Principal Software Engineer serves as a senior technical leader responsible for advancing... ...applications. Define architectural standards for reliability, scalability, resiliency, observability, and operational...PrincipalFull timeTemporary workRemote work- ...responsibility, and professionalism. V2X is seeking a Senior Principal Systems Engineer to join our Engineering team in Indianapolis, IN... ...implementing performance requirements (i.e. -ilities) such as Reliability, Testability, Maintainability, Producibility, and System/...Principal
- ...Sr. Principal Engineer Our client is internationally recognized due to their commitment to and delivery of quality products to consumers... ..., robust process design, effective training, and reliable operations. • Drive knowledge transfer and development of...Principal
- ...The Principal Mechanical Design Engineer leads the design and integration of next-gen defense platforms, including armament systems and flight hardware. They are responsible for guiding cross-functional teams from concept through field deployment and generating technical...Principal
$296.3k - $423.9k
...future of transportation on a global scale.? We are looking for a Principal Technical Lead Manager (TLM) to lead the Trajectory Generation... ...real-world scenarios. You will lead a high-performing team of engineers building ML-driven trajectory generation systems, and drive...PrincipalFull timeLocal areaRemote workWork from homeRelocationRelocation packageFlexible hours$128.6k - $184.9k
...to keep their digital systems secure and reliable. Come help organizations be their best, while... ...team! We are looking for an experienced engineer to support the next generation of our... ...insurance. Please see the Cisco careers site to discover more benefits and perks. Employees...Permanent employmentFull timeTemporary workLocal areaRemote workFlexible hoursShift workNight shiftWeekend work$126.2k - $264.1k
Job Description Manage the development and implementation process of a specific company product. Responsibilities Manage the development and implementation process of a specific company product involving departmental or cross-functional teams focused on the delivery...PrincipalTemporary workFlexible hours- ...Principal, National Healthcare About the Company Technology consulting and engineering firm specializing in integrated technology systems across complex building environments Industry Information Technology and Services Type Privately Held About the...Principal
$140k - $200k
...– Speechify has no office. These include frontend and backend engineers, AI research scientists, and others from Amazon, Microsoft, and... ...→ testing → release → maintenance. Ensure quality, reliability, and consistency across releases. Identify, diagnose, and resolve...Remote jobFull timeWork at office- Join BSA as a Registered Architect Are you a Registered Architect passionate about creating inspired solutions that improve lives? BSA, a 100% employee-owned firm, is looking for talented individuals like you to join our team. Contribute to transformative healing, learning...PrincipalWork at officeFlexible hours
- ...security focused agencies within the Department of Defense and U.S. Intelligence Community. We are currently seeking a Software Engineer (Senior) to join our team. Position Summary The Software Engineer (Senior) is responsible for conducting and/or...Full timeWork at office
- ...challenges with integrity, respect, responsibility, and professionalism.To lead the charge, we seek an exceptional Principal Mechanical Design Engineer—someone ready to innovate, execute, and lead from the front. The Principal Mechanical Design Engineer is a senior technical...PrincipalWork experience placement
$51 - $61 per hour
...onsite at the project, significantly reducing and/or eliminating the demands to travel. Key Responsibilities:As a Release Train Engineer, you will be responsible for facilitating Agile Release Train events and processes including communicating with stakeholders...Hourly payLive inWork at officeLocal areaImmediate startFlexible hoursShift work$146k - $171k
...Location: Carmel, IN or Eagan, MN As a Principal Engineer – Planning R&D , you’ll help MISO evaluate the technologies, resources, and... ...not only what is changing, but what those changes mean for reliability, markets, planning, and the future of the power system—and...PrincipalFull timeLocal area$118.4k - $196.8k
...Less than 25% PHYSICAL, MENTAL DEMANDS and WORKING CONDITIONS Position Type Office-Based or Remote Position Physical work site required Frequently Disclaimer: The job description has been designed to indicate the general nature and essential duties and...PrincipalFull timeFor contractorsWork at officeLocal areaRemote work- ...Vice President, Principal Owner Development & Growth About the Company Top-tier mutual life insurance company Industry Financial Services Type Privately Held, Private Equity-backed Founded 1860 Employees 5001-10,000 Categories Financial...PrincipalHome office
- ...Workplace type: Hybrid As a Principal Mechanical Engineer , you will lead across a broad range... ...meet manufacturing, quality, safety, reliability, and cost goals. Considered Subject Matter... ...complex problems across multiple sites and implements sustainable solutions...PrincipalLocal areaWorldwide
$140k - $200k
...exponential growth. Overview We're looking for a Senior Software Engineer to join our Core Experiences Team. This team builds and... ...thinks strategically, and is passionate about designing clear, reliable APIs and simple systems that directly enhance the user experience...Full timeRemote work- ...Responsibilities And Requirements Our client seeks an experienced Principal Technical Artist to join their team and play a crucial role in... ...technical art team and collaborating closely with artists and engineers to develop innovative solutions that enhance our games' visual...Principal
$95k - $125k
...generated static pages to a live, interactive product including charts, maps, filtering, cross-chart interaction. A dedicated front-end engineer will own that work. Your job is to make sure the API gives them everything they need, and to be comfortable enough in the front-...Full timeContract workWork at officeFlexible hours- The Principal Scientist is responsible for upstream process development activities, including cell line development, process optimization... ...must hold a BS or MS degree in a relevant scientific or engineering field with 3 to 6+ years of industrial experience. Strong expertise...PrincipalFlexible hours
- ...seeking an Autonomous Systems Technical Lead with a passion for autonomous vehicle system integration and test to join our Systems Engineering team in Indianapolis, IN. This position involves creating advanced autonomous technologies as part of a fast moving, cross...Principal
$112.5k - $195.8k
Overview Sr Principal Scientist - Global Technical Services Molecule Steward - Dry Products... ...Development, CM&C teams, Manufacturing sites, and applicable functional areas to commercialize... ...integrate different disciplines such as engineering and analytical science on technical...PrincipalFull timeFlexible hours- ...resume and a cover letter explaining why you are interested in future opportunities with Mesh Systems as a Software or Senior Software Engineer. We are always looking for top talent! Your application will be reviewed, and we will contact you should a suitable position...Work at officeRemote workWork from homeFlexible hours
$132k - $186k
...needs. It's our driving force to help patients live longer and healthier lives. Join us and be part of our inspiring journey. The Principal Biostatistician will apply advanced biostatistical expertise to support clinical research and inform strategic decision-making...PrincipalFull timeRemote work- ...Principal Product Manager (Cloud Media Solutions)Contract Location: Indianapolis, IN (Remote... ...with business leaders, customers, and engineering teams. You will drive product... ...carrier interoperability, with a focus on reliability, scalability, and global reachOwn the strategy...PrincipalContract workRemote workWork from home
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Principal Site Reliability Engineer. Be the first to apply!
- senior chief engineer Indianapolis, IN
- general engineer Indianapolis, IN
- chief engineer Indianapolis, IN
- principal developer Indianapolis, IN
- senior principal engineer Indianapolis, IN
- engineering director Indianapolis, IN
- senior civil engineer project manager Indianapolis, IN
- senior director engineering Indianapolis, IN
- data center chief engineer Indianapolis, IN
- hotel chief engineer Indianapolis, IN




