Senior Systems Reliability Engineer
$109k - $150kThe Electric Reliability Council of Texas (ERCOT)
At ERCOT, our diverse and dynamic work environment provides a platform on which employees can work together to build the future of the Texas power grid and wholesale market utilizing the latest technologies and resources. We encourage you to join our talented, dedicated workforce to develop world-class solutions for today and tomorrow's energy challenges while learning new skills and growing your career.
ERCOT is committed to fostering inclusion at all levels of our company. It is the cornerstone of our corporate values of accountability, leadership, innovation, trust, and expertise. We know that individuals with a wide variety of talents, ideas, and experiences propel the innovation that drives our success. An inclusive and diverse workforce strengthens us and allows for a collaborative environment to solve the challenges that face our industry today and in the future. JOB SUMMARY The Senior Systems Reliability Engineer applies software engineering discipline to reliability problems - designing, building, and operating the systems that make production software measurable, scalable, and self-healing. This role treats operational challenges as engineering problems: when a process is manual, it gets automated; when a failure mode is unknown, it gets instrumented; when a system degrades, the degradation is understood before it recurs. At this level, the specialist owns SLO and error budget frameworks for assigned systems, architects the observability stack that the team relies on, leads engineering-driven incident response, and holds NERC/CIP compliance responsibility for assigned systems. This role partners directly with Software Engineers as a technical peer - participating in design reviews, influencing architecture decisions for reliability, and building the production readiness standards that govern how software ships. Advancement to Lead is based on demonstrated ability to define reliability engineering standards at the platform level, influencing practice across multiple teams and portfolios. JOB DUTIES- Performs complex reliability engineering work autonomously; recognized subject matter expert within the team and adjacent teams.
- Designs and builds production software systems, reliability tooling, and automation frameworks; treats operational problems as engineering problems to be solved through code.
- Owns SLO governance, error budget management, and observability architecture for assigned systems; leads engineering-driven incident response including failover scenarios.
- Holds NERC/CIP compliance responsibility for assigned systems; formally mentors less experienced specialists; may coordinate team delivery and on-call activities.
- Engineer reliability solutions: when a process is manual and repeatable, automate it; when a failure mode is opaque, instrument it; when a system is fragile, redesign the failure boundary.
- Define and own SLIs and SLOs for assigned systems; treat error budgets as a shared engineering contract with development teams, not an operations metric.
- Respond to production incidents as an engineer: form a hypothesis, isolate the failure, resolve it, and close the loop with a post-mortem that addresses root cause.
- Instrument systems so that on-call responders have sufficient telemetry to diagnose and act without tribal knowledge.
- Participate in 24/7 on-call rotation; treat every alert as signal - either actionable or worth eliminating.
- Write production-quality code: reliability tooling, automation frameworks, and operational software are held to the same engineering standards as application code.
- Partner with development teams as a peer in design reviews; reliability is designed in, not bolted on after deployment.
- Design, build, and maintain reliability tooling: automated remediation systems, self-healing infrastructure components, and operational software that reduces human intervention in production.
- Own SLO and error budget definitions for assigned systems; review error budget consumption with development teams and drive engineering decisions based on budget status.
- Architect and implement chaos engineering programs: define failure injection scenarios, automate resilience tests, and validate recovery behavior against defined SLOs.
- Build and maintain CI/CD reliability gates: automated canary analysis, progressive delivery validation, and rollback triggers based on SLI thresholds.
- Design capacity planning models for assigned systems; build tooling to project resource needs and surface capacity risks before they affect availability.
- Contribute to production readiness reviews: define and enforce the engineering criteria that a system must meet before it ships to production.
- Reduce operational toil through engineering: measure toil, track reduction targets, and build the automation that eliminates it.
- Architect MLTP (Metrics, Logs, Traces, Profiling) observability solutions using the Grafana LGTM stack (Loki, Grafana, Tempo, Mimir), Dynatrace APM, Splunk, and Datadog.
- Define and enforce instrumentation standards: structured logging schemas, metric naming conventions, trace context propagation, and continuous profiling configuration for assigned systems.
- Build distributed tracing coverage across service boundaries; identify and close observability gaps that produce blind spots during incidents.
- Design SLI instrumentation: translate user-facing reliability requirements into specific, measurable signals that accurately represent system health from the user's perspective.
- Build and maintain alerting frameworks: alerts must be actionable, calibrated to SLO burn rate, and free of noise; own alert quality as an engineering output.
- Correlate application performance data - JVM heap behavior, GC pressure, thread contention - with infrastructure events to enable root cause analysis across layers.
- Lead high-severity incident response for assigned systems, including dual-datacenter failover execution; own the technical resolution from detection through remediation.
- Apply structured root cause analysis: distinguish symptoms from causes, identify contributing factors across system layers, and drive remediation that addresses root cause rather than surface behavior.
- Author post-mortems that produce actionable engineering work items - not process improvements alone; track remediation to completion and validate effectiveness.
- Diagnose complex cross-layer failures: Java/JVM application failures, distributed system race conditions, database connection pool exhaustion, messaging system backpressure, and cross-datacenter synchronization issues.
- Build and maintain incident response runbooks as engineering artifacts: automated where feasible, version-controlled, and validated during chaos engineering exercises.
- Participate in blameless post-mortem facilitation; model the engineering culture that treats incidents as system failures, not human failures.
- Diagnose and resolve Java application performance problems in production: heap memory pressure, garbage collection tuning, thread pool exhaustion, connection leak detection, and class loading anomalies.
- Perform JVM performance analysis using heap dumps, thread dumps, and continuous profiling; translate findings into engineering recommendations for development teams.
- Instrument Spring Boot applications with production-grade observability: Micrometer metrics, structured logging with correlation IDs, and distributed trace integration.
- Diagnose failures across the Java application stack: Spring Boot service behavior, PostgreSQL and Oracle query performance, Kafka and ActiveMQ messaging reliability, and REST/SOAP API integration failures.
- Contribute to Java application design reviews with a reliability lens: identify failure modes, single points of failure, and observability gaps before code ships to production.
- Design and operate Kubernetes and OpenShift workloads for reliability: resource quotas, pod disruption budgets, horizontal pod autoscaling, and liveness and readiness probe engineering.
- Build infrastructure-as-code for reliability infrastructure: Terraform modules, Ansible/AAP playbooks, and Azure Resource Manager templates that are tested, version-controlled, and peer-reviewed.
- Own dual-datacenter reliability architecture for assigned systems: synchronization validation, automated failover triggering, traffic management, and recovery time objective verification.
- Design and automate environment promotion pipelines: ensure that configuration, secrets, and infrastructure state are consistent and validated across development, test, staging, and production.
- Build and maintain automated patch compliance workflows; integrate CVE remediation into CI/CD pipelines rather than treating it as a manual operational process.
- Own NERC/CIP compliance for assigned systems: interpret applicable reliability standards, implement required controls, maintain evidence documentation, and prepare for regulatory audit.
- Engineer compliance controls into the platform where possible: automated hardening scripts, configuration drift detection, access control validation, and audit log integrity verification.
- Maintain currency on applicable NERC/CIP standards and ERCOT-specific regulatory requirements; escalate emerging compliance risks to the Lead or Manager.
- Participate in regulatory audit preparation: produce control evidence, respond to auditor inquiries, and coordinate with compliance stakeholders on findings remediation.
- Hold formal mentoring responsibility for Systems Reliability Specialist I and II team members: structured coaching on SRE practices, code review for reliability tooling, and career development conversations.
- Serve as the recognized technical authority on reliability engineering and Java application operations for the team; adjacent teams and development engineers seek out this specialist for guidance.
- Lead design reviews for systems within the team's scope; identify reliability risks and observability gaps before systems reach production.
- Set engineering standards for the team: post-mortem quality, observability instrumentation, chaos engineering practices, and on-call readiness.
- Contribute to the broader engineering organization: internal technical talks, SRE practice documentation, and shared tooling that other teams can adopt.
- Minimum 5 years of progressive experience in systems reliability, software with an SRE focus, or a closely related discipline.
- Demonstrated experience building and operating reliability engineering systems in production: SLO frameworks, observability platforms, chaos engineering programs, and automated remediation tooling.
- Strong software engineering fundamentals: proficiency in Python and Java with experience writing production-quality reliability tooling and automation.
- Deep Java/Spring Boot and JVM performance expertise: heap analysis, GC tuning, thread profiling, and application instrumentation.
- Python, Bash, Linux operating system, CI Pipelines, Kubernetes. Open Telemetry, basic networking knowledge.
- Expert knowledge of MLTP observability tooling: Grafana LGTM stack (Loki, Grafana, Tempo, Mimir), Dynatrace, and Splunk.
- Experience with Kubernetes and OpenShift: workload design, autoscaling, pod reliability, and container networking troubleshooting.
- Experience with infrastructure-as-code: Terraform, Ansible/AAP, or equivalent.
- Experience leading high-severity incident response and driving blameless post-mortem programs.
- Experience with dual-datacenter or hybrid cloud reliability architecture preferred.
- NERC/CIP compliance experience: control implementation, audit preparation, and regulatory engagement preferred.
- Bachelor's Degree: Computer Science, Software Engineering, MIS, or related field (Required)
- Master's Degree: Computer Science, Software Engineering, or related field (Preferred)
- A combination of education and experience that provides equivalent knowledge to a major in such fields is required.
- Azure - Preferred
- Certified Kubernetes Administrator (CKA) - Preferred
- ITIL Foundation or Managing Professional - Preferred
$109,000 - $150,000
$84k - $135k
...Position Summary In your role as a SME Senior Electrical Engineer, you'll use your detailed knowledge... ...electrical troubleshooting, control systems and advanced programming skills.... ...SFC), ensuring robust, efficient, and reliable code. Testing and validation: Perform...SeniorHourly payFull timeWork experience placementImmediate start$50 - $60 per hour
...Duration: 12 months Job Summary Provides electrical engineering analysis and technical support for the interconnection of new... ...generation resources and future planning of the electric power system. Implements appropriate system modeling, develops tools and procedures...SuggestedHourly payTemporary workWork experience placement$131k - $174.5k
...are Innovating Today to Power the Devices of Tomorrow. Come innovate with us! Position Summary Staff-level reliability engineer will care semiconductor reliability of at Samsung Foundry's new technology introduction (NTI)/new product introduction (NPI)...SuggestedHourly payFull timeLocal areaImmediate startOverseas- Austin Industrial is seeking a Millwright Senior Technician in Taylor, TX to perform... ...maintenance supporting industrial waste treatment systems and facility utilities. The role leads... ..., maintenance strategy, system reliability, and continuous improvement with minimal...Senior
$60 - $90 per hour
...Provides support for Market Management Systems (MMS) applications portfolio such as Security... ...(SCED), Day-Ahead Market (DAM), Reliability Unit Commitment (RUC), Congestion Revenue... ...Member of the 24/7 Market Applications Engineering support on call team and supports,...SuggestedHourly payTemporary workWork experience placement- ...customer relationships, achieving sales goals, and supervising Team Members. The ideal candidate should possess expert knowledge of automotive systems and parts, with responsibilities also including dispatching drivers for timely deliveries. #J-18808-Ljbffr Advance Auto PartsSeniorFull time
- ...Description Job Description: Title: Senior Account Associate - Commercial Lines Work Mode: Remote: Eastern and Central... ...delinquent accounts, collecting outstanding balances. ~ System Maintenance: Maintain agency management systems and carrier/vendor...SeniorContract workFor contractorsRemote work
- ...Staff Electrical Engineer - Power System Protection JOB-10047166 Anticipated Start Date August 10, 2026 Location... ...analysis using industry-standard software to deliver accurate and reliable protection solutions. Job Description Develop protective...Full timeContract workLocal area
$70.48k - $150.76k
...Summary We are looking for experienced engineers to join HVAC Engineering Team that will... ...but not limited to: chilled water systems, cooling towers, process cooling systems... ...responsible for overall system quality, reliability, capacity and air permit compliance. Drives...Hourly payFull timeImmediate start- ...Instrument & Electrical Reliability Engineer JOB-10047109 Anticipated Start Date July 27, 2026 Location Big Spring... ...performance, and improvement of plant instrumentation and electrical systems. You will collaborate closely with operations, maintenance,...Full timeContract workLocal area
- ...Senior Project ManagerDoota Industrial America, a leading electrical/electronic manufacturing company, is seeking an experienced Senior... ...management and clients.RequirementsBachelor's degree in engineering, project management, or a related field.Minimum of 7 years of...Senior
- ERCOT seeks a senior leader to manage the Grid Implementation team, directing reliability analyses of upcoming grid changes and leading cross-functional coordination.... ...license preference, a Bachelor's in Electrical Engineering, and 8-10+ years of relevant experience,...SeniorRemote job
- Electric Reliability Council of Texas is seeking a Supply Chain Coordinator Sr in Taylor, TX to manage the flow of goods across the supply chain. You will coordinate shipments, negotiate with carriers, and ensure accurate inventory records while supporting receiving and...Senior
- Compunnel, Inc. is looking for a Project Manager to manage IT-related projects, ensuring adherence to project governance and delivering on goals. You will lead multiple medium to high complexity projects, collaborate with business owners and technical leads, and effectively...Senior
- Samsung Austin Semiconductor seeks a seasoned Contracts Manager to lead contractor agreements within the Infra Innovation Ops team. You will coordinate negotiations, maintain contract registers, track obligations, and ensure compliance across FS, Procurement, and vendors...SeniorContract workFor contractors
- ...Senior Electrical Engineer JOB-10047239 Anticipated Start Date August 10, 2026 Location... ...projects that improve operational reliability, safety, and efficiency. This role... ...supports maintenance, capital projects, PLC systems, power distribution, and facility...SeniorFull timeContract workLocal area
$80k - $100k
...Systems Protection Engineer – Technical Services JOB-10047348 Anticipated Start Date Sept. 1, 2026 Location Houston, TX... ...monitoring, and control. It is recognized for delivering reliable, high-performance solutions that ensure power system stability...Full timeLocal areaRemote work- ...Senior Scheduler Taylor, Texas, United States About the Job Senior Scheduler Role & Responsibility [Schedule Management... ...communicate orally and in writing Key Notes / Requirements - Engineering base preferred - Construction Manager experienced preferred -...SeniorContract workFor contractorsFor subcontractor
- ...achieve individual sales goals to support the store's sales and profit objectives, provide superior customer service, and take on other senior-level responsibilities within a store. Essential Functions (not all-inclusive): Generate sales to exceed personal...SeniorWork experience placementLocal area
$100 per hour
...Senior Controls Engineer - PLC/DCS/SCADA JOB-10047436 Anticipated Start Date September 9, 2026 Location Fremont, CA... ...startup, and troubleshooting of automated industrial control systems. The Engineer will have deep and broad expertise in process...SeniorFull timeContract workTemporary workLocal area- ...Description Description: Doota US is seeking a Design Electrical Engineer - Construction to support engineering and project coordination... ...Review and develop drawings related to: Power distribution systems Conduit and raceway routing Cable tray systems...For contractors
$99.11k - $168.45k
...future. Job Summary Provides engineering analysis and technical support to ensure continuing reliable operations of the electric power... ...of the electric power system. Implements appropriate system... ...Powerworld) under the direction of a senior level engineer or management...SeniorLocal area- ...Senior Electrical Engineer - Data Center JOB-10047615 Anticipated Start Date October 5, 2026 Location Overland Park, KS... ...Autodesk Revit Familiarity with Data Center tier rating systems Strong knowledge of Autodesk Revit Experience utilizing...SeniorFull timeContract workWork experience placementLocal areaHome officeFlexible hours
$18 - $21.5 per hour
...Code of Conduct. Performs duties as assigned by Pharmacy Manager, Staff Pharmacist and Store Manager including utilizing pharmacy systems to enter patient and drug information, ensuring information is entered correctly, filling prescriptions by retrieving, counting and...SeniorHourly payWork experience placementLocal areaImmediate startFlexible hoursAfternoon shift$38.62 - $42.62 per hour
...Job Type: Full Time Department: Engineering Division: Construction Inspection... ...Position Overview Title: Senior Construction Inspector Department: Public... ...Inspect various structures such as street systems, drainage facilities, water and sewer systems...SeniorHourly payFull timeContract workTemporary workFor contractorsFlexible hours- ...Industrial is seeking a Sr. Millwright. The Senior Facilities Technician performs advanced... ...supporting industrial waste treatment systems, facility utilities, chemical delivery... ...troubleshooting, maintenance strategy, system reliability, process improvement, and technical...SeniorPermanent employmentLocal areaImmediate startFlexible hoursShift work
- ...Job Description Job Description The Senior Estimator leads complex estimating and... ...evaluation, constructability review, value engineering, risk assessment, and schedule analysis... ...planning for complex electrical systems · Coordinate and lead subcontractor and...SeniorContract workFor subcontractorWork at officeLocal area
$60k - $100k
A Team Home Services in Hutto, TX is seeking a skilled Service Carpenter & Handyman with extensive residential carpentry, repair, maintenance, and installation experience to perform finish work and responsive repairs. This full-time, W-2 position offers commission-based...SeniorFull time$112k - $154k
...sponsors, steering committees, and senior leadership. Manages... ...dependencies across teams, systems, and delivery timelines to maintain... ...: Business, Management, Engineering, IT, or related field (... ...,000 - $154,000 The Electric Reliability Council of Texas (ERCOT)SeniorContract workWork experience placementLocal area- ...Join our team today and help make a difference in the lives of those we serve. Must have reliable transportation and flexible availability with 1 year of verifiable senior care experience. Benefits Flexible Schedule In-home and facility shifts available (vary...SeniorWeekly payLocal areaImmediate startFlexible hoursShift work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Systems Reliability Engineer. Be the first to apply!
- senior operations technician Taylor, TX
- senior Taylor, TX
- senior performance engineer Taylor, TX
- senior application security Taylor, TX
- srs Taylor, TX
- senior manager diversity & inclusion Taylor, TX
- remote senior project manager Taylor, TX
- systems installation engineer
- healthcare systems engineer
- systems engineer stf




