SRE
LTM
Role description Location:Plano, TX Key Responsibilities
- Design and maintain highly available scalable and faulttolerant systems
- Implement and manage monitoring ing and observability tools Grafana Prometheus etc
- Participate in incident management RCA and postmortems
- Automate operational tasks using scripting and tools Python Bash etc
- Collaborate with development infrastructure and support teams to improve system reliability
- Drive adoption of SRE practices like SLIs SLOs and error budgets
- Ensure performance optimisation capacity planning and system stability
- Build and maintain automation pipelines and infrastructure as code
- Strong experience in SRE Production Support DevOps environment
- Handson with LinuxUnix systems
- Experience with cloud platforms AWS Azure GCP
- Expertise in Docker Kubernetes
- Knowledge of monitoring tools Grafana Prometheus OpenSearch Instana
- Strong troubleshooting incident management skills
- Experience with Infrastructure as Code Terraform Ansible
- Programmingscripting using Python Bash PowerShell
- Knowledge of Kafka databases SQLNoSQL
- Understanding of ITIL processes and SRE concepts Toil SLOs
- Exposure to CICD tools Jenkins
- Experience with multicloud or hybrid environments
Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the SRE in Plano, TX vacancy
$155.4k - $261.1k
...incidentsStrong communication skills to explain complex issues clearlyRequired:7+ years in Systems Engineering, ITSM, RM/CMBackground in SRE, Support or QAOne or more of the following SRE Tools: T-APM, T-Trace, CatchPoint, GrafanaHands-on experience and understanding of...SuggestedPermanent employmentFull timeTemporary workWork at officeLocal areaRelocationShift work- We are seeking a Delivery SRE leader who will ensure security applications are delivered with strong SDLC discipline and measurable reliability. This role partners closely with Product Owners and engineering leadership to challenge assumptions, sharpen the Definition of...Suggested
- Job Title SRE Lead Location Plano, TX - Onsite Job Description We are seeking an experienced 13 to 18 years of experience to join our team. The ideal candidate will have expertise in AWS, SRE, and Datadog, and a background in the automotive industry is a plus. This hybrid...SuggestedDay shift
- Job DescriptionYou belong to the top echelon of talent in your field. At JPMorganChase, infrastructure isn't just a foundation — it's a competitive advantage. This is your opportunity to bring deep storage expertise to a team that operates at global scale, where your contributions...SuggestedShift work
- JPMorganChase & Co. is seeking a Delivery SRE leader to ensure security applications are delivered with strong SDLC discipline and measurable reliability. The role partners with Product Owners and engineering leadership to bake SRE requirements into design and build phases...Suggested
$69.99 per hour
...Job Description Job Description Job Title: Site Reliability Engineer (SRE) Location: Plano, TX / Pennington, NJ / Charlotte, NC Duration: Contract - 12 months Pay Range: $69.99/hr (W2) Job ID: 409158 About BCforward BCforward is a leading global...Contract work- ...Python Bash etc Collaborate with development infrastructure and support teams to improve system reliability Drive adoption of SRE practices like SLIs SLOs and error budgets Ensure performance optimisation capacity planning and system stability Build and maintain...Permanent employmentTemporary workWork experience placement
- ...SRE Production Support Engineer Location: Plano, TX Duration: 6 Months (Contract to hire) Interview Process: 1st round - Zoom 2nd round – In Person Role Overview: Position is part of the Central Site Reliability Engineering (SRE) Team. Looking for...Contract workShift work
$197.3k - $225.1k
...Overview Lead Software Engineer, Full Stack (Java, Python, SRE, AWS, AI) (Cloud Operations Resilience Engineering) Do you love building and pioneering in the technology space? Do you enjoy solving complex business problems in a fast-paced, collaborative, inclusive...Full timePart timeInternshipH1bLocal area- ...Site Reliability Engineer (Sre) Our client, an IT Services and Consulting company, is looking for a Site Reliability Engineer (SRE) for their Plano, TX/Atlanta, GA/Middletown, NJ location. Responsibilities: Design and execute performance, load, stress, failover...
- ...and tools Python Bash etcCollaborate with development infrastructure and support teams to improve system reliabilityDrive adoption of SRE practices like SLIs SLOs and error budgetsEnsure performance optimisation capacity planning and system stabilityBuild and maintain...Full time
- ...knowledge of Linux administrationExperience in production support troubleshooting and incident managementGood to HaveExperience with SRE practices observability SLAs reliability engineeringExposure to monitoring tools Prometheus Grafana etcKnowledge of data replication...
$106k - $142k
...Experience with software application testing approaches and tools like jUnit, Cucumber, Selenium.- Experience with Chaos Testing and other SRE practices for building highly resilient cloud-based software applications- Experience designing and developing applications in...Full timeTemporary workWork at officeRemote workFlexible hours- ...Center, Major Incident Management, Monitoring & Observability, Event Management, and the evolution of Site Reliability Engineering (SRE). The Director will also manage the third-party managed services provider responsible for delivering 24x7 Digital Operations capabilities...Full time
- ...recommendations before use, escalating when uncertain and following data handling expectations.Must have a background in development or SRE Advanced expertise in stakeholder management, with the ability to establish productive working relationships and influence decision-...
- ...Python, Java/Spring Boot, .NetDemonstrated experience using enterprise-authorized AI capabilities within the work environment to improve SRE workflows (e.g., incident investigation support and knowledge capture) with strong validation habits and awareness of data...
$96.8k - $145.2k
...infrastructure, and CI/CD layersDrive continuous reliability and performance improvementsLeverage or support adoption of AI/automation in SRE workflows (alerting, incident triage, or productivity tools like Copilot)Basic Qualifications:5+ years of Hands-on experience with...Full timeTemporary workWork at officeRemote workFlexible hours- ...Troubleshoot and resolve application and production issues. Manage and elevate incidents, problems, and changes. Collaborate with Development, SRE, Infrastructure, AIOps, and business teams. Support trading, client onboarding, and operational applications. Communicate issues and...Work at office
- ...Maintain and improve monitoring, alerting, and observability for owned applications (logging, tracing, dashboards). Coordinate with SRE, platform, and DevOps teams on deployment pipelines, on-call schedules, and incident response. Work closely with business stakeholders...Work experience placement
- ...Ansible, CloudFormation, ARM) for networking configurations, and contribute to GitOps and pipeline automation workflows. Collaborate with SRE, Platform Architecture, Security, and Product Engineering teams to enforce compliance, improve network security posture, and support...Work at officeLocal areaImmediate startRemote work3 days per week
- ...environment. Operational Frameworks: Advanced understanding and application of ITIL v4 principles, FinOps, and Site Reliability Engineering (SRE) concepts. Automation/IaC: Strong practical experience with: IaC: Terraform, GCP instance Templates, Azure ARM/Bicep, CloudFormation....
- ...practicesCreating of APIs and Dashboards to determine and report data on environment health and application quality.Partnership with SRE team to enable and improve automation with tools and processes in non-prod environments.Interact and communicate with technical and non...Contract workWork experience placementRemote work3 days per week
- ...Description Job Description Responsibilities: \t3-4 years of experience in production engineering and site reliability engineering (SRE) to design, implement, and maintain highly available, scalable, and resilient systems. \tOwn end-to-end operational...
- ...ART model, including PI Planning and execution. Experience supporting production readiness and operational handoffs with Operations/SRE. A hybrid schedule requirement of a minimum of 3 days per week onsite. Technical Skills: Strong working knowledge of Jira and Jira Align...For contractorsWork experience placement3 days per week
- ...Architectures Cloud Engineering & Infrastructure as Code (IaC) DevOps, CI/CD & Engineering Productivity Site Reliability Engineering (SRE) & Operational Excellence Monitoring, Observability & Platform Reliability Security, Governance & Engineering Controls Why Should You...Work experience placement
$8,896.99 per month
...GitOps pipelines using Cloud Build/GitHub Actions/Artifact Registry and integrate IaC with policy-as-code. Establish observability and SRE practices: Cloud Monitoring, Logging, Trace, Error Reporting, SLOs/SLIs, incident runbooks, and game days. Define cost governance and...Work experience placement- Technology Operation Center (TOC) - Production Support And Site Reliability | DevOps Engineer SRE engineers will be engaged in automation and development work, reducing toil, developing self-service capabilities, automating manual tasks and develop support tools, utilizing...
- ...limits, N+1 mitigation (DataLoader). Proven delivery of API/schema governance, versioning/deprecation, and CI policy gates. Strong SRE practices: SLIs/SLOs, error budgets, OpenTelemetry, data-driven post-incident improvements. Developer productivity: time-to-first-hello...Contract work
- ...Architecture Cloud Engineering & Cloud-Native Architecture Infrastructure as Code (IaC) DevOps & CI/CD Site Reliability Engineering (SRE) Monitoring & Observability Platform Reliability & Operational Excellence Developer Experience (DevEx) Self-Service Platforms APIs, SDKs...Contract workTemporary work2 days per week1 day per week
$125.3k - $167.9k
...monitoring capabilities outlined in the application and system monitoring designs put forward by the Senior Site Reliability Engineer (SRE)Develops and maintains reliability scripts, tools and libraries and leverages them for common instrumentation, automation, and...Full timeWork at officeFlexible hoursDay shift
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to SRE. Be the first to apply!



