SRE
3B Staffing LLC
Title: ECO Event Management Integrate & Sustain: Site Reliability
Position Type: Contract
Location: Remote, United States
Use modern system monitoring tools to improve VA enterprise reliability and improve the quality of services provided to veterans.
Work with system and application owners to obtain existing design and functionality, leverage comprehension of workflow systems and applications processes within multiple system environments and work across technology and development teams to diagnose outages and recommend changes to increase reliability.
Use your hardware and software experience to help strengthen the systems the VA relies on. Your primary focus will be investigation, working with event management, application owners, DevOps teams, and system and network administrators to examine issues across enterprise applications and technology stacks.
Partner with system and application owners to understand their platform designs and how they operate across different environments. This insight will help you diagnose outages, trace workflow issues, and recommend changes that enhance stability.
Collaborate with developers and identity and access teams when deeper technical investigations are needed.
You'll gain hands-on experience with enterprise-level triage and incident analysis, which will deepen your understanding of the VA's infrastructure. Tools like SolarWinds, Dynatrace, and Splunk will be part of your daily workflow, giving you the visibility needed to identify reliability concerns and support improvements to the services delivered to veterans.
Must Have:
Deep expertise (3+ years) in two or more of the following tools used for troubleshooting application logging in an enterprise environment (Dynatrace, Splunk, SolarWinds, ServiceNow Operator Workspace)
Extensive experience in one or more Technology Areas (Network, Windows, Desktop, Unix/Linux, AWS or Azure Cloud, WebSphere Middleware, Java/JS Development, Microsoft or Oracle Database)
8+ years of experience working with key indicators for IT system operability, reliability, application performance and code quality
8+ years of experience deploying, maintaining and troubleshooting complex applications at an enterprise scale while working with cross-functional teams
1+ years of experience in service virtualization, AWS or Azure Cloud technologies, and SaaS and PaaS implementation.
Experience with using Microsoft Office, including Word, Excel, and PowerPoint
2+ years independently leading a team to solve difficult technical challenges
HS diploma or GED and 20+ years of relevant professional experience or MA or MS degree in computer science, electronics engineering, or other engineering or technical discipline with 10+ years of relevant professional experience Nice to Have:
Experience with test-driven development, distributed systems, microservices and cloud-native application implementation
Experience with the following tools: Oracle Enterprise Manager, Riverbed - Aternity, and ServiceNow VTBs
Possession of excellent written and verbal communication skills
Possession of strong critical thinking and error assessment capabilities
Virtual team management
Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the SRE in Murphy, TX vacancy
- ...GEICO is seeking an experienced Senior SRE Software Engineer to build and operate enterprise-grade platforms for incident management. You will design automation, dashboards, and data pipelines to support on-call, paging, and troubleshooting across high-availability services...Suggested
- ...GEICO is seeking a Senior SRE Software Engineer to design, build, and operate high-performance distributed platforms with zero-downtime reliability. You will own incident management tooling, lead on-call operations, and help scale automation across complex systems....Suggested
- ...pipelines across incident response, on-call, and runbooks. This role demands deep technical depth, strong leadership, and collaboration with SRE, platform, product, security, and business partners to raise reliability and reduce incidents. #J-18808-Ljbffr Jobleads-USSuggested
- ...GEICO is seeking a Staff Engineer in SRE to design, build, and operate high-performance, zero-downtime distributed platforms. You will lead incident management tooling, craft runbooks, and improve detection, troubleshooting, and recovery across critical services. The...Suggested
- JPMorgan Chase & Co. is recruiting a Lead Site Reliability Engineer to define the future of reliability on the AI ML and Data platform. You will lead a team, drive resiliency reviews, and mentor engineers while advocating for secure, observable, scalable systems across...Suggested
- ...designs, develops, and maintains scalable full-stack enterprise software solutions with user-centric interfaces. The team implements SRE principles to automate workflows, monitor reliability via SLIs/SLOs, and manage cloud infrastructure. Requires a bachelor’s degree or...
- ...collaborating with J.P. Morgan to connect them with exceptional professionals for this role. JOB DESCRIPTION We are seeking a Delivery SRE leader who will ensure security applications are delivered with strong SDLC discipline and measurable reliability. This role partners...
$68.31 per hour
...Job Description Job Description Job Title: Senior Linux Platform / SRE Infrastructure Engineer Location: Plano, TX or Jersey City, NJ or Charlotte, NC Duration: Contract - 11 months Pay Range: $68.31/hr (W2) Job ID: 410801 Work Arrangement: Hybrid...Contract work3 days per week- ...Develop unit, integration, and end-to-end tests and participate in code reviews. Collaborate with Infrastructure, Platform, and SRE teams on scheduling, resource allocation, quotas, and placement policies. Document architecture, APIs, operational procedures, and...
- JPMorgan Chase & Co. is seeking a Lead Software Engineer within the Consumer and Community Banking technology -Deposits Platform to lead resiliency and DR efforts across the product portfolio. The role involves cross-team collaboration, development of failover frameworks...
- ...SRE Production Support Engineer Location: Plano, TX Duration: 6 Months (Contract to hire) Interview Process: 1st round - Zoom 2nd round – In Person Role Overview: Position is part of the Central Site Reliability Engineering (SRE) Team. Looking...Contract workShift work
- Job Title Responsibilities Talents with Big data expert with 5+ years' experience in Hadoop Big data Required Skill Sets Working knowledge of AWS Kubernetes/EKS Middleware technologies Jenkins, JIRA, Ansible and expertise on DevOps tools Istio...
- ...SRE MAHIN-JOB-31492 Location: Plano TX Skill: Web application designing basics-1 8+ years of professional experience processing a culture of learning through the development and sharing of skills, knowledge, process and tools 2. A driving passion for...
$229.9k - $262.4k
...Overview Full-Stack Engineer 5 (Java, Python, SRE, AWS, AI) (Cloud Operations Resilience Engineering) Do you love building and pioneering in the technology space? Do you enjoy solving complex business problems in a fast-paced, collaborative, inclusive and iterative...Full timePart timeInternshipLocal area- Job Title Mandatory skills: Java, Perl, Python 8 Years: Relevant Experience Experience utilizing Java, Perl, Python, Go and scripting experience in Shell and Perl to automate reports and monitor enterprise business transactions under finance or billing domain for mobile...
- ...Role : SRE Engineer Location : Plano, Texas Job Summary: We are looking for a highly motivated Site Reliability Engineer (SRE) to improve system reliability, scalability, and performance of mission-critical applications. The ideal candidate should have strong...
$87.5k - $125k
Observability Engineer DISH is transforming the future of connectivity. We're doing it by building the country's first virtualized, standalone 5G wireless network from scratch. The foundation of a connected world, it's a network free of the limitations of the past, ...Flexible hoursNight shift- Duration: Long Term ContractPay Rate: $40/Hr. W2Experience: 3-5 YearsOverviewWe are seeking a remote Junior SRE/DevOps Engineer role. The ideal candidate has foundational knowledge of Site Reliability Engineering (SRE) and Kubernetes, and is enthusiastic about growing in...Remote work
$40 per hour
A technology solutions provider is seeking a remote Junior SRE/DevOps Engineer. The ideal candidate should have foundational knowledge of Site Reliability Engineering (SRE) and Kubernetes. Responsibilities include gaining experience in a DevOps-driven environment. Applicants...Long term contractInternshipRemote work- ...architecture patterns, DevOps practices, and migration execution. Cross-Functional Leadership Serve as a trusted advisor to engineering, SRE, product, and executive leadership. Partner with AWS teams and vendors to drive best practices and architectural decisions....Permanent employmentContract workLocal area
$100k - $215k
...incident management processes and builds platforms that allow GEICO to manage and recover from incidents.GEICO is seeking an experienced SRE Software Engineer with a passion for building, operating and troubleshooting high-performance, low-maintenance, zero-downtime complex...Hourly payWork experience placementLocal area- ...management processes and builds platforms that allow GEICO to manage and recover from incidents. GEICO is seeking an experienced SRE Software Engineer with a passion for building, operating and troubleshooting high-performance, low-maintenance, zero-downtime complex...Local area
- ...integration, and end-to-end tests for control plane components; participate in code reviews Collaborate with infrastructure, platform, and SRE teams to define scheduling policies, resource quotas, and placement constraints Document architecture decisions, APIs, and...
$147.25k - $190k
...storage, load balancing, DNS, certificates/TLS), managing scope, dependencies, risk, and status. Partner with security, IAM, network, SRE/operations, and vendors to architect and implement scalable, resilient solutions and platform modernization. Automate and...Full timeWork at office$229.9k - $262.4k
...incident response efforts and ensure root‑cause analysis drives durable reliability improvements Partner with Cloud, Security, and SRE teams to define cross‑platform observability, telemetry, and self‑healing automation Mentor and upskill engineers across the...Full timePart timeH1bLocal area- ...Cloud-Native Architecture & Platform Modernization* DevOps, CI/CD & Engineering Productivity Platforms* Site Reliability Engineering (SRE) & Operational Excellence* Security, Governance, Compliance & Engineering Controls* LLM Workflows, Prompt Engineering & Context...Hourly payContract workWork experience placement
- ...limits, N+1 mitigation (DataLoader). Proven delivery of API/schema governance, versioning/deprecation, and CI policy gates. Strong SRE practices: SLIs/SLOs, error budgets, OpenTelemetry, data-driven post-incident improvements. Developer productivity: time-to-first-hello...Contract work
$54.68 - $64.68 per hour
...OpenSSH,Oracle Enterprise Linux,performance tuning,performance optimization,postfix,Python,system reliability,scripting,sendmail,SMTP,SRE Practices,vulnerability remediation,high availability,TCP/IP,analytical,communication,organizational skills,Leadership,Troubleshoot,...Hourly payContract workTemporary workWork experience placement$150k - $190k
...modeling, pipeline reliability, observability, and performance optimization at enterprise scale Partner with product, backend, cloud/SRE, and AI/ML teams to ensure cohesive platform evolution Guide senior engineers, promote reusable patterns, and foster a culture of...Permanent employmentContract work- ...reliability, and scalability; follow Agile practices such as Scrum and Continuous Delivery. Support Site Reliability Engineering (SRE) practices to ensure excellent user experience and system performance. Required qualifications, capabilities, and skills: Formal...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to SRE. Be the first to apply!


