SRE
3B Staffing LLC
Title: ECO Event Management Integrate & Sustain: Site Reliability
Position Type: Contract
Location: Remote, United States
Use modern system monitoring tools to improve VA enterprise reliability and improve the quality of services provided to veterans.
Work with system and application owners to obtain existing design and functionality, leverage comprehension of workflow systems and applications processes within multiple system environments and work across technology and development teams to diagnose outages and recommend changes to increase reliability.
Use your hardware and software experience to help strengthen the systems the VA relies on. Your primary focus will be investigation, working with event management, application owners, DevOps teams, and system and network administrators to examine issues across enterprise applications and technology stacks.
Partner with system and application owners to understand their platform designs and how they operate across different environments. This insight will help you diagnose outages, trace workflow issues, and recommend changes that enhance stability.
Collaborate with developers and identity and access teams when deeper technical investigations are needed.
You'll gain hands-on experience with enterprise-level triage and incident analysis, which will deepen your understanding of the VA's infrastructure. Tools like SolarWinds, Dynatrace, and Splunk will be part of your daily workflow, giving you the visibility needed to identify reliability concerns and support improvements to the services delivered to veterans.
Must Have:
Deep expertise (3+ years) in two or more of the following tools used for troubleshooting application logging in an enterprise environment (Dynatrace, Splunk, SolarWinds, ServiceNow Operator Workspace)
Extensive experience in one or more Technology Areas (Network, Windows, Desktop, Unix/Linux, AWS or Azure Cloud, WebSphere Middleware, Java/JS Development, Microsoft or Oracle Database)
8+ years of experience working with key indicators for IT system operability, reliability, application performance and code quality
8+ years of experience deploying, maintaining and troubleshooting complex applications at an enterprise scale while working with cross-functional teams
1+ years of experience in service virtualization, AWS or Azure Cloud technologies, and SaaS and PaaS implementation.
Experience with using Microsoft Office, including Word, Excel, and PowerPoint
2+ years independently leading a team to solve difficult technical challenges
HS diploma or GED and 20+ years of relevant professional experience or MA or MS degree in computer science, electronics engineering, or other engineering or technical discipline with 10+ years of relevant professional experience Nice to Have:
Experience with test-driven development, distributed systems, microservices and cloud-native application implementation
Experience with the following tools: Oracle Enterprise Manager, Riverbed - Aternity, and ServiceNow VTBs
Possession of excellent written and verbal communication skills
Possession of strong critical thinking and error assessment capabilities
Virtual team management
Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the SRE in Murphy, TX vacancy
$56 - $80 per hour
AWS SRE/Platform EngineerContract to HireAddison, TX (Hybrid)ROLE OVERVIEW Seeking a highly experienced Technical Lead - SRE and Platform Engineering to provide technical leadership for the reliability, performance, security, observability, and operational management of...Suggested- Job Title SRE Lead Location Plano, TX - Onsite Job Description We are seeking an experienced 13 to 18 years of experience to join our team. The ideal candidate will have expertise in AWS, SRE, and Datadog, and a background in the automotive industry is a plus. This hybrid...SuggestedDay shift
- We have an opportunity to impact your career and provide an adventure where you can push the limits of what's possible.As a Software Engineer III at JPMorgan Chase, you serve as a seasoned member of an agile team to design and deliver trusted market-leading technology products...Suggested
- ...Engineering (SPE) team within the Toyota Financial Services CTO PM organization is seeking a Senior Product Manager to own and transform the SRE & Observability product area with an expanded charter to elevate the space and shape broader AIOps strategy for TFS Team Members. You...SuggestedH1bShift work
- Duration: Long Term Contract Pay Rate: $40/Hr. W2 Experience: 3-5 Years Overview We are seeking a remote Junior SRE/DevOps Engineer role. The ideal candidate has foundational knowledge of Site Reliability Engineering (SRE) and Kubernetes, and is enthusiastic about growing...SuggestedLong term contractContract workInternshipRemote work
- Toyota Financial Services seeks a Senior Product Manager to own the SRE & Observability space, driving a product-driven model with reliability outcomes and business impact. You will lead roadmap prioritization, VOC synthesis, acceptance criteria, and metrics across SRE,...
- T-Mobile USA, Inc. is seeking a senior DevOps/SRE leader to drive globally distributed teams (US and India) with end-to-end accountability for workforce planning, performance, and development of top talent. You will oversee reliability, scalability, and cost efficiency...
$155.4k - $261.1k
...communication skills to explain complex issues clearly Required: ~7+ years in Systems Engineering, ITSM, RM/CM ~ Background in SRE, Support or QA ~ One or more of the following SRE Tools: T-APM, T-Trace, CatchPoint, Grafana ~ Hands-on experience and...Permanent employmentFull timeTemporary workWork at officeLocal areaRelocationShift work$40 per hour
A technology solutions provider is seeking a remote Junior SRE/DevOps Engineer. The ideal candidate should have foundational knowledge of Site Reliability Engineering (SRE) and Kubernetes. Responsibilities include gaining experience in a DevOps-driven environment. Applicants...Remote jobLong term contractInternship- ...Window OS, SQL/Oracle DB Unix/Linux. ~ Excellent knowledge of Identity, Authentication and Access Management (IAM) domain including SRE and DevOps space. ~ Must have senior level production support experience and troubleshooting skills in IAM technologies. ~...3 days per week
- ...to enterprise expectations for reliability and risk management. • Collaborate with cross-functional teams (engineering, product, SRE, operations, architecture, controls) to identify high-value automation opportunities and deliver outcomes that can be adopted at scale...
$197.3k - $225.1k
...Overview Lead Software Engineer, Full Stack (Java, Python, SRE, AWS, AI) (Cloud Operations Resilience Engineering) Do you love building and pioneering in the technology space? Do you enjoy solving complex business problems in a fast-paced, collaborative, inclusive...Full timePart timeInternshipH1bLocal area- ...Sre Position 8+ years of professional experience processing a culture of learning through the development and sharing of skills, knowledge, process and tools. A driving passion for finding solutions to hard problems at scale and operationalizing them. Exceptional critical...
- Job Title Responsibilities Talents with Big data expert with 5+ years' experience in Hadoop Big data Required Skill Sets Working knowledge of AWS Kubernetes/EKS Middleware technologies Jenkins, JIRA, Ansible and expertise on DevOps tools Istio...
- ...SRE Production Support Engineer Location: Plano, TX Duration: 6 Months (Contract to hire) Interview Process: 1st round - Zoom 2nd round – In Person Role Overview: Position is part of the Central Site Reliability Engineering (SRE) Team. Looking for...Contract workShift work
$87.5k - $125k
Observability Engineer DISH is transforming the future of connectivity. We're doing it by building the country's first virtualized, standalone 5G wireless network from scratch. The foundation of a connected world, it's a network free of the limitations of the past, ...Flexible hoursNight shift$66.73 per hour
...Description We are seeking a Site Reliability Engineer (SRE) to help establish and scale our client's Google Cloud Platform (GCP) SRE practice. This is a unique opportunity to join a team during its formative stage and play a key role in building the operational foundation...Contract workTemporary work- ...Site Reliability Engineer (SRE) We are seeking an experienced Site Reliability Engineer (SRE) to join a newly forming team. This is an opportunity to establish SRE practices from the ground up, influencing the culture and approach for reliability. The ideal candidate...
- ...Dynatrace with ServiceNow Splunk PagerDuty Jira or similar platformsHandson scripting experience with Python Shell Script or PowerShellFamiliarity with DevOps CICD Infrastructure as Code and Site Reliability Engineering SRE practicesITIL Foundation certification preferred
- ...knowledge of Linux administrationExperience in production support troubleshooting and incident managementGood to HaveExperience with SRE practices observability SLAs reliability engineeringExposure to monitoring tools Prometheus Grafana etcKnowledge of data replication...
$128.6k - $184.9k
...position may also perform work that the U.S. government has specified can only be performed by a U.S. citizen on U.S. soil.Meet the TeamThe SRE Fleet team is responsible for maintaining the stability, scalability, and efficiency of the infrastructure that powers our global...Permanent employmentFull timeTemporary workLocal areaWorldwideFlexible hours- ...storage, load balancing, DNS, certificates/TLS), managing scope, dependencies, risk, and status.Partner with security, IAM, network, SRE/operations, and vendors to architect and implement scalable, resilient solutions and platform modernization.Automate and standardize...Work at office
$61.37 per hour
DescriptionWe are seeking a Technical Product Owner to partner closely with Engineering and Site Reliability Engineering (SRE) teams in driving infrastructure and cloud platform initiatives. This is a highly technical, engineering-facing role focused on translating business...Contract workTemporary work- ...how we’re UNSTOPPABLE for our employees!Are you ready for the next chapter in your Uncarrier journey? The System Reliability Engineer (SRE) improves and protects the software and systems behind all of T-Mobile's IT services, including management of scalability,...Full timeTemporary workPart timeWork experience placementLocal areaFlexible hours
- ...writing database scripts using DDL or queries using DML• CI/CD skills to build automated pipelines using ADO• Good understanding of SRE skillset and ability to design the services considering observability, traceability and monitoring aspects.• Experience in Agile scrum...Full timeTemporary workRelocation
- ...flexibility for diverse business needs.Serve as the primary observability subject matter expert and trusted advisor to engineering, product, SRE, Cloud Operations, FinOps, governance, and executive leadership teams.Evaluate, recommend, and guide the implementation of...Shift work
- ...Infrastructure as Code tools (Terraform, AWS CDK, CloudFormation).• Experience with observability platforms (Datadog, CloudWatch) and SRE practices.• Knowledge of security frameworks and compliance standards (SOC2, PCI-DSS, NIST).• Financial services or banking domain...Full timeTemporary workRelocation
- ...root cause analysis, identifying systemic risks, and implementing preventative measures to reduce production incidents.Familiarity with SRE practices such as service level indicators, service level objectives, error budgets, toil reduction, blameless post-incident reviews,...
- ...query limits, N+1 mitigation (DataLoader).Proven delivery of API/schema governance, versioning/deprecation, and CI policy gates.Strong SRE practices: SLIs/SLOs, error budgets, OpenTelemetry, data-driven post-incident improvements.Developer productivity: time-to-first-...Contract work
- ...environments. Oversee cloud engineering, platform engineering, and infrastructure modernization initiatives. Implement modern DevOps, SRE, automation, and continuous delivery practices. Establish engineering standards, governance, and reusable frameworks that improve...Full time
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to SRE. Be the first to apply!

