SRE
TriOptus LLC
Site Reliability Engineer (SRE)
As a part of the FRDC Site Reliability Engineer (SRE) team, you will help identify resilience challenges, build reusable, foundational software and infrastructure components to improve, influence, and validate the resilience and reliability for technologies that move trillions of dollars per day. Responsibilities include, but are not limited to:
- Participate in the design of build, refactor major software components that improve the availability, resilience, performance of our system
- Design, code, test, and deliver software to automate manual operational work
- Support incident responses, blameless postmortem, design and implement the product improvement to prevent incident reoccurring
- Implement application patterns in support of better service level objectives
- Implement self-healing, resiliency patterns
- Exercise failure cases regularly to validate resilience assumptions
- Engage with development teams throughout the life cycle of incident, ensure lessons learned are translated into automated or process adjust responses to help develop software for reliability and scale, ensuring minimal refactoring or changes
- Code, test and deliver software to automate manual operational work
- Troubleshoot incidents, participate in blameless post-incident evaluations and ensure permanent closure of incidents
- Identify application patterns and analytics in support of better service level objectives
- Analyze self-healing and resiliency patterns and contribute to software which can use these outcomes
- Implement best in class monitoring frameworks to accomplish end to end flow monitoring and noiseless alerting
Requirements & Qualifications:
- Bachelor’s degree or equivalent experience in a software engineering discipline
- 2+ years of hands-on software engineer experience
- Curious about solving resilience problems in run time at scale
- Expertise in at least one technology stack designing, coding, testing, and delivering software
- Knowledge in a few of infrastructure components (e.g. routers, load balancers, cloud products, container systems, compute, storage, and networks)
- Experience in cloud native, distributed application design and implementation
- Demonstrated communication and ownership skills
- Debugging and trouble shooting skills
- Collaboration with a diversified high-performing multi-location team
- Excellent analytical, interpersonal and communication skills
- Understanding of SRE methodologies/practices
Required Skills: Technical expertise of 4+ years the below areas, overall IT experience of 6+ years:
- Proficiency in Java / JVM based system design & implementation
- Infrastructure knowledge required including Unix, Windows, networking, and scripting (e.g. Perl / Python)
- Experience with orchestration tools like Jenkins CI/CD, or Jules
- Experience following source control best practices: Git/bitbucket
- Experience with database development (MySQL / Oracle)
- Understanding of architecture and design across distributed systems
Prefer Skills:
- Knowledge of SpringBoot / Microservices architecture
- Experience using Pivotal Cloud Foundry
- Experience with Public Cloud: AWS
- Enterprise platforms using Big Data tools and technologies (e.g. Hadoop, Spark, Hive, Impala, Dremio, Nifi, Ignite)
- Experience setting up & building solutions for Containers e.
Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the SRE in San Antonio, TX vacancy
- ...SLOs and engage with exception processes when technical limitations exist. Work with dependent process (Site Reliability Management, SRE, Event Management and Incident Management, CMDB) to create enhancement stories that will improve SLO efficacy and value when passing...Suggested
- Job Title Responsibilities: Develop, test, and debug automated tasks (Apps, Systems, Infrastructure) Troubleshoot minor incidents and contribute to resolution through post-mortems Participate in the application or service development lifecycle through code ...Suggested
$80k - $133k
...to Obtain Public Trust What You Will Do : * Collaborate with team members and cross-departmental partners to establish and maintain SRE practice in an Agile Scrum framework. * Participate in system design reviews to identify points of failure, promote automation and...SuggestedPermanent employmentContract workTemporary workFlexible hours- ..., CISSP, CISM, CRISC, CISA, Cisco CCNP (Enterprise/Security), VMware Certified Professional (VCP), and Site Reliability Engineering (SRE) Foundation.SkillsAdaptability - Maintaining effectiveness when experiencing major changes in work responsibilities or environment (e...SuggestedFull timeTemporary workLocal area
- ...degree in Computer Science or related field or equivalent practical experience4+ years of experience in DevOps, Cloud Engineering, or SRE rolesStrong expertise in:CI/CD tools (Azure DevOps, GitHub Actions, Jenkins)Cloud platforms (Azure strongly preferred; AWS/GCP...SuggestedFull time
- ...· AWS DevOps Engineer Professional· AWS Solutions Architect Professional· Terraform Associate· FinOps Practitioner· ITIL Foundation· SRE Practitioner· Datadog CertificationAdditional Required Qualifications• Bachelor’s degree or foreign equivalent required from an accredited...Full timeTemporary workRelocation
$131.3k - $237.35k
...implementing Kubernetes-based developer platforms or Internal Developer Platforms (IDPs).Experience with Site Reliability Engineering (SRE) practices and operational excellence programs.Professional certifications such as:AWS Certified Solutions ArchitectAWS Certified...Full timeWork at office- ...about what we do!What We Need:iHeartMedia Entertainment, Inc. seeks candidates for the position of Senior Site Reliability Engineer (SRE), responsible for leading a talented team of SREs/DevOps Engineers across a wide variety of Cloud Services to ensure the reliability,...Full timeFlexible hours
- ...technical support services to specific Enterprise customers and serving as a subject-matter expert on Network security products. The SRE relies on experience and judgment to plan and perform a variety of critical tasks/responsibilities pertaining to the customer’s deployment...
- ...lead the design and implementation of Infrastructure Automation, CI/CD for infrastructure operations, Site Reliability Engineering (SRE), DevSecOps, and Software Development Lifecycle (SDLC) processes. Provide technical leadership in Software-Defined Networking (SDN),...Full timeLocal area
- ...Proficient with ITSM process- improve data and process that enable improved IT Business Service Management. Proficient with ServiceNow, SRE, Configuration Management Databases (CMDB), and Common Service Data Models (CSDM). Strong verbal and technical writing skills are...
- ...Village ] (No Remote/Hybrid please) ABOUT OODLE Oodle’s mission is to build an iconic observability company! Oodle enables developer and SRE teams to handle 10x more data at a tenth of the price. Oodle is disrupting Observability just like snowflake has disrupted data...Work experience placementRemote work
- ...infrastructure using AWS / Azure services. Responsibilities Work across multiple teams to define and implement end-end DevOps /SRE strategy Expert in automation by writing Shell, Perl & Python scripts to monitor Production Applications Expertise in...Work experience placement
- ...networking (VPCs, Route53, etc), GoLang Terraform, Understanding regulatory compliance (e.g, SOC 2, ISO 27001), Datadog, Elastic, SRE or DevOps methodologies Key Responsibilities Monitoring & Alerting • Continuously monitor AWS Resilience Hub to track resilience...
$148k - $160k
...faster. About you Required 2+ years in a technical role where hands-on troubleshooting was a core part of the work — engineering, DevOps, SRE, QA, or similar. Customer-facing experience is a plus, not a requirement. Strong debugging instincts — you dig into logs, configs,...Local areaImmediate startRemote workFlexible hoursShift work$131.5k
...systems Remediate infrastructure and/or container security vulnerabilities What You'll Need: U.S citizen required 2+ years of related SRE experience Apply core software engineering fundamentals* (e.g., data structures, algorithms, concurrency, system design) to solve...Work experience placementH1bLocal areaRemote workFlexible hours- ..., including API development, data management, and system integrations. Leadership in observability and Site Reliability Engineering (SRE) practices, including defining SLOs, monitoring, and incident management. Oversight of operational security, including secrets management...Full timeWork at officeRelocation packageFlexible hours3 days per week
- ...within the architecture. Build and maintain strong cross-functional relationships with product managers, site reliability engineers (SRE), quality assurance (QA) teams, integration managers, and gaming studios to ensure seamless collaboration. Provide constructive feedback...Full timeLocal areaVisa sponsorshipFlexible hours
- ...AWS/Azure/GCP) bility to interpret architecture diagrams, data flows, and system design Knowledge of DevOps, CI/CD, monitoring, SRE fundamentals Familiarity with data engineering concepts (ETL, data models, analytics). Product visioning, roadmap planning, and...
- ...by building self-service capabilities, golden paths, and automation for common workflows, while ensuring platform reliability through SRE practices and compliance automation. The ideal candidate will have experience with Atlassian Compass (preferred) or Backstage for...
- ...work from idea to stable production: Drive technical discovery and spike work when requirements are ambiguous. Collaborate with QA and SRE to ensure readiness for peak events (launches, BFCM, campaigns). Support on-call rotations by helping debug, fix, and prevent...Temporary workWork at officeLocal areaRemote workFlexible hours
$265k
...developer platforms, internal tooling, or paved-road infrastructure at scale Strong understanding of observability practices and SRE principles (SLOs, error budgets, incident management) Hands-on experience with AWS, Azure, or GCP at production scale Strong understanding...Full timeWork at officeLocal areaRemote workWork from homeFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to SRE. Be the first to apply!

