Senior Site Reliability Engineer
Spectraforce Technologies Inc
Title: Senior Site Reliability Engineer Duration: 6 Months (Could extend upto 18 months) Location: Austin, TX - Hybrid 4 days weekly onsite Qualifying Reason for Opening? support for moving from on-prem to cloud, infra support etc Why is this role important to your team/the project/the company? We are looking for a skilled engineer with disciplines that incorporate aspects of software systems engineering and operations. We are combining these skills to come up with better ways of managing and operating applications - including AI/ML-driven approaches to observability and reliability. Our Opportunity:
We are looking for a skilled engineer with disciplines that incorporate aspects of software systems engineering and operations. We are combining these skills to come up with better ways of managing and operating applications - including AI/ML-driven approaches to observability and reliability.
What you'll do:
* Evangelize SRE mindset and solve problems through systematization.
* Identify opportunities to build innovative tools and solve unique operations problems on large enterprise and mission-critical applications.
* Create scripts to automate operational tasks and incorporate solutions into infrastructure; architect and own production automation solutions that measurably reduce manual toil and improve operational throughput.
* Design and implement AI/ML-driven automation pipelines, observability enhancements, and proactive operational response systems - including anomaly detection and predictive alerting to improve platform reliability.
* Lead expansion of automation coverage across deployment, monitoring, alerting, and self-healing workflows for Cloud and Login Platforms.
* Collaborate with Engineering, Scrum, and Ops resources to provide technical expertise and support on key initiatives for system availability and reliability.
* Triage alerts and diagnose/resolve critical issues; manage implementation of changes with clear communication and minimal risk.
* Develop tools, frameworks, and instrumentation to validate and increase rollout success for applications; leverage AI/ML capabilities to enhance operational visibility and rollout validation at scale.
* Champion AIOps platform adoption and ML-assisted observability practices across the team.
* Coordinate capacity planning using data-driven trend analysis and ML-informed forecasting.
* Develop CI/CD orchestration systems to reduce friction for software delivery to production; drive adoption of GitOps concepts and AI-assisted pipeline optimization.
* Real-time troubleshooting of mission-critical application workflows and incorporate feedback into product development.
* Participate in on-call support. What do you have:
Required Skills:
* 6-8 years of experience with enterprise-level administration and support.
* 6-8 years of experience writing automation scripts, building application dashboards for proactive monitoring, and setting up alerts for early issue determination.
* 6-8 years practicing SDLC, process improvements.
* Hands-on enterprise systems administration, monitoring, and deployment activities.
* Experience with Windows 2019/2022 and Linux hosted via Virtual Machine.
* Experience in Cloud application configuration, deployment, support, and migration - GCP/PCF is a plus.
* Knowledge of IP networking including DNS, DHCP, firewalls, IP routing, etc.
* Familiarity with large-scale distributed systems and high-availability architecture.
* Linux and Windows system administration, troubleshooting, and tuning.
* Development experience in one or more programming languages: .NET, PowerShell, Java, Python, Bash.
* Knowledge of one or more of SQL, Oracle, MongoDB databases.
* Working knowledge of Actimize.
* Knowledge of one or more Message Brokers: Solace, RabbitMQ, IBM MQ, Kafka.
* Knowledge of Splunk, AppDynamics, or similar observability tools.
* Demonstrated experience applying AI/ML or AIOps approaches (e.g., anomaly detection, predictive alerting, ML-assisted observability) in production environments.
* Bachelor's degree in computer science or related discipline. Helpful Skills:
* Financial services industry experience.
* Agile methodologies.
* Hands-on experience with AIOps platforms or ML-driven observability tooling.
* Experience integrating AI/ML capabilities into CI/CD or operational automation workflows.
* Familiarity with CI/CD tools (Harness, Jenkins, GitHub Actions) or GitOps concepts.
* Exposure to container orchestration (Kubernetes, OpenShift) or cloud platforms (AWS, Azure, GCP). Personal Skills:
* Strong customer orientation with an affinity to proactively own, communicate, and follow through on projects and issues.
* Extreme sense of ownership to resolve problems in a distributed environment.
* Gritty resolve to dig deeper into technical issues in a complex login ecosystem.
* A self-starter with the ability and confidence to independently resolve issues and bring results back to the team.
Must Have
Notes from my Call: This is an IC role
Responsible for virtual application resiliency and sustainability 'expected to be working on all Ai assisted automation and AI observability building , defend system and production and platform issues
Handel cross platform communication
Cloud - GCp exp or at least exp with 2 other cloud platforms (azure/Aws etc)
Test , setup up and take it into production
And AI hands on
SRE proactive and knowledge
4 main areas they will work on:
Work ranges from - production availably stability and resiliency
Always monitoring 24/7 monitoring
Take action on alerts
Make sure application is stable
be proactive not reactive
Monitor signals and take action before it happens
Reposting to originality is priority
Monitoring is AI driven and continuously and issues are caught before issue happens Observability of application Production issues
AI - 50/50: As and when needed jump in a resolve production issues
Remining time will be infra and automation
QUESTIONS FROM TEAM:
We are looking for a skilled engineer with disciplines that incorporate aspects of software systems engineering and operations. We are combining these skills to come up with better ways of managing and operating applications - including AI/ML-driven approaches to observability and reliability.
What you'll do:
* Evangelize SRE mindset and solve problems through systematization.
* Identify opportunities to build innovative tools and solve unique operations problems on large enterprise and mission-critical applications.
* Create scripts to automate operational tasks and incorporate solutions into infrastructure; architect and own production automation solutions that measurably reduce manual toil and improve operational throughput.
* Design and implement AI/ML-driven automation pipelines, observability enhancements, and proactive operational response systems - including anomaly detection and predictive alerting to improve platform reliability.
* Lead expansion of automation coverage across deployment, monitoring, alerting, and self-healing workflows for Cloud and Login Platforms.
* Collaborate with Engineering, Scrum, and Ops resources to provide technical expertise and support on key initiatives for system availability and reliability.
* Triage alerts and diagnose/resolve critical issues; manage implementation of changes with clear communication and minimal risk.
* Develop tools, frameworks, and instrumentation to validate and increase rollout success for applications; leverage AI/ML capabilities to enhance operational visibility and rollout validation at scale.
* Champion AIOps platform adoption and ML-assisted observability practices across the team.
* Coordinate capacity planning using data-driven trend analysis and ML-informed forecasting.
* Develop CI/CD orchestration systems to reduce friction for software delivery to production; drive adoption of GitOps concepts and AI-assisted pipeline optimization.
* Real-time troubleshooting of mission-critical application workflows and incorporate feedback into product development.
* Participate in on-call support. What do you have:
Required Skills:
* 6-8 years of experience with enterprise-level administration and support.
* 6-8 years of experience writing automation scripts, building application dashboards for proactive monitoring, and setting up alerts for early issue determination.
* 6-8 years practicing SDLC, process improvements.
* Hands-on enterprise systems administration, monitoring, and deployment activities.
* Experience with Windows 2019/2022 and Linux hosted via Virtual Machine.
* Experience in Cloud application configuration, deployment, support, and migration - GCP/PCF is a plus.
* Knowledge of IP networking including DNS, DHCP, firewalls, IP routing, etc.
* Familiarity with large-scale distributed systems and high-availability architecture.
* Linux and Windows system administration, troubleshooting, and tuning.
* Development experience in one or more programming languages: .NET, PowerShell, Java, Python, Bash.
* Knowledge of one or more of SQL, Oracle, MongoDB databases.
* Working knowledge of Actimize.
* Knowledge of one or more Message Brokers: Solace, RabbitMQ, IBM MQ, Kafka.
* Knowledge of Splunk, AppDynamics, or similar observability tools.
* Demonstrated experience applying AI/ML or AIOps approaches (e.g., anomaly detection, predictive alerting, ML-assisted observability) in production environments.
* Bachelor's degree in computer science or related discipline. Helpful Skills:
* Financial services industry experience.
* Agile methodologies.
* Hands-on experience with AIOps platforms or ML-driven observability tooling.
* Experience integrating AI/ML capabilities into CI/CD or operational automation workflows.
* Familiarity with CI/CD tools (Harness, Jenkins, GitHub Actions) or GitOps concepts.
* Exposure to container orchestration (Kubernetes, OpenShift) or cloud platforms (AWS, Azure, GCP). Personal Skills:
* Strong customer orientation with an affinity to proactively own, communicate, and follow through on projects and issues.
* Extreme sense of ownership to resolve problems in a distributed environment.
* Gritty resolve to dig deeper into technical issues in a complex login ecosystem.
* A self-starter with the ability and confidence to independently resolve issues and bring results back to the team.
Must Have
- AI tools - Claud, co pilot (none specific)
- Must have strong understanding of AI ops and hands on exp
- Tools
- GCP or other cloud platforms (more than 1)
- Kubernetes
- Terraform
- Python
Notes from my Call: This is an IC role
Responsible for virtual application resiliency and sustainability 'expected to be working on all Ai assisted automation and AI observability building , defend system and production and platform issues
Handel cross platform communication
Cloud - GCp exp or at least exp with 2 other cloud platforms (azure/Aws etc)
Test , setup up and take it into production
And AI hands on
SRE proactive and knowledge
4 main areas they will work on:
- production issue is primary
- Observability
- infra set up
- work on infra set up and building that infra
- Moving to the could 'Kubernetes terraform and moving to GCP cloud
- Moving from on prem to cloud
- Team owns setting up the infra
- Anything/everything manual they aim to automate it
Work ranges from - production availably stability and resiliency
Always monitoring 24/7 monitoring
Take action on alerts
Make sure application is stable
be proactive not reactive
Monitor signals and take action before it happens
Reposting to originality is priority
Monitoring is AI driven and continuously and issues are caught before issue happens Observability of application Production issues
AI - 50/50: As and when needed jump in a resolve production issues
Remining time will be infra and automation
QUESTIONS FROM TEAM:
- Could you please clarify how critical AIOps platform experience is for this role? Is it a must-have requirement , or would you consider it nice to have if the candidate has strong experience in the core areas of the role?
- yes this is must have skill
Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Senior Site Reliability Engineer in Austin, TX vacancy
- ...and foster a dynamic work environment where new ideas thrive. Are you ready to join our team and make an impact?As a Senior Site Reliability Engineer at TeamViewer, you’ll be a key player in ensuring the reliability, scalability, and performance of our Azure-based SaaS...SeniorTemporary workCasual workWorldwide
- ...importance of in-office collaboration and fully intend for the selected candidate for this role to work on site in the specified location(s). As a Senior Site Reliability Engineer within the CET SAvE organization, you will play a critical leadership role advancing the...SeniorFull timeWork at office
$127k - $249k
The TeamPlatform Engineering is the department within SRE that is responsible for a range of critical infrastructure and operational functions... ...fleet, alongside the critical components that ensure cluster reliability and security (e.g., CoreDNS, cert-manager, and Gatekeeper)....SeniorWork at officeLocal areaRemote workWorldwideFlexible hours$127k - $249k
The TeamPlatform Engineering sits within SRE and builds the core infrastructure powering MongoDB... ...a pivotal role in engineering the reliable, globally connected, multi-cloud... ...Role OverviewWe are seeking a talented Senior Site Reliability Engineer (SRE) with a strong...SeniorLocal areaRemote workWorldwideFlexible hours- ...the rapidly evolving state of the art to engineer scalable, innovative, and research... ...business. About the Role:We are looking for a Senior SRE to serve as the operations owner for... ...tooling ecosystemsOwn the operational reliability of developer tooling ecosystems, including...SeniorFull timeLocal area
- ...importance of in-office collaboration and fully intend for the selected candidate for this role to work on site in the specified location(s).As a Senior Reliability Engineer, you will help shape the reliability, scalability, and operational excellence of mission-critical...SeniorFull timeWork at office
$152k - $241.5k
...artificial intelligence.We’re looking for a Senior SRE to join our Compute Farm team and... ...host lifecycle management, fleet reliability/auto-healing, E2E observability or data-... ...Python, Go, Perl, or Ruby.Mentored other engineers and influenced technical direction through...SeniorFull time$127k - $249k
We are looking for an experienced Senior or Staff Engineer for our SRE, InfraSec team, to guide the security of our cloud-based infrastructure. As a Staff SRE, you will be very hands-on technically while also mentoring a small team of SREs.The InfraSec team collaborates...SeniorLocal areaRemote workWorldwideFlexible hours$127k - $249k
...Eastern or Central time zones. We are looking for an experienced Senior Engineer for our SRE, Atlas team to support, maintain and grow the... ...crucial workloads. Role OverviewWe are seeking a talented Site Reliability Engineer (SRE) with a strong infrastructure background....SeniorLocal areaRemote workWorldwideFlexible hours$192.4k - $275.8k
...the world's most demanding enterprise customers, blending Site Reliability Engineering, Systems Engineering, and Service Engineering disciplines... ...this is the team for you Your ImpactYou will be the most senior technical individual contributor on the team — setting the...SeniorFull timeTemporary workLocal areaFlexible hours- ...Schwab. We are an integrated product, engineering, strategy and risk team, all based in San... ...how we serve our clients. As a Senior Engineer on AI.x, you will play a key role... ...areas of technology today.As a Senior AI Site Reliability Engineer you will support reliability efforts...SeniorFull time
- ...and best in class outcomesVisionary in future focused problem-solvingExceptional in execution and impactThe RoleAs a Senior Site Reliability Engineer at Proofpoint you will develop a deep understanding of the various servicesand applications that come together to deliver...SeniorFull timeFlexible hours
$136.2k - $214.01k
...outcomes Visionary in future focused problem-solving Exceptional in execution and impact The Role As a Senior Site Reliability Engineer at Proofpoint you will develop a deep understanding of the various services and applications that come together to...SeniorFull timeFlexible hours$75 - $80 per hour
...Location: Austin, TX Salary: $75.00 USD Hourly - $80.00 USD Hourly Description: Role: Senior Site Reliability Engineer / DevOps Engineer (AIOps & Observability) Work Arrangement: Hybrid (4 days weekly on-site in Austin, TX) Position Type: 6+ Month Contract...SeniorHourly payContract workWork experience placement- Job Description:About the Role:We are looking for a Senior SRE to join our Platform Engineering team where you’ll own the reliability, scalability, and operational excellence of our workflow orchestration platforms - primarily Apache Airflow and Broadcom Automic/UC4. This...SeniorFull time
- ...: 2026-09-21Location: Austin, TXCompany: Realtor.comSr Site Reliability EngineerLocation: Austin, TX, United StatesCategory: TechnologyJob... ...confidence through expert guidance.We are seeking a Senior Site Reliability Engineer to join our newly formed Operations Excellence...SeniorWork at officeLocal area
- ...Sr. Software Engineer - Site Reliability About ShipperHQ: ShipperHQ is a trusted leader in the e-commerce shipping space, with over 15 years... ...logistics. Position Overview: We’re seeking a Senior Site Reliability Engineer to join our fast-paced Engineering...SeniorFull timeWork at office
$140k - $215k
...intersection of our Core Platform and Embedded Reliability charters: building the foundational... ...while embedding directly with product engineering teams and their leadership to drive... ...eliminated manual deployment processes.At the Senior Engineer level, your influence is...SeniorFull timeWork experience placementWork at officeLocal area2 days per week3 days per week$28 - $38 per hour
DescriptionKforce has a client seeking a remote Site Reliability Engineer to join their team. We are seeking a highly motivated Site Reliability Engineer (SRE) to help build, scale, and maintain cloud infrastructure, CI/CD pipelines, and deployment automation. This role...Work experience placementRemote work$109.65k - $182.76k
...encrypt data to make the connected world more secure.Austin, TX - Hybrid (3 days a week)Position SummaryWe are seeking a Site Reliability Engineer to ensure the high level of service and operation excellence for the development of the innovative and ambitious Telecommunication...Full timeLocal area3 days per week$198.24k - $272.58k
We’re looking for a Principal Site Reliability Engineer to join Procore’s Compute Division to work on our FedRAMP initiative. In this role, you’ll help build Procore’s next-generation construction compute platform for others to build upon, including Procore developers,...Full timeContract workWork at officeLocal areaImmediate start$98.58k - $138.02k
...Northern California / Silicon Valley Region / Denver, COProduct Engineering - DevOps /Full Time /HybridRestaurant365 is a SaaS company... ...office locations: Austin, TX; Irvine, CA; or Akron, OH. The Site Reliability Engineer II will be responsible for supporting, enhancing,...Full timeWork at office$128.6k - $184.9k
...to keep their digital systems secure and reliable. Come help organizations be their best, while... ...team! We are looking for an experienced engineer to support the next generation of our... ...insurance. Please see the Cisco careers site to discover more benefits and perks. Employees...Permanent employmentFull timeTemporary workLocal areaRemote workFlexible hoursShift workNight shiftWeekend work- ...Infrastructure Code. Builds reliability into the ecosystem by applying... ...best practices in resiliency engineering and observability by developing... ...engineering techniques with site reliability engineering... ...repeatable business processes.Advises senior management on technical...Full time
$84.9k - $209.5k
...Help ensure healthcare professionals can reliably access the applications they depend on... ...Oracle Health is seeking a Principal Site Reliability Engineer to strengthen the reliability,... ...valuable insights and information with senior team members, management, and beyond to...Temporary workImmediate startFlexible hoursShift work$152k - $241.5k
...the world.Join the Simulation Software team at NVIDIA as a Senior System Software Engineer! This role offers an outstanding opportunity to work on... ...development by enabling Chips Simulation as a trusted and reliable virtual platform.What you will be doing:Drive early...SeniorFull time$172k - $300k
Job DescriptionGM Vehicle Autonomy is forming a centralized Site Reliability Engineering team to make reliability a measurable, engineered property of the systems used to build, validate, release, and operate autonomous-vehicle software.As one of our founding SREs, you...Full timeWork at officeLocal areaRemote workWork from homeRelocationRelocation packageFlexible hours- ...Role: Site Reliability Engineer Location: Southlake / Austin, TX - Onsite 4 days weekly Duration: 12 Months Job Summary We are seeking a motivated Site Reliability Engineer (Contractor) with 3 to 5 years of experience in automation, cloud infrastructure, and production...For contractors
- ...selected candidate for this role to work on site in the specified location(s).The Client... ...team is responsible for ensuring the reliability, scalability, and operational excellence... ...around the clock. As a Site Reliability Engineer, you will partner across application engineering...Full timeWork at office
$48.48 - $53.87 per hour
...you a highly motivated and experienced reliability professional passionate about automation... ...Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical... ...3 to 5 years of hands-on experience in Site Reliability Engineering, Production...Hourly payTemporary work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!
Related searches
- site reliability engineer sre Austin, TX
- site reliability engineer Austin, TX
- site reliability engineer remote Austin, TX
- senior operations technician Austin, TX
- senior operations associate Austin, TX
- senior cloud service delivery manager Austin, TX
- senior it service manager Austin, TX
- senior project engineer Austin, TX
- senior chief engineer Austin, TX
- senior manufacturing supervisor Austin, TX


