SRE
Saxon Global Inc
Responsibilities:
*5-7 years of experience
• Gathers and analyzes metrics from monitoring platforms to assist in performance tuning
and fault tolerance.
• Participates in system design, platform management and capacity planning.
• Balances feature development speed and reliability with service-level objectives.
• Works closely with the incident response team and restoring service to normal operation.
• Understands debugging and applying troubleshooting skills.
• Investigates, blocks and rate-limits unwanted traffic.
• Utilizes monitoring systems and dashboards for proactive changes and alerting.
• Establishes continuous process improvement cycles where the process, performance,
and supporting technologies are reviewed and enhanced where applicable.
• Partners with development teams to improve services through testing and release
procedures.
KNOWLEDGE, SKILLS, ABILITIES
• Understanding of Kubernetes, containers, clusters and elastic scalability.
• Expertise in SRE principles.
• Mindset of continually finding ways to drive scalability, stability and performance.
• Cloud Services experience with Google Cloud Platform (GCP).
• Experience with API, service-based or microservice-based architecture.
• Proficiency in infrastructure, network, database, operating systems or security
troubleshooting and remediation.
• Architecture-level knowledge of Windows and Linux and Infrastructure systems.
• Experience with production deployment, monitoring and operational support for enterprise-class applications (Dynatrace a plus).
• Experience working with Continuous Integration/ Continuous Deployment tools.
• Experience in performance diagnostics, capacity planning, performance architecture
design, performance tuning and performance monitoring.
• Experience with Azure DevOps (ADO), Dynatrace, Prometheus, Terraform and Grafana
*5-7 years of experience
• Gathers and analyzes metrics from monitoring platforms to assist in performance tuning
and fault tolerance.
• Participates in system design, platform management and capacity planning.
• Balances feature development speed and reliability with service-level objectives.
• Works closely with the incident response team and restoring service to normal operation.
• Understands debugging and applying troubleshooting skills.
• Investigates, blocks and rate-limits unwanted traffic.
• Utilizes monitoring systems and dashboards for proactive changes and alerting.
• Establishes continuous process improvement cycles where the process, performance,
and supporting technologies are reviewed and enhanced where applicable.
• Partners with development teams to improve services through testing and release
procedures.
KNOWLEDGE, SKILLS, ABILITIES
• Understanding of Kubernetes, containers, clusters and elastic scalability.
• Expertise in SRE principles.
• Mindset of continually finding ways to drive scalability, stability and performance.
• Cloud Services experience with Google Cloud Platform (GCP).
• Experience with API, service-based or microservice-based architecture.
• Proficiency in infrastructure, network, database, operating systems or security
troubleshooting and remediation.
• Architecture-level knowledge of Windows and Linux and Infrastructure systems.
• Experience with production deployment, monitoring and operational support for enterprise-class applications (Dynatrace a plus).
• Experience working with Continuous Integration/ Continuous Deployment tools.
• Experience in performance diagnostics, capacity planning, performance architecture
design, performance tuning and performance monitoring.
• Experience with Azure DevOps (ADO), Dynatrace, Prometheus, Terraform and Grafana
Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the SRE in Vestavia Hills, AL vacancy
- ...Job Description Job Description Job Title: Senior AWS Site Reliability Engineer (SRE) Location: Birmingham, Alabama Type: Contract To Hire Work Model: Onsite – onsite Hours: 40.0 Security Clearance: Overview Responsibilities Implement and improve...SuggestedContract workLocal area
$22 per hour
...This CO-OP opportunity is at our research facility (Kratos SRE) located in Birmingham, AL. The CO-OP requires a 3-semester commitment - rotating working a semester then returning to school for a semester. CO-OPs support a department of engineers in test preparation and...SuggestedFull timeWork at officeImmediate start$22 per hour
...This CO-OP opportunity is at our research division (Kratos SRE) located in Birmingham, AL. The CO-OP requires a 3 semester commitment -- rotating working a semester, then returning to school for a semester. CO-OPs support a department of engineers in test preparation...SuggestedImmediate start- ...JOB SUMMARY: This is a full-time position that spans two complementary technical areas within the Hypersonics department at Kratos SRE. Approximately half of the candidate's effort (~20 hours per week) will be devoted to supporting the design, development, and commissioning...SuggestedFull timeWork at office
- ...position serves as a technical team lead and developing subject matter expert within the Hypersonic Structures department at Kratos SRE, one of the nation's leading organizations in the evaluation of thermal protection systems and refractory materials for ballistic and...SuggestedFull timeWork at office
$124.5k - $170k
...the platform standards the pod builds on.QUALIFICATIONSREQUIRED QUALIFICATIONS• 8+ years in data, ML, or platform engineering, or in SRE/DevOps, including several years operating production data and/or ML systems.• Demonstrated technical leadership — setting standards,...Full timeTemporary work- Reference: 85347-370314This CO-OP opportunity is at our research facility (Kratos SRE) located in Birmingham, AL. The CO-OP requires a 3-semester commitment - rotating working a semester then returning to school for a semester.CO-OPs support a department of engineers in...Temporary work
- ...position serves as a technical team lead and developing subject matter expert within the Hypersonic Structures department at Kratos SRE, one of the nation's leading organizations in the evaluation of thermal protection systems and refractory materials for ballistic and...Temporary work
- ...; Hire and Develop the Best.; Be Curious and Learn.; Win as a Team Job Summary We’re looking for a Site Reliability Engineer (SRE) who’s passionate about building resilient, high-performing systems that our customers can depend on every day. In this role, you’ll...Local areaRemote work
- ...reliability and resilience. This role focuses on building automation to reduce manual effort and prevent service-impacting incidents. The SRE combines software and systems engineering to build and support large-scale, distributed, fault-tolerant systems. This role ensures...
- ...automated builds for O365 and cloud platforms company-wide, upholding standards via Azure Policy. Drive Site Reliability Engineering (SRE) Practices : Team with Operations to define SLOs/SLIs, using tools like Azure Monitor and Application Insights for advanced self-...Full time
- ...with ACH systems, especially PEP+ (FiServ Solution/Product) Experience with ACI UPF (Universal Payments Framework) platform for Wires Experience with ISO 20022 data and upgrades Experience with both mainframe and distributed systems Cloud or SRE or DevSecOps experience
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to SRE. Be the first to apply!


