Site Reliability Engineer
Input Technology Solutions
Job Title: Site Reliability Engineer (SRE)
Location: Washington, DC (Onsite)
Clearance: TS/SCI
• Monitor system health, availability, and performance using enterprise observability tools
• Analyze metrics and logs to proactively detect and remediate issues
• Tune alerting to reduce noise and prioritize mission impact Incident Management & Reliability
• Respond to and resolve production incidents across distributed environments
• Perform root cause analysis and lead post-incident reviews
• Implement corrective and preventive actions to improve resilience
• Participate in on-call rotation for outages, upgrades, and urgent activities Automation & DevOps Enablement
• Automate repetitive operational tasks to improve efficiency and reduce human error
• Support CI/CD pipelines and automated deployment workflows
• Develop scripts and tooling to improve reliability and repeatability Platform & Infrastructure Support
• Maintain Linux/Unix systems and containerized workloads
• Support Kubernetes/Docker environments and microservices architectures
• Assist with configuration management and environment standardization
• Ensure secure and compliant system configurations Collaboration & Continuous Improvement
• Partner with development teams to improve service reliability and performance
• Support backlog refinement and reliability engineering initiatives
• Document runbooks, procedures, and knowledge articles
• Contribute to continuous service improvement efforts Required Qualifications Education & Experience
• Bachelor's degree in Computer Science, Engineering, or related technical field
• Minimum 5 years of relevant technical experience
• At least 3 years of systems programming or SRE/DevOps experience Technical Skills
• Strong proficiency in Python, Bash, or similar scripting languages
• Hands-on experience with Linux/Unix administration
• Experience with Kubernetes and Docker
• Familiarity with cloud platforms (AWS, Azure, or Google Cloud)
• Experience with monitoring and logging tools (e.g., Grafana, Kibana, Prometheus, ELK)
• Working knowledge of CI/CD tools (e.g., GitLab, Jenkins, ArgoCD)
• Understanding of microservices architecture and DevOps practices
• Experience with Git-based workflows Infrastructure & Networking
• Knowledge of networking fundamentals, load balancers, and firewalls
• Experience with identity and access management (IAM, SSH, VPN, security groups)
• Experience deploying to on-premises or data center environments Professional Skills
• Strong analytical and troubleshooting abilities
• Excellent time management and ability to work independently
• Effective written and verbal communication skills
• Experience using Jira and Confluence in an Agile environment Preferred Qualifications
• Experience defining or working with SLIs, SLOs, and error budgets
• Familiarity with Helm and Kubernetes deployment pipelines
• Experience supporting high-availability or mission-critical systems
• Knowledge of security best practices and compliance frameworks
Vacancy posted more than 2 months ago
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
Related searches
- site reliability engineer Washington DC
- site reliability engineer remote Washington DC
- site reliability engineer sre Washington DC
- construction site safety Washington DC
- site recruiter Washington DC
- on site coordinator Washington DC
- website content developer Washington DC
- website coordinator Washington DC
- on-site clinical research associate (traveling/remote) Washington DC
- site safety Washington DC
