Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Sr. Site Reliability Engineer

Tiger Analytics

Role Overview

We are seeking a high-caliber Site Reliability Engineer (SRE) to join our Forward Engineering team. You will be the guardian of our production ecosystems, ensuring that our complex, data-driven AI platforms remain resilient, scalable, and highly performant. This role is a hybrid of software engineering and systems architecture, with a specialized focus on MLOps —bridging the gap between model development and production-grade reliability.

Key Responsibilities
1. Reliability & Performance Engineering
  • SLA/SLO Management: Define, monitor, and maintain Service Level Objectives (SLOs) and Service Level Indicators (SLIs) for critical AI/ML services.
  • Error Budgeting: Manage error budgets to balance the velocity of feature releases from the ML team with the stability of the production environment.
  • Scalability: Architect and manage auto-scaling strategies for Kubernetes (GKE) to handle fluctuating workloads during model training and high-volume inference.
2. MLOps & AI Infrastructure
  • Model Serving Reliability: Ensure the high availability of Vertex AI endpoints and custom inference services.
  • GPU/TPU Optimization: Monitor and optimize compute resource utilization (accelerators) to ensure cost-efficient performance for Large Language Models (LLMs).
  • Pipeline Resilience: Support and stabilize ML pipelines (Vertex AI Pipelines/Kubeflow) to ensure seamless data flow from ingestion to model retraining.
3. Automation & Orchestration (Eliminating "Toil")
  • Infrastructure as Code (IaC): Use Terraform or Pulumi to provision and manage consistent, version-controlled cloud environments.
  • CI/CD & GitOps: Design and optimize robust deployment pipelines for both application code and ML models using GitHub Actions, Cloud Build, or ArgoCD.
  • Task Automation: Develop custom Python or Go scripts to automate repetitive operational tasks, self-healing mechanisms, and resource cleanup.
4. Monitoring, Alerting & Incident Response
  • Observability: Build and manage comprehensive dashboards using Prometheus, Grafana, or Google Cloud Operations Suite (Stackdriver) .
  • Incident Management: Act as a primary responder in on-call rotations, leading the technical resolution of production outages.
  • Blameless Post-Mortems: Conduct deep-dive root cause analysis (RCA) to ensure systemic issues are identified and permanently remediated through code.

Orchestration: Expert-level knowledge of Kubernetes (K8s) and Docker.

MLOps Stack: Familiarity with tools such as Kubeflow, Vertex AI, MLflow, or DVC .

Scripting: Strong proficiency in Python (for automation) and Bash; knowledge of Go is a plus.

Data Systems: Experience managing the reliability of data-heavy services (BigQuery, Pub/Sub, or Vector Databases like Pinecone/Milvus).

Networking: Solid understanding of VPCs, Load Balancers, DNS, and secure service mesh (Istio/Anthos).

Benefits

Significant career development opportunities exist as the company grows. The position offers a unique opportunity to be part of a small, fast-growing, challenging and entrepreneurial environment, with a high degree of individual responsibility.

Tiger Analytics provides equal employment opportunities to applicants and employees without regard to race, color, religion, age, sex, sexual orientation, gender identity/expression, pregnancy, national origin, ancestry, marital status, protected veteran status, disability status, or any other basis as protected by federal, state, or local law.

#J-18808-Ljbffr
Vacancy posted 5 hours ago
Similar jobs that could be interesting for youBased on the Sr. Site Reliability Engineer in Washington DC vacancy
  • $106.3k - $221.1k

     ...Senior Site Reliability Engineer At Accenture Federal Services, nothing matters more than helping the US federal government make the nation stronger and safer and life better for people. Our 13,000+ people are united in a shared purpose to pursue the limitless potential... 
    Senior

    Accenture Federal Services

    Arlington, VA
    1 day ago
  •  ...Site Reliability Engineer (SRE) Dexian is seeking a savvy Site Reliability Engineer (SRE) who will play a key role in building a sustainable platform by developing systems for analyzing environments, predicting, and resolving issues, and supporting the production environment... 
    Senior
    Work experience placement

    Samprasoft

    Washington DC
    2 days ago
  • $175k - $250k

     ...Senior Cloud Infrastructure Engineer Location: San Francisco, CA. Remote unavailable. Modality: On‑Site only. Must live within commuting distance of San Francisco or...  ...while ensuring scalability, performance, and reliability across environments. What You’ll Do Design, build... 
    Senior
    Full time
    Remote work
    Relocation
    Relocation package

    The Recruiting Guy

    Washington DC
    1 day ago
  • $168k - $200k

     ...is passionate about creating transformative change in healthcare. What We're Looking For We're looking for a Senior Site Reliability Engineer to join our Data & ML Platform team. You'll be at the forefront of building and operating a resilient, observable, and... 
    Senior
    Remote work

    Datavant

    Washington DC
    1 day ago
  • $166k - $220k

     ...requirements and customer expectations. Our systems integration engineers internalize the nuances of each deployment, ensuring the...  ...solutions we ship. ABOUT THE JOB We are looking for a Site Reliability Engineer (SRE) to join AGD, our rapidly growing team in Irvine... 
    Senior
    Full time
    Work experience placement

    Mosaic

    Washington DC
    5 hours ago
  •  ...new job is posted. Sign in to set job alerts for “Senior Site Reliability Engineer” roles. Bellevue, WA $204,000.00-$259,000.00 1 day ago...  ...weeks ago Seattle, WA $151,300.00-$261,500.00 1 hour ago Sr. Software Engineer (TS/SCI Clearance Required) Sr Software... 
    Senior
    Contract work
    Remote work

    Signature IT World Inc

    Washington DC
    1 day ago
  •  ...solutions using a tailored Agile methodology. We are seeking a highly motivated and intellectually curious Senior Site Reliability Engineer to join our team working with a Federal client. The position will be a remote role open to US citizens residing in the... 
    Senior
    Remote work

    Elevate Government Solutions

    Washington DC
    3 days ago
  • $210k - $230k

     ...GovCIO is currently hiring for a Senior Site Reliability Engineer (SRE) to design, implement, and maintain highly available, scalable, and resilient infrastructure systems. The ideal candidate will bridge the gap between development and operations, focusing on automation... 
    Senior
    Currently hiring
    Remote work

    GovCIO

    Arlington, VA
    1 day ago
  • $149.4k - $202k

     ...Senior Software Engineer- Site Reliability Engineering (SRE) DC, MD, VA, CA The Site Reliability Engineering discipline at Noctua Technology, LLC is a strategic force driving digital transformation. We treat operations as a software engineering challenge, focusing... 
    Senior
    Remote work

    Noctua Technology

    Washington DC
    3 days ago
  • $106.3k - $221.1k

     ...more. Join us to drive positive, lasting change that moves missions and the government forward! Job Description The Site Reliability Engineer will ensure the reliability, performance, and scalability of the Client System. The engineer will define and track Key... 
    Senior
    Live in
    Work at office
    Local area

    Accenture

    Arlington, VA
    2 days ago
  • $121.4k - $218.6k

     ...solve complex challenges? Do you have a passion for automation and building systems that scale? Join our highly skilled Site Reliability Engineering team! Our team designs, develops, and manages applications and infrastructure that support Akamai Cloud's products and... 
    Senior
    Work experience placement
    Work at office

    Akamai

    Washington DC
    5 days ago
  • $136.2k - $214.01k

     ...outcomes Visionary in future focused problem-solving Exceptional in execution and impact The Role As a Senior Site Reliability Engineer at Proofpoint you will develop a deep understanding of the various services and applications that come together to... 
    Senior
    Full time
    Flexible hours

    Proofpoint

    Laurel, MD
    2 days ago
  • $150k - $180k

     ...what’s possible in remote sensing, you belong here at Umbra. About the Job We are seeking an experienced Senior Site Reliability Engineer to help design, build, operate, and scale the mission- and business-critical infrastructure that powers Umbra's systems.... 
    Senior
    Permanent employment
    Work at office
    Local area
    Remote work
    Worldwide
    Flexible hours

    Umbra

    Arlington, VA
    2 days ago
  • $81.1k - $187k

     ...Job Description We are looking for a Site Reliability Engineer 3 to support mission-critical cloud services and production operations. The role focuses on improving service reliability, reducing operational risk, automating repetitive tasks, and driving faster detection... 
    Senior
    Temporary work
    Immediate start
    Flexible hours
    Shift work

    Oracle

    Washington DC
    1 day ago
  •  ...Job Description Job Description Description: Onsite in Washington, DC   our client seeks a Sr. Site Reliability Engineer III to design, automate, and operate mission-critical systems for federal environments. The role focuses on Kubernetes or VMWare platforms,... 
    Senior
    Hourly pay
    Permanent employment
    Full time
    Local area
    Immediate start

    Eliassen Group

    Washington DC
    a month ago
  • $120k

     ...Principal Site Reliability Engineer location- Washington DC Remote- No Salary- $120K/Y Tech M/ Amtrak Job Summary we're seeking a seasoned Principal Site Reliability Engineer with a strong focus on pipelines as code, CI... 
    Remote work

    Yochana

    Washington DC
    4 days ago
  •  ...Principal Site Reliability Engineer The Principal Site Reliability Engineer will be a critical technical leader responsible for driving the operational excellence, resilience, and security of our core systems for a key Randstad client in the Washington D.C. area. This... 

    Software Technology Inc

    Washington DC
    1 day ago
  • $169.3k - $304.7k

     ...in building and maintaining fast, efficient, scalable, and reliable routing software and infrastructure that is responsible...  ...growth and stability of our global platform. As a Principal Site Reliability Engineer - Network, you will be responsible for: Architecting,... 
    Work experience placement
    Work at office

    Akamai

    Washington DC
    1 day ago
  • $165k - $265k

     ...SpaceX is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.SR. SITE RELIABILITY ENGINEER (STARSHIELD) At SpaceX we’re leveraging our experience in building rockets and spacecraft to deploy the Starshield... 
    Senior
    Permanent employment
    Temporary work
    Immediate start
    Weekend work

    SpaceX

    Washington DC
    11 hours ago
  • $84.9k - $209.5k

     ...About the Opportunity Help ensure healthcare professionals can reliably access the applications they depend on to deliver patient care. Oracle Health is seeking a Principal Site Reliability Engineer to strengthen the reliability, performance, security, and... 
    Temporary work
    Immediate start
    Flexible hours
    Shift work

    Oracle

    Washington DC
    1 day ago
  •  ...Site Reliability Engineer III (AI Platform) Location: Mount Laurel, NJ (Onsite) Duration: Contract Experience: 4+ years About the Role We are seeking a Site Reliability Engineer (SRE) III to support a cutting-edge AI Platform Engineering team responsible... 
    Contract work

    GCS Recruitment

    Laurel, MD
    5 hours ago
  • $135k - $154k

     ...where you matter.Your ImpactAs a contributor in the APX platform engineering organization on the CloudNet team, you are passionate about...  .... You are also obsessed about achieving the high quality and reliability our customers demand. You will work closely with sovereign... 
    Work experience placement
    Work at office
    Remote work

    Axon

    Washington DC
    1 day ago
  • $95k - $171k

     .... Opportunities exist to focus on GPU infrastructure, Kubernetes, and ensuring reliability for AI workloads within Akamai's serverless inference platform. As an Site Reliability Engineer II, you will be responsible for: Building and maintaining dashboards, alerts... 
    Permanent employment
    Work experience placement
    Work at office
    Remote work
    Work from home
    Worldwide
    Flexible hours

    Akamai

    Washington DC
    5 days ago
  • $230k - $250k

     ...GovCIO is hiring a Site Reliability Engineer with an active Secret clearance to ensure reliability, scalability, performance, and availability of mission-critical systems by combining software engineering practices with infrastructure operations expertise. This role is... 
    Remote work

    GovCIO

    Arlington, VA
    1 day ago
  • $112k - $179k

     ...system, network, software, and security solutions. About The Role Peraton is seeking a self-driven and resourceful Site Reliability Engineer to join our dynamic of Network and UC engineers in Washington, DC. This position combines software engineering and systems... 
    Contract work
    Worldwide
    Shift work

    Peraton

    Washington DC
    2 days ago
  • $107k - $220k

     ...The Site Reliability Engineer (SRE) will ensure the reliability, performance, and scalability of the WDP System. This person will define and track Key Performance Indicators (KPIs) and Service Level Objectives (SLOs), identify and resolve performance bottlenecks, and perform... 
    Full time
    Contract work
    Temporary work
    Work at office
    Visa sponsorship
    Work visa

    Avalore, LLC

    Arlington, VA
    4 days ago
  •  ...ears, and hands on the ground at a government customer site, ensuring the reliability and performance of Twenty's mission-critical platform running...  ...of deep technical ownership and customer-facing engineering: you'll define how we measure reliability, lead incident... 
    Full time
    Contract work
    Remote work
    Flexible hours

    Twenty Inc.

    Arlington, VA
    1 day ago
  •  ...Site Reliability Engineer Qualifications: ~10+ years of overall experience in IT including, with hands-on Development and Systems engineering background ~3-5 years of experience in a Site Reliability Engineering role ~ Experience with Enterprise Cloud transformation... 
    Temporary work
    Immediate start

    Samprasoft

    Washington DC
    2 days ago
  • $160k - $180k

     ...Site Reliability Engineer Location: Hybrid – Washington DC/Virginia/Maryland metro with the ability to travel to Patuxent River, MD, as needed (up to 20% of the time). Compensation: $160,000 - 180,000 per year, depending on experience and qualifications. Employment... 
    Full time
    Temporary work
    Local area
    Remote work
    Flexible hours

    RiseMe

    Washington DC
    5 hours ago
  •  ...Site Reliability Engineer Location- Wilmington De, Washington DC, Dallas, TX (Onsite Position) Full time position Minimum Qualifications Bachelor’s degree in computer science, Engineering, or a related technical field. Minimum of 5 years of experience... 
    Full time

    Yochana

    Washington DC
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Sr. Site Reliability Engineer. Be the first to apply!