Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Staff Software Engineer, Reliability

$203.5k - $248.5k

Metropolis

Job Description

Job Description

Who we are

The real world is the next frontier, and at Metropolis, we are creating the artificial intelligence to make it responsive. We are pioneering the Recognition Economy — a future where mundane repetition disappears and being known unlocks access, comfort, and belonging everywhere you go. From transforming parking into a seamless drive-in, drive-out experience for millions of Members to expanding our intelligence layer across retail and hospitality, we are building a world that feels instinctive and magical. The future isn't coming; it's here, and we need builders, innovators, and problem solvers to help us create it.

Who you are

Metropolis is seeking a Staff Software Engineer focused on Reliability to own reliability across the entire Metropolis platform and drive the comprehensive practices that ensure system availability, resilience, and observability for our mission-critical mobility infrastructure. In this role, you will build reliability from first principles, architecting failover systems, implementing chaos engineering, and improving our observability foundation to maintain 99.9%+ uptime as we scale to new markets. 
As the technical owner of our reliability posture, you will tackle challenges like external service failover, dependency mirroring, and database replication, working alongside highly technical teams across the organization to influence architecture decisions and establish company-wide reliability standards. You will join the Product Foundations team, playing a key role in building the foundational infrastructure that powers the future of mobility commerce.

What you'll do
  • Own the overall reliability posture for the Metropolis platform, establishing practices, metrics, and systems that ensure 99.9%+ uptime across all services
  • Design and implement automatic failover mechanisms for critical external dependencies like Twilio for SMS/voice and Stripe for payments with circuit breakers, retry policies, and degraded mode operations
  • Architect and build active-passive or active-active regional deployment strategies with database replication, automated failover, and DNS-based traffic routing including disaster recovery planning and testing
  • Establish comprehensive monitoring using Datadog for APM, logs, and metrics correlation
  • Implement synthetic monitoring, SLO-based alerting, on-call rotation, and escalation policies while building service health dashboards that show customer impact
  • Own the incident management process including workflows, tooling, post-mortem culture, runbook automation, and MTTR reduction initiatives to drive down mean time to recovery from detection to resolution
  • Drive adoption of resilience patterns across all services including health checks, graceful degradation, feature flags, rate limiting, backpressure mechanisms, and chaos engineering practices
  • Build and maintain local mirrors for critical dependencies with artifact caching, dependency pinning, and vulnerability scanning to prevent build failures from upstream outages
What we're looking for
  • 8+ years of engineering experience including software engineering, reliability engineering, SRE practices, or production operations at scale
  • Demonstrate expert-level reliability engineering skills including hands-on experience with multi-region architectures, failover automation, circuit breakers, chaos engineering, and disaster recovery
  • Utilize production observability expertise with deep experience implementing monitoring, alerting, tracing, and logging systems at scale – specifically Datadog or similar APM platforms in high-load environments
  • Apply strong systems thinking with proven ability to design resilient distributed systems that gracefully handle failures, network partitions, and external dependency outages
  • Demonstrate database and data systems knowledge including replication strategies, backup/restore procedures, connection pooling, query optimization, and experience with both relational and NoSQL databases
  • Leverage cloud platform expertise with production experience operating and ensuring reliability of systems on AWS including multi-region deployments, load balancing, and DNS-based failover
  • Possess experience with AI-powered development tools such as Claude Code, GitHub Copilot, or similar agentic coding tools for enhanced productivity – context engineering in particular
  • Exhibit excellent technical communication with ability to influence technical decisions across teams, document complex systems, conduct post-mortems, and establish reliability standards organization-wide
  • Demonstrate expert-level Java and/or Scala proficiency with strong understanding of JVM performance, concurrency, and operational characteristics
While not required, these are a plus:
  • Bring Scala experience
  • Possess SRE or Reliability Engineering experience at companies known for operational excellence such as Google, Amazon, Netflix or high-growth startups where you built reliability practices from the ground up
  • Demonstrate incident response leadership including experience building incident management processes, conducting blameless post-mortems, and driving MTTR reduction initiatives in production environments
  • Utilize chaos engineering experience with tools like Chaos Monkey, Gremlin, or similar, including designing and executing game days and failure injection testing
  • Exhibit performance optimization experience with profiling, benchmarking, capacity planning, and system tuning at hyperscale including experience optimizing for high-throughput, low-latency systems
  • Show open source contributions or technical blog writing that demonstrates depth of expertise in reliability engineering, distributed systems, or production operations
Our Stack
  • Languages + Frameworks: TypeScript, React, Scala (principally), Java (limited)
  • Datastores: MySQL, PostgreSQL, Snowflake
  • Cloud: AWS
  • Version control: Git & GitHub
  • AI Tooling: Copilot on GitHub
  • Observability: Datadog

4 Days in Office: Metropolis values in-person collaboration to drive innovation, strengthen culture, and enhance the Member experience. Our corporate team members hold to our office-first model, which requires employees to be on-site at least four days a week, fostering organic interactions that spark creativity and connection

When you join Metropolis, you'll join a team of world-class product leaders and engineers, building an ecosystem of technologies at the intersection of parking, mobility, and real estate. Our goal is to build an inclusive culture where everyone has a voice and the best idea wins. You will play a key role in building and maintaining this culture as our organization grows. The anticipated base salary for this position is $203,500 USD to $248,500 USD annually. The actual base salary offered is determined by a number of variables, including, as appropriate, the applicant's qualifications for the position, years of relevant experience, distinctive skills, level of education attained, certifications or other professional licenses held, and the location of residence and/or place of employment. Base salary is one component of Metropolis' total compensation package, which may also include access to or eligibility for healthcare benefits, a 401(k) plan, short-term and long-term disability coverage, basic life insurance, a lucrative stock option plan, bonus plans, and more. #LI-LR1 #LI-Onsite

Metropolis may utilize an automated employment decision tool (AEDT) to assess or evaluate your candidacy for employment or promotion. AEDTs are used to assist in assessing a candidate's application relative to the required job qualifications and responsibilities listed in the job posting.

As part of this process, Metropolis retains data relevant to your candidacy, including personal information, for a period that is reasonably necessary for the use of the tool. If you are hired for the position, your data may become part of your employee records.

Metropolis Technologies is an equal opportunity employer. We make all hiring decisions based on merit, qualifications, and business needs — without regard to race, color, religion, sex (including gender identity, sexual orientation, and pregnancy), national origin, disability, veteran status, or any other protected characteristic under federal, state, or local law.

We are committed to providing a welcoming, accessible hiring experience. If you need a reasonable accommodation for any part of the application or interview process, please reach out to us at View email address on us.fitly.work.

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Staff Software Engineer, Reliability in Washington DC vacancy
  • $189k - $303k

     ...accessible for all. We are searching for an exceptional Staff-level Backend Software Engineer to join the Aurora Services Engineering team and take...  ...infrastructure to scale our products with high availability and reliability. Collaborate with stakeholders including Security,... 
    Suggested
    Full time
    Remote work

    Aurora Innovation

    Washington DC
    16 hours ago
  •  ...users. But it didn’t stop there. They engineered Snowflake to power the Data Cloud, where...  ...and maintaining Snowflake’s security, reliability and performance. The team culture is...  ...mentorship from Principal engineers. AS A STAFF SOFTWARE ENGINEER - IDENTITY & ACCESS... 
    Suggested
    Full time

    Snowflake

    Washington DC
    16 hours ago
  •  ...About the Team Join the engineering teams that bring OpenAI’s ideas...  ...About the Role We’re seeking Software Engineers who can solve...  ...needed to deliver a scalable, reliable platform. You’ll also partner...  ...internal title Member of Technical Staff . We use Senior Staff... 
    Suggested
    Full time

    OpenAI

    Washington DC
    16 hours ago
  •  ...home day is currently Tuesday.   About the Role As a Staff Software Engineer for the Compute pillar, you will play a critical role in...  ...and low-level semiconductor architecture to enable seamless, reliable cloud provisioning and lifecycle management of a heterogeneous... 
    Suggested
    Full time
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda

    Washington DC
    16 hours ago
  • $200k - $275k

     ...internationally.  Our Team As an engineering team, we believe strongly that empathy...  ...serves. We’re looking for a Staff Software Engineer to join our core engineering teams...  ...synthesize product requests into strong and reliable software components What we look for... 
    Suggested
    Full time
    Work at office
    Local area

    Peregrine Technologies, Llc

    Washington DC
    16 hours ago
  • $160.2k - $246.3k

     ...pipelines that enable and quantify the accuracy, reliability, and efficiency of simulation tests used for autonomous vehicle software validation. Lead cross-functional initiatives with Autonomy, Systems Engineering, Simulation, and Data teams to tightly integrate team... 
    Full time
    Local area
    Remote work
    Work from home
    Relocation
    Relocation package
    Flexible hours

    General Motors

    Washington DC
    2 days ago
  • $220k - $275k

     ...passionate about creating transformative change in healthcare. Staff Software Engineer The Role As a Staff Software Engineer, you will...  ...initiatives in data architecture, security, and reliability. Advise leadership on technical direction and major tradeoffs... 

    Datavant

    Washington DC
    1 day ago
  • $254k - $336k

     ...ABOUT THE JOB At Anduril, our Software Engineers are at the forefront of defense technology, crafting high-impact, cutting-edge solutions...  ...challenging real-world problems by building scalable and reliable software solutions for our next-generation robotic platforms... 
    Full time
    Work experience placement
    Immediate start

    Anduril Industries

    Washington DC
    16 hours ago
  • $150k - $203k

     ...build-many model, we're accelerating the deployment of safe, reliable, and affordable nuclear energy. We operate with an AI-...  ...unwavering commitment to our mission. About the role As a Staff Software Engineer, you will serve as a technical leader responsible for... 
    Bi-weekly pay

    The Nuclear Company

    Washington DC
    4 days ago
  • $260.1k

     ...surges.”learn more about working at Coinbase. As a Senior Staff Software Engineer on thePlatform Payments team, you'll define the...  ...quarter technical strategy, architect for global scale and reliability, and drive platform-level improvements that directly impact... 
    Local area

    Coinbase

    Washington DC
    16 hours ago
  •  ...Overview: The Staff Software Engineer L5 works across all service aspects of high-throughput and multi-tenant systems, with the ability...  ...Kubernetes; apply cloud-native patterns to support platform reliability and scalability; Leverage observability and monitoring... 
    H1b

    Inovalon

    Bowie, MD
    1 day ago
  • $100 per hour

     ...Staff Software Engineer Washington, D.C We're on a mission to bring joy to the business of events. Weddings, galas, festivals... the moments...  ...of AI, and you'll accelerate how the team keeps systems reliable alongside building net new product. Who you are You... 
    Work at office
    Work from home

    Goodshuffle Pro

    Washington DC
    2 days ago
  •  ...Responsibilities Own the reliability outcomes for customer-facing calls, voicemail, and activation, including connection, reachability...  ...with the reliability PM to turn diagnosis into an actionable engineering backlog. Requirements ~5+ years building and operating... 
    Full time

    Tin Can

    Washington DC
    26 days ago
  •  ...deployment pipelines for releasing new models reliably. Provide high-performance inference...  .... Requirements Significant software engineering experience, particularly with...  ...Benefits Hybrid work policy requiring staff to work from an Anthropic office at least... 
    Full time
    Work at office
    Worldwide
    Visa sponsorship
    Flexible hours

    Anthropic

    Washington DC
    26 days ago
  • $300k - $380k

     ...Golang), AWS, Rust​, FastAPI The Role We're seeking a Staff Backend Engineer to lead our core product development. Our backend serves a...  ...founders to set standards and build a highly performant, reliable, and secure app from the ground up. This role may evolve... 
    Flexible hours

    AHU Technologies Inc

    Washington DC
    25 days ago
  • $209.1k - $282.9k

    As a Software Engineer on our AI Inference Runtime team, you will set technical direction for critical components of distributed Inference...  ...validation, and safe-rollout systems to improve latency, throughput, reliability, and resource efficiency.Partner with cloud, framework,... 
    Work at office
    Local area

    ARM

    Washington DC
    6 days ago
  • $174k - $299k

     ...unparalleled reputation for being leading and reliable force in South Korean commerce. We are...  ...You will be a part of the Observability Engineering team at Coupang to build and maintain...  ...OSS solutions, as well as building software components from scratch. You would work... 
    Temporary work
    Work experience placement
    Flexible hours

    Coupang

    Washington DC
    29 days ago
  •  ...with data scientists and machine learning engineers to integrate models and algorithms into...  ...for data pipelines to support reliability, performance, and high availability....  ...role. Three years of experience with software development in Python, Java, or Scala;... 
    Full time
    Temporary work
    Work at office
    Remote work

    The Trade Desk

    Washington DC
    3 days ago
  • $208k - $260k

     ...providing secure, simple, and reliable ways to manage their money,...  ...the moments that matter most. Engineers on these teams apply a full-stack...  ...Canada send business. This Staff Engineer will work across...  ...systems concepts.10+ years of software development experience, including... 
    Full time
    Work at office
    Worldwide
    Flexible hours

    Remitly

    Washington DC
    a month ago
  •  ...recommendations that help customers run workloads reliably at scale. Beyond query observability,...  ...across all these surfaces, raising the engineering bar of the combined team, and shaping...  ....Champion reliable, high-quality software and the operational practices that let a... 
    Worldwide

    DataBricks

    Washington DC
    a month ago
  • $177.19k - $364.8k

     ...offsite conversions in a privacy-preserving way. We are hiring a Staff Software Engineer to lead the backend architecture and implementation of a...  ..., add evaluation and safety checks, and make the agent reliable enough for always-on monitoring and decision support and make... 
    Work at office
    Local area
    Relocation
    Relocation package

    Pinterest

    Washington DC
    a month ago
  • $250k - $300k

     ...build with us at Crusoe. About This Role As a Senior Staff Software Engineer for SDN Architecture, you will be a driving technical force...  ...Product, Hardware, and Infrastructure teams to deliver a reliable, scalable networking fabric that aligns with Crusoe's mission... 
    Temporary work

    Crusoe

    Washington DC
    25 days ago
  • $131.6k - $210.3k

     .... Progress starts with you. Job Description The Staff Software Engineer is a senior technical leader responsible for designing, building...  ..., SAML, mTLS, IAM).  ~ Experience with observability and reliability (Splunk, Datadog, Grafana, OpenTelemetry, SLOs).  ~... 
    Full time
    Contract work
    Work experience placement
    Work at office
    Local area

    VISA

    Washington DC
    4 days ago
  • $240k - $285k

     ...management, bill pay, and travel software, Brex enables founders and...  ...need to grow your career.Engineering at BrexEngineering at Brex is...  ...leaders.What you’ll doAs a Staff Software Engineer in Banking...  ...infrastructure, driving architecture, reliability, and execution across that... 
    Bank staff
    Work at office
    Remote work
    Work from home
    3 days per week

    Brex

    Washington DC
    a month ago
  •  ...rapidly growing healthcare technology company is seeking a Staff Software Engineer to help architect and scale a mission-critical platform...  ...practices around performance, observability, maintainability, and reliability Evaluate technical tradeoffs and make sound... 
    Remote work

    Prevail Recruiting

    Washington DC
    25 days ago
  • $140.6k - $173.1k

     ....About the Team/RoleThe Data Platform Engineering Team acts as the backbone of our enterprise...  ...-grade platform.We are seeking a Staff Software Engineer (Semantic Foundationa) with 6...  ...solutions to ensure high platform uptime and reliability.Central Data Platform, Tools &... 
    Full time
    Remote work
    Flexible hours

    WEX

    Washington DC
    2 days ago
  •  ...looking for a talented Distributed Systems Engineer to join our core payroll team and play a...  ...systems with high availability and reliability, targeting four or five 9s uptime.What you...  ...+ years of professional experience as a software engineerProficiency in a modern programming... 
    Work at office
    3 days per week

    Rippling

    Washington DC
    6 days ago
  •  ...transportation and ground-up build autonomous robotaxis that are safe, reliable, clean, and enjoyable for everyone. We are still in the...  ...ML use cases. You will work alongside a team of strong software engineers and act as a force multiplier for our internal customers.... 

    Zoox

    Washington DC
    23 days ago
  •  ...millions of patients and members they serve. Overview: The Staff Software Development Engineer L5 works across all service aspects of high-throughput...  ...patterns such as circuit breakers to ensure platform reliability; Design and implement API Gateway configurations and... 
    H1b

    RXinsider LTD.

    Bowie, MD
    5 days ago
  • $174k - $238k

     ...talk.The Federal SRE TeamWe are looking for an experienced Staff Site Reliability Engineer to join Okta's Federal SRE team for the Emerging Products...  ...within the EPG SRE organization, partnering closely with software engineers, architects, and product teams to design, build... 
    Local area
    Worldwide
    Flexible hours

    Okta, Inc.

    Washington DC
    22 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Staff Software Engineer, Reliability. Be the first to apply!