Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Staff Software Engineer, Reliability

$203.5k - $248.5k

Metropolis

Job Description

Job Description

Who we are

The real world is the next frontier, and at Metropolis, we are creating the artificial intelligence to make it responsive. We are pioneering the Recognition Economy — a future where mundane repetition disappears and being known unlocks access, comfort, and belonging everywhere you go. From transforming parking into a seamless drive-in, drive-out experience for millions of Members to expanding our intelligence layer across retail and hospitality, we are building a world that feels instinctive and magical. The future isn't coming; it's here, and we need builders, innovators, and problem solvers to help us create it.

Who you are

Metropolis is seeking a Staff Software Engineer focused on Reliability to own reliability across the entire Metropolis platform and drive the comprehensive practices that ensure system availability, resilience, and observability for our mission-critical mobility infrastructure. In this role, you will build reliability from first principles, architecting failover systems, implementing chaos engineering, and improving our observability foundation to maintain 99.9%+ uptime as we scale to new markets. 
As the technical owner of our reliability posture, you will tackle challenges like external service failover, dependency mirroring, and database replication, working alongside highly technical teams across the organization to influence architecture decisions and establish company-wide reliability standards. You will join the Product Foundations team, playing a key role in building the foundational infrastructure that powers the future of mobility commerce.

What you'll do
  • Own the overall reliability posture for the Metropolis platform, establishing practices, metrics, and systems that ensure 99.9%+ uptime across all services
  • Design and implement automatic failover mechanisms for critical external dependencies like Twilio for SMS/voice and Stripe for payments with circuit breakers, retry policies, and degraded mode operations
  • Architect and build active-passive or active-active regional deployment strategies with database replication, automated failover, and DNS-based traffic routing including disaster recovery planning and testing
  • Establish comprehensive monitoring using Datadog for APM, logs, and metrics correlation
  • Implement synthetic monitoring, SLO-based alerting, on-call rotation, and escalation policies while building service health dashboards that show customer impact
  • Own the incident management process including workflows, tooling, post-mortem culture, runbook automation, and MTTR reduction initiatives to drive down mean time to recovery from detection to resolution
  • Drive adoption of resilience patterns across all services including health checks, graceful degradation, feature flags, rate limiting, backpressure mechanisms, and chaos engineering practices
  • Build and maintain local mirrors for critical dependencies with artifact caching, dependency pinning, and vulnerability scanning to prevent build failures from upstream outages
What we're looking for
  • 8+ years of engineering experience including software engineering, reliability engineering, SRE practices, or production operations at scale
  • Demonstrate expert-level reliability engineering skills including hands-on experience with multi-region architectures, failover automation, circuit breakers, chaos engineering, and disaster recovery
  • Utilize production observability expertise with deep experience implementing monitoring, alerting, tracing, and logging systems at scale – specifically Datadog or similar APM platforms in high-load environments
  • Apply strong systems thinking with proven ability to design resilient distributed systems that gracefully handle failures, network partitions, and external dependency outages
  • Demonstrate database and data systems knowledge including replication strategies, backup/restore procedures, connection pooling, query optimization, and experience with both relational and NoSQL databases
  • Leverage cloud platform expertise with production experience operating and ensuring reliability of systems on AWS including multi-region deployments, load balancing, and DNS-based failover
  • Possess experience with AI-powered development tools such as Claude Code, GitHub Copilot, or similar agentic coding tools for enhanced productivity – context engineering in particular
  • Exhibit excellent technical communication with ability to influence technical decisions across teams, document complex systems, conduct post-mortems, and establish reliability standards organization-wide
  • Demonstrate expert-level Java and/or Scala proficiency with strong understanding of JVM performance, concurrency, and operational characteristics
While not required, these are a plus:
  • Bring Scala experience
  • Possess SRE or Reliability Engineering experience at companies known for operational excellence such as Google, Amazon, Netflix or high-growth startups where you built reliability practices from the ground up
  • Demonstrate incident response leadership including experience building incident management processes, conducting blameless post-mortems, and driving MTTR reduction initiatives in production environments
  • Utilize chaos engineering experience with tools like Chaos Monkey, Gremlin, or similar, including designing and executing game days and failure injection testing
  • Exhibit performance optimization experience with profiling, benchmarking, capacity planning, and system tuning at hyperscale including experience optimizing for high-throughput, low-latency systems
  • Show open source contributions or technical blog writing that demonstrates depth of expertise in reliability engineering, distributed systems, or production operations
Our Stack
  • Languages + Frameworks: TypeScript, React, Scala (principally), Java (limited)
  • Datastores: MySQL, PostgreSQL, Snowflake
  • Cloud: AWS
  • Version control: Git & GitHub
  • AI Tooling: Copilot on GitHub
  • Observability: Datadog

4 Days in Office: Metropolis values in-person collaboration to drive innovation, strengthen culture, and enhance the Member experience. Our corporate team members hold to our office-first model, which requires employees to be on-site at least four days a week, fostering organic interactions that spark creativity and connection

When you join Metropolis, you'll join a team of world-class product leaders and engineers, building an ecosystem of technologies at the intersection of parking, mobility, and real estate. Our goal is to build an inclusive culture where everyone has a voice and the best idea wins. You will play a key role in building and maintaining this culture as our organization grows. The anticipated base salary for this position is $203,500 USD to $248,500 USD annually. The actual base salary offered is determined by a number of variables, including, as appropriate, the applicant's qualifications for the position, years of relevant experience, distinctive skills, level of education attained, certifications or other professional licenses held, and the location of residence and/or place of employment. Base salary is one component of Metropolis' total compensation package, which may also include access to or eligibility for healthcare benefits, a 401(k) plan, short-term and long-term disability coverage, basic life insurance, a lucrative stock option plan, bonus plans, and more. #LI-LR1 #LI-Onsite

Metropolis may utilize an automated employment decision tool (AEDT) to assess or evaluate your candidacy for employment or promotion. AEDTs are used to assist in assessing a candidate's application relative to the required job qualifications and responsibilities listed in the job posting.

As part of this process, Metropolis retains data relevant to your candidacy, including personal information, for a period that is reasonably necessary for the use of the tool. If you are hired for the position, your data may become part of your employee records.

Metropolis Technologies is an equal opportunity employer. We make all hiring decisions based on merit, qualifications, and business needs — without regard to race, color, religion, sex (including gender identity, sexual orientation, and pregnancy), national origin, disability, veteran status, or any other protected characteristic under federal, state, or local law.

We are committed to providing a welcoming, accessible hiring experience. If you need a reasonable accommodation for any part of the application or interview process, please reach out to us at View email address on us.fitly.work.

Vacancy posted 2 days ago
Similar jobs that could be interesting for youBased on the Staff Software Engineer, Reliability in Seattle, WA vacancy
  • $129.5k - $214.02k

     ...together. Because at UKG, your work matters—and so do you.   Staff Software Engineer-ENG: We are seeking a highly experienced Staff Software...  ...ensuring high standards of performance, scalability, and reliability. Collaborate with architects on mid-level and high-level... 
    Suggested
    Full time
    Worldwide

    Ukg

    Seattle, WA
    16 hours ago
  • $405k

     ...Anthropic’s mission is to create reliable, interpretable, and steerable AI systems...  ...growing group of committed researchers, engineers, policy experts, and business leaders working...  ...is seeking talented and experienced Staff+ Software Engineers to join our Continuous Integration... 
    Suggested
    Full time
    Work at office
    Visa sponsorship
    Flexible hours

    Anthropic

    Seattle, WA
    16 hours ago
  • $189k - $303k

     ...accessible for all. We are searching for an exceptional Staff-level Backend Software Engineer to join the Aurora Services Engineering team and take...  ...infrastructure to scale our products with high availability and reliability. Collaborate with stakeholders including Security,... 
    Suggested
    Full time
    Remote work

    Aurora Innovation

    Seattle, WA
    16 hours ago
  •  ...users. But it didn’t stop there. They engineered Snowflake to power the Data Cloud, where...  ...and maintaining Snowflake’s security, reliability and performance. The team culture is...  ...mentorship from Principal engineers. AS A STAFF SOFTWARE ENGINEER - IDENTITY & ACCESS... 
    Suggested
    Full time

    Snowflake

    Bellevue, WA
    16 hours ago
  •  ...home day is currently Tuesday.   About the Role As a Staff Software Engineer for the Compute pillar, you will play a critical role in...  ...and low-level semiconductor architecture to enable seamless, reliable cloud provisioning and lifecycle management of a heterogeneous... 
    Suggested
    Full time
    Work at office
    Local area
    Work from home
    Flexible hours

    Lambda

    Bellevue, WA
    16 hours ago
  •  ...uses simulation to develop our driving software, validate our safety, and analyze our real...  ...queues ~ More Info:   Software Engineer - Simulation Data Platform ~8+ years of...  ...new versions of the simulation framework reliably and efficiently ~8+ years of experience... 
    Full time
    Temporary work
    Relocation package

    Zoox

    Seattle, WA
    16 hours ago
  • $210k - $240k

     ...Staff Software Engineer Palo Alto, CA About Typeface We help the world's biggest brands move from brief to fully personalized campaigns...  ...related field ~12+ years of experience in developing scalable, reliable, performant, and secure full-stack applications. ~... 
    Work at office
    Flexible hours
    3 days per week

    Typeface

    Seattle, WA
    1 day ago
  • $207k - $275k

     ...Staff Software Engineer Livingston, NJ / New York, NY / Sunnyvale, CA / San Francisco, CA / Bellevue, WA CoreWeave is The Essential Cloud...  ...Infrastructure organization, responsible for the availability, reliability, scalability, and security of the company's data platform.... 
    Permanent employment
    Full time
    Temporary work
    Casual work
    Work at office
    Flexible hours

    CoreWeave

    Bellevue, WA
    16 hours ago
  • $320k - $405k

     ...Staff Software Engineer, Product San Francisco, CA | New York City, NY | Seattle, WA About Anthropic Anthropic's mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole... 
    Work at office
    Visa sponsorship
    Flexible hours
    Shift work

    Anthropic

    Seattle, WA
    2 days ago
  • $160k - $288k

     ...more information, visit QXO.com. What you'll do: As a Staff Software Engineer on the QXO eCommerce team, you will lead the design,...  ...across engineering, product, and design teams to build secure, reliable, and extensible services and user interfaces that serve millions... 

    QXO

    Seattle, WA
    2 days ago
  • $254k - $336k

     ...ABOUT THE JOB At Anduril, our Software Engineers are at the forefront of defense technology, crafting high-impact, cutting-edge solutions...  ...challenging real-world problems by building scalable and reliable software solutions for our next-generation robotic platforms... 
    Full time
    Work experience placement
    Immediate start

    Anduril Industries

    Bellevue, WA
    16 hours ago
  • $209.7k - $283.8k

     ...To support this growth, we need strong technical ownership to ensure our ML pipelines remain reliable, scalable, and architecturally sound. We are seeking a staff ML engineer to design and evolve the large-scale offline platform. This role focuses on building... 
    Work at office
    Worldwide
    Relocation package

    Unity

    Bellevue, WA
    3 days ago
  • $228k - $240k

     ...exceptional service to seller and buyer clients. As a Staff Software Engineer in the Transaction Journey (TJ) organization, you will use...  ...Champion a culture of operational excellence by defining reliability KPIs, establishing accountability for service health, and... 
    Minimum wage
    Flexible hours

    Compass

    Seattle, WA
    2 days ago
  • $170k - $200k

     ...Staff Software Engineer Over 50,000 customers globally trust our end-to-end, cloud-driven networking solutions. They rely on our top-rated...  ...experiences across Extreme. These systems must be scalable, secure, reliable, observable, adaptable, and safe for enterprise use cases.... 
    Local area
    Work from home
    Flexible hours
    Shift work

    Extreme Networks

    Seattle, WA
    1 day ago
  • $165.69k - $307.71k

     ...celebrated, here you can thrive. Software development teams in WBD's...  ...within Release and Delivery Engineering owns the final leg of WBD's...  ..., providing a simple, reliable path to production built on modern...  ...You are a technically broad Staff engineer who takes complex, ambiguous... 
    Temporary work
    Local area
    Immediate start
    Remote work

    Warner Bros. Discovery

    Bellevue, WA
    2 days ago
  •  ...future. DataRobot’s Fleet team is the engine behind how our platform runs across...  ...velocity. That’s where you come in. As a Staff Software Engineer, you’ll be responsible for...  .... Champion best practices in CI/CD, reliability, container lifecycle management, and dev... 
    Full time
    Local area
    Remote work
    Worldwide
    Flexible hours

    DataRobot

    Seattle, WA
    3 days ago
  • $200k - $260k

     ...Staff Software Engineer (Scopely Explore, Inc.; fka Niantic, Inc., San Francisco, California): Research, design, and develop computer and...  ...utility programs to build AR technology. Build and design reliable, high-throughput, low latency and scalable server and networking... 
    Local area
    Immediate start
    Remote work
    Worldwide

    Scopely

    Bellevue, WA
    4 days ago
  •  ...Responsibilities Own the reliability outcomes for customer-facing calls, voicemail, and activation, including connection, reachability...  ...with the reliability PM to turn diagnosis into an actionable engineering backlog. Requirements ~5+ years building and operating... 
    Full time

    Tin Can

    Seattle, WA
    26 days ago
  • $240k - $285k

     ...management, bill pay, and travel software, Brex enables founders and...  ...need to grow your career.Engineering at BrexEngineering at Brex is...  ...leaders.What you’ll doAs a Staff Software Engineer in Banking...  ...infrastructure, driving architecture, reliability, and execution across that... 
    Bank staff
    Work at office
    Remote work
    Work from home
    3 days per week

    Brex

    Seattle, WA
    a month ago
  •  ...recommendations that help customers run workloads reliably at scale. Beyond query observability,...  ...across all these surfaces, raising the engineering bar of the combined team, and shaping...  ....Champion reliable, high-quality software and the operational practices that let a... 
    Worldwide

    DataBricks

    Bellevue, WA
    a month ago
  • $131.6k - $210.3k

     ...connect the world through the most innovative, convenient, reliable, and secure payments network, enabling individuals, businesses...  ...: We are looking for Versatile, curious, and energetic Software Engineers who embrace solving complex challenges on a global scale. As... 
    Work experience placement
    Work at office
    Local area
    Relocation package

    Visa

    Bellevue, WA
    3 days ago
  • $250k - $300k

     ...build with us at Crusoe. About This Role As a Senior Staff Software Engineer for SDN Architecture, you will be a driving technical force...  ...Product, Hardware, and Infrastructure teams to deliver a reliable, scalable networking fabric that aligns with Crusoe's mission... 
    Temporary work

    Crusoe

    Seattle, WA
    25 days ago
  •  ...with data scientists and machine learning engineers to integrate models and algorithms into...  ...for data pipelines to support reliability, performance, and high availability....  ...role. Three years of experience with software development in Python, Java, or Scala;... 
    Full time
    Temporary work
    Work at office
    Remote work

    The Trade Desk

    Bellevue, WA
    3 days ago
  •  ...transportation and ground-up build autonomous robotaxis that are safe, reliable, clean, and enjoyable for everyone. We are still in the...  ...ML use cases. You will work alongside a team of strong software engineers and act as a force multiplier for our internal customers.... 

    Zoox

    Seattle, WA
    23 days ago
  • $213k - $256k

     ...Staff Software Engineer - Backend Seattle, Washington, United States LVT is redefining how businesses operate in the physical world, moving...  ...to transform complex business requirements into elegant, reliable cloud services while elevating the technical output of the... 
    Full time
    Work at office
    Flexible hours
    3 days per week

    LVT Corp

    Seattle, WA
    2 days ago
  • $182.4k - $247k

     ...growing SaaS companies in the world. Our engineering teams build highly technical products...  ...and operate one of the largest scale software platforms. The fleet consists of millions...  ...machines. Data Plane Storage : Deliver reliable and high performance services and client... 
    Work at office
    Local area
    Worldwide
    Flexible hours

    Databricks

    Seattle, WA
    3 days ago
  • $140k - $193k

     ...Staff Software Engineer - Backend Hybrid - San Francisco, California; Hybrid - Seattle Motive empowers the people who run physical operations...  ...scale. What You'll Do Develop highly observable, reliable, and scalable backend services that process telemetry from... 
    Temporary work
    Work experience placement
    Work at office
    2 days per week

    Motive

    Seattle, WA
    4 days ago
  • $117k - $171k

     ...Staff Software Engineer, Backend At Bayer we're visionaries, driven to solve the world's toughest challenges and striving for a world where...  ...is to ensure those platforms are architected for scale, reliability, and the demands of AI and machine learning at a global level... 

    Bayer Global

    Seattle, WA
    1 day ago
  • $174k - $299k

     ...unparalleled reputation for being leading and reliable force in South Korean commerce. We are...  ...You will be a part of the Observability Engineering team at Coupang to build and maintain...  ...OSS solutions, as well as building software components from scratch. You would work... 
    Temporary work
    Work experience placement
    Flexible hours

    Coupang

    Seattle, WA
    29 days ago
  • $208k - $260k

     ...providing secure, simple, and reliable ways to manage their money,...  ...the moments that matter most. Engineers on these teams apply a full-stack...  ...Canada send business. This Staff Engineer will work across...  ...systems concepts.10+ years of software development experience, including... 
    Full time
    Work at office
    Worldwide
    Flexible hours

    Remitly

    Seattle, WA
    a month ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Staff Software Engineer, Reliability. Be the first to apply!