Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Software Architect

Andromeda

Full-Time Location: North America Remote/SF-Hybrid Full-Time About Andromeda Andromeda Cluster was founded by Nat Friedman and Daniel Gross to give early-stage startups access to the kind of scaled AI infrastructure once reserved only for hyperscalers. We began with a single managed cluster — but it filled almost instantly. Since then, we’ve been quietly building the systems, network, and orchestration layer that makes the world’s AI infrastructure more accessible. Today, Andromeda works with leading AI labs, data centers, and cloud providers to deliver compute when and where it’s needed most. Our platform routes training and inference jobs across global supply, unlocking flexibility and efficiency in one of the fastest-growing markets on earth. Our long-term vision is to build the liquidity layer for global AI compute. We are expanding to new frontiers to find the brightest that work in AI infrastructure, research and engineering. The Role This role owns our provider relationships on the technical side. You're the person who decides whether a provider's cluster is good enough to join the Andromeda network, and the person who helps them get there when it isn't. That means vetting prospective providers against our quality bar and then working alongside their engineers to bring their clusters onto the network cleanly. You’ll work closely with our compute procurement team to identify and qualify new providers, and you’ll be the standing technical relationship with the providers we already have. When procurement finds a promising provider, you're the one who validates the claims. When a provider says their fabric is non-blocking and their nodes are burn-in tested, you're the one who verifies it by reviewing their evidence, and sometimes by running the tests yourself on their hardware. The qualification bar you'll hold providers to mainly doesn't exist yet in written form. Some of it lives informally in how our SREs evaluate clusters today; much of it hasn't been defined at all. You'd be the first person in this role, so a large part of the job is building that bar: the acceptance test suite, the quality thresholds, and the technical standards a provider must meet. Then making them rigorous enough that we can trust them and clear enough that providers can build to them. This role is the technical counterpart to our Provider Technical Program Manager, who owns delivery timelines, escalations, and incident command. You own the technical judgment: is this cluster ready, what's wrong with it, and what will it take to fix. During provider incidents, command and coordination sit with the program manager and technical response sits with our SREs. Your involvement is upstream, making sure clusters that would have caused those incidents never make it onto the network. What You’ll Do Vet prospective compute providers: assess cluster architecture, GPU hardware, network fabric, storage, and orchestration against Andromeda's quality metrics, and make the call on whether a cluster qualifies. Define the qualification bar itself. Build the acceptance test suite, benchmark methodology, and quality thresholds from scratch, formalizing what currently exists only as SRE tribal knowledge. Own and evolve these standards as the fleet and the market change. Run validation hands-on where it matters: burn-in testing, fabric validation (InfiniBand/RoCE), NCCL and application-level benchmarks, storage performance testing. For routine or repeat validation, define the methodology and review results rather than executing everything yourself. Guide providers through technical onboarding: work directly with their data‑center and platform engineers to remediate gaps, tune configurations, and bring clusters up to the bar on a predictable path. Partner with compute procurement to identify and qualify new providers. Perform technical due diligence during sourcing, and a clear read on how much remediation a candidate cluster needs before procurement commits. Build and maintain technical relationships with existing providers: you're the engineer their engineers call, and the one who spots architectural or quality drift before it becomes a customer problem. Feed what you learn back into provider-facing standards and internal documentation, so each qualification is faster and more consistent than the last. What We’re Looking For Deep HPC experience: you've designed, built, or operated GPU clusters at meaningful scale and understand what makes them fast, stable, and debuggable. Strong fabric knowledge, ideally both InfiniBand and RoCE. You can evaluate a topology, interpret fabric diagnostics, and identify why a fabric underperforms, not just that it does. Experience with distributed orchestration and the HPC software stack: Slurm, Kubernetes, OpenMPI or equivalent, and the Linux systems engineering underneath all of it. Data-center literacy: power, cooling, cabling, and physical‑layer realities. You can walk a provider's facility and know what questions to ask. Benchmarking judgment. You know which numbers matter for large‑scale training workloads, how providers game them, and how to design tests that can't be gamed. The ability to write standards others can build against: precise, testable, and usable by a provider's engineering team without you in the room. Credibility in the room with external engineering teams, including the ability to deliver a failing grade to a provider who wants your business, and keep the relationship intact. Comfort with ambiguity. The qualification framework you'll apply mostly doesn't exist yet; you'll write it. Strong Candidates May Have Experience inside a neocloud, hyperscaler, colocation, or data‑center provider. NVIDIA data‑center GPU depth: DGX/HGX platforms, NVLink/NVSwitch, GPU health and RAS behavior at fleet scale. Experience supporting AI research labs or other large‑scale training customers, and familiarity with what their workloads punish: stragglers, fabric jitter, storage stalls. Background in cluster acceptance testing or site bring‑up. Ideally you've taken a cluster from delivery to production before. What Success Looks Like A qualification bar that’s written down and trusted. Andromeda's acceptance criteria, benchmark suite, and quality thresholds exist as documents and tooling — not tribal knowledge. SREs, procurement, and providers all work from the same bar, and a passing grade from you means the cluster performs in production. Providers who build to our standards before we ever test them. Prospective providers get our technical requirements upfront and arrive closer to qualified, because the standards are clear enough to engineer against. Time‑from‑sourcing‑to‑qualified drops with each provider. Onboarding that's technically boring. Clusters that join the network have been validated the same way every time, and the surprises that used to appear in a customer's training run get caught in acceptance instead. A technical relationship providers value. Provider engineering teams treat you as the authority on what Andromeda needs and the first call when they're planning a build‑out so that we hear about architecture decisions early enough to influence them. A procurement partnership that changes what we buy. Procurement's provider pipeline reflects your technical due diligence, and deals are shaped by a realistic view of remediation cost. Why You’ll Love It Here High‑growth environment: Get in early at a company at the center of the AI infrastructure boom Ownership: First HPC Architect for the solutions engineering team, you’ll get to build this function from the ground up Competitive compensation: + meaningful equity Comprehensive benefits: for you and your dependents, including healthcare, dental, and vision coverage, 401(k), and unlimited PTO Andromeda Cluster is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. We do not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status. #J-18808-Ljbffr

Vacancy posted 11 hours ago
Similar jobs that could be interesting for youBased on the Software Architect in San Francisco, CA vacancy
  • $263k - $353k

     ...building a world where Identity belongs to you. The Engineering Architect Team Auth0 is an easy-to-implement authentication and...  ...provide guidance, build proof of concepts, and deliver production software implementations that help Auth0 Engineering teams move faster... 
    Suggested
    Full time
    Flexible hours

    Okta, Inc.

    San Francisco, CA
    2 days ago
  •  ...Okta is seeking a Senior Software Architect (P6) in San Francisco, California. The role is hybrid, requiring office presence three times a week. Responsibilities include implementing new product demonstrations and leading industry-wide identity protocols. The ideal candidate... 
    Suggested
    Work at office

    Okta, Inc.

    San Francisco, CA
    12 hours ago
  • A leading AI infrastructure startup is seeking a Compiler Tech Lead to drive technical and team leadership. You'll manage a team of engineers, guide architecture decisions, and influence core compiler systems. The ideal candidate has over 8 years of compiler development...
    Suggested

    Amadeus Search

    San Francisco, CA
    12 hours ago
  •  ...views, 20K+ users within two weeks. We’re backed by Y Combinator and angels such as Dane Knecht (Cloudflare). The Context Aside is software people live in every day, so design is core to the product, not a layer added at the end. Our challenge is to combine familiar... 
    Suggested
    Live in
    Work at office
    Immediate start

    Socket

    San Francisco, CA
    5 days ago
  • Nixon Peabody LLP in San Francisco is seeking a Senior Full Stack Developer to lead the development of enterprise-grade web and desktop applications. The role requires extensive experience with the Microsoft technology stack including C#, ASP.NET, and SQL Server while ...
    Suggested
    Remote work

    Nixon Peabody

    San Francisco, CA
    12 hours ago
  • $197.3k - $365.2k

     ...months to ensure you are not duplicating efforts. Job Category Software Engineering Job Category Software Engineering Job Details...  ..., and you are the future of Salesforce. Overview of the Role: Architecting the Core: Drive the end-to-end architecture for mission-critical... 

    Salesforce.Com Inc

    San Francisco, CA
    3 days ago
  •  ...Note: By applying to the architect posting, recruiters and hiring managers across the organization hiring architects will review your...  ...AI/ML Architecture (this can change frequently) Salesforce has Software Architect opportunities throughout the company! These positions... 

    Salesforce

    San Francisco, CA
    12 hours ago
  • $130k - $150k

     ...Job Title Senior Software Architect – Data Center Infrastructure Management (DCIM) Locations Dallas, TX | Austin, TX | Boston, MA | Ashburn, VA | San Francisco, CA | Charlotte, NC | New York, NY | Miami, FL | Denver, CO About the Role As a Senior Software Architect –... 

    Digital Realty

    San Francisco, CA
    2 days ago
  •  ...Software Architect Location: San Francisco Bay Area Hybrid Duration: 6 months Visa: Any Visa (except H1B and CPT) Interview: Phone/Zoom Need two references Duties & Responsibilities: Provide leadership, mentorship, and strategic engineering guidance Hands... 
    H1b

    ShiftCode Analytics

    San Francisco, CA
    5 days ago
  • Principal Software Architect - Ground Automation & Distributed Systems Partnering with an innovative aerospace technology organisation, I am supporting the search for a Principal Software Architect to define and lead the software architecture underpinning next-generation... 

    Strativ Group

    San Francisco, CA
    2 days ago
  •  ...expertise but also bold vision and an unwavering commitment to excellence. Position Summary: Sceye is seeking an experienced Software Architect to define and evolve the software architecture powering our stratospheric High-Altitude Platform Station (HAPS) fleet. This role... 
    Temporary work
    Work at office
    Remote work

    Sceye

    San Francisco, CA
    3 days ago
  • NScale is seeking a Strategic Copilot to the GTM organization to build the RevOps engine from the ground up, aligning sales, marketing, and customer success data into a single, fluid machine. In our hyper-growth environment, you will establish strict data hygiene, build...

    Nscale

    San Francisco, CA
    11 hours ago
  • Sentry is seeking a Deal Desk Analyst to strengthen our Revenue Operations pillar. You will own end-to-end deal processing, optimize CPQ and Salesforce workflows, and drive automation to reduce manual steps for reps across a fast-paced B2B SaaS environment. You’ll work...

    Segment (Twilio)

    San Francisco, CA
    12 hours ago
  •  ...people across the globe. Join us on this journey to redefine resource management—and change lives along the way. The Role As a Software Architect / Solutions Architect at Air Apps, you will be responsible for defining and overseeing the overall system architecture for... 
    Temporary work
    Worldwide

    Air Apps

    San Francisco, CA
    3 days ago
  • Principal Software Architect - Ground Systems & Fleet Automation Location: Moriarty, NM or North Bay, CA Job Type: Full-Time, Exempt Reports to: Director of Operations Who Are We: Sceye is a high-growth technology company, building the next generation of instant communications... 
    Full time
    Remote work
    Flexible hours

    Sceye

    San Francisco, CA
    4 days ago
  •  ...departments and operational autonomy. Candidates should possess deep expertise in Salesforce and have a proven track record of architecting revenue systems in high-growth B2B environments. Excellent benefits, including premium healthcare and a wellness stipend, are provided... 

    Exa Corporation

    San Francisco, CA
    12 hours ago
  •  ...founders from Oculus and Ubiquity6, alongside proven leaders from Meta, Google, and Apple, with deep expertise spanning hardware and software. Join us in shaping a future where computers truly come alive. About the Role We are seeking an engineer living at the... 
    Full time
    Contract work
    Flexible hours

    Sesame

    San Francisco, CA
    11 hours ago
  • $180k - $200k

     ...bottlenecks, and perform low-level debugging. Work with firmware teams to integrate sensor algorithms / ML models into system software. Develop monitoring and observability systems to track model performance, data drift, data quality, and overall system health.... 
    Full time

    Gridware

    San Francisco, CA
    11 hours ago
  • $100 - $120 per hour

     ...your recruiter to learn more. Base pay range $100.00/hr - $120.00/hr Location: Newark CA Overview We are seeking a Salesforce Architect with deep expertise in Sales Cloud and Service Cloud and strong hands‑on development capability. This role will lead solution architecture... 
    Contract work

    DynPro

    San Francisco, CA
    4 days ago
  •  ...pipelines that feed training. Run training and evaluation cycles against partner requirements on a continuous loop through major software releases. Make fast-moving work presentable: dashboards, analyses, documentation, and polished partner-facing deliverables.... 
    Full time
    Internship
    Shift work

    Liquid Ai Inc

    San Francisco, CA
    11 hours ago
  •  ...ID.me is seeking a Senior Software Engineer to focus on client-side experiences for the digital identity wallet. You will design and ship features for both mobile and web clients, working with advanced standards like W3C Verifiable Credentials. The position is based in... 
    Full time
    Work at office

    ID.me

    San Francisco, CA
    11 hours ago
  •  ...and Android engineers while collaborating with product, design, and hardware teams. The ideal candidate will have over 8 years of software development experience and 5 years in engineering management to ensure high-performance mobile solutions for residential solar systems... 

    Qcells North America

    San Francisco, CA
    12 hours ago
  • $70 - $85 per hour

     ...Base pay range $70.00/hr - $85.00/hr Software Engineer Location: Bay Area, CA or Charlotte, NC Start Date: Tentative January start (approximately 10 working days) Pay Range: $70-85/hour Onsite Expectation: Negotiable; ideally Tuesday and Thursday About the Role LHH is... 
    Contract work
    Temporary work
    January start
    Local area

    LHH

    San Francisco, CA
    5 days ago
  •  ...Oura is seeking an experienced Staff Enterprise Data Architect to define and implement the enterprise data vision. You will align technology strategy with business goals, lead data mesh initiatives, and guide senior architects across domains to build scalable, governed... 

    Socket

    San Francisco, CA
    12 hours ago
  •  ...ensure that our team members have what they need to do their best work — both in and out of the office. The Staff Enterprise Data Architect defines the overarching architectural vision and framework for Oura\'s enterprise data ecosystem, ensuring total alignment with long... 
    Work at office
    Local area
    Remote work
    Flexible hours

    Socket

    San Francisco, CA
    12 hours ago
  •  ...deliver immediate emissions and fuel-cost reduction — without replacing engines. We are seeking an Embedded Controls Engineer Lead to architect, implement, and validate the real-time ECU controlling hydrogen-diesel combustion at sea. You will own the firmware stack and... 
    Immediate start
    Worldwide

    New Light

    San Francisco, CA
    2 days ago
  • Multiply Labs in San Francisco, CA, is applying AI to solve complex manipulation challenges in robotic cell therapy manufacturing. We're seeking a Research Scientist specializing in robotic manipulation to advance our intelligent systems. You will collaborate with a multi...

    Multiply Labs

    San Francisco, CA
    3 days ago
  • $100k - $200k

     ...hired directly by an Insight Global client Job Summary: Embedded Software Engineer - San Francisco, CA You will be responsible for...  ...engineer, your breadth of experience should allow you to both architect the high-level system and implement low-level modules. In addition... 
    Permanent employment
    Full time
    Work at office

    Insight Global

    San Francisco, CA
    5 days ago
  • $500 per month

     ...you! Voice AI startup Giga raises $61M Series A DoorDash and Giga Partnership The Role We're looking for a Staff Software Engineer to own our iOS and Android SDKs what our customers use to embed Giga's chat and voice agents into their mobile apps. You'... 
    Immediate start

    Giga AI Inc

    San Francisco, CA
    1 day ago
  • Ando is a messaging platform where AI agents can take on work alongside human. We're rebuilding Slack from the ground up around two core ideas: durable memory and agents as first-class participants. While we are pushing the frontier, we still have classic messaging ...
    Work from home

    Ando

    San Francisco, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Software Architect. Be the first to apply!