Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Member of Technical Staff - AI Infrastructure Reliability

Jobleads-US

About Us:

Fireworks is the platform for specialized intelligence, enabling companies to build, train, and serve AI models tailored to their own data, workflows, and products. Founded by the team behind PyTorch and backed by AMD, Atreides, Benchmark Capital, Index Ventures, Lightspeed, NVIDIA, Sequoia Capital, and TCV, Fireworks powers production AI with hundreds of state-of-the-art open models across text, image, embedding, audio, and multimodal workloads. Today, Fireworks is a Series D company valued at $17.5 billion, bringing together an ambitious, collaborative team that's building the future of enterprise AI.

About the Role

Fireworks AI is one of the industry leaders in inference and training for open models. Open models are how the rest of the world gets to build on frontier AI without handing the keys to a single vendor, and our job is to make them fast, cheap, and dependable enough that this is a real choice. That work is systems work: GPU scheduling, kernel and runtime performance, networking, storage, Linux. We serve over 40 trillion tokens a day doing it.

Reliability Engineering makes sure that platform runs dependably as it grows. You will work across cloud infrastructure, AI systems, and product teams to make sure the pieces fit together, fail gracefully, and hold up under load.

How We Think About Ownership

  • You own the bar. You define what "reliable" means at Fireworks: SLOs, error budgets, production readiness, on-call expectations. Then you drive adoption across engineering.

  • You own the process and the tooling. Incident management, postmortems, observability standards, failure testing, guardrails, and automation are yours end to end.

  • Every team owns the reliability of what they build. You make that ownership practical. Structured logging and aggregation, metrics and tracing that work the same way everywhere, alerting that routes to the right owner, dashboards that answer "why is this slow."

  • You choose where the leverage is. You have a wide view of the platform and the latitude to spend your time where it changes outcomes most.

Responsibilities

  • Define reliability standards: SLOs, error budgets, production readiness criteria. Not written in a vacuum: you instrument the systems and read the real telemetry the numbers come from.

  • Own the reliability toolchain: Logging and telemetry pipelines, alerting standards, failure injection, load testing, self-healing automation, and AI-assisted investigation tooling.

  • Keep customer experience from falling through the cracks: Per-service reliability is necessary but not sufficient. A customer can hit a bad experience while every system sits inside its SLO. You make sure those failures get an owner and a fix.

  • Own the seams: The hardest failures live between systems: retries that amplify load, timeouts that do not compose, dependencies nobody mapped. You find them before customers do and drive fixes through the teams that own them.

  • Run incident management: Coordinate live production issues, run blameless postmortems, and track follow-ups to completion.

  • Reduce toil: Automate repetitive operational work so growth does not turn into an unsustainable on-call load.

  • Partner across the org: Cloud infrastructure on capacity and multi-region risk, inference and training on failure modes in the serving and training stacks, performance on zero-downtime rollouts, product and control plane on customer-facing reliability.

Qualifications

  • Systems fundamentals: 5+ years with Linux internals, system performance troubleshooting, and networking fundamentals (TCP/IP, gRPC).

  • Software engineering: 5+ years in Python, Go, C++, or Rust, writing production-grade tools and systems code.

  • Cloud-native operations: Operating and debugging Kubernetes, Terraform, and Docker in high-throughput production.

  • Distributed systems: High-throughput control planes, microservices, or multi-region setups.

  • Reliability fundamentals: Fault-tolerant design, SLO/SLA management, automated failover, high-availability architecture.

  • Influence without authority: You can get other teams to adopt a standard through credibility and useful tooling rather than mandate.

  • Breadth over comfort: Willingness to dig into unfamiliar parts of the stack when a problem crosses boundaries.

  • Education: Bachelor's or Master's in Computer Science, Computer Engineering, or equivalent practical experience.

Preferred Qualifications

  • Observability tooling: Prometheus, Grafana, OpenTelemetry, and alerting people actually act on.

  • GPU and ML infrastructure exposure: GPUs, inference serving, or distributed training.

  • AI-assisted operations: Building agents or LLM-based tooling for investigation, triage, or automation.

  • Open source background: Contributions to infrastructure, systems, or ML serving projects.

  • Startup agility: Comfortable where pragmatism and teamwork matter more than process.

Why Fireworks?

  • Solve Hard Problems: Tackle challenges at the forefront of AI infrastructure, from low-latency inference to scalable model serving.

  • Build What's Next: Work with bleeding-edge technology that impacts how businesses and developers harness AI globally.

  • Ownership & Impact: Join a fast-growing, passionate team where your work directly shapes the future of AI -no bureaucracy, just results.

  • Learn from the Best: Collaborate with world-class engineers and AI researchers who thrive on curiosity and innovation.

Fireworks AI is an equal-opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all innovators.

#J-18808-Ljbffr Jobleads-US
Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Member of Technical Staff - AI Infrastructure Reliability in New York, NY vacancy
  • $120k - $200k

     ...Member of Technical Staff — Infrastructure San Francisco / New York · $120,000 – $200,000 + equity AI has advanced by expanding what machines can represent. Deep learning enabled...  ..., while preserving performance, reliability, and security across all deployments.... 
    Suggested

    Jobleads-US

    New York, NY
    2 days ago
  • $160k - $300k

    ⚡ Member of Technical Staff (QA Engineer - Agentic Systems) AI-Powered Pharmaceutical Marketing New York City (on-site) $160k - $300k...  ...build the testing, evaluation, and quality infrastructure that keeps the platform reliable, compliant, and production-ready.... 
    Suggested
    Full time
    Relocation package

    Storm3

    New York, NY
    3 days ago
  •  ...Finch, we're building the infrastructure to make justice radically...  ...operators and purpose-built AI work together – which is...  ...software to fix it. As a Member of the Technical Staff at Finch, you'll own critical...  ...release. Build reliable, performant systems across... 
    Suggested
    Work at office
    Remote work
    Flexible hours
    1 day per week

    Finch Services

    New York, NY
    3 days ago
  •  ...Responsibilities Deploy and scale our AI-agent infrastructure. Own observability end to end. Own CI/CD. Make shipping fast and safe. Treat developer experience as a product. Build for the product, not just the platform. Requirements ~4+ years of... 
    Suggested

    Jobtailor

    New York, NY
    3 days ago
  •  ...Member of Technical Staff Shared Context is building adaptive personal AI that understands the texture of real life: our relationships...  ...core agent, product, and infrastructure systems. Create...  ...long-horizon agent behavior, reliability, and safety. Design systems... 
    Suggested

    Shared Context Lab

    New York, NY
    4 days ago
  • $220k - $405k

     ...Perplexity is hiring a Member of Technical Staff (Software Engineer, Enterprise...  ...software engineering and applied AI, focused on making the...  ...operations. Engineer scalable infrastructure: Make knowledge stores,...  ...process needs into simple, reliable product experiences in... 
    Full time
    Work at office

    Jobleads-US

    New York, NY
    3 days ago
  •  ...Job Description The role As a Member of Technical Staff focused on Applied AI, you'll own our AI stack end to...  ...logic into something coherent and reliable. Generative UI and human-in-the...  ...just take turns. Evaluation infrastructure that holds two bars at once: high... 
    Work at office
    1 day per week

    NxT Level

    New York, NY
    a month ago
  •  ...DoorDash, and Ramp. About the Role Members of Technical Staff (MTS) are the senior engineers who...  ...at its core. Multi-tenant data infrastructure across very different portcos. Event-...  ...developer tooling projects. How We Use AI in Our Hiring Process: To ensure transparency... 

    BEACON SOFTWARE COMPANY

    New York, NY
    1 day ago
  • $300 per month

     ...Delangue and many other operators/technical leaders. _"Basis is on the...  ...." — Prashant Mital, Applied AI Lead, OpenAI_ The Work Being a Member of Technical Staff at Basis means you'll face...  ...) ~ Platform & Infrastructure ~ Agent Data (e.g... 
    Work at office
    Shift work

    Basis

    New York, NY
    4 days ago
  • $200k - $250k

     ...digital assets and data center infrastructure, delivering solutions that...  ...infrastructure to power AI and high-performance computing...  ...ll move fluidly between deep technical exploration and shipping production...  ...operating systems with real reliability, security, and auditability... 
    Local area
    Flexible hours

    Galaxy Digital

    New York, NY
    18 hours ago
  •  ...At Inductive, we're building the AI and machine learning tools to radically...  ...drug discovery. We are seeking a Member of Technical Staff, Machine Learning to join our talented...  ...industry Build and optimize scalable infrastructure for model training, deployment, and... 

    Inductive Bio

    New York, NY
    3 days ago
  •  ...Member of Technical Staff: Software Engineering Asymmetric is the first AI-native Digital Forensics and Incident Response company....  ...artefacts from customer tenants reliably and within their API limits....  ...and least-privilege infrastructure for a security product that... 

    Jobleads-US

    New York, NY
    1 day ago
  • $120k - $250k

     ...,000 - $250,000 + equity AI has advanced by expanding what...  ...building the missing layer: infrastructure that makes human experience...  ...systems, make foundational technical decisions, and help establish...  ...methods for evaluating the reliability under real-world conditions... 
    Immediate start

    Jobleads-US

    New York, NY
    1 day ago
  • $90.9k - $254.1k

     ...Job Description About the Role As a Member of Technical Staff focused on AI, you will design applied AI systems and build the products around...  ...powered applications, agent systems, and the evaluation and reliability systems required to run them in production. The role... 
    Work at office

    Jobleads-US

    New York, NY
    2 days ago
  • $180k - $300k

     ...builds security into a complex AI-native product, not one who...  ...Tenant Isolation. ListenLabs staff, customer admins, and study...  ..., sessions, RBAC), cloud and infrastructure security basics, secrets management...  ...Room to grow : As an early member of the team, you'll own end-... 
    Work at office
    Flexible hours
    Shift work

    Jobleads-US

    New York, NY
    9 hours ago
  • $185k - $250k

     ...labor-intensive manual processes. Stuut’s AI agent works across order management,...  .... The Role We’re hiring a Member of Technical Staff – Full Stack, Credit to help build Stuut...  ...decisions explainable, auditable, and reliable as we expand the platform. What You... 
    Full time
    Worldwide
    Flexible hours

    Stuut

    New York, NY
    9 days ago
  •  ...About the Role As a Member of Technical Staff - Applied AI at Entendre, you will design and ship user-facing products that combine cutting-edge AI...  ...models, and frontend components that are understandable and reliable. Integrate large language model (LLM) capabilities—... 
    Full time

    Entendre Finance

    New York, NY
    12 days ago
  • $150k - $300k

     ...About the Role Join an applied AI engineering team building systems that change...  ...across business workflows. Develop the infrastructure, interfaces, and operating approaches...  ...through launch, ensuring the systems work reliably in practice. Understand the underlying... 
    Remote work
    Visa sponsorship

    Jobleads-US

    New York, NY
    1 day ago
  • $150k - $220k

     ...effective mental healthcare. Our AI-native care platform is...  ...from day one and shape both technical decisions and product...  ...continuously improve AI quality, reliability, and performance Work across...  ...systems, and application infrastructure to solve real-world product... 
    Work at office

    Blossom Health

    New York, NY
    23 days ago
  •  ...Member of Technical Staff Location: NYC (onsite only – not remote) Alliance is the leading accelerator for crypto & AI founders. Since 2020 we've backed 300+ startups (Rain, Pump, Synthetix...  ...quality : code reviews, CI/CD, reliability, and accessibility. What we're... 
    Full time
    Temporary work
    Relocation

    DeFi Alliance

    New York, NY
    2 days ago
  • $150k - $300k

     ...Job Description Job Description Member of Technical Staff - Atlas Company: Basis Location:...  ...Basis started from the belief that AI agents would become integral to knowledge...  ...take over entire jobs. Invent new infrastructure, interfaces and operating models for... 
    Full time
    H1b
    Work at office
    Visa sponsorship

    Transparent Search Group

    New York, NY
    16 days ago
  •  ...Member of Technical Staff: Agents, Evaluations, and Environments Asymmetric is the first AI-native Digital Forensics and Incident Response company. Our platform has helped...  ...in Python and comfortable building the infrastructure your experiments need. Think quantitatively... 

    Jobleads-US

    New York, NY
    1 day ago
  • $150k - $300k

     ...Job Description Job Description Member of Technical Staff – AI Agents & Internal Platforms Location: New York, NY – Flatiron Employment...  ...Tech Stack: TypeScript, Python, AI Agents, Infrastructure About the Role What does knowledge work look like when... 
    Full time
    H1b
    Work at office
    Immediate start
    Visa sponsorship

    Morgan Pinnacle Group

    New York, NY
    a month ago
  • $200k - $250k

     ...Job Title: Member of Technical Staff, Medical Research Job Type: Full-time Location: Remote...  ...Research, where you'll help advance AI systems capable of supporting healthcare...  ...for healthcare AI evaluation and reliability. What We're Looking For... 
    Full time
    Remote work

    Pro Integrate

    New York, NY
    2 days ago
  •  ...Member of Technical Staff, Machine Learning Pace is an AI-native business process outsourcer for insurers. We combine the speed of AI agents with expert review by our insurance operations team to automate insurance tasks. Almost $400bn per year is spent on outsourcing... 

    Pace

    New York, NY
    18 hours ago
  •  ...unstructured data at scale including calls, chats, emails, and AI agent responses to build autonomous systems that adapt, detect,...  ...Ruby or Python, with experience in data pipelines, AWS or GCP infrastructure, and orchestration frameworks. Experience deploying production... 

    Rulebase: The voice fraud defense system for financial servi...

    New York, NY
    18 hours ago
  •  ...OffDeal is the world’s first AI-native investment bank for small businesses. Instead of selling AI to Goldman Sachs, we’...  ...started. Read more about our story: The Role As a Member of Technical Staff, you'll work directly with the CTO to build the AI-native platform... 
    Work experience placement
    Relocation package

    OffDeal, Inc.

    New York, NY
    3 days ago
  •  ...deeper understanding in healthcare. Our AI-powered platform was purpose-built for medical...  .... The ideal candidate will bring technical mastery, fluency with foundation models,...  ...new things and recognize that great team members might not perfectly match a job description... 
    Hourly pay
    Full time
    Currently hiring
    Work at office
    Relocation package
    Flexible hours

    Abridge Al, Inc

    New York, NY
    4 days ago
  • $150k - $300k

     ...Venture Partners to acquire and modernize CPA firms using Agentic AI. We streamline financial audits from the ground up — starting...  ...pipelines, document parsing, browser automation, and backend infrastructure. You’ll shape the product and the company. What you’ll do... 
    Work at office

    Modus

    New York, NY
    4 days ago
  •  ...As a Member of Technical Staff at Quadrillion Labs, you'll build the systems that power Qualia, our research agent. This...  ...tooling researchers actually touch, and scaling the infrastructure that keeps all of it fast and reliable as we grow - and the product work of turning a... 
    Work at office
    Local area

    Quadrillion Labs, Inc

    New York, NY
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Member of Technical Staff - AI Infrastructure Reliability. Be the first to apply!