Member of Technical Staff - AI Infrastructure Reliability
Jobleads-US
About Us:
Fireworks is the platform for specialized intelligence, enabling companies to build, train, and serve AI models tailored to their own data, workflows, and products. Founded by the team behind PyTorch and backed by AMD, Atreides, Benchmark Capital, Index Ventures, Lightspeed, NVIDIA, Sequoia Capital, and TCV, Fireworks powers production AI with hundreds of state-of-the-art open models across text, image, embedding, audio, and multimodal workloads. Today, Fireworks is a Series D company valued at $17.5 billion, bringing together an ambitious, collaborative team that's building the future of enterprise AI.
About the Role
Fireworks AI is one of the industry leaders in inference and training for open models. Open models are how the rest of the world gets to build on frontier AI without handing the keys to a single vendor, and our job is to make them fast, cheap, and dependable enough that this is a real choice. That work is systems work: GPU scheduling, kernel and runtime performance, networking, storage, Linux. We serve over 40 trillion tokens a day doing it.
Reliability Engineering makes sure that platform runs dependably as it grows. You will work across cloud infrastructure, AI systems, and product teams to make sure the pieces fit together, fail gracefully, and hold up under load.
How We Think About Ownership
You own the bar. You define what "reliable" means at Fireworks: SLOs, error budgets, production readiness, on-call expectations. Then you drive adoption across engineering.
You own the process and the tooling. Incident management, postmortems, observability standards, failure testing, guardrails, and automation are yours end to end.
Every team owns the reliability of what they build. You make that ownership practical. Structured logging and aggregation, metrics and tracing that work the same way everywhere, alerting that routes to the right owner, dashboards that answer "why is this slow."
You choose where the leverage is. You have a wide view of the platform and the latitude to spend your time where it changes outcomes most.
Responsibilities
Define reliability standards: SLOs, error budgets, production readiness criteria. Not written in a vacuum: you instrument the systems and read the real telemetry the numbers come from.
Own the reliability toolchain: Logging and telemetry pipelines, alerting standards, failure injection, load testing, self-healing automation, and AI-assisted investigation tooling.
Keep customer experience from falling through the cracks: Per-service reliability is necessary but not sufficient. A customer can hit a bad experience while every system sits inside its SLO. You make sure those failures get an owner and a fix.
Own the seams: The hardest failures live between systems: retries that amplify load, timeouts that do not compose, dependencies nobody mapped. You find them before customers do and drive fixes through the teams that own them.
Run incident management: Coordinate live production issues, run blameless postmortems, and track follow-ups to completion.
Reduce toil: Automate repetitive operational work so growth does not turn into an unsustainable on-call load.
Partner across the org: Cloud infrastructure on capacity and multi-region risk, inference and training on failure modes in the serving and training stacks, performance on zero-downtime rollouts, product and control plane on customer-facing reliability.
Qualifications
Systems fundamentals: 5+ years with Linux internals, system performance troubleshooting, and networking fundamentals (TCP/IP, gRPC).
Software engineering: 5+ years in Python, Go, C++, or Rust, writing production-grade tools and systems code.
Cloud-native operations: Operating and debugging Kubernetes, Terraform, and Docker in high-throughput production.
Distributed systems: High-throughput control planes, microservices, or multi-region setups.
Reliability fundamentals: Fault-tolerant design, SLO/SLA management, automated failover, high-availability architecture.
Influence without authority: You can get other teams to adopt a standard through credibility and useful tooling rather than mandate.
Breadth over comfort: Willingness to dig into unfamiliar parts of the stack when a problem crosses boundaries.
Education: Bachelor's or Master's in Computer Science, Computer Engineering, or equivalent practical experience.
Preferred Qualifications
Observability tooling: Prometheus, Grafana, OpenTelemetry, and alerting people actually act on.
GPU and ML infrastructure exposure: GPUs, inference serving, or distributed training.
AI-assisted operations: Building agents or LLM-based tooling for investigation, triage, or automation.
Open source background: Contributions to infrastructure, systems, or ML serving projects.
Startup agility: Comfortable where pragmatism and teamwork matter more than process.
Why Fireworks?
Solve Hard Problems: Tackle challenges at the forefront of AI infrastructure, from low-latency inference to scalable model serving.
Build What's Next: Work with bleeding-edge technology that impacts how businesses and developers harness AI globally.
Ownership & Impact: Join a fast-growing, passionate team where your work directly shapes the future of AI -no bureaucracy, just results.
Learn from the Best: Collaborate with world-class engineers and AI researchers who thrive on curiosity and innovation.
Fireworks AI is an equal-opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all innovators.
#J-18808-Ljbffr Jobleads-US$120k - $200k
...Member of Technical Staff — Infrastructure San Francisco / New York · $120,000 – $200,000 + equity AI has advanced by expanding what machines can represent. Deep learning enabled... ..., while preserving performance, reliability, and security across all deployments....Suggested$160k - $300k
⚡ Member of Technical Staff (QA Engineer - Agentic Systems) AI-Powered Pharmaceutical Marketing New York City (on-site) $160k - $300k... ...build the testing, evaluation, and quality infrastructure that keeps the platform reliable, compliant, and production-ready....SuggestedFull timeRelocation package- ...Finch, we're building the infrastructure to make justice radically... ...operators and purpose-built AI work together – which is... ...software to fix it. As a Member of the Technical Staff at Finch, you'll own critical... ...release. Build reliable, performant systems across...SuggestedWork at officeRemote workFlexible hours1 day per week
- ...Responsibilities Deploy and scale our AI-agent infrastructure. Own observability end to end. Own CI/CD. Make shipping fast and safe. Treat developer experience as a product. Build for the product, not just the platform. Requirements ~4+ years of...Suggested
- ...Member of Technical Staff Shared Context is building adaptive personal AI that understands the texture of real life: our relationships... ...core agent, product, and infrastructure systems. Create... ...long-horizon agent behavior, reliability, and safety. Design systems...Suggested
$220k - $405k
...Perplexity is hiring a Member of Technical Staff (Software Engineer, Enterprise... ...software engineering and applied AI, focused on making the... ...operations. Engineer scalable infrastructure: Make knowledge stores,... ...process needs into simple, reliable product experiences in...Full timeWork at office- ...Job Description The role As a Member of Technical Staff focused on Applied AI, you'll own our AI stack end to... ...logic into something coherent and reliable. Generative UI and human-in-the... ...just take turns. Evaluation infrastructure that holds two bars at once: high...Work at office1 day per week
- ...DoorDash, and Ramp. About the Role Members of Technical Staff (MTS) are the senior engineers who... ...at its core. Multi-tenant data infrastructure across very different portcos. Event-... ...developer tooling projects. How We Use AI in Our Hiring Process: To ensure transparency...
$300 per month
...Delangue and many other operators/technical leaders. _"Basis is on the... ...." — Prashant Mital, Applied AI Lead, OpenAI_ The Work Being a Member of Technical Staff at Basis means you'll face... ...) ~ Platform & Infrastructure ~ Agent Data (e.g...Work at officeShift work$200k - $250k
...digital assets and data center infrastructure, delivering solutions that... ...infrastructure to power AI and high-performance computing... ...ll move fluidly between deep technical exploration and shipping production... ...operating systems with real reliability, security, and auditability...Local areaFlexible hours- ...At Inductive, we're building the AI and machine learning tools to radically... ...drug discovery. We are seeking a Member of Technical Staff, Machine Learning to join our talented... ...industry Build and optimize scalable infrastructure for model training, deployment, and...
- ...Member of Technical Staff: Software Engineering Asymmetric is the first AI-native Digital Forensics and Incident Response company.... ...artefacts from customer tenants reliably and within their API limits.... ...and least-privilege infrastructure for a security product that...
$120k - $250k
...,000 - $250,000 + equity AI has advanced by expanding what... ...building the missing layer: infrastructure that makes human experience... ...systems, make foundational technical decisions, and help establish... ...methods for evaluating the reliability under real-world conditions...Immediate start$90.9k - $254.1k
...Job Description About the Role As a Member of Technical Staff focused on AI, you will design applied AI systems and build the products around... ...powered applications, agent systems, and the evaluation and reliability systems required to run them in production. The role...Work at office$180k - $300k
...builds security into a complex AI-native product, not one who... ...Tenant Isolation. ListenLabs staff, customer admins, and study... ..., sessions, RBAC), cloud and infrastructure security basics, secrets management... ...Room to grow : As an early member of the team, you'll own end-...Work at officeFlexible hoursShift work$185k - $250k
...labor-intensive manual processes. Stuut’s AI agent works across order management,... .... The Role We’re hiring a Member of Technical Staff – Full Stack, Credit to help build Stuut... ...decisions explainable, auditable, and reliable as we expand the platform. What You...Full timeWorldwideFlexible hours- ...About the Role As a Member of Technical Staff - Applied AI at Entendre, you will design and ship user-facing products that combine cutting-edge AI... ...models, and frontend components that are understandable and reliable. Integrate large language model (LLM) capabilities—...Full time
$150k - $300k
...About the Role Join an applied AI engineering team building systems that change... ...across business workflows. Develop the infrastructure, interfaces, and operating approaches... ...through launch, ensuring the systems work reliably in practice. Understand the underlying...Remote workVisa sponsorship$150k - $220k
...effective mental healthcare. Our AI-native care platform is... ...from day one and shape both technical decisions and product... ...continuously improve AI quality, reliability, and performance Work across... ...systems, and application infrastructure to solve real-world product...Work at office- ...Member of Technical Staff Location: NYC (onsite only – not remote) Alliance is the leading accelerator for crypto & AI founders. Since 2020 we've backed 300+ startups (Rain, Pump, Synthetix... ...quality : code reviews, CI/CD, reliability, and accessibility. What we're...Full timeTemporary workRelocation
$150k - $300k
...Job Description Job Description Member of Technical Staff - Atlas Company: Basis Location:... ...Basis started from the belief that AI agents would become integral to knowledge... ...take over entire jobs. Invent new infrastructure, interfaces and operating models for...Full timeH1bWork at officeVisa sponsorship- ...Member of Technical Staff: Agents, Evaluations, and Environments Asymmetric is the first AI-native Digital Forensics and Incident Response company. Our platform has helped... ...in Python and comfortable building the infrastructure your experiments need. Think quantitatively...
$150k - $300k
...Job Description Job Description Member of Technical Staff – AI Agents & Internal Platforms Location: New York, NY – Flatiron Employment... ...Tech Stack: TypeScript, Python, AI Agents, Infrastructure About the Role What does knowledge work look like when...Full timeH1bWork at officeImmediate startVisa sponsorship$200k - $250k
...Job Title: Member of Technical Staff, Medical Research Job Type: Full-time Location: Remote... ...Research, where you'll help advance AI systems capable of supporting healthcare... ...for healthcare AI evaluation and reliability. What We're Looking For...Full timeRemote work- ...Member of Technical Staff, Machine Learning Pace is an AI-native business process outsourcer for insurers. We combine the speed of AI agents with expert review by our insurance operations team to automate insurance tasks. Almost $400bn per year is spent on outsourcing...
- ...unstructured data at scale including calls, chats, emails, and AI agent responses to build autonomous systems that adapt, detect,... ...Ruby or Python, with experience in data pipelines, AWS or GCP infrastructure, and orchestration frameworks. Experience deploying production...
- ...OffDeal is the world’s first AI-native investment bank for small businesses. Instead of selling AI to Goldman Sachs, we’... ...started. Read more about our story: The Role As a Member of Technical Staff, you'll work directly with the CTO to build the AI-native platform...Work experience placementRelocation package
- ...deeper understanding in healthcare. Our AI-powered platform was purpose-built for medical... .... The ideal candidate will bring technical mastery, fluency with foundation models,... ...new things and recognize that great team members might not perfectly match a job description...Hourly payFull timeCurrently hiringWork at officeRelocation packageFlexible hours
$150k - $300k
...Venture Partners to acquire and modernize CPA firms using Agentic AI. We streamline financial audits from the ground up — starting... ...pipelines, document parsing, browser automation, and backend infrastructure. You’ll shape the product and the company. What you’ll do...Work at office- ...As a Member of Technical Staff at Quadrillion Labs, you'll build the systems that power Qualia, our research agent. This... ...tooling researchers actually touch, and scaling the infrastructure that keeps all of it fast and reliable as we grow - and the product work of turning a...Work at officeLocal area
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Member of Technical Staff - AI Infrastructure Reliability. Be the first to apply!
- application support technician New York, NY
- junior IT service management analyst New York, NY
- infrastructure support analyst New York, NY
- operations support technician New York, NY
- technical solutions specialist New York, NY
- help desk technical support New York, NY
- desktop support analyst New York, NY
- it technical specialist New York, NY
- customer support technician New York, NY
- technical support representative New York, NY



