Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Member of Technical Staff, Model Efficiency (San Francisco)

Full-time

Cohere

Member of Technical Staff, Model Efficiency

Who are we?

Our mission is to scale intelligence to serve humanity. We’re training and deploying frontier models for developers and enterprises who are building AI systems to power magical experiences like content generation, semantic search, RAG, and agents. We believe that our work is instrumental to the widespread adoption of AI.

We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. We like to work hard and move fast to do what’s best for our customers.

Cohere is a team of researchers, engineers, designers, and more, who are passionate about their craft. Each person is one of the best in the world at what they do. We believe that a diverse range of perspectives is a requirement for building great products.

Join us on our mission and shape the future!

Why this role?

Our team is a fast-growing group of researchers and engineers focused on building reliable ML systems and pushing the boundaries of LLM inference efficiency. We develop techniques that improve how models execute in production, driving lower latency, higher throughput, and consistent quality across diverse workloads.

As an engineer on this team, you’ll work across the inference stack to improve core performance metrics by diving deep into model execution, identifying bottlenecks, and developing innovative optimizations. You’ll collaborate closely with modeling and systems teams to experiment, measure, and ship improvements that meaningfully accelerate inference. As the team evolves, you’ll have opportunities to build expertise in advanced performance techniques, including GPU/CUDA optimizations, kernel-level improvements, and model execution strategies for MoE and large‑scale architectures.

We have offices in Toronto, Montreal, San Francisco, New York, Paris, Seoul, and London. Remote‑friendly environment, with preferred locations in EST and PST time zones.

You may be a good fit for the Model Efficiency team if you have:

  • 5+ years of experience writing high‑performance, production‑quality code
  • Strong programming skills in C++ or Python (Rust/Go also welcome)
  • Experience working with large language models and familiarity with the LLM inference ecosystem (e.g., vLLM, SGLang, etc.)
  • Ability to diagnose and resolve performance bottlenecks across the model execution stack
  • A strong bias for action — you ship fast, measure impact, and iterate

It’s a big plus if you have experience with:

  • GPU programming, CUDA, or low‑level systems optimization
  • Language modeling with transformers (MoE, speculative decoding, KV‑cache optimizations)
  • Scaling performance‑critical distributed systems (e.g., computation, search, storage)

If some of the above doesn’t line up perfectly with your experience, we still encourage you to apply!

We value and celebrate diversity and strive to create an inclusive work environment for all. We welcome applicants from all backgrounds and are committed to providing equal opportunities. Should you require any accommodations during the recruitment process, please submit an Accommodations Request Form, and we will work together to meet your needs.

Full‑time employees at Cohere enjoy these perks

  • An open and inclusive culture and work environment
  • Work closely with a team on the cutting edge of AI research
  • Weekly lunch stipend, in‑office lunches & snacks
  • Full health and dental benefits, including a separate budget to take care of your mental health
  • 100% parental leave top‑up for up to 6 months
  • Personal enrichment benefits towards arts and culture, fitness and well‑being, quality time, and workspace improvement
  • Remote‑flexible, offices in Toronto, New York, San Francisco, London, and Paris, as well as a co‑working stipend
  • 6 weeks of vacation (30 working days!)

Seniority level

Mid‑Senior level

Employment type

Full‑time

Job function

Engineering and Information Technology

Industries: Software Development

Referrals increase your chances of interviewing at Cohere by 2×.

#J-18808-Ljbffr
Vacancy posted 3 hours ago
Similar jobs that could be interesting for youBased on the Member of Technical Staff, Model Efficiency (San Francisco) in San Francisco, CA vacancy
  • A leading AI research firm in San Francisco is seeking a Member of Technical Staff specialized in Model Efficiency. In this role, you will enhance LLM inference systems by tackling performance issues and collaborating with cross-functional teams. Ideal candidates have... 
    Suggested
    Full time
    Remote work

    Cohere

    San Francisco, CA
    3 hours ago
  • $240k - $280k

     ...Cabana Senior Engineer / Member of Technical Staff @ AI Healthcare Startup $2...  ...cutting-edge treatments more efficiently Direct collaboration with...  ...-person collaboration in San Francisco with mission-driven teammates...  ...Software Engineer, GenAI Model Quality Senior Software... 
    Suggested
    Full time
    Remote work
    Worldwide
    Relocation

    Cabana

    San Francisco, CA
    3 hours ago
  • $170k - $220k

     ...Member of Technical Staff – Infrastructure & LLMs Location: San Francisco, CA (Hybrid) Compensation: $170,000 – $220,000 base + 1–3% equity Work Authorization:...  ...Kubernetes (or equivalent) Focus: Batch inference, model distillation, low‑latency pipelines Soft... 
    Suggested
    Full time
    Temporary work
    Immediate start
    Visa sponsorship
    Work visa

    Amadeus Search

    San Francisco, CA
    3 hours ago
  • $130k - $200k

    Member of Technical Staff, Founding Frontend Engineer Join to apply for the Member...  ...the team. Location: San Francisco, CA, In-Person\...  ...handle millions of entities efficiently. Design Collaborative Canvas...  ...Member of Technical Staff, Model Serving We’re unlocking community... 
    Suggested
    Full time
    Work at office

    Ocular AI (YC W24)

    San Francisco, CA
    3 hours ago
  • $256k - $276k

     ...Overview Member of Technical Staff, AI Reliability & Monitoring Engineering Lead — Postman...  ..., particularly GPU/accelerator efficiency, ensuring cost-effective AI...  ..., we embrace a hybrid work model. For roles based in the San Francisco Bay Area, Boston, Bangalore, Hyderabad... 
    Suggested
    Full time
    Work at office
    Flexible hours
    3 days per week

    Postman

    San Francisco, CA
    3 hours ago
  • $100k - $300k

     ...early team is fully in-person in San Francisco and New York and consists of...  ...ambitious Backend Senior and Staff Engineers who are excited to...  ...and uplevel future team members Participate in, provide feedback...  ...as a hands‑on engineer and technical leader, overseeing and... 
    Full time

    Cogent Security

    San Francisco, CA
    3 hours ago
  • $256k - $276k

     ...Member of Technical Staff, AI Agent Development Lead Who Are We? Postman is the world’s leading...  .... The company is headquartered in San Francisco and has offices in Boston, New York,...  ...leveraging state‑of‑the‑art language models and associated technologies. Collaborate... 
    Full time
    Work at office
    Flexible hours
    3 days per week

    Postman

    San Francisco, CA
    3 hours ago
  • $300k

     ...We're hiring Senior Engineers that will be part of a deeply technical, fully hands‑on engineering team, tackling problems that if solved...  ...of interviewing at Amigos by 2x Get notified about new Senior Software Engineer jobs in San Francisco Bay Area . #J-18808-Ljbffr... 
    Full time

    Amigos

    San Francisco, CA
    3 hours ago
  • $200k

     ...Overview Tzafon is a foundation model lab building scalable compute systems and advancing machine intelligence, with offices in San Francisco, Zurich & Tel Aviv. We’ve raised over $12m in funding to advance our mission of expanding the frontiers of machine intelligence... 
    Full time
    Work at office
    Visa sponsorship

    Tzafon

    San Francisco, CA
    3 hours ago
  • $150k - $200k

     ...Member of Technical Staff, Founding Backend Engineer Join to apply for the Member of Technical Staff...  ..., Computer Vision, and Enterprise AI models. We help companies transform...  ...Staff, DevSecOps / Infrastructure San Francisco, CA $136,947.00 - $239,699.00 5... 
    Full time
    Work at office
    Shift work

    Ocular AI (YC W24)

    San Francisco, CA
    3 hours ago
  •  ...Member of Technical Staff – Machine Learning I’m partnering with a rapidly scaling healthtech startup that has just raised a $40M Series A to...  ...Work across the full ML lifecycle: from data pipelines, to model training, to deployment Adapt and fine-tune large foundation... 
    Full time
    Work at office

    Quantix Search

    San Francisco, CA
    3 hours ago
  • $200k

     ...Join to apply for the Member of Technical Staff role at Listen Labs . TL;DR: We are seeing strong market demand and an aggressive 6‑month...  ...—layers of consulting intelligence. Customer Preference Model & Synthetic Personas: Building profound understanding of... 
    Full time
    Flexible hours

    Listen Labs

    San Francisco, CA
    3 hours ago
  • $100k - $200k

     ...Member of Technical Staff, Computer Vision / Graphics Join to apply for the Member of Technical Staff, Computer Vision / Graphics role at Outerport...  ...Collect data, annotate, and train computer vision models / VLMs Write beautiful visualization code for both debugging... 
    Full time

    Outerport

    San Francisco, CA
    3 hours ago
  • $120k - $180k

     ...Software Engineer to build systems that leverage physics-informed models. The ideal candidate will have strong software engineering...  ...ownership and innovation. This is a full-time position located in San Francisco, California, with a competitive salary ranging from $120,000... 
    Full time

    Godela (YC X25)

    San Francisco, CA
    3 hours ago
  •  ...Crusoe in San Francisco is seeking a Senior Director for the Model LifeCycle team to establish a managed platform for the development lifecycle of Machine Learning models. The ideal candidate will have over 10 years of experience in AI, strong leadership skills, and expertise... 
    Full time

    Crusoe

    San Francisco, CA
    3 hours ago
  • Vapi (/ˈwɑːpi/): We’re creating the shift to voice as humanity’s default interface. We’re the most configurable platform for deploying voice agents. We’re grown to 400,000 developers in 20 months, adding 2,000+ every day. Try talking to Vapi now! Why We’re...
    Full time
    Flexible hours
    Shift work

    Vapi Inc.

    San Francisco, CA
    3 hours ago
  •  ...engineer the next interface for intelligence. you are either the best in the world at a certain function or a brilliant 15 to 20-something-year-old with a steep growth curve you want to explore the edge of whats technically possible and whats useful #J-18808-Ljbffr
    Part time

    Attention Engineering

    San Francisco, CA
    3 hours ago
  • $120k - $220k

     ...recruiter to learn more. Base pay range $120,000.00/yr - $220,000.00/yr Responsibilities Work on in-house protein language models — design, build, & test; you will be implementing your own ideas. Work with the lab team to produce the best proteins possible.... 
    Full time
    Relocation

    Anthrogen

    San Francisco, CA
    3 hours ago
  • A leading AI technology firm in San Francisco seeks a Machine Learning Engineer to join their AI/ML team. The role involves creating innovative AI experiences, training models, and developing unique features to wow customers. Candidates should have extensive experience... 
    Full time

    Lightfield

    San Francisco, CA
    3 hours ago
  • $130k - $200k

     ...technology. The Role Being a Member of Technical Staff at SketchPro means the...  ...agent represents a Revit model in context, then shift to...  ...problems and craft reliable and efficient solutions. Thrives in...  ...experience preferred In-person in San Francisco, 5 days a week (our office... 
    Work at office
    Shift work

    SketchPro.ai

    San Francisco, CA
    1 day ago
  • $70k - $110k

     ...engineering team. As a key member of our team, you will be responsible...  ...looking to apply their technical skills and knowledge in a...  ...highest standards of safety and efficiency. 2. Conduct regular...  ...Initiative for Hiring and the San Francisco Fair Chance Ordinance. Information... 
    Temporary work
    Local area

    Jobot

    San Francisco, CA
    5 days ago
  • $227.5k - $401k

     ...individuals who tackle unique technical challenges at scale...  ...engineering team in San Francisco to drive our next...  ...sector. As a Member of Technical Staff, you will operate with...  ...of foundation models. For instance, you might...  ...record of writing clean, efficient, and scalable code... 
    Work at office
    Immediate start
    Relocation
    Flexible hours

    Adyen

    San Francisco, CA
    5 days ago
  •  ...Pixeltable Inc. Member of Technical Staff San Francisco, CA·Full time Apply for Member of Technical Staff As...  ...engineers to focus on building impactful models, not wrangling with complex data...  ...compute resources (CPU and GPU) efficiently? What data model will enable us to... 
    Full time
    Part time
    Work at office
    Work from home
    Flexible hours
    2 days per week

    Pixeltable, Inc.

    San Francisco, CA
    1 day ago
  • $250k

     ...San Francisco, CA · On-site · Full-time Compensation: $...  .... The team is small, technical, and moving fast, with...  ...: AI Tools. The Role Member of Technical Staff who can handle everything from modeling to systems to product...  ...latency, throughput, cost efficiency, and reliability of... 
    Full time

    David Joseph & Company

    San Francisco, CA
    1 day ago
  • $150k - $250k

     ...organization as well as serve as a member of our IT engineering team....  ...based in Sydney, Melbourne, San Francisco or Singapore....  ...serve Airwallex, building a model to improve and scale their function...  ...applications and serve as the technical owner or primary IT engineer... 
    Full time
    Worldwide

    Airwallex

    San Francisco, CA
    3 hours ago
  • $301.75k - $355k

     ...This Role The Senior Director for the Model LifeCycle team will undertake a pivotal...  ...checkpointing, failure recovery, and cost‑efficient scaling. Implement and maintain end‑...  ...AI products and solving challenging technical problems. Bonus Points PhD in Machine... 
    Full time
    Temporary work

    Crusoe Energy Systems LLC

    San Francisco, CA
    3 hours ago
  •  ...Member of Technical Staff – Full Stack / AI Systems Company : AdsGency AI Relocation : San Francisco City Required Authorization : Applicants must be permanently authorized to work...  ...knowledge of distributed systems, data models, and service orchestration. Ability to design... 
    Full time
    Work experience placement
    Relocation
    Visa sponsorship

    AdsGency AI

    San Francisco, CA
    5 days ago
  •  ...reshape how people discover and buy online. Role As a Member of Technical Staff, you will ship core systems, set engineering culture, and...  ...Want high ownership in a small, talent-dense team in San Francisco. Example problems Design and ship agentic-... 
    Work at office

    Catalog.com

    San Francisco, CA
    4 days ago
  •  ...researcher with strong experience in generative modeling. You will join an interdisciplinary...  ...culture of trust across our London and San Francisco sites. We’re looking for innovators...  ...setting where goals must be achieved efficiently and urgently. What sets you apart (preferred... 
    Flexible hours

    Latent Labs Ltd.

    San Francisco, CA
    2 days ago
  •  ...mighty team with an office in San Francisco and others distributed...  ...to production Contribute to technical discussions and help improve...  ...MongoDB Have experience with AI models or are excited to learn how...  ...engineering fundamentals, write efficient code, and have a clear... 
    Work at office

    Mixpeek

    San Francisco, CA
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Member of Technical Staff, Model Efficiency (San Francisco). Be the first to apply!