Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Member of Technical Staff, Model Efficiency

Cohere

Member of Technical Staff, Model Efficiency

Who are we?

Our mission is to scale intelligence to serve humanity. We’re training and deploying frontier models for developers and enterprises who are building AI systems to power magical experiences like content generation, semantic search, RAG, and agents. We believe that our work is instrumental to the widespread adoption of AI.

We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. We like to work hard and move fast to do what’s best for our customers.

Cohere is a team of researchers, engineers, designers, and more, who are passionate about their craft. Each person is one of the best in the world at what they do. We believe that a diverse range of perspectives is a requirement for building great products.

Join us on our mission and shape the future!

Why this role?

Our team is a fast-growing group of researchers and engineers focused on building reliable ML systems and pushing the boundaries of LLM inference efficiency. We develop techniques that improve how models execute in production, driving lower latency, higher throughput, and consistent quality across diverse workloads.

As an engineer on this team, you’ll work across the inference stack to improve core performance metrics by diving deep into model execution, identifying bottlenecks, and developing innovative optimizations. You’ll collaborate closely with modeling and systems teams to experiment, measure, and ship improvements that meaningfully accelerate inference. As the team evolves, you’ll have opportunities to build expertise in advanced performance techniques, including GPU/CUDA optimizations, kernel-level improvements, and model execution strategies for MoE and large‑scale architectures.

We have offices in Toronto, Montreal, San Francisco, New York, Paris, Seoul, and London. Remote‑friendly environment, with preferred locations in EST and PST time zones.

You may be a good fit for the Model Efficiency team if you have:

  • 5+ years of experience writing high‑performance, production‑quality code
  • Strong programming skills in C++ or Python (Rust/Go also welcome)
  • Experience working with large language models and familiarity with the LLM inference ecosystem (e.g., vLLM, SGLang, etc.)
  • Ability to diagnose and resolve performance bottlenecks across the model execution stack
  • A strong bias for action — you ship fast, measure impact, and iterate

It’s a big plus if you have experience with:

  • GPU programming, CUDA, or low‑level systems optimization
  • Language modeling with transformers (MoE, speculative decoding, KV‑cache optimizations)
  • Scaling performance‑critical distributed systems (e.g., computation, search, storage)

If some of the above doesn’t line up perfectly with your experience, we still encourage you to apply!

We value and celebrate diversity and strive to create an inclusive work environment for all. We welcome applicants from all backgrounds and are committed to providing equal opportunities. Should you require any accommodations during the recruitment process, please submit an Accommodations Request Form, and we will work together to meet your needs.

Full‑time employees at Cohere enjoy these perks

  • An open and inclusive culture and work environment
  • Work closely with a team on the cutting edge of AI research
  • Weekly lunch stipend, in‑office lunches & snacks
  • Full health and dental benefits, including a separate budget to take care of your mental health
  • 100% parental leave top‑up for up to 6 months
  • Personal enrichment benefits towards arts and culture, fitness and well‑being, quality time, and workspace improvement
  • Remote‑flexible, offices in Toronto, New York, San Francisco, London, and Paris, as well as a co‑working stipend
  • 6 weeks of vacation (30 working days!)

Seniority level

Mid‑Senior level

Employment type

Full‑time

Job function

Engineering and Information Technology

Industries: Software Development

Referrals increase your chances of interviewing at Cohere by 2×.

#J-18808-Ljbffr
Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Member of Technical Staff, Model Efficiency in San Francisco, CA vacancy
  •  ...Face) and other industry leaders. The Opportunity Language model evaluation is the sharpest question in AI: what can these systems...  ..., are the reference the industry uses. We’re hiring Members of Technical Staff to build the next generation of them. This is a role for people... 
    Suggested

    Artificial Analysis

    San Francisco, CA
    4 days ago
  • We're partnering with a frontier AI research company on a search for a Member of Technical Staff focused on AI Safet y. The company is building next-generation open-weight foundation models with a mission to make advanced AI broadly accessible. Their team includes researchers... 
    Suggested

    Xcede

    San Francisco, CA
    23 hours ago
  • Job You will own the training pipeline behind the models that power both Parallel’s search stack and Parallel’s agents. On the search side, that means the rankers, classifiers, and query models that surface the right information. On the agent side, that means the models... 
    Suggested
    Work at office
    Visa sponsorship

    Parallel Web Systems

    San Francisco, CA
    2 days ago
  •  ...assembling a founding core engineering team to build and train models that understand these systems, optimize operations, anticipate...  ...collaborate across the hardware and software stack. Want to build the technical DNA of a new applied research org from the ground up. #J-18808... 
    Suggested

    Meter

    San Francisco, CA
    1 day ago
  • A leading AI research firm in San Francisco is seeking a Member of Technical Staff specialized in Model Efficiency. In this role, you will enhance LLM inference systems by tackling performance issues and collaborating with cross-functional teams. Ideal candidates have over... 
    Suggested
    Remote work

    Cohere

    San Francisco, CA
    1 day ago
  • $227.5k - $401k

     ...individuals who tackle unique technical challenges at scale and...  ...financial technology sector. As a Member of Technical Staff , you will operate with a...  ...development of foundation models . For instance, you might...  ...record of writing clean, efficient, and scalable code suitable... 
    Work at office
    Immediate start
    Relocation
    Flexible hours

    Adyen

    San Francisco, CA
    23 hours ago
  • $240k - $280k

     ...poster from Cabana Senior Engineer / Member of Technical Staff @ AI Healthcare Startup $240,000 - $2...  ...deliver cutting-edge treatments more efficiently Direct collaboration with healthcare...  ...Manager Senior Software Engineer, GenAI Model Quality Senior Software Quality... 
    Full time
    Remote work
    Worldwide
    Relocation

    Cabana

    San Francisco, CA
    23 hours ago
  •  ...breakthrough AI application, from foundation models to autonomous vehicles, relies on...  ...district office. Your Role As a Member of Technical Staff, you will be responsible for building...  ...level primitives when performance and efficiency demand it. We look for engineers with... 
    Work at office
    Immediate start
    Flexible hours
    Night shift

    Socket

    San Francisco, CA
    2 days ago
  • $200k - $300k

     ...autonomous supply chain to power fast, efficient and economical commerce. We’re...  ...agency. The Role Nimble is looking for a Member of Technical Staff to help us advance our robotics moonshot...  ...and implementing robotic foundation models, Vision‑Language‑Action Models,... 
    Local area
    Immediate start
    Flexible hours
    Weekend work

    Nimble Robotics

    San Francisco, CA
    23 hours ago
  •  ...Pixeltable Inc. Member of Technical Staff San Francisco, CA·Full time Apply for Member of Technical...  ...engineers to focus on building impactful models, not wrangling with complex data...  ...heterogeneous compute resources (CPU and GPU) efficiently? What data model will enable us to... 
    Full time
    Part time
    Work at office
    Work from home
    Flexible hours
    2 days per week

    Pixeltable, Inc.

    San Francisco, CA
    1 day ago
  • $180k - $300k

     ...world creates images and video. We’re creating the generative models that power how people make images and video—tools used by millions...  ...depth over noise, collaboration over hero culture, and honest technical conversations over hype. Our models have been downloaded... 
    Remote work
    Worldwide
    2 days per week

    Black Forest Labs

    San Francisco, CA
    4 days ago
  • $70k - $110k

     ...our dynamic engineering team. As a key member of our team, you will be responsible for...  ...for individuals looking to apply their technical skills and knowledge in a challenging and...  ...to the highest standards of safety and efficiency. 2. Conduct regular inspections of HVAC... 
    Temporary work
    Local area

    Jobot

    San Francisco, CA
    10 hours ago
  •  ...In-person collaboration. Join a lean, staff-level team (ex-Affirm, Uber, DoorDash,...  ...About the Role We’re looking for a Member of Technical Staff to help reshape how the insurance...  ...application. Write clean, maintainable, and efficient code with appropriate test coverage.... 
    Work at office
    Immediate start
    Relocation

    Fulcrum

    San Francisco, CA
    3 days ago
  •  ...person can. We're looking for founding members of technical staff: engineers who will co-own Hivemind's...  ...§3 — What You'll Do. Build fast, efficient infrastructure that enables social...  ...scale. Own context management, model orchestration, and systems that let one... 
    Full time

    Hivemind

    San Francisco, CA
    1 day ago
  •  ...precedents to copy from. About the Role Members of Technical Staff (MTS) are the senior engineers who...  ...internal tools. The canonical data model that survives contact with very different...  ...made by humans. We use AI to support efficiency and consistency, not to replace human... 

    Beacon Software

    San Francisco, CA
    2 days ago
  • $150k - $300k

     ...stack - from frontier agentic models to the infra that enables...  ...infrastructure to serve LLMs efficiently at scale. Optimization and integration...  ...our RL training stack. Core Technical Responsibilities LLM Serving...  ...and encourage team members to contribute to the broader... 
    Work at office
    Remote work
    Visa sponsorship
    Relocation package
    Flexible hours
    Shift work

    Prime Intellect

    San Francisco, CA
    1 day ago
  •  ...multi-silicon neocloud designed for fast, efficient inference. As AI workloads become more...  ...systems research problems grounded in frontier models, cutting-edge production workloads, and...  ...build and run this company. As an early member of the team, you will have significant... 

    The Consensus

    San Francisco, CA
    4 days ago
  • $150k - $300k

    United States Digital Space LLC is hiring a Member of Technical Staff to work onsite in New York City or San Francisco. This full-time role offers...  ...data sources and leverage AI to enhance operational efficiency. The position demands a hands-on approach to problem-solving... 
    Full time

    United States Digital Space LLC

    San Francisco, CA
    2 days ago
  •  ...Token Company trains machine learning models to compress raw LLM inputs before...  ...a research and product focus. As a Member of Technical Staff on our infrastructure team, you'll own...  ...and scaling to reliability and cost-efficiency. This is a very high ownership role where... 
    Visa sponsorship

    The Token Company

    San Francisco, CA
    23 hours ago
  •  ...hardware that best fits its performance and efficiency needs. This approach enables...  ...datacenters. Gimlet Labs is seeking a Member of Technical Staff focused on ML systems and inference....  ...inference systems that execute full models end-to-end under real production constraints... 

    Gimlet Labs

    San Francisco, CA
    4 days ago
  •  ...funded company pioneering a new model of acquisition-led growth....  ...company-building, with each member having previously steered AI...  ...designing platforms that unlock efficiency, scale, and profitability....  ...practical systems. Act as the technical lead for this research direction... 

    Enam, Inc.

    San Francisco, CA
    23 hours ago
  •  ...funded company pioneering a new model of acquisition-led growth....  ...AI and company-building, with members who have helped build AI...  ...building platforms that unlock efficiency, scale, and profitability. What...  ...Computer Science or a related technical field strongly preferred.... 

    Enam, Inc.

    San Francisco, CA
    23 hours ago
  •  ...about our vision, team, and backers at . About the Role As a Member of Technical Staff, you will help invent and build the next generation of...  ...systems performance. Experience with power management, energy‑efficient computing, sustainability, or data center infrastructure.... 
    Work from home
    Flexible hours
    2 days per week

    Emerald AI

    San Francisco, CA
    4 days ago
  •  ...We build cutting‑edge foundation AI models and end‑to‑end products that are designed...  ...matter, and join the team. As a member of technical staff with a focus on multimodal AI, you will...  .... Bonus: Experience in writing efficient GPU kernels using CUDA, optimising performance... 
    Full time
    Work at office
    Local area
    Remote work
    Home office

    Cohere

    San Francisco, CA
    2 days ago
  •  ...hardware that best fits its performance and efficiency needs. This approach enables...  ...datacenters. Gimlet Labs is seeking a Member of Technical Staff focused on compilers. In this role, you...  ..., and translating emerging AI models and execution patterns into production... 

    Gimlet Labs

    San Francisco, CA
    4 days ago
  • Careers / Member of Technical Staff (AI research) Member of Technical Staff (AI research) You will...  ...s platform from deploying our custom models to building frontend experiences....  ...About our interview process. We run an efficient and thorough interview process to ensure... 
    Full time
    Work at office

    Kindredventures

    San Francisco, CA
    3 days ago
  •  ...customize, and build on. We build open models that let anyone control their...  ...Overview Reflection AI is looking for a Member of Technical Staff - IT Engineer. In this role, you’ll be...  ...will help the IT function operate more efficiently Act as the on‑prem IT lead at our New... 
    Work at office
    Visa sponsorship

    Reflection

    San Francisco, CA
    1 day ago
  • $160k - $250k

     ...trading to optimize ad spend, delivering 20-50% upside in spend efficiency. We are managing spend for global category leaders like Cider...  ...We are looking for a Machine Learning Engineer to build the models, optimization systems and algorithms that drive our autonomous... 
    Full time
    Immediate start
    Relocation
    Relocation package

    Pepr Ai

    San Francisco, CA
    23 hours ago
  • $175k - $240k

     ...biology, physics, chemistry, and AI. The Role As a Member of Technical Staff, Infrastructure Engineer, you'll play a key role in designing...  ...(agents, jobs, services) with high availability and efficient resource utilization. Drive the strategy for cluster scaling... 
    Full time
    Work at office

    Edison Scientific

    San Francisco, CA
    2 days ago
  • $160k - $220k

     ...hard real-time engineering problems. Founded 2025 · ~10 people · Industry: AI Tools / Voice AI infrastructure The Role As a Member of Technical Staff, you\'ll architect and build the systems that simulate, analyze, and evaluate conversational AI agents across voice and... 
    Full time

    David Joseph & Company

    San Francisco, CA
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Member of Technical Staff, Model Efficiency. Be the first to apply!