Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Software Engineer, Inference

Pulse Software Corp

Overview

Pulse is tackling one of the most persistent challenges in data infrastructure: extracting accurate, structured information from complex documents at scale. We have a breakthrough approach to document understanding that combines intelligent schema mapping with fine-tuned extraction models where legacy OCR and other parsing tools consistently fail.

We are a small, fast-growing team of engineers in San Francisco powering Fortune 100 enterprises, YC startups, public investment firms, and growth-stage companies. We are backed by tier 1 investors and growing quickly.

What makes our tech special is our multi-stage architecture:
  • Layout understanding with specialized component detection models
  • Low-latency OCR models for targeted extraction
  • Advanced reading-order algorithms for complex structures
  • Proprietary table structure recognition and parsing
  • Fine-tuned vision-language models for charts, tables, and figures
If you are passionate about the intersection of computer vision, NLP, and data infrastructure, your work at Pulse will directly impact customers and shape the future of document intelligence.

What we are looking for
  • 5 days in-office at our San Francisco office
  • Eager to learn and adapt quickly
  • Prior startup or founding experience is a plus
What we are looking for
  • 5 days in-office at our San Francisco office
  • Eager to learn and adapt quickly
  • Prior startup or founding experience is a plus
About the Role Specialize in low-latency, high-throughput inference for OCR and multimodal models. Own profiling, batching, and autoscaling across single-tenant and multi-tenant environments.

Responsibilities
  • Build inference services with smart batching and caching
  • Optimize kernels, tokenization, and model graphs
  • Evaluate vLLM, TensorRT LLM, and Triton tradeoffs
  • Implement autoscaling and admission control with clear SLOs
  • Own performance dashboards and capacity planning
Requirements
  • 3+ years in performance engineering or ML systems
  • Strong Python, plus C++ or CUDA exposure
  • Experience with GPU profiling and model serving
Nice to have
  • Experience reducing p95 and cost in production ML systems

Sponsorship Sponsorship available.

Compensation and benefits Competitive base salary plus equity, performance-based bonus, relocation assistance for Bay Area moves, daily meal stipend, medical, vision, and dental coverage.
Vacancy posted 10 hours ago
Similar jobs that could be interesting for youBased on the Software Engineer, Inference in San Francisco, CA vacancy
  •  ...About the Team We’re hiring a Developer Productivity engineer to support OpenAI’s Inference Runtime teams. These teams own the systems responsible for serving models reliably, efficiently, and safely across Codex, ChatGPT, API, and internal research workloads. We’re... 
    Suggested
    Full time

    OpenAI

    San Francisco, CA
    10 hours ago
  •  ...BASETEN Baseten powers mission-critical inference for the world's most dynamic AI...  ...Conviction. Join us and help build the platform engineers turn to to ship AI products. THE ROLE...  ..., reliability, and ease of use. As a Software Engineer on the Inference Stack team,... 
    Suggested
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    1 day ago
  •  ...About the Team OpenAI’s Inference team powers the deployment of our most advanced models...  ...world. We're a small, fast-moving team of engineers focused on delivering a world-class...  ...About the Role We’re looking for a software engineer to help us serve OpenAI’s multimodal... 
    Suggested
    Full time

    OpenAI

    San Francisco, CA
    10 hours ago
  •  ...About the Team Our Inference team brings OpenAI’s most capable research and technology to the world through our products. We empower...  ...via model inference. About the Role We’re hiring engineers to scale and optimize OpenAI’s inference infrastructure across... 
    Suggested
    Full time

    OpenAI

    San Francisco, CA
    10 hours ago
  • $320k

     ...growing group of committed researchers, engineers, policy experts, and business leaders working...  ...the Role Our mandate is to make inference deployment boring and unattended....  ...deployment continuous and unattended. As a Software Engineer on the Launch Engineering team,... 
    Suggested
    Full time
    Work at office
    Visa sponsorship
    Flexible hours
    Shift work

    Anthropic

    San Francisco, CA
    1 day ago
  •  ...About the Team Our team analyzes inference stack performance across the application, model, and fleet layers to identify bottlenecks...  ..., analysis, and optimization. Enjoy collaborating with engineering and research teams to improve real production systems. About... 
    Full time

    OpenAI

    San Francisco, CA
    10 hours ago
  •  ...ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence...  ...Greylock, and Conviction. Join us and help build the platform engineers turn to to ship AI products. THE ROLE: Voice is becoming... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    10 hours ago
  •  ...tools consistently fail. We are a small, fast-growing team of engineers in San Francisco powering Fortune 100 enterprises, YC startups...  ...About the Role Specialize in low-latency, high-throughput inference for OCR and multimodal models. Own profiling, batching, and... 
    Visa sponsorship
    Relocation package

    PULSE

    San Francisco, CA
    4 days ago
  • $170k - $216k

     ...products that evaluate the Waymo Driver's software stack at a massive scale. We solve...  ...for a broad range of customers Software Engineers, Product, Data Science, System Engineering...  ...You will: Build and evolve ML inference infrastructure for simulations. Be responsible... 
    Full time
    Remote work

    Waymo

    San Francisco, CA
    10 hours ago
  •  ...About the Team Our Inference team brings OpenAI’s most capable research and technology to...  ...About the Role We are looking for an engineer who wants to take the world's largest and...  ...Have at least 5 years of professional software engineering experience. Have or can quickly... 
    Full time

    OpenAI

    San Francisco, CA
    10 hours ago
  • $300k

     ...growing group of committed researchers, engineers, policy experts, and business leaders working...  .... About the Role The Cloud Inference team scales and optimizes Claude to serve...  ...Fit If You: Have significant software engineering experience, with a strong background... 
    Full time
    Work at office
    Visa sponsorship
    Flexible hours

    Anthropic

    San Francisco, CA
    10 hours ago
  • $125k - $160k

     ...Role Overview We are seeking a versatile Full Stack Software Engineer to join our engineering team. Reporting to the Software Engineering...  ...-Augmented Generation) architectures, or local model inference (Ollama). Experience in automated testing at multiple levels... 
    Full time
    Local area
    Visa sponsorship
    Work visa
    Shift work

    Cala Health

    San Francisco, CA
    10 hours ago
  • $150k - $180k

     ...Capital , and JFF Ventures , and are now hiring a Full Stack Engineer to help build the product that institutions use to interact...  ...application layer up to our data stack (Postgres + DuckDB) and model inference, and keep query and inference latency low enough that the... 
    Full time
    Work at office
    Immediate start

    Straia

    San Francisco, CA
    10 hours ago
  • $215k - $260k

     ...reliably in production. That means owning the inference stack end to end: profiling where time...  ...will also work directly with customer engineering teams to tailor deployments to their...  ...service.Build and support the software and product features around the inference... 
    Temporary work

    Crusoe

    San Francisco, CA
    3 days ago
  •  ...video. Our team also manages large-scale inference and platform infrastructure that...  ...over unchecked growth. Within Applied Engineering, the Ads Monetization team in Financial...  ...Possess a minimum of 5 years of professional software engineering experience. Bring... 
    Full time

    OpenAI

    San Francisco, CA
    10 hours ago
  •  ...ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence...  ...Greylock, and Conviction. Join us and help build the platform engineers turn to to ship AI products. THE ROLE As a Senior Enterprise... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    10 hours ago
  • $190.9k - $232.8k

    P-1285About This RoleAs a staff software engineer for GenAI inference, you will lead the architecture, development, and optimization of the inference engine that powers Databricks Foundation Model API.. You’ll bridge research advances and production demands, ensuring high... 
    Local area
    Worldwide

    DataBricks

    San Francisco, CA
    10 hours ago
  •  ...s best for our customers. Cohere is a team of researchers, engineers, designers, and more, who are passionate about their craft. Each...  ...), especially how they influence latency and throughput of inference. ~ Strong understanding or working experience with distributed... 
    Full time
    Work experience placement
    Work at office
    Remote work
    Flexible hours

    Cohere

    San Francisco, CA
    10 hours ago
  • $165k - $190k

     ...The RoleWe're hiring fullstack Applied AI Engineers to help us build AI agents that power...  ...mission to make intelligent, autonomous software a reality both internally and for our customers...  ...: prompting, retrieval, orchestration, inference APIs, and model selection across... 
    Work at office
    Flexible hours

    Langchain

    San Francisco, CA
    4 days ago
  • $160k - $230k

     ...the full generative AI lifecycle, combining the fastest LLM inference engine with state-of-the-art AI cloud infrastructure. As a Senior...  ...business needs Write clear, well-tested, and maintainable software and IaC for both new and existing systems Conduct design and... 
    Full time
    Remote work

    Together Ai

    San Francisco, CA
    10 hours ago
  • $190k - $265k

     ...use deep data insights to improve their business. Founded by engineers — and customer-obsessed — we leap at every opportunity to solve...  ...serving layer for large language models across real-time and batch inference, powering model inference at enterprise scale. We are looking... 
    Local area
    Worldwide

    DataBricks

    San Francisco, CA
    3 days ago
  •  ...tools they need.About the Position:We’re looking for a seasoned software engineer to join Parafin’s Infrastructure team and lead the...  ...lifecycle (training, deployment, monitoring, retraining), real-time inference.Contributions to internal tooling or open-source projects in... 
    Flexible hours

    Parafin

    San Francisco, CA
    10 hours ago
  • $238k - $290k

     ...started.Role OverviewAs a Backend Platform Engineer at Harvey, you will help build and...  ...such as supporting high-throughput model inference, managing streaming and long-running API...  ...with confidenceWhat You Have5+ years of software engineering experience (post-BS/MS), including... 
    Flexible hours
    Shift work

    Harvey

    San Francisco, CA
    10 hours ago
  •  ...About the Team The Kernels team at OpenAI builds the low-level software that accelerates our most ambitious AI research. We work at...  ..., and runtime improvements to make large-scale training and inference more efficient. Our work enables OpenAI to push the limits... 
    Full time

    OpenAI

    San Francisco, CA
    10 hours ago
  • $150k - $237.5k

     ...What Our Team Does We build the software platform that powers optimized control...  ...value collaboration, trust, learning, and engineering excellence, and we genuinely enjoy building...  ...or employment information, and inferences drawn from your PI. We collect your PI... 
    Full time
    Flexible hours

    Redwood Materials

    San Francisco, CA
    10 hours ago
  •  ...for GPT-4, GPT-3, embeddings, and fine-tuning. We also operate inference infrastructure at scale. There's a lot more on the immediate...  ...features that were never before possible.  About the Role The Engineering Acceleration team designs, builds and maintains the... 
    Full time
    Immediate start
    Relocation package

    OpenAI

    San Francisco, CA
    10 hours ago
  •  ...About the Role We’re hiring three exceptional Founding Software Engineers to help us scale the computational biology platform that...  ...Computational Biology. We deal with problems ranging from scaling ML inference on AWS for hundreds of GPUs to dissecting pdb files with... 
    Full time
    Relocation

    Tamarind Bio

    San Francisco, CA
    10 hours ago
  •  ...latency voice interactions. We partner closely with Research and Inference to bring frontier model capabilities to developers and use...  ...feedback to improve our models. About the Role As a software engineer on API Multimodal, you will build and operate the products... 
    Full time
    Internship

    OpenAI

    San Francisco, CA
    10 hours ago
  •  ...ABOUT BASETEN Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence...  ...Greylock, and Conviction. Join us and help build the platform engineers turn to to ship AI products. THE ROLE We’re seeking a... 
    Full time
    Flexible hours

    Baseten

    San Francisco, CA
    10 hours ago
  • $150k - $250k

     ...administrative burden. This isn’t just better software. It’s infrastructure that wasn’t...  ...They are hiring a full-stack software engineer to help shape the future of healthcare....  ...been leveraged to integrate between the inference layer, transcription (text to speech), voice... 
    Full time
    Night shift

    Rockstar

    San Francisco, CA
    10 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Software Engineer, Inference. Be the first to apply!