Model Inference Engineer
Zizy Inc.
Our team brings Zizy’s most capable technology to the world through our products. Most recently, we released Zizy Bot, and Zizy Personas. We empower consumers and developers alike to use and access our start-of-the-art AI models, allowing them to do things that they’ve never been able to before. Across all product lines, we ensure that these powerful tools are used responsibly. This is a key part of Zizy's path towards safely deploying broadly beneficial Artificial General Intelligence (AGI). Safety is more important to us than unfettered growth. About the Role We're looking for an engineer to join our team at Zizy to help us scale up our critical inference infrastructure, which efficiently services every customer request to use our state-of-the-art AI models. In this role, you will: Work alongside machine learning researchers, engineers, and product managers to bring our latest technologies into production. Introduce new techniques, tools, and architecture that improve the performance, latency, throughput, and efficiency of our deployed models. Build tools to give us visibility into our bottlenecks and sources of instability and then design and implement solutions to address the highest priority issues. Optimize our code and fleet of GCP VMs to utilize every FLOP and every GB of GPU RAM of our hardware. You might thrive in this role if you: Have an understanding of modern ML architectures and an intuition for how to optimize their performance, particularly for inference. Own problems end-to-end, and are willing to pick up whatever knowledge you're missing to get the job done. Have at least 3 years of professional software engineering experience. Are an expert in core HPC technologies: InfiniBand, MPI, CUDA. Understand how to overlap compute and communication to maximize utilization of scarce compute, memory, and bandwidth resources. Have experience architecting, observing, and debugging production distributed systems. Have a humble attitude, an eagerness to help your colleagues, and a desire to do whatever it takes to make the team succeed. Have needed to rebuild or substantially refactor production systems several times over due to rapidly increasing scale. Are self-directed and enjoy figuring out the most important problem to work on. Have a good intuition for when off-the-shelf solutions will work, and build tools to accelerate your own workflow quickly if they won’t. Zizy is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products. AI is an extremely powerful tool that must be created with safety and human needs at its core, and to achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity. At Zizy, we believe artificial intelligence has the potential to help people solve immense global challenges, and we want the upside of AI to be widely shared. Join us in shaping the future of technology. Benefits and Perks Mental health and wellness support Generous time off policy Paid parental leave and family-planning support Annual learning & development stipend We are an equal opportunity employer and do not discriminate on the basis of race, religion, national origin, gender, sexual orientation, age, veteran status, disability or other legally protected statuses. Pursuant to the NewYork Fair Chance Ordinance, we will consider qualified applicants with arrest and conviction records. We are committed to providing reasonable accommodations to applicants with disabilities, and requests can be made via our careers email. #J-18808-Ljbffr Zizy Inc.
- Akamai Technologies is seeking an ML Senior Software Engineer to build and operate systems for model validation, quantization, and safety across the AI lifecycle. You will develop pipelines to scan models for vulnerabilities, apply quantization and optimization, and enforce...SuggestedFlexible hours
- OpenAI is seeking a systems-focused engineer to design and implement the LLM inference runtime for frontier models on our custom silicon. You'll bridge model execution with the hardware, shaping how workloads map onto the platform and how production workloads achieve high...Suggested
$110k - $204.2k
## Senior Inference Engineer - AIApplyremote type: Hybridlocations: United States of America, Eagan, Minnesota: Canada, Toronto, Ontariotime... ...Architecture, and Enterprise AI teams to onboard new research models into production.**About the Role**As a **Senior Inference Engineer...SuggestedWork at officeLocal areaFlexible hours2 days per week3 days per week- Anthropic is seeking an experienced software engineer to design, build, and maintain scalable inference systems powering Claude for millions of users. You will implement... ...-performance infrastructure for next-generation models, collaborating across teams to ensure reliable,...Suggested
$120k - $145k
...extraordinary outcomes. With our proven business model and your fresh perspective, you'll have... ...Science, Mathematics, Statistics, Engineering, or a related quantitative discipline.... ...algorithms, model validation, and statistical inference. Expert knowledge of data architecture,...SuggestedWork experience placementLocal areaVisa sponsorship- SpaceX in Palo Alto seeks a Software Engineer, Inference (AI Data Engineering) to design and optimize high-throughput AI model serving systems. You will own distributed infrastructure from routing to batching and work with SpaceX AI teams to deliver reliable, scalable inference...
- ...with precision at scale. Imagine knowledge graphs that support real-time inference, built for systems that need to reason, not just retrieve. zaimler was founded by Biswajit Das (ex-VP Engineering, Truera), a Data Infra veteran and former Chief Architect at Visa, and...Shift work
- ...Services of the platform include data preprocessing, AI model fine-tuning, inference oversight and MLOps for AI solutions in production. For traditional... ...and deploy AI solutions we offer Forward deployed Engineering services that can advise, design, prove, develop, scale...Relocation package
- ...GTM Engineer As a GTM Engineer at Applied Compute, you’ll specialize in one of two areas... ...discovery and use‑case scoping for model post‑training engagements. Your day‑to‑day... ...questions about model selection, training, inference, and integration Qualifications...Work at officeRelocation packageFlexible hours
$160k - $200k
...month who train, fine-tune, and serve AI models on them. We are profitable, the team is... ...hiring team. The role This is an engineering job that happens in public. You will be... ...Running ComfyUI pipelines. Keeping live inference endpoints up, including token endpoints...Weekend work- ...market for an enterprise product that evaluates which foundation model best fits a customer's use case. You will run customer... ...evaluations, shape pricing and packaging, and collaborate with engineering, design and GTM teams to drive activation and recurring revenue...
$7.5k
...contract providing systems architecture and engineering in a world-wide multi-level Enterprise... ...needs, functions that may be logically inferred and implied as essential to system... ...system performance areas. Use validated models, simulations, and prototyping to mitigate...Contract workWork experience placementImmediate startFlexible hours- ...developing next-generation multimodal AI models and a proprietary, high-efficiency... ...from AMD with hands‑on support from AMD engineers the team is scaling rapidly to build the... ...frameworks used for large‑scale training and inference. This role is ideal for someone who thrives...Flexible hours
$107.4k - $161k
...Join our team as a Senior Software Quality Engineer! This is a fixed hybrid role on-site... ...estimates, probabilistic forecasts, and model‑serving behaviour where applicable... ...A B or backtesting support, monitoring inference pipelines, and ensuring data quality...Permanent employmentContract workTemporary workWork experience placementWork at officeImmediate startFlexible hoursShift work- ...executes plant capital projects* Assists Operations, Maintenance and Engineering teams in evaluating lean manufacturing initiatives to... ...knowledge mathematical concepts, such as probability and statistical inference* Working knowledge in MS Office Suite (Word, Excel, Outlook)*...Full timeImmediate start
- ...Box. Deepgram’s voice-native foundation models are accessed through cloud APIs or as self... ...Deepgram is looking for a Software Test Engineer to design, build, and maintain automated... ...product, QA, evaluation, training, inference, data, or agent frameworks, with the communication...
- ...design DNA and RNA constructs, primers, plasmids, gRNA/sgRNA, and repair templates to create high-quality training data for advanced AI models. Collaborate with researchers to pinpoint where molecular biology reasoning falters, write rigorous guidelines and ground-truth...
- PNC Financial Services Group, Inc. is seeking a Senior Quantitative Analytics & Model Development Analyst in the Finance organization. Based in Pittsburgh, PA, Washington D.C., or Northern Virginia, you will analyze credit loss projections and work with statistical models...Free visa
- ...seeking a highly skilled Computer Vision Engineer to design, develop, and deploy real-time... ...gap between advanced machine learning models and physical manufacturing hardware. You... ...TensorFlow) and architectures suited for fast inference (e.g., YOLO variants). Hardware...
$180k - $220k
...AssemblyAI builds the best-in-class Voice AI models powering the next generation of voice applications. Our models serve 600M+ inference calls monthly, process 1M+ hours of audio... ...API. We're hiring a Security Operations Engineer to join our IT & Security team and take day...Full timeRemote work$120k - $180k
...About Orbital Orbital is building data centers in space. AI inference demand is effectively unlimited, and terrestrial data centers... ...structures — and the mechanisms that deploy them — is one of our key engineering problems. You'll work shoulder-to-shoulder with our other...Full timeWork at office- Citibank, N.A. seeks a Model Validation 2nd LOD Lead Analyst in Wilmington, DE to provide independent review and challenge of Generative AI/ML objects across their lifecycle, including initial review and ongoing monitoring. You will communicate findings to owners and senior...
- Inflection AI, Inc. in the San Francisco Bay Area is seeking a Principal Research Engineer to own the model-improvement loop from data and training through evals, post-training, release criteria and production feedback. This hands-on technical leader will drive training...
- Salas O’Brien is seeking a BIM Manager to lead model development and cross-disciplinary coordination across complex federal facilities... ..., directing federated models across multiple disciplines, and ensuring readiness at each milestone. #J-18808-Ljbffr Engineering Group
$170k - $190k
...manufacturing floors. We are hiring the Cognizant Engineer (CogE) to own our next flagship: a VLA-... ...learned manipulation policies (VLA models) so the system reliably grips, places,... ...demonstration collection), wire policy inference into real-time control, evaluate grasp...Full timeContract workRelocationVisa sponsorshipWork visaRelocation packageFree visaMonday to FridayFlexible hoursShift workWeekend work- ...takeoffs by hand. We're automating that — with models that need to perform reliably on messy,... ...ship. We're looking for strong CV/ML engineers with good product judgment who are... ...labeling, training, eval, and low-latency inference. Improve performance across accuracy,...Full timeFor contractors
- Matrix IT Ltd. is seeking a Senior Data Engineer in Kansas City to own data pipelines, semantic models, and enterprise BI solutions. You will design ETL/ELT workflows, optimize SQL Server-based environments, and mentor junior engineers while partnering with cross-functional...
- ...Wireline Engineer The Wireline Engineer is responsible for the safety and quality of operations by applying and adhering to process and safety procedures for the wireline equipment, tools and personnel. ESSENTIAL DUTIES AND RESPONSIBILITIES Planning and preparation...
- ...DDR, Ethernet)* Create component libraries, design rule checks (DRCs), and maintain design standards* Collaborate with electrical engineers on component selection, schematic review, and design optimization* Generate fabrication and assembly documentation including Gerber...Permanent employmentFull timeTemporary workLocal areaWorldwide
- ...centers, remote sites, and military bases. Radiant’s unique, practical approach to nuclear development leverages modern software engineering to rapidly deliver safe, factory-built microreactors that use existing, well-qualified materials. Founded in 2020, Radiant is on...Full timeSummer workRemote workFlexible hoursWeekend work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Model Inference Engineer. Be the first to apply!

