Data Engineer
$157.5k - $231kEli Lilly and Company
Job ID: R-110976Company: LillyLocation: San Francisco, California, United States of AmericaJob Type: Full TimeCategory: Information TechnologyPosted Date: 2026-09-16At Lilly, the work is demanding because patients are waiting. We unite caring with discovery to help make life better for people around the world, knowing that every decision, every detail, and every day matters. Headquartered in Indianapolis, Indiana, our over 50,000 employees around the globe take on complex challenges to discover and deliver life-changing medicines, strengthen how health is understood and managed, and support the communities we serve. This is hard, urgent, selfless work—but it’s work worth doing. If you’re driven by purpose and ready to bring your best to work that truly matters for patients, we invite you to join us. Where AI Meets Medicine: Build the Future of Drug Discovery in the Heart of Silicon Valley!Making medicine that’s never been made means doing what’s never been done. If you’re an engineer, scientist, or builder who thrives on problems no one has solved before, this is your invitation; we want you on the team. We are ready to challenge the status quo and push medicine forward, all in the name of health. Are you up for the challenge? If so, join us! About the Lilly and NVIDIA PartnershipLilly and NVIDIA are launching a new AI co-innovation lab in the heart of Silicon Valley — an up-to-$1 billion, multi-year commitment to solve drug discovery’s toughest challenges. The lab brings Lilly scientists, technologists, chemists and biologists together with NVIDIA engineers under one roof. Together, we are building purpose-built foundation and frontier AI models trained on Lilly data at scale, tightening the feedback loop between automated wet labs and computational dry labs, designing the next generation of medicines for millions of patients across the globe.What You’ll Be DoingAs a Data Engineer, you will build and maintain the data platforms that power AI-driven research and discovery. You will develop scalable pipelines that ingest, transform, and deliver chemical, biological, and experimental data for machine learning and scientific workflows. Partnering with AI Scientists, AI Engineers, and laboratory researchers, you will ensure that data is accurate, traceable, and accessible at scale. Your work will provide the trusted data foundation behind next-generation AI models and experiments.How You’ll SucceedEngineer datasets in large language environment for model training specifically efficient formats and storage layout (Parquet, Zarr, Arrow) and delivery fast enough that GPU clusters are never left waiting on data.Design, develop, and maintain scalable and efficient data pipelines to support data analytics, reporting, and machine learning initiatives.Ensure seamless data flow between systems and applications, optimizing data transfer and transformation processes for performance and scalability.Build the ingestion path from the automated lab, so experimental results reach the models in hours rather than weeks, closing the loop between what a model proposes and what the next model learns from.Own the correctness of what models train on completeness, sound joins across experimental sources, and validation that catches a bad dataset before it reaches a training run rather than after.Build dataset versioning, lineage, and reproducibility into the platform, so any model can be traced to the exact data it was trained on months or years later.Work with the laboratory, instrument, and external teams producing the data so that a change upstream does not quietly corrupt a training run downstream.What You Should BringStrong Python, or equivalent experience building data-intensive software systems.Strong SQL and data modeling experience including designing schemas that hold up as scientific data grows and diversifies, with expert knowledge of Postgres or a comparable enterprise database.Distributed data processing (Spark, Ray, or Dask) and pipeline orchestration (Airflow or Dagster) at scale.Experience with cloud platforms — AWS and Azure preferred — and with high-performance and object storage feeding large-scale compute environments.A track record of building data systems that other people depend on, and of taking responsibility for them when they broke.Strong testing practices and test automation, with solid CI/CD and Git fundamentals.Adaptability and a collaborative mindset, with the ability to translate complex scientific questions into data solutions that accelerate experimentation and decision-making.Experience streaming and event-driven integration (Kafka, MQTT, or AMQP), including instrument and laboratory data capture.Cheminformatics or scientific data experience — compound registration, structure notation (SMILES, InChI, HELM), RDKit, multi-omics, assay, or sequencing data — is a strong plus.Prior experience across the following: data modeling, ETL/ELT at scale, ontology development, semantic graph construction and linked data, or relational schema design.Experience standing up, migrating, or consolidating databases and data platforms, including production cutover of systems in active use.Your Basic QualificationsBachelor’s degree in Computer Science, Data Science, Engineering, Mathematics, or a related technical field.5+ years of data engineering experience building and operating production data systems.Location & Work FlexibilityThis role is based at our Silicon Valley Hub. We offer a flexible hybrid work model, with three days onsite and two days working remotely each week, supporting both collaboration and work‑life balance.Lilly is dedicated to helping individuals with disabilities to actively engage in the workforce, ensuring equal opportunities when vying for positions. If you require accommodation to submit a resume for a position at Lilly, please complete the accommodation request form () for further assistance. Please note this is for individuals to request an accommodation as part of the application process and any other correspondence will not receive a response.Lilly is proud to be an EEO Employer and does not discriminate on the basis of age, race, color, religion, gender identity, sex, gender expression, sexual orientation, genetic information, ancestry, national origin, protected veteran status, disability, or any other legally protected status.Our employee resource groups (ERGs) offer strong support networks for their members and are open to all employees. Our current groups include: Africa, Middle East, Central Asia (AMECA), Black Employees at Lilly (View email address on click.appcast.io), Chinese Culture Network (CCN), EnAble, Evolve, Lilly Indian Network (LIN), Organization of Latinx at Lilly (OLA), Pride (LGBTQ+ Allies), Veterans Leadership Network (VLN) and Women’s Initiative for Leading at Lilly (WILL).Actual compensation will depend on a candidate’s education, experience, skills, and geographic location. The anticipated wage for this position is$157,500 - $231,000Full-time equivalent employees also will be eligible for a company bonus (depending, in part, on company and individual performance). In addition, Lilly offers a comprehensive benefit program to eligible employees, including eligibility to participate in a company-sponsored 401(k); pension; vacation benefits; eligibility for medical, dental, vision and prescription drug benefits; flexible benefits (e.g., healthcare and/or dependent day care flexible spending accounts); life insurance and death benefits; certain time off and leave of absence benefits; and well-being benefits (e.g., employee assistance program, fitness benefits, and employee clubs and activities).Lilly reserves the right to amend, modify, or terminate its compensation and benefit programs in its sole discretion and Lilly’s compensation practices and guidelines will apply regarding the details of any promotion or transfer of Lilly employees.#WeAreLillySkillspythonsqldata modelingdistributed data processingsparkdata pipeline orchestrationpostgresraydaskairflowawsazureci/cdgitetl/eltcheminformaticsdata ingestiondata validationsemantic graph constructionlinked datarelational schema designdatabase migrationdata platform consolidationkafkamqttamqprdkitinchihelmassay data
$160k - $220k
...Lead Data Engineer Location: San Francisco, CA, United States Location Type: On-site Salary Range: 160000 - 220000 USD Annually We are seeking a Lead Data Engineer to architect, build, and lead the development of scalable, cloud-based data platforms that support...Suggested- ...Afresh is hiring a Staff Data Engineer to lead the analytics platform that powers our AI-driven grocery analytics. You will evolve data models, strengthen pipelines, and set data governance for internal and customer analytics. Collaborate with data scientists, product...Suggested
$2,000 per month
...About Build AI Build AI is the data hyperscaler for Physical AI. We co-design hardware, collection, infrastructure, and research... ...Eventual/Daft. Demo-scale ETL is not this job Strong software engineering. Python and at least one systems language. Linux You measure...SuggestedContract workWork at officeRelocation package- ...Thoughtspot and other BI tools • Write SQL for processing raw data, kafka ingestions, adf pipelines, data validation and QA •... ...other Big Data related technologies Work with product and engineering team to understand requirements, evaluate new features and architecture...Suggested
- ...Lead Data Engineer with MarTech Location: SFO, CA (Hybrid 2 days a week) Key Responsibilities Lead end-to-end MarTech engineering initiatives across orchestration, data processing, and activation pipelines. Architect scalable, event-driven systems that...SuggestedRemote work2 days per week
- ...Lead Data Engineer RADIUMONE IS A GLOBAL PROGRAMMATIC AD BUYING PLATFORM RadiumOne is the 6th largest web property in the U.S. according to comScore We build intelligent software that automates media buying, making big data actionable for marketers and connects...
- ...Job Title Mandatory Skills: (Oracle or PostgreSQL) and ETL Pipelines and Big Data and AWS Responsibilities • Uses structured tools for analysis and presentation of concepts and models to enhance the BRD • Develop, maintain and deliver training materials to the...Work experience placement
- ...Collect metrics based on user interactions. Visualize data for business teams. Develop and redesign data pipelines using... ...business partners, and product managers. Balance between hands-on engineering (50%) and team leadership (50%). Collaboration Structure:...Local area
- ...Lead Data Engineer The Office of Information Technology (IT) is responsible for enabling State Bar's internal and external stakeholders by the management, implementation, and maintenance of an organization's technology to support of State Bar's mission and goals. The...Work at office
- ...Downtown San Francisco | Hybrid (4 days onsite) What You'll Do Design, build, and maintain data pipelines from ingestion through transformation and delivery Own data modeling and transformation across the firm's core data sources Build and maintain a reliable, scalable...Work at office
- ...Stripe and Metronome billing, Hex on top. This is a hands-on contract role with real authority to own the dbt project, review and merge data-platform pull requests, and ensure pipeline reliability. You will turn hand-built SQL into tested dbt models, build the marts...Contract work
- ...software agents. We're the makers of Devin, the first AI software engineer. Our team is extremely talent-dense. Among our founding team,... ...apply to join us. About the Role We're hiring a technical Data Engineer to own our full data stack – from database architecture...
$150k - $250k
...Gradient is everything. Torus is building the agentic engineering tool that physical engineering firms use to build infrastructure 10x... ...multiple Fortune 500s and major EPCs on critical infrastructure like data centers and energy installations. Gradient is everything....- ...Job Posting Lead Data Engineer - Data & AI, Supply Chain Company Description/Details We are seeking an experienced Lead Data Engineer to join the Data & AI - Supply Chain organization. This role will lead the design, development, and implementation of enterprise...Contract workLocal area
$148k - $185k
...Data Engineer III Los Angeles, California, United States; San Francisco, CA, United States About Crunchyroll Founded by fans, Crunchyroll delivers the art and culture of anime to a passionate community. We super-serve over 100 million anime and manga fans across...Flexible hours$140k - $180k
...Data Engineer, Data Platform About the Role We are building out our Data Platform team at Sigma, with a relentless focus on developing data models that fuel trusted insights across the company. As a Data Platform Engineer, you will be responsible for the underlying...Full timeWork at officeFlexible hours- ...Data Engineer RadiumOne is seeking a Data Engineer to join our Data Technology team with the responsibility of designing and implementing scalable data infrastructure, as well as implementing data transformations developed by data scientists as scalable and robust processes...Local area
- ...About The Role The Data Engineer designs and operates the data infrastructure that powers product analytics, business reporting, and operational decision-making. The role owns reliable batch and streaming pipelines across application, transactional, and third-party data...
$144k - $153k
...Data Engineer Position The Golden State Warriors are looking for an experienced Data Engineer to join the growing Business Strategy and Analytics team. In this role, you will help us achieve our mission of becoming a global leader in experiences and entertainment by...Full timeWork experience placementSummer work$180k - $250k
...Fluency in San Francisco is hiring a full-time AI Engineer to create LLM-powered features and improve AI output quality. You will build new features end to end, partner with product engineers, and evaluate models. Applicants should have experience in TypeScript or Python...Full time- ...Data Engineer As a Data Engineer at Regard, you will help build and maintain the data pipelines and infrastructure that turn raw data into the metrics and insights that drive our product decisions and research. We run an engineering-first stack that prioritizes transparent...Work at officeLocal areaHome officeVisa sponsorshipRelocation package
- ...Data Engineer We are looking for a talented, highly enthusiastic team player to join a team of data enthusiasts on the Reporting Team. The ideal candidate would be interested in data and look for different ways to tell the story with data, and possess excellent analytical...
- ...Salesforce is seeking a Senior Data Engineer to build a trusted data foundation powering decision-making across the organization. You will design and scale data models and pipelines, champion data governance, and provide technical leadership across Product, Data Science...
$200.4k - $300.6k
## Staff, Data PlatformApplylocations: San Francisco, CA, USAtime type: Full timeposted on: Posted Yesterdayjob requisition id: JOBREQ-2616119**The opportunity** Unity is looking for a Staff Engineer, Data Platform to help build and scale the data infrastructure powering...Work at officeWorldwideRelocation package$150k - $250k
...the boundaries of AI-driven communication. If you're ready to own data strategy at a high-growth AI company, this is your chance to... ...our AI—it’s the foundation of everything we build. As Senior Data Engineer, you won’t just clean datasets and maintain pipelines. You’ll own...Full timeRemote workFlexible hours- ...Data Engineer Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. By uniting applied AI research, flexible infrastructure, and seamless developer tooling, we enable...Flexible hours
- ...years. But that only happens when clean, structured scientific data and AI are built into how science gets done. Benchling is the... ...software to modern science. Benchling is building AI & Data Engineering (AIDE), a small, autonomous team within our Security & IT...Work at officeLocal areaFlexible hours3 days per week
- ...Job Title: Lead Data Engineer - Data & AI, Supply Chain Location: San Francisco, CA 94105 (Onsite) Duration: 6 Months Contract About the Role Client is seeking an experienced Lead Data Engineer to join the Supply Chain Data & AI Organization. This role will...Contract work
- ...Data Engineer Responsibilities: Develop and automate large scale, high-performance data processing systems (batch and/or streaming) to drive Airbnb business growth and improve the product experience. Build scalable Spark data pipelines leveraging Airflow scheduler...
- ...veterans and innovative thinkers. We don't believe culture can be engineered - but when it falls into place, it's a once-in-a-lifetime... ...felt so present. Position Overview We're looking for a data engineer to help us turn raw driving data into the well-structured...Local areaFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Data Engineer. Be the first to apply!
- aws data engineer San Francisco, CA
- director data engineering San Francisco, CA
- data platform engineer San Francisco, CA
- data engineer machine learning San Francisco, CA
- data science developer San Francisco, CA
- senior data engineer San Francisco, CA
- finance data engineer San Francisco, CA
- principal data engineer San Francisco, CA
- big data devops engineer San Francisco, CA
- data engineer analytics San Francisco, CA

