Data Engineer
Soteris
ABOUT SOTERIS
Soteris is a YC-backed AI company building the future of pricing and product management for the
insurance industry. Our mission is to infuse the $5 trillion P&C insurance industry with best-in-class,
proprietary, AI-driven data analytics. We’ve spent years building our own proprietary AI models on
personal auto claims and exposure data to help insurers improve their loss ratios, with over 100 million
submissions and $180 billion in premium scored to date.
Each year, roughly $750 billion in insurance policies are written in the United States. Our machine
learning platform helps insurers evaluate policies at a granular level, moving beyond broad segmentation approaches that often lead to risks being over- or underpriced. Our modeling approach incorporates multiple model families and calibration methods to rank policy risk within a book of business.
We are a team of 10 and growing quickly. As our second Data Engineer, you will own the data layer that turns customer policy, claims, quote, and financial data into trusted inputs for actuarial analysis, model development, and production scoring. This is a builder role: you will work directly with customer data teams, create repeatable ingestion and transformation pipelines, reconcile outputs to source-of-truth control totals, and make the platform easier to operate as we add customers and products.
WHAT YOU'LL BE DOING
Customer Data Onboarding and Integration
- Leading the technical data workstream for new customer implementations by understanding policy, claims, rating, quote, and financial systems and establishing secure access to the data.
- Building reusable extraction and synchronization workflows for databases, backups, secure file transfer, APIs, and other delivery methods while preserving source lineage and supporting backfills.
- Mapping customer data into Soteris’s internal ontology and working with customer technical teams and internal project leads to resolve definitions, transformations, data-quality issues, and onboarding blockers.
Lakehouse and Pipeline Engineering
- Owning Databricks and AWS pipelines that move customer data from raw and Bronze ingestion through standardized Silver tables and curated Gold or model-ready datasets.
- Designing idempotent, incremental, observable workflows that handle schema evolution, late-arriving data, backfills, orchestration, performance, and cost.
- Developing shared components and configuration-driven patterns, then publishing well-defined datasets for actuarial analysis, backtesting, model training, production scoring, reporting, and monitoring.
Data Quality and Modeling Readiness
- Building automated quality gates for completeness, uniqueness, referential integrity, valid ranges, freshness, balance, schema drift, and other customer-specific controls.
- Reconciling written and earned premium, exposure, incurred and ultimate loss, claim counts, fee income, and other economic drivers to customer control statistics at the required state, program, year, and coverage levels.
- Partnering with data science to produce leakage-resistant, point-in-time-correct datasets and productionize approved actuarial reference data
Production Platform and Operational Ownership
- Owning data flows for quote requests, model inputs, and bound policy outcomes, including pre-live comparisons that confirm production request fields match corresponding policy data and post-launch drift checks.
- Operating pipelines with monitoring, alerting, recovery behavior, runbooks, and clear incident diagnostics across customer synchronization, Databricks processing, and downstream model- serving dependencies.
- Testing, reviewing, and deploying data code through GitHub and GitHub Actions, with automated tests and deployment controls for production data assets.
- Managing Databricks permissions, Unity Catalog controls, sensitive data, and least-privilege access in support of Soteris’s security and SOC 2 requirements.
OUR CURRENT STACK
- Python, SQL, PySpark, pandas, and related data-engineering libraries
- Databricks, Delta Lake, Databricks Workflows, and Unity Catalog
- AWS, including S3, Lambda, EC2, ECS, SageMaker, and secure customer file transfer
- Customer databases, backups, SFTP, APIs, and file-based ingestion
- Terraform, GitHub, GitHub Actions, automated testing, and infrastructure as code
- MLflow, SageMaker, and production scoring APIs at the model handoff boundary
- Modern generative AI development tools, including Claude, ChatGPT, Codex, or similar models
ABOUT YOU
You must have the following:
- Strong Python and SQL skills and the ability to write production-quality transformations, tests, utilities, and operational tooling.
- Hands-on experience building and operating production pipelines using Spark, Databricks, or a comparable distributed data platform.
- Strong understanding of data modeling, lakehouse or warehouse design, incremental processing, schema evolution, idempotency, backfills, and lineage.
- Experience designing data-quality controls and reconciling complex datasets to source systems or independent control totals.
- Experience with AWS, Git-based development, automated testing, CI/CD, and practical tradeoffs involving reliability, security, performance, and cost.
- The ability to work directly with customer technical teams, understand unfamiliar schemas, ask precise questions, and document decisions clearly.
- Comfort operating with significant ownership and ambiguity where customer implementation, platform development, security, and production operations overlap.
- The judgment to use AI development tools effectively while verifying generated code, tests, and transformations against source evidence.
You’d be a great fit if you also have:
- Experience with P&C insurance data, including policy transactions, coverages, claims, premium, exposure, rating, or underwriting data.
- Experience integrating with policy administration, claims management, rating, or other operational source systems.
- Deep experience with Databricks, Delta Lake, Unity Catalog, Databricks Workflows, or configuration-driven data pipelines.
- Experience with Terraform, secure file transfer, database replication, or customer-specific ingestion infrastructure.
- Experience supporting or building machine-learning feature pipelines, point-in-time datasets, model monitoring, MLflow, SageMaker, or production scoring systems.
$134.5k - $265.1k
Position Summary Deloitte is seeking a Lead Data Engineer- Databricks to support the design, build, and delivery of modern data and analytics solutions for clients across industries. In this role, you will help teams translate business needs into scalable data platform...SuggestedLocal areaVisa sponsorship$105.4k - $207.8k
Position Summary Our Deloitte AI & Engineering team works to transform technology platforms, drive innovation, and help make a significant... ..., and fuel growth through innovation.Work you'll do As a Data Engineer III on the AI & Data team, you will be responsible for...SuggestedLocal area- ***MUST HAVE CURRENT HEALTHCARE EXPERIENCE*** As a Lead Data Engineer, you will play a hands-on technical leadership role in evolving the data capabilities that power our client's solutions. You will lead complex initiatives across two closely connected areas: onboarding...Suggested
$200k
...Lead Data Engineer - Remote 100% Remote (U.S. Based) Up to $200,000 Base Salary U.S. Citizens & Green Card Holders Only I'm partnering with a rapidly growing healthcare technology company that is looking to hire a Lead Data Engineer to help drive the evolution...SuggestedRemote work- ...We are seeking a Senior Data Engineer to design, develop, and optimize scalable data integration pipelines and data solutions. This is a hands-on engineering role with a strong focus on Databricks, Spark/PySpark, Python, SQL, Delta Lake, and Unity Catalog. The ideal...Suggested3 days per week
$102.3k - $209.5k
Designs and builds proper data pipelines required for optimal data collection and processing from a variety of data sources. Implements... ...of data validation and governance.Data Pipeline and Solutions Engineering - Pipeline Design: Independently designs, develops, and...Temporary workFlexible hoursShift work- ...Position Overview: Describes the overall purpose of the position or why the position existsThe Data Engineer designs, develops, and maintains secure, scalable data infrastructure that integrates clinical, operational, financial, quality, and value-based care data. This...
- ...Brooklyn, OH, with a three-day per week onsite requirement. Candidates must be located in the Cleveland/Brooklyn area. Job Summary: Data engineer responsible for the design and development of data integration pipelines. Build data solutions for business problems and support...Work experience placement3 days per week
$95k - $135k
...health benefits, a 401k with company match, and generous time off to recharge, an employee SHARE program. JOB SUMMARY: The Data Engineer designs, builds, and maintains the data platform infrastructure that powers Healthfuse's internal and client-facing operations....Permanent employmentH1b- ...Job Description Hands on Experience in Azure Synapse Analytics, Azure Data Factory and Data Bricks, Azure Storage, Azure Key Vault, SQL Pools CI/CD Pipeline Designing and other Azure services like functions, logic apps - Linked services, Various Runtimes, Datasets...
- ...Senior Data Engineer With Databricks Remote, Anywhere in the Continental US Due to the nature of the role, this position requires U.S. Citizenship. Data Surge is disrupting the services industry with cutting-edge technology that brings together the very best...Full timeImmediate startRemote work
- ...Location: Remote We are seeking an experienced Data Engineer to support challenging data engineering, data management, and analytics initiatives. In this role, you will develop and maintain data pipelines and data structures that enable advanced analytics and data...Temporary workLocal areaRemote work
- ...The Data Engineer role supports Pharmacy Benefit Management clients by maintaining and improving complex data systems. The position focuses on building scalable data solutions for analytics and business intelligence while adhering to healthcare data compliance requirements...Full timeTemporary workRemote workFlexible hours
- ...Data EngineerWe are seeking a Data Engineer to design, develop, and support data integration and transformation solutions that enable data-driven decision-making across the organization. This role is responsible for building scalable data pipelines, ensuring data quality...
- ...Data Engineer – Remote Remote | Full-time | Data Engineering We're looking for an experienced Data Engineer to join a growing data team and help build the pipelines, models and infrastructure that power data-driven decision making across a global travel business....Full timeRemote work
- ...Location: Remote (United States) Employment Type: contract About the Role As a Data Engineer, you’ll build and operate the pipelines, tables, and services that power our hybrid data platform across on-premises and AWS. You’ll implement lakehouse patterns, productionize...Contract workRemote work
- ...Inclusive Care for the Elderly. With over 1,400 team members, we care for more than 100,000 people across all 50 states.Job DetailsThe Data Engineer maintains complex data systems while delivering improvements to existing processes and building new solutions for our Pharmacy...Full timeTemporary workLocal areaRemote workFlexible hours
- ...IDR is seeking a Data Engineer to join one of our top clients for a remote opportunity. This role involves designing and implementing scalable data solutions, supporting AI-ready data processing, and collaborating with cross-functional teams to enable data-driven decision...Remote work
$100k - $120k
...Data Engineering - DevOps Fractal Analytics is a strategic AI partner to Fortune 500 companies with a vision to power every human decision in the enterprise. Fractal is building a world where individual choices, freedom, and diversity are the greatest assets. An ecosystem...Hourly payFull timeLocal areaRemote work$116.2k - $229.1k
Position Summary Join our AI & Engineering team in transforming technology platforms, driving innovation, and helping make a significant... ...operate integrated/verticalized sector solutions in software, data, AI, network, and hybrid cloud infrastructure. These solutions...Local areaVisa sponsorship- ...solutions that enable seamless communication and collaboration across the supply chain. DESCRIPTION: Transflo is seeking a Senior Data Engineer to architect and own our enterprise data platform — from raw ingestion through curated, analytics-ready data products. You will...Remote work
$86.7k - $170.9k
Position Summary Join Deloitte’s Core AI & Data practice and help organizations modernize data platforms, strengthen enterprise... ...and artificial intelligence capabilities. As a Databricks Data Engineer, you will support the design, build, and optimization of cloud-based...Local areaVisa sponsorship- ...a premier Google Cloud partner helping organizations modernize data ecosystems, build real-time analytics capabilities, and responsibly... ...Foundation Models, and Gemini Enterprise.You AreA hands-on Engineer with foundational experience in Data Engineering, Analytics, or...Full timeWork experience placementLive inWork at officeLocal area
$116.2k - $229.1k
Position Summary Join Deloitte’s AI & Engineering practice and help organizations transform enterprise technology platforms, modernize data environments, and unlock value through innovation. As a Databricks Engineer in our AI & Data practice, you will design, build...Local areaVisa sponsorship- ...Quantiphi: Quantiphi is an award-winning, AI-First digital engineering and consulting company focused on delivering high-impact Services... ...by combining deep industry expertise, disciplined cloud and data engineering practices, and cutting-edge applied AI research. Our...Full timeRemote work
- ...SENIOR DATA ENGINEER We are looking for a highly skilled and strategic Senior Data Engineer to lead in our data engineering consulting team. In this role, you will serve as the technical cornerstone for our clients, designing and deploying the sophisticated data architectures...
- ...Responsibilities: Build and operate production-grade batch and streaming data pipelines across SQL Server, cloud applications, device/event... ...on-call activities. Work closely with Product, Application Engineering, Quality Assurance, DevOps/Site Reliability Engineering,...
$82.6k - $162.8k
...business rules (calculations, validations, allocations, and transformations) using OneStream development tools and best practices.Support data management activities including data loads, mappings, transformations, reconciliations, and error handling within OneStream (and/or...Local areaRemote work- ...Datacenter Engineer The engineer in this position should have a wealth of experience engineering/installing datacenter equipment. The... ...cohesive package. This would include the complete engineering of data centers including, but not limited to, infrastructure, overhead...For contractorsWork at officeLocal area
$127k - $150k
Position Summary Healthcare Data Engineer Sr. Consultant Deloitte Consulting's technology professionals help clients identify and solve their most critical information and technological challenges. We provide advisory through end-to-end implementation services as...Local areaVisa sponsorship
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Data Engineer. Be the first to apply!
- data engineer machine learning Cleveland, OH
- aws data engineer Cleveland, OH
- sr data engineer Cleveland, OH
- finance data engineer Cleveland, OH
- entry level data engineer Cleveland, OH
- sr information security engineer Cleveland, OH
- data center engineer Cleveland, OH
- senior data integration developer Cleveland, OH
- senior cloud data engineer Cleveland, OH
- data engineer Cleveland, OH



