Data Engineer
Soteris
ABOUT SOTERIS
Soteris is a YC-backed AI company building the future of pricing and product management for the
insurance industry. Our mission is to infuse the $5 trillion P&C insurance industry with best-in-class,
proprietary, AI-driven data analytics. We’ve spent years building our own proprietary AI models on
personal auto claims and exposure data to help insurers improve their loss ratios, with over 100 million
submissions and $180 billion in premium scored to date.
Each year, roughly $750 billion in insurance policies are written in the United States. Our machine
learning platform helps insurers evaluate policies at a granular level, moving beyond broad segmentation approaches that often lead to risks being over- or underpriced. Our modeling approach incorporates multiple model families and calibration methods to rank policy risk within a book of business.
We are a team of 10 and growing quickly. As our second Data Engineer, you will own the data layer that turns customer policy, claims, quote, and financial data into trusted inputs for actuarial analysis, model development, and production scoring. This is a builder role: you will work directly with customer data teams, create repeatable ingestion and transformation pipelines, reconcile outputs to source-of-truth control totals, and make the platform easier to operate as we add customers and products.
WHAT YOU'LL BE DOING
Customer Data Onboarding and Integration
- Leading the technical data workstream for new customer implementations by understanding policy, claims, rating, quote, and financial systems and establishing secure access to the data.
- Building reusable extraction and synchronization workflows for databases, backups, secure file transfer, APIs, and other delivery methods while preserving source lineage and supporting backfills.
- Mapping customer data into Soteris’s internal ontology and working with customer technical teams and internal project leads to resolve definitions, transformations, data-quality issues, and onboarding blockers.
Lakehouse and Pipeline Engineering
- Owning Databricks and AWS pipelines that move customer data from raw and Bronze ingestion through standardized Silver tables and curated Gold or model-ready datasets.
- Designing idempotent, incremental, observable workflows that handle schema evolution, late-arriving data, backfills, orchestration, performance, and cost.
- Developing shared components and configuration-driven patterns, then publishing well-defined datasets for actuarial analysis, backtesting, model training, production scoring, reporting, and monitoring.
Data Quality and Modeling Readiness
- Building automated quality gates for completeness, uniqueness, referential integrity, valid ranges, freshness, balance, schema drift, and other customer-specific controls.
- Reconciling written and earned premium, exposure, incurred and ultimate loss, claim counts, fee income, and other economic drivers to customer control statistics at the required state, program, year, and coverage levels.
- Partnering with data science to produce leakage-resistant, point-in-time-correct datasets and productionize approved actuarial reference data
Production Platform and Operational Ownership
- Owning data flows for quote requests, model inputs, and bound policy outcomes, including pre-live comparisons that confirm production request fields match corresponding policy data and post-launch drift checks.
- Operating pipelines with monitoring, alerting, recovery behavior, runbooks, and clear incident diagnostics across customer synchronization, Databricks processing, and downstream model- serving dependencies.
- Testing, reviewing, and deploying data code through GitHub and GitHub Actions, with automated tests and deployment controls for production data assets.
- Managing Databricks permissions, Unity Catalog controls, sensitive data, and least-privilege access in support of Soteris’s security and SOC 2 requirements.
OUR CURRENT STACK
- Python, SQL, PySpark, pandas, and related data-engineering libraries
- Databricks, Delta Lake, Databricks Workflows, and Unity Catalog
- AWS, including S3, Lambda, EC2, ECS, SageMaker, and secure customer file transfer
- Customer databases, backups, SFTP, APIs, and file-based ingestion
- Terraform, GitHub, GitHub Actions, automated testing, and infrastructure as code
- MLflow, SageMaker, and production scoring APIs at the model handoff boundary
- Modern generative AI development tools, including Claude, ChatGPT, Codex, or similar models
ABOUT YOU
You must have the following:
- Strong Python and SQL skills and the ability to write production-quality transformations, tests, utilities, and operational tooling.
- Hands-on experience building and operating production pipelines using Spark, Databricks, or a comparable distributed data platform.
- Strong understanding of data modeling, lakehouse or warehouse design, incremental processing, schema evolution, idempotency, backfills, and lineage.
- Experience designing data-quality controls and reconciling complex datasets to source systems or independent control totals.
- Experience with AWS, Git-based development, automated testing, CI/CD, and practical tradeoffs involving reliability, security, performance, and cost.
- The ability to work directly with customer technical teams, understand unfamiliar schemas, ask precise questions, and document decisions clearly.
- Comfort operating with significant ownership and ambiguity where customer implementation, platform development, security, and production operations overlap.
- The judgment to use AI development tools effectively while verifying generated code, tests, and transformations against source evidence.
You’d be a great fit if you also have:
- Experience with P&C insurance data, including policy transactions, coverages, claims, premium, exposure, rating, or underwriting data.
- Experience integrating with policy administration, claims management, rating, or other operational source systems.
- Deep experience with Databricks, Delta Lake, Unity Catalog, Databricks Workflows, or configuration-driven data pipelines.
- Experience with Terraform, secure file transfer, database replication, or customer-specific ingestion infrastructure.
- Experience supporting or building machine-learning feature pipelines, point-in-time datasets, model monitoring, MLflow, SageMaker, or production scoring systems.
$105.4k - $207.8k
Position Summary Our Deloitte AI & Engineering team works to transform technology platforms, drive innovation, and help make a significant... ..., and fuel growth through innovation.Work you'll do As a Data Engineer III on the AI & Data team, you will be responsible for...SuggestedLocal area$80k - $133k
Job Family:Data Science & AnalysisTravel Required:Up to 10%Clearance Required:Ability to Obtain Public TrustWhat You Will Do:Guidehouse seeks a Senior Data Engineer to design, develop, and optimize modern data platforms, pipelines, and cloud-based analytics solutions. The...SuggestedFull timeContract workFlexible hours- ...DescriptionArchitects, develops, tests, and maintains scalable data pipelines to ingest structured and unstructured data from disparate... ...CoE, translating data requirements from data scientists, ML engineers, and analysts into technical specifications for logistics...Suggested
- ...Job Functions: Provide PIR Time Constrained, Content Centric (TC3) and Content Dominant (CD) in-depth analytical support. Improve raw data and PED systems supporting Requests for Information (RFI), rapid scripting, process improvement, technique discovery, and raw data...SuggestedRemote work
- Job SummarySee below for important information regarding this job. Position will be filled at any of the locations listed below. Site specific salary information as follows:Battle Creek, MI: $125,776- $163,514Columbus, OH: $131,245- $170,624Dayton, OH: $130,461 - $169,...Suggested
- Position Summary Join our Deloitte AI & Engineering team to transform technology platforms, drive innovation, and help make a significant... ...fuel growth through innovation.Work you'll doAs a Lead AI and Data Science Engineer II on the AI & Data team, you will be...Local area
- ...Senior Ai Data EngineerJoin a team where innovation meets mission. Our AI, cloud, cyber, and modernization solutions save agencies thousands... ...challenges.Credence has an immediate need for a Senior AI Data Engineer to join our growing AI and Automation practice. You will be a...Immediate startWorldwide
$140k - $190k
...The Data Engineer will be writing code that moves data through a pipeline, fixing data pipeline issues, optimizing data systems, and collaborating with stakeholders, day to day. The Capital Software Project is building a data fabric system. The data fabric system includes...Contract workWork experience placementWork at office- ...Responsibilities: Build and operate production-grade batch and streaming data pipelines across SQL Server, cloud applications, device/event... ...on-call activities. Work closely with Product, Application Engineering, Quality Assurance, DevOps/Site Reliability Engineering,...
- ...is currently a hybrid position that includes participation at an in-person workshop. We are seeking a skilled, customer-focused Data Engineer with experience in Databricks, Apache Spark, Python, AWS, Microsoft 365, and SharePoint. Ideal candidates will have a solid grasp...For contractorsFlexible hours
- ...Systems EngineeringTravel Required:NoneClearance Required:Ability to Obtain Public TrustWhat You Will DoSupport data analytics, reporting, and data engineering activities across enterprise data lakes, data warehouses, and business intelligence platforms.Design, develop,...Full timeWork at officeFlexible hours
$150k - $180k
...Lead Control Engineer // Data Center Developer Location: Austin, Texas (On-site)Relocation Needed Compensation: $150,000–$180,000 base About the Company A leading developer of hyperscale data centers, delivering scalable, efficient, and sustainable solutions...RelocationFlexible hours- Job Family:Data Science & AnalysisTravel Required:Up to 10%Clearance Required:Ability to Obtain Public TrustWhat You Will Do:Guidehouse seeks a Data Engineer to support the development, maintenance, and enhancement of data pipelines, cloud-based data platforms, and analytics...Permanent employmentFull timeInternshipWork at officeFlexible hours
$135k - $170k
...Enterprise Data Engineer Location: US-OH-Dayton ID: 2026-4370 Category: Engineering Position Type: Full Time Salary Riverside Overview Riverside Research is an independent National Security Nonprofit dedicated to research and development in the national...Civilian ContractorFull timeLocal area- ...Data EngineerReporting to the AI & Data Solutions Manager, the Data Engineer will design, build, and maintain scalable data solutions that support reporting, analytics, data governance, and informed decision-making across Covation Global / United Wheels Inc. This role...
- ...A technology firm supporting national defense is looking for a Database Engineer in Dayton, OH. The role involves maintaining and enhancing database systems, ensuring compliance with security requirements, and directly supporting intelligence operations. Candidates should...
- ...ARCTOS a DCS company seeks a Technical Leader to establish and lead the Integrated Data and Engineering Analytics (IDEA) team within the Aerospace Structures & Materials Department. The role will define the architectural roadmap and build a multidisciplinary team to transform...
- ...Job Description Summary Responsible for designing, developing, testing and implementing data engineering processes to generate analytical and reporting solutions. Responsible for analyzing and preparing the data needed for data science based outcomes. Also responsible...Permanent employmentFull timeVisa sponsorshipWork visaRelocation package
$110k - $135k
Job DescriptionEverforth ECS is seeking a Data Analyst, GEOINT Technical Subject Matter Expert (SME), to work at our Dayton, OH customer site. Note: This position is contingent upon contract award and final acceptance by government customer.Data Analyst - GEOINT Technical...Contract workRemote work- ...Our work includes enterprise architectural assessments, systems engineering and integration, test, planning and execution, cost estimating... ...solutions for our National Defense!SPA has near-term need for a Data Scientist.ResponsibilitiesThe Data Scientist will support NASIC...Full timeWork at office
$103.47k - $132.5k
...Our capabilities in cybersecurity, network architecture, reverse engineering, software and hardware development uniquely enable us to... ...Technologies is leading the next evolution of national defense - the data evolution - by accelerating a breadth of national security...Full timeWork at officeLocal areaWorldwide$99k - $225k
Data ScientistThe Opportunity:As a data scientist, you’re excited at the prospect of unlocking the secrets held by a data set, and you’re fascinated by the possibilities presented by IoT, machine learning, and artificial intelligence. In an increasingly connected world...Full timeContract workPart timeWork at officeLocal areaRemote work- ...Join to apply for the Database Engineer role at Evans & Chambers Technology Evans & Chambers Technology is seeking a highly motivated Database... ...We work with our customers on everything from conquering their data to improving and safeguarding IT infrastructure. Our ultimate...Full timeWork experience placement
- ...Database Engineer Belong. Connect. Grow. with KBR! KBR's National Security Solutions team provides high-end engineering and advanced... ...access to compute power, modeling and simulation tools, and data sharing across organizations to the Department of the Air Force...Local areaFlexible hours
$140k - $160k
...from secure cloud architectures and enterprise infrastructure to data center operations, scientific analysis, cutting-edge cyber... ....Mission Focused Expertise: From veteran leadership to cleared engineers, our people understand both the technology and the mission.Summary...Work experience placementRelocationFlexible hours- ...Winsupply organization, contributing to innovative solutions and long-term success. Position Summary We are seeking a detail-oriented Data Analyst Intern to join our team. This role is focused on supporting Support Services. As an intern, you will assist in maintaining...InternshipLocal area
- ...last five years. Learn more at CFD Research is seeking an AI/ML Engineer to build and deploy advanced artificial intelligence and machine... ...security ecosystem. The candidate will contribute end-to-end—spanning data processing, model development, MLOps, and integration of...Permanent employmentFull timeTemporary workPart timeWork at office
- ...Data Scientist Remote, Anywhere in the Continental US Due to the nature of the role, this position requires that you are a U... ...technical and non-technical audiences. Collaborate with policy, engineering, and operations teams to ensure analytics solutions align with...Full timeImmediate startRemote work
- ...Great Place to Work® certification year after year. Principal Data Scientist Job requirements ~ Experience Range: 15 - 18... ...PySpark, and R, ensuring efficient data processing and feature engineering Develop, validate, and maintain probabilistic graph models and...
- ...Data Scientist - Mid Dayton, Ohio, United States Clearance Required: TS/SCI Dayton, Ohio - 100% onsite Position Overview... ...-day implementation of basic automation routines. The mid-level engineer supports the enterprise software pipeline by preparing high-quality...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Data Engineer. Be the first to apply!
- data engineer machine learning Dayton, OH
- finance data engineer Dayton, OH
- data center engineer Dayton, OH
- senior cloud data engineer Dayton, OH
- data engineer Dayton, OH
- data engineer analytics Dayton, OH
- senior data center engineer Dayton, OH
- data science developer Dayton, OH
- data developer Dayton, OH
- ai data Dayton, OH



