Data Engineer
Soteris
ABOUT SOTERIS
Soteris is a YC-backed AI company building the future of pricing and product management for the
insurance industry. Our mission is to infuse the $5 trillion P&C insurance industry with best-in-class,
proprietary, AI-driven data analytics. We’ve spent years building our own proprietary AI models on
personal auto claims and exposure data to help insurers improve their loss ratios, with over 100 million
submissions and $180 billion in premium scored to date.
Each year, roughly $750 billion in insurance policies are written in the United States. Our machine
learning platform helps insurers evaluate policies at a granular level, moving beyond broad segmentation approaches that often lead to risks being over- or underpriced. Our modeling approach incorporates multiple model families and calibration methods to rank policy risk within a book of business.
We are a team of 10 and growing quickly. As our second Data Engineer, you will own the data layer that turns customer policy, claims, quote, and financial data into trusted inputs for actuarial analysis, model development, and production scoring. This is a builder role: you will work directly with customer data teams, create repeatable ingestion and transformation pipelines, reconcile outputs to source-of-truth control totals, and make the platform easier to operate as we add customers and products.
WHAT YOU'LL BE DOING
Customer Data Onboarding and Integration
- Leading the technical data workstream for new customer implementations by understanding policy, claims, rating, quote, and financial systems and establishing secure access to the data.
- Building reusable extraction and synchronization workflows for databases, backups, secure file transfer, APIs, and other delivery methods while preserving source lineage and supporting backfills.
- Mapping customer data into Soteris’s internal ontology and working with customer technical teams and internal project leads to resolve definitions, transformations, data-quality issues, and onboarding blockers.
Lakehouse and Pipeline Engineering
- Owning Databricks and AWS pipelines that move customer data from raw and Bronze ingestion through standardized Silver tables and curated Gold or model-ready datasets.
- Designing idempotent, incremental, observable workflows that handle schema evolution, late-arriving data, backfills, orchestration, performance, and cost.
- Developing shared components and configuration-driven patterns, then publishing well-defined datasets for actuarial analysis, backtesting, model training, production scoring, reporting, and monitoring.
Data Quality and Modeling Readiness
- Building automated quality gates for completeness, uniqueness, referential integrity, valid ranges, freshness, balance, schema drift, and other customer-specific controls.
- Reconciling written and earned premium, exposure, incurred and ultimate loss, claim counts, fee income, and other economic drivers to customer control statistics at the required state, program, year, and coverage levels.
- Partnering with data science to produce leakage-resistant, point-in-time-correct datasets and productionize approved actuarial reference data
Production Platform and Operational Ownership
- Owning data flows for quote requests, model inputs, and bound policy outcomes, including pre-live comparisons that confirm production request fields match corresponding policy data and post-launch drift checks.
- Operating pipelines with monitoring, alerting, recovery behavior, runbooks, and clear incident diagnostics across customer synchronization, Databricks processing, and downstream model- serving dependencies.
- Testing, reviewing, and deploying data code through GitHub and GitHub Actions, with automated tests and deployment controls for production data assets.
- Managing Databricks permissions, Unity Catalog controls, sensitive data, and least-privilege access in support of Soteris’s security and SOC 2 requirements.
OUR CURRENT STACK
- Python, SQL, PySpark, pandas, and related data-engineering libraries
- Databricks, Delta Lake, Databricks Workflows, and Unity Catalog
- AWS, including S3, Lambda, EC2, ECS, SageMaker, and secure customer file transfer
- Customer databases, backups, SFTP, APIs, and file-based ingestion
- Terraform, GitHub, GitHub Actions, automated testing, and infrastructure as code
- MLflow, SageMaker, and production scoring APIs at the model handoff boundary
- Modern generative AI development tools, including Claude, ChatGPT, Codex, or similar models
ABOUT YOU
You must have the following:
- Strong Python and SQL skills and the ability to write production-quality transformations, tests, utilities, and operational tooling.
- Hands-on experience building and operating production pipelines using Spark, Databricks, or a comparable distributed data platform.
- Strong understanding of data modeling, lakehouse or warehouse design, incremental processing, schema evolution, idempotency, backfills, and lineage.
- Experience designing data-quality controls and reconciling complex datasets to source systems or independent control totals.
- Experience with AWS, Git-based development, automated testing, CI/CD, and practical tradeoffs involving reliability, security, performance, and cost.
- The ability to work directly with customer technical teams, understand unfamiliar schemas, ask precise questions, and document decisions clearly.
- Comfort operating with significant ownership and ambiguity where customer implementation, platform development, security, and production operations overlap.
- The judgment to use AI development tools effectively while verifying generated code, tests, and transformations against source evidence.
You’d be a great fit if you also have:
- Experience with P&C insurance data, including policy transactions, coverages, claims, premium, exposure, rating, or underwriting data.
- Experience integrating with policy administration, claims management, rating, or other operational source systems.
- Deep experience with Databricks, Delta Lake, Unity Catalog, Databricks Workflows, or configuration-driven data pipelines.
- Experience with Terraform, secure file transfer, database replication, or customer-specific ingestion infrastructure.
- Experience supporting or building machine-learning feature pipelines, point-in-time datasets, model monitoring, MLflow, SageMaker, or production scoring systems.
- Be part of a dynamic team where your distinctive skills will contribute to a winning culture and team.As a Lead Data Engineering at JPMorgan Chase within the Consumer and Community Banking team, you design, develop, and maintain robust data pipelines and architectures....Suggested
- We are looking for a Data Engineer to join a Financial Services team in Plano, Texas on a contract basis with the potential for a permanent role. This role is focused on designing and delivering reliable data pipelines in a cloud-first environment, with Snowflake serving...SuggestedPermanent employmentContract work
$82.42k - $126.6k
...driving digital transformation for financial institutions. We specialize in leveraging advanced technologies such as AI, cloud, and data-led innovation to help our clients accelerate growth and unlock business value. Our AI-driven solutions empower financial institutions...SuggestedFull timeTemporary workRelocation- ...capture, manage, store and utilize structured and unstructured data from internal and external sources. Establishes and builds processes... ...to systems and storage to accommodate ongoing needs.Why TI?Engineer your future. We empower our employees to truly own their career...SuggestedInternshipLocal area
- Req ID:373994NTT DATA strives to hire exceptional, innovative and passionate individuals who want to grow with us. If you want to be... ...thinking organization, apply now.We are currently seeking a Data Engineer to join our team in Plano, Texas (US-TX), United States (US).General...SuggestedWork at officeRemote workFlexible hours
- ...Lead Data EngineerLocations: Richmond - VA / McLean - VA/ Plano - TX / Chicago - IL / NYC - NY / Wilmington - DE (Preferred Location... ...:Lead the design, development, and deployment of scalable data engineering solutions that deliver operational and analytical data to third...Work at office
- ...Lead Data EngineerJoin us as we embark on a journey of collaboration and innovation, where your unique skills and talents will be valued... ...brighter future and make a meaningful difference.As a Lead Data Engineer at JPMorganChase within the Corporate Sector, you are an...
- ...Be part of a dynamic team where your distinctive skills will contribute to a winning culture and team. As a Lead Data Engineering at JPMorgan Chase within the Consumer and Community Banking team, you design, develop, and maintain robust data pipelines and architectures...
$200k
...Lead Data Engineer - Remote 100% Remote (U.S. Based) Up to $200,000 Base Salary U.S. Citizens & Green Card Holders Only I'm partnering with a rapidly growing healthcare technology company that is looking to hire a Lead Data Engineer to help drive the evolution...Remote work- ...Data Engineering LeadLead end-to-end delivery of large-scale data engineering initiatives, ensuring high-quality and timely outcomes. Architect, design, and implement scalable, cloud-native data platforms using Databricks, Apache Spark (Scala), and AWS services. Develop...
$102.68k - $171.13k
NTT DATA strives to hire exceptional, innovative and passionate individuals who want to grow with us. If you want to be part of an inclusive... ..., apply now.We are currently seeking a Senior Oracle Data Engineer - Onsite to join our team in Plano, Texas (US-TX), United States...Full timeTemporary workWork at officeRemote workFlexible hours$179.4k - $204.7k
...Lead Data Engineer (Java, Python, AWS) Do you love building and pioneering in the technology space? Do you enjoy solving complex business problems in a fast-paced, collaborative,inclusive, and iterative delivery environment? At Capital One, you'll be part of a big group...Full timePart timeInternshipH1bLocal area$50 per hour
...Data EngineerWork Location: Plano, TX or Atlanta, GAMaximum Pay Rate: $50/hr (to me) Client: BOFACandidate’s Experience: 5 Years of ExperienceCandidates will also need to go through a tech screening.We are looking for candidates with the following primary skills for this...- ...The data engineer will be responsible for designing, building, and maintaining the company's data infrastructure and analytics pipelines to support the business. About Digicode Digicode is a leading global iGaming company that provides cutting-edge technology and...
- ...Job TitleStrong in SQL and Python, 8+ years' experience.Experience with big data frameworks (i.e. Hadoop and Spark), 5+ years' experience.Experience building automated data pipelines, 5+ years of experience.Experience performing data analysis and data exploration, 5+...
$100k - $120k
...Data Engineering - DevOps Fractal Analytics is a strategic AI partner to Fortune 500 companies with a vision to power every human decision in the enterprise. Fractal is building a world where individual choices, freedom, and diversity are the greatest assets. An ecosystem...Hourly payFull timeLocal areaRemote work- ...Location: Remote We are seeking an experienced Data Engineer to support challenging data engineering, data management, and analytics initiatives. In this role, you will develop and maintain data pipelines and data structures that enable advanced analytics and data...Temporary workLocal areaRemote work
- ...TX (Hybrid) 3+ years About This Role Build and maintain data pipelines that power analytics and AI initiatives for our clients... ...Requirements ~3+ years of experience in data engineering ~ Proficiency in SQL and Python ~ Experience with data warehousing...
- ...Data EngineerWe are seeking a skilled Data Engineer to design, build, and maintain scalable data pipelines and analytics solutions within our Azure Databricks environment. The ideal candidate will have hands-on experience with Apache Spark, Delta Lake, and Azure cloud...Full timeTemporary workRemote work
- ...Data EngineerTalent should have 7 years of experience in software analysis, design and development with deep expertise in Java for backend and distributed systems. Strong background in streaming technologies (Dataflow, Kafka, Flink, or equivalent). Proven experience in...
- ...Job Description Job Description We are seeking a detail-oriented Data Engineer to design, develop, and optimize scalable data solutions and SQL Server databases. This role focuses on building efficient data pipelines, developing high-performance T-SQL solutions,...Local area
- ...Lead Data Platform EngineerLocation: PLANO, TX- 75024 Duration: 0-4+ months Shift: M-F 8am-4pm Hybrid Pay Rate: $51.09-$72.99/HR on... ...data protection, and compliance controls.Collaborate with data engineering, analytics and cross-functional teams to enable efficient data...Shift work
- ...Data Engineer Opportunity At AscenttAscentt is transforming the future of manufacturing through advanced Data Analytics, AI/ML, and Generative AI solutions. We partner with global manufacturing enterprises to convert complex industrial data into actionable, real-time business...
$60 per hour
...Senior/Lead Data EngineerTrident Consulting is seeking a Senior/Lead Data Engineer for one of our clients. A global leader in business and technology services.Job Title: Senior/Lead Data EngineerLocation: RemoteJob Type: ContractPay Rate: $60Position SummarySeeking a...Contract work- ...Sr. Snowflake Data EngineerWe are seeking a skilled Data Engineer with a strong background in Snowflake to join our dynamic team. The ideal candidate will have 7 to 10 years of experience in data engineering, with a proven track record of designing, building, and maintaining...
$63 per hour
...Data EngineerBFS -Domain preferred***Max rate $63***Hybrid in Plano TXExperience: 5 to 7 Years Required Skills: Unix, Java, Spring Boot, Spark, SQL, AWSAs a Data Engineer within our multinational corporation, you will leverage your expertise in Java, SpringBoot, Spark...- ...Data EngineerWe are seeking a skilled professional with over 5 years of experience in AWS, Python, Scala, Spark, and SQL. The ideal candidate will have a strong understanding of data processing, API development, and data pipeline construction.Key responsibilities include...Immediate start
- ...Engineer IIThe Engineer II role plans, designs, develops and tests high-quality, high-performance, innovative, and fully functioning products and live services related to Toyota vehicle data. The scope of the projects includes ingestion of data from multiple types of...
- ...-on experience; expert level - 1-2 years of hands-on experience; complex problem solving - Training and small amount of hands-onBig Data/Hadoop - Working experience with Hadoop stack dealing huge volumes of data. - 3-4 experience working with Hadoop / EMR or any other...Work experience placementWork at officeRemote work
- ...Job TitleTechnical hands on skill and minimum 10+ years of on job experience on below tech stack:Data factoryAzurePysparkAzure DatabricksDelta lakePythonSynpaseGood to have:Ability to analyze and understand data, data flows, end to end reconciliation of data should be...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Data Engineer. Be the first to apply!
- data engineer machine learning Plano, TX
- aws data engineer Plano, TX
- big data developer Plano, TX
- sr data engineer Plano, TX
- big data cloud engineer Plano, TX
- finance data engineer Plano, TX
- junior big data engineer Plano, TX
- sr information security engineer Plano, TX
- data center engineer Plano, TX
- senior data integration developer Plano, TX


