Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Data Platform Engineer

Take2 Consulting LLC

Take2 has proven experience bridging the intersection of technology and people solutions. As a proven, trusted provider for our Federal and commercial clients, we provide the right solutions, at the right time through trusted partnerships, customized to solve our client’s unique business challenges. Take2 invests time, discipline, and rigor into our technology and people solutions, as well as utilizes our proprietary People Cloud. Whether we are bridging the gap between IT talent and our customers’ business challenges, Take2 will work as a partner to best resolve client needs.

Take2 is hiring an AWS Lakehouse Data Engineer who is eligible to be sponsored for a Public Trust Clearance . This position is Remote , but it will require you to work East Coast Hours while being located in the United States.

Job Description:

Take2 is seeking an AWS Lakehouse Data Engineer to design, implement, and operate the cloud-native data platform that powers AI/ML, analytics, reporting, and data visualization. You will build a modern lakehouse on Amazon S3 using AWS-native services and open table formats, providing Databricks-like capabilities while maintaining portability, strong governance, cost efficiency, and operational control. You will also develop scalable batch and streaming ingestion, Python and PySpark ETL/ELT pipelines, metadata and governance services, and automated cloud provisioning and CI/CD across environments.

This role is ideal for an engineer who enjoys platform building, automation, performance optimization, and enabling advanced analytics through trusted, secure, and well-governed data.

What You Will Do

Build and Operate Data Pipelines (Batch and Streaming)

  • Design and implement batch and streaming ingestion from APIs, relational databases, file drops, event streams, and external partners.
  • Implement, test, and optimize ETL/ELT pipelines using Python and PySpark to produce curated, analytics-ready datasets for reporting, visualization, and machine learning.
  • Implement incremental processing, change data capture (CDC), data contracts, schema validation, and reusable transformation frameworks.
  • Improve pipeline reliability through automated testing, orchestration, monitoring, retry handling, and operational runbooks.

Deliver an AWS-Native Lakehouse Data Platform

  • Design and implement a Delta Lakehouse-style data platform using AWS-native services to provide Databricks-like capabilities for data engineering, analysis, and data visualization.
  • Build and manage a scalable lakehouse on Amazon S3 using Apache Iceberg and open columnar formats such as Apache Parquet.
  • Implement SQL-like table reliability for data stored in Amazon S3, including ACID transactions, schema evolution, partition evolution, snapshot isolation, time travel, and rollback capabilities using Apache Iceberg.
  • Enable fast, interactive querying of lakehouse data using AWS-native query and compute services such as Amazon Athena, Amazon EMR, AWS Glue, and Amazon Redshift where appropriate.
  • Optimize performance and cost through partitioning, compaction, file sizing, statistics, caching, lifecycle policies, and efficient separation of compute and storage.
  • Establish standardized development, test, and production environments with consistent configuration and controlled promotion across stages.

Metadata, Governance, Access Control, Lineage, and Quality

  • Implement data governance and fine-grained access control using AWS-native services, including AWS Lake Formation, AWS Glue Data Catalog, AWS Identity and Access Management (IAM), AWS Key Management Service (KMS), and related security services.
  • Implement a managed metadata repository for dataset cataloging, ownership, business definitions, tagging, classification, and discoverability.
  • Enable end-to-end lineage from source through transformation and consumption to support auditability, impact analysis, and regulatory requirements.
  • Apply policy-based access, least-privilege permissions, row-, column-, and cell-level controls where required, data classification, retention, encryption, and secure data handling.
  • Build operational data quality checks for freshness, completeness, uniqueness, validity, consistency, and anomaly detection, and publish measurable SLAs/SLOs.

AWS Automation, CI/CD, and Operations

  • Implement automated AWS provisioning using Infrastructure as Code (IaC) to create consistent environments and secure-by-default baselines.
  • Build and enhance CI/CD for data pipelines and lakehouse components, including automated tests, security checks, validation gates, packaging, deployment, promotion, and rollback strategies.
  • Implement observability with centralized metrics, logs, traces, alerts, dashboards, runbooks, and incident-response procedures.
  • Continuously evaluate platform performance, scalability, reliability, security, and cost, and implement measurable improvements.

Cross-Team Collaboration and Documentation

  • Work closely with data, application, analytics, AI/ML, security, networking, and cloud platform teams to support mission needs and delivery timelines.
  • Maintain high-quality engineering documentation, including architecture diagrams, data models, SOPs, interface specifications, operational runbooks, and secure configuration baselines.
  • Present technical findings, trade-offs, risks, and recommendations clearly to technical and non-technical stakeholders.

What You Will Need

  • Bachelor's degree in Engineering, Information Technology, Computer Science, Data Engineering, or a related field, or FOUR (4) years equivalent practical experience in leu of degree.
  • SIX (6) years of relevant experience.
  • Hands-on experience implementing AWS-native data lake or lakehouse architectures using Amazon S3 and services such as AWS Glue, Amazon Athena, Amazon EMR, AWS Lake Formation, and Amazon Redshift.
  • Strong experience developing production ETL/ELT pipelines using Python and PySpark, including data modeling, transformation, testing, performance tuning, and error handling.
  • Hands-on experience with Apache Iceberg, including ACID transactions, snapshots, schema and partition evolution, time travel, table maintenance, and query optimization.
  • Advanced SQL skills and experience supporting analytical queries, semantic layers, reporting tools, and data visualization workloads.
  • Experience implementing metadata management and governance capabilities, including cataloging, lineage, ownership, classification, policy enforcement, and fine-grained access controls.
  • Experience with AWS security fundamentals, including IAM and least privilege, KMS encryption, secrets management, network security, logging, and secure SDLC practices.
  • Experience provisioning AWS resources using IaC and operating data platforms across multiple environments.
  • Experience building or operating CI/CD pipelines for data workflows, including testing, packaging, deployment automation, environment promotion, and rollback.
  • Ability to troubleshoot distributed data-processing workloads and optimize performance, reliability, and cost.

What Would Be Nice to Have

  • Hands-on experience with Databricks, Delta Lake, or migrating Databricks workloads to AWS-native services and Apache Iceberg.
  • Experience with AWS Step Functions, Amazon Managed Workflows for Apache Airflow (MWAA), Amazon Kinesis, AWS Database Migration Service (DMS), AWS Lambda, Amazon MSK, or similar ingestion and orchestration services.
  • Experience with modern DevOps practices and tools such as Git, Terraform, AWS CloudFormation or AWS CDK, Jenkins, AWS CodePipeline, GitHub Actions, and Docker.
  • Experience integrating lakehouse data with business intelligence and visualization tools such as Amazon QuickSight, Tableau, or Power BI.
  • Experience using AI-assisted coding tools, such as GitHub Copilot, ChatGPT, Cursor, or Kiro, to accelerate implementation while maintaining code quality, testing, review, privacy, and security controls.
  • Knowledge graph and Graph RAG experience, including graph modeling, ontology and taxonomy alignment, entity resolution, relationship extraction, and hybrid retrieval that combines graph traversal with semantic or vector search.
Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Data Platform Engineer in Santa Clara, CA vacancy
  • $70 - $75 per hour

     ...Meghana GorusuCompany: SRI Tech SolutionsData Platform EngineerLocation - San Jose, CA ( 4 Days...  ...berribot coding test within 24-48hr. Data & Analytics Technologies SQL - Advanced...  ...warehousesSnowflake, RedshiftData Engineering & QualityUnderstanding of: Data models,... 
    Suggested
    Hourly pay

    SRI Tech

    San Jose, CA
    3 days ago
  • $240.1k - $420.2k

    Company DescriptionIt all started when engineer Fred Luddy wrote code that automated a tedious task for his coworker,...  ...control tower for business reinvention. Our ServiceNow AI platform brings together any AI, any data, and any workflow— helping 85% of the Fortune 500 work... 
    Suggested
    Work at office
    Immediate start
    Remote work
    Flexible hours

    ServiceNow

    Santa Clara, CA
    1 day ago
  • $176.1k - $308.2k

    Company DescriptionIt all started when engineer Fred Luddy wrote code that automated a tedious task for his coworker,...  ...control tower for business reinvention. Our ServiceNow AI platform brings together any AI, any data, and any workflow— helping 85% of the Fortune 500 work... 
    Suggested
    Work experience placement
    Work at office
    Immediate start
    Remote work
    Flexible hours
    2 days per week

    ServiceNow

    Santa Clara, CA
    1 day ago
  • $176.1k - $308.2k

     ...enterprises face: who can and should take what action on what data.Veza's Access Graph platform maps an organization's entire identity ecosystem across...  ..., data, cloud environments, and AI agents. For engineers joining Veza today, this means the scale and resources of... 
    Suggested
    Work at office
    Remote work
    Flexible hours

    ServiceNow

    Santa Clara, CA
    1 day ago
  •  ...AI systems that automate complex business workflows across platforms such as SAP, Salesforce, Workday, Snowflake, MuleSoft, and...  ...serious, AI forward. About the role We’re looking for a Data Platform Engineer to build the software that connects enterprise systems and... 
    Suggested
    Remote work

    Tessera Labs

    San Jose, CA
    4 days ago
  •  ...Overview Senior Data Platform Engineer - Direct-Hire/FTE - Remote (US) This is a hands-on Senior Data Platform Engineer role that will require strong Python/PySpark coding skills. This is NOT an ETL/ELT development role. Responsibilities Contribute to the enhancements... 
    Full time
    Work experience placement
    Local area
    Remote work
    Flexible hours

    INSPYR Solutions

    San Jose, CA
    4 days ago
  •  ...A technology solutions company seeks a skilled Back-End Engineer in Sunnyvale, California. You will design and maintain scalable data pipelines and back-end systems using Spark, Python, and Java. The role demands collaboration with data scientists and knowledge of cloud... 

    Robotics Prcocess Automation, LLC

    Sunnyvale, CA
    3 days ago
  •  ...Amazon.com Services LLC in Sunnyvale is seeking a Software Development Engineer II to design and build scalable data processing components powering data catalog, discovery, and transformation. You will own end-to-end delivery across design, implementation, testing, deployment... 

    Jobleads-US

    Sunnyvale, CA
    2 days ago
  •  ...Adobe is seeking a Senior Software Engineer for its Real-Time Customer Data Platform (RTCDP) in California. You will architect globally distributed systems for real-time identity resolution, consent management, and highly scalable customer profiles, while integrating AI... 

    Jobleads-US

    San Jose, CA
    2 days ago
  •  ...ServiceNow is seeking a Principal Data Platform Software Engineer (RaptorDB) to help architect, build, and operate large-scale database tooling across our hybrid data centers and cloud environments. You will lead architectural discussions, drive best practices, and... 

    Jobleads-US

    Santa Clara, CA
    2 days ago
  • $201.3k - $352.3k

     ...It all started when engineer Fred Luddy wrote code that automated a tedious task for his coworker, Phyllis. She...  ...control tower for business reinvention. Our ServiceNow AI platform brings together any AI, any data, and any workflow— helping 85% of the Fortune 500® work... 
    Full time
    Work experience placement
    Work at office
    Immediate start
    Remote work
    Flexible hours
    2 days per week

    ServiceNow

    Santa Clara, CA
    10 days ago
  •  ...Technical Staff 3 to design and implement a next-generation Data Platform for on-prem and cloud deployments. You will help shape data models...  ...buses, and databases. This role emphasizes hands-on data engineering and collaboration with cross-functional teams. The team... 

    Jobleads-US

    San Jose, CA
    3 days ago
  •  ...direction and architecture strategy for the data access layer. Build and improve core...  ..., performance, and multiple database engines. Own features from design through production...  ...design. ~ Advanced-to-expert backend platform development experience. ~ Experience... 
    Full time
    Work experience placement
    Work at office
    Flexible hours
    2 days per week

    ServiceNow

    Santa Clara, CA
    10 days ago
  • $240.1k - $420.2k

     ...with database development, infrastructure automation, systems engineering, security, architecture, and customer support teams. Requirements...  ...+ years focused on database systems, infrastructure, or cloud platforms. ~ Deep expertise in database operations, operational... 
    Full time
    Flexible hours

    ServiceNow

    Santa Clara, CA
    4 days ago
  • $240.1k - $420.2k

     ...functional architecture and design sessions. Requirements ~12+ years of software engineering experience, including 8+ years focused on database systems, infrastructure, or cloud platforms. ~ Deep expertise in database operations, operational management, and backup... 
    Full time
    Flexible hours

    ServiceNow

    Santa Clara, CA
    17 days ago
  • $201.3k - $352.3k

     ...direction and architecture strategy for the data access layer. Design and improve core...  ..., and intermittent failures. Mentor engineers and raise the technical bar across the organization...  ...design. ~ Advanced-to-expert backend platform development experience. ~ Knowledge of... 
    Full time
    Work experience placement
    Work at office
    Flexible hours
    2 days per week

    ServiceNow

    Santa Clara, CA
    10 days ago
  •  ...NVIDIA Corporation seeks a Manager of Data & AI Engineering to lead and grow a high-performing team building the next-generation autonomous data intelligence platform for Supply Chain Operations. The role blends hands-on architecture with delivery leadership, partnering... 

    Jobleads-US

    Santa Clara, CA
    20 hours ago
  •  ...Roche, where every voice matters.The PositionAs a Senior Data Scientist - ML Engineering, you will join the Global Analytics and Technology Center...  ...across the engineering team.The OpportunityML Systems & Platform Engineering: Architect, deploy, and maintain end-to-end production... 
    Full time
    Relocation package

    Roche

    San Jose, CA
    2 days ago
  •  ...Walmart is seeking a Senior Data Engineer in Sunnyvale, CA to design and deploy scalable data solutions powering agent-enabled AI for Sam’s Club. Collaborate with AI/ML, product, and engineering teams to ensure data services are trustworthy, discoverable, and consumable... 

    Jobleads-US

    Sunnyvale, CA
    20 hours ago
  •  ...Enigma's Tolias Lab in Palo Alto seeks a Data Engineer to build and maintain the data platform supporting Enigma’s experimental workflows. You will design, implement, and operate ETL/ELT pipelines moving raw data from acquisition to downstream stores, with emphasis on... 

    Jobleads-US

    Palo Alto, CA
    2 days ago
  •  ...Mindlance seeks a Senior Data Engineer to design and build infrastructure supporting Finance and Business Model Strategy. You will architect data models, integrate diverse sources, and manage a data warehouse to unlock insights for internal users and external customers... 

    Jobleads-US

    San Jose, CA
    1 day ago
  •  ...SLAC invites applications for a Data Engineer (Software Developer 2) to design, build, and maintain a data platform supporting Enigma’s experimental workflows. You will own end-to-end data pipelines, implement durable data models, and partner with researchers to enable... 

    Jobleads-US

    Palo Alto, CA
    2 days ago
  • $96.49k - $144.74k

     ...NTT DATA strives to hire exceptional, innovative and passionate individuals who want to grow with us. If you want to be part...  ...organization, apply now.   We are currently seeking a Data Engineer - Data Platform (Spark/Kafka/Flink/Scala/Java) - Onsite Hybrid to join our... 
    Temporary work
    Work experience placement
    Work at office
    Remote work
    Flexible hours

    NTT DATA, Inc.

    Cupertino, CA
    a month ago
  •  ...for managing an organization.Job DescriptionPrimary Skills:Big Data experience8+ years exp in Java, Python and ScalaWith Spark and...  ..., calling APIs, write SQL queries, etc.).Work closely with our engineering team to integrate and build algorithmsProcess unstructured data... 

    Comtech

    San Jose, CA
    5 days ago
  •  ...4-06-04Company Name:HITACHI ENERGY USA INCProfession (Job Category):Engineering & ScienceJob Schedule: Full timeRemote:NoJob Description:General information:Hitachi Energy is seeking for a Senior Data Management Engineer for it's San Jose, CA location.Your Responsibilities... 
    Full time

    Hitachi

    Santa Clara, CA
    3 days ago
  •  ...is complex, and we need to build a robust data layer to be able to answer core questions around...  ....We’re looking for a startup-minded data engineer who can wear a lot of hats, work with multiple teams, and build the data platform needed to support answering these questions... 
    Work at office
    Home office
    Flexible hours
    3 days per week

    Gridmatic

    Cupertino, CA
    3 days ago
  •  ...We are seeking an experienced Kinaxis Data Migration Engineer with 8–10 years of experience in data migration, supply chain systems, data analysis...  ...RapidResponse/Maestro or comparable supply chain planning platforms. ~ Strong understanding of master data management and... 
    Full time

    Long Finch Technologies

    San Jose, CA
    11 days ago
  •  ...801708Reference26-07957Remote100% Remote Job Title: Contract Data Engineer (API & Database Focus) Role Summary: We are seeking an experienced...  ...and loading processes (ETL/ELT) within the data platform. Manage and optimize data storage and schemas in Snowflake and... 
    Contract work
    Remote work

    Mindlance

    Sunnyvale, CA
    1 day ago
  • Position: Data Engineer IILocation: Sunnyvale, CaliforniaDuration: ContractJob ID: 175305Job Overview: We are seeking a skilled and detail...  ...knowledge of SQL and database systems.Familiarity with cloud platforms such as AWS, Azure, or Google Cloud.Excellent problem-solving... 
    Full time

    Pinnacle Group

    Sunnyvale, CA
    5 days ago
  •  ...doYou will work in cross-functional agile teams alongside data scientists, machine learning engineers, product managers, and industry experts to innovate...  ...the development of core data frameworks and technical platforms that power advanced analytics engagements across data... 
    Apprenticeship
    Work at office
    Local area
    Easy work
    Shift work

    McKinsey & Company

    San Jose, CA
    5 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Data Platform Engineer. Be the first to apply!