Apache Spark Developer
Bright Vision Technologies
Role Description
We are seeking an experienced Apache Spark Developer to design, develop, and optimize large-scale distributed data processing applications supporting enterprise analytics, machine learning, real-time reporting, and cloud-based data platforms. This role focuses on building high-performance Spark applications capable of processing billions of records across structured and semi-structured data sources while delivering scalable, reliable, and cost-efficient data pipelines.
You will work closely with data architects, data engineers, cloud platform teams, machine learning engineers, and business intelligence developers to build modern data processing solutions leveraging Apache Spark, cloud-native technologies, and distributed computing frameworks. The ideal candidate possesses deep expertise in Spark architecture, distributed systems, performance optimization, and cloud-based big data ecosystems.
Key Responsibilities
- Design, develop, and maintain high-performance distributed data processing applications using Apache Spark.
- Build scalable batch and real-time ETL/ELT pipelines processing large volumes of enterprise data.
- Develop Spark applications using PySpark, Scala, or Spark SQL for data transformation, aggregation, and analytics.
- Optimize Spark jobs for memory utilization, partitioning strategies, shuffle performance, and execution efficiency.
- Process structured, semi-structured, and streaming data from enterprise databases, APIs, Kafka, cloud storage, and data lakes.
- Develop reusable Spark libraries, data processing frameworks, and metadata-driven ingestion pipelines.
- Collaborate with cloud engineering teams to deploy Spark workloads on Databricks, EMR, Azure Synapse, or Kubernetes.
- Implement data quality validation, reconciliation, monitoring, and automated error handling across distributed pipelines.
- Integrate Spark applications with enterprise data warehouses, lakehouses, and reporting platforms.
- Participate in architecture reviews, code reviews, technical design discussions, and Agile development activities.
- Troubleshoot production issues involving distributed processing, cluster performance, resource utilization, and data quality.
- Support cloud migration initiatives by modernizing legacy ETL workloads into Spark-based architectures.
Qualifications
- Six or more years of professional software or data engineering experience.
- Four or more years of hands-on Apache Spark development experience in enterprise production environments.
- Strong proficiency in PySpark, Scala, or Spark SQL for distributed data processing.
- Deep understanding of Apache Spark architecture including RDDs, DataFrames, Datasets, Catalyst Optimizer, DAG execution, and Tungsten engine.
- Strong experience with distributed computing concepts including partitioning, shuffling, caching, broadcast joins, and fault tolerance.
- Advanced SQL skills with databases such as SQL Server, Oracle, PostgreSQL, Snowflake, or Teradata.
- Experience working with Hadoop ecosystem technologies including Hive, HDFS, YARN, and Parquet.
- Experience processing streaming data using Spark Structured Streaming, Apache Kafka, or Event Hubs.
- Hands-on experience with cloud platforms including Azure Databricks, AWS EMR, AWS Glue, Azure Synapse Analytics, or Google Dataproc.
- Experience integrating Spark applications with Delta Lake, Apache Iceberg, or Apache Hudi.
- Strong understanding of data warehousing concepts, dimensional modeling, and data lake architecture.
- Experience using Git, CI/CD pipelines, Azure DevOps, GitHub Actions, or Jenkins.
- Strong debugging, troubleshooting, and Spark performance tuning skills.
- Experience working in Agile Scrum development environments.
Preferred Qualifications
- Experience building enterprise Lakehouse architectures using Databricks or Delta Lake.
- Familiarity with Apache Airflow, Azure Data Factory, AWS Step Functions, or Control-M for workflow orchestration.
- Experience with machine learning workflows using Spark MLlib, MLflow, or feature engineering pipelines.
- Knowledge of Kubernetes, Docker, and containerized Spark deployments.
- Experience implementing Data Quality frameworks using Great Expectations or Deequ.
- Familiarity with Apache NiFi, Apache Flink, Trino, or Presto.
- Experience working with cloud object storage including Amazon S3, Azure Data Lake Storage (ADLS Gen2), or Google Cloud Storage.
- Knowledge of Infrastructure as Code using Terraform or ARM templates.
- Experience with enterprise monitoring tools including Prometheus, Grafana, Datadog, or OpenTelemetry.
- Cloud certifications in Azure, AWS, Databricks, or Apache Spark-related technologies are highly desirable.
Project Environment
You will be joining a modern data engineering team responsible for building cloud-native big data platforms supporting enterprise analytics, AI, and business intelligence initiatives. Current projects include:
- Enterprise data lakehouse implementation using Databricks and Delta Lake
- Real-time streaming analytics processing billions of daily events
- Large-scale customer analytics and behavioral data platforms
- Financial risk modeling and fraud detection pipelines
- Healthcare clinical and operational analytics solutions
- Cloud migration of legacy Hadoop and ETL workloads
- Machine learning feature engineering and model training pipelines
- Enterprise reporting platforms supporting executive dashboards and self-service analytics
- Distributed data processing infrastructure deployed on Azure and AWS
How to Apply
Would you like to know more about this opportunity? For immediate consideration, please send your resume to [email protected] or contact us at View phone number on us.fitly.work. Learn more about Bright Vision Technologies at .
$181.2k - $317.1k
...solutions . Strong expertise in designing and operating data ingestion pipelines using: Apache Iceberg (tables, catalogs, schema evolution, metadata management) Apache Spark (batch & streaming jobs, optimization, partitioning) Expertise in data formats such as...SuggestedPermanent employmentFull timeWork experience placementWork at officeImmediate startRemote workFlexible hours2 days per week$130k - $270k
...team of engineers take pride in what they develop and constantly innovate to provide the... ...workflows and automation pipelines using Apache Airflow. This role focuses on building reliable... ...Data processing engines including Apache Spark Experience with containerization...SuggestedHourly payFull timeTemporary work$184k - $230k
...Engineer with deep expertise in distributed systems to join the Apache Spark Team. You will be at the forefront of innovation, building our... ...in the open-source community. Build with Modern Stacks: Develop high-performance features using Scala, Java, and Python on modern...SuggestedRemote workWork from homeFlexible hours$130k - $270k
...timely manner. Our team of engineers take pride in what they develop and constantly innovate to provide the best solution. Captivation... ...with Distributed Big Data processing engines including Apache Spark Experience using Jupyter Notebook Experience with data wrangling...SuggestedHourly payFull timeTemporary work- About the TeamThe Spark Platform team owns and operates DoorDash's Apache Spark ecosystem — the execution runtime, remote shuffle service, cluster scheduler, and reliability tooling that powers the company's data, analytics, and ML workloads. We run Spark across the company...SuggestedHourly payWork at officeLocal areaRemote workRelocationFlexible hours
- About the TeamThe Spark Platform team owns and operates DoorDash's Apache Spark ecosystem — the execution runtime, remote shuffle service, cluster scheduler, and reliability tooling that powers the company's data, analytics, and ML workloads. We run Spark across the company...Hourly payWork at officeLocal areaRemote workRelocationFlexible hours
- ...build and maintain reusable frameworks, libraries and internal developer tooling to standardise and accelerate data pipeline... ...Demonstrated experience with enterprise data pipeline tooling (Apache Spark, dbt, Airflow), providing guidance on best practices and helping...Permanent employmentFull timePart timeWork at officeFlexible hours
- ...solutions using technologies such as Kafka, Spark, Elasticsearch, and cloud-based platforms... ...| AWS | Linux What You’ll Do Design, develop, and maintain enterprise data ingestion,... ...solutions. Experience building and supporting Apache Kafka producers, consumers, and streaming...Full timeRemote workFlexible hours
- ...your time and have a great week ahead!!! Role: Senior Java/Spark Developer - Remote (Should be inside US to apply for this role) Need... ...and data processing solutions using Java, Kotlin, Scala, and Apache Spark. ~Design and implement data loading and transformation...Full timeRemote work
- Role Description Como Senior Backend Spark Developer, serás una pieza clave en el diseño, desarrollo y optimización de soluciones de procesamiento... ...pipelines de Big Data robustos y escalables utilizando Apache Spark. ~Optimizar el rendimiento de las soluciones de...Full time
$135k - $150k
...AWS. This role centers on hardening and operating Amazon EMR (Spark) and OpenSearch workloads, building secure CI/CD pipelines, managing... ...Git Data & Development Foundation ~ Working knowledge of Apache Spark core concepts: RDDs, DataFrames, Spark SQL ~...Full time- About the job Tech Lead (Spark and Python expertise) Position: Tech Lead (Spark and Python expertise) Location... ...working in heavy data background needed. Develop, program, and maintain applications using the Apache Spark open-source framework. Work with different...Remote work
- ...ETL Engineer / Java Spark/ Ab Initio Wilmington DE (3 days WFO, 2 days WFH) Look for candidate who can be onsite from day 1 Abinitio - is needed (if not abinitio then any other ETL tools such as Informatica, Talend etc. ) Spark AWS...Work from homeFlexible hours
$96.49k - $144.74k
...seeking a Data Engineer - Data Platform (Spark/Kafka/Flink/Scala/Java) - Onsite Hybrid... ...at enterprise scale. You will design and develop streaming and batch frameworks, Kafka/Flink... ...real-time and batch data pipelines using Apache Spark, Apache Flink, Apache Kafka, and...Temporary workWork experience placementWork at officeRemote workFlexible hours- ...Job Description The Senior Front-End Developer will be part of a team supporting established projects and creating products from the... ...process, such as Jira. - Familiarity with web servers such as Apache, Nginx, etc. Bonus/Nice To Have - Interest in design and...Full time
$100k - $150k
...Apache Kafka Developer - Remote Bright Vision Technologies is a technology consulting and software development company delivering cloud... ...). Familiarity with stream processing frameworks (Flink, Spark Streaming). Experience with data governance and lineage...Full timeH1bLocal areaImmediate startRemote workVisa sponsorship- ...ad platform teams to drive advertisers' value. The role requires 6+ years building production data or backend systems, strong Scala or JVM background, and experience with Spark, Delta Lake, Parquet, or Iceberg on Databricks. Remote options available. #J-18808-Ljbffr FOXRemote job
- ...prior to applying Description: Looking for a full stack developer supporting a backend development team focused on middleware and... ...in a big data environment using tools such as Hadoop, Pyspark, Spark, and Hbase (prefer at least two of the list) 3. Minimum of 5 years...Full time
- ...Migration, Continuous Delivery, Big Data, Apache, Web Logic, Jenkins) in Reston, VA... ...NoSQL, HADOOP, NOSQL, Apache Kafka, Apache Spark, WebLogic, Oracle, Linux, Unix, Windows... ...will take ownership of conceptualizing, developing, standardizing and driving the adoption of...Permanent employmentFull timeRemote work
- ...Description Job Title: Technical Lead – Data Engineering (Palantir, Spark, PySpark, Python) Please dont apply if you dont have PALANTIR... ...~Strong hands-on experience with: ~ Python & PySpark ~ Apache Spark ~Advanced SQL ~Experience with Palantir Foundry or...Work from homeFlexible hours
- ...Job Description Hi, Role : Sr. Apache Druid Administrator Location: Irving, TX Duration: 12 Months Contract MOI... ...Configure and maintain Druid metadata stores using MySQL. Develop operational automation using Ansible and Infrastructure-as-Code...Full timeContract work
- ...We are hiring a Sr. Software Engineer (Apache, Python, Trino, Kubernetes) to work in... ...Description: The Software Engineer develops, maintains, and enhances complex and diverse... ...degree. Apache AirFlow Python Apache Spark Trino Kubernetes Themis Insight...Full timeLocal areaWork from homeFlexible hours
$111k - $114k
...Innovations is seeking Mid-Level Software Developers to provide remote support for a federal... ...cloud platforms.Hands-on experience with Apache Kafka for event-driven architectures and... ...Experience with distributed data systems such as Spark, HBase, or Solr.Quality &...Work experience placementLocal areaRemote work$180.5k - $225.6k
...production scale. This greenfield provisioning layer will power all non-Spark compute workloads on Serverless (Notebooks, AI Agents, Remote... ...globe and was founded by the original creators of Lakehouse, Apache Spark, Delta Lake and MLflow. To learn more, follow Databricks...Local areaRemote workWorldwide$85 - $90 per hour
...Experience working on Microservices.• Good experience on AWS Cloud, MSK, Kinesis, Lamda• Good hands-on experience on AWS Glue using Apache Spark, Amazon EMR• Good experience on container orchestration system like Kubernetes, ECS• Good working knowledge in SPRING Framework...Hourly payFull timeRemote work- ....Research industry best practices, evaluate new technologies, develop standards and engineering best practices and recommend innovative... ...Azure/API GatewaysExperience with data processing technology (Apache Spark etc.)Experience with data virtualization technology (Tibco DV,...Work at officeRemote workFlexible hours
$158k - $197k
...datastores like DyanmoDB, Postgres, MySQL, etc.Distributed systems design & distributed processing through technologies such as Apache Spark or Databricks / Snowflake is a plusKnowledge on messaging technologies like Kafka / AWS Kinesis or similar is a plusOur Values Act...H1bRemote workWorldwideVisa sponsorshipWork visa- ...remote option.)Job SummaryThis Engineer 2 is responsible for developing, enhancing, and supporting software applications and data platforms... ...; hands on knowledge of Scala required.Experience with Apache Spark or other distributed data processing technologies.Experience...Full timeWork at officeRemote workWorldwideNight shiftWeekend work
$300k - $425k
...helping you.What you’ll be doingDesign and Development:Architect, develop, and maintain scalable backend systems and APIs using Java and... ....Big Data Expertise:Leverage big data technologies such as Apache Spark, Kafka, Flink, and related tools to build high-performance...Work at officeLocal areaRemote workMonday to ThursdayFlexible hours$94k - $120k
...frameworks, open-source libraries, and APIs to develop basic application solutions.Learn and... ...technologies (e.g. MS Cosmos DB, Apache Cassandra, Amazon DynamoDB)Understanding... ...distributed computing (MS HPC, Sagemaker, Spark)Two years of experience with integration...Full timeContract workWork at officeLocal areaRemote workWorldwideWork visaRelocation packageFlexible hours3 days per week
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Apache Spark Developer. Be the first to apply!


