SDET – PySpark Automation Engineer
Location: Jersey City, NJ
Experience: 8+ Years
Employment Type: W2 Contract
Interview: Onsite Interview Required
Job Summary
We are seeking an experienced SDET – PySpark Automation Engineer with 8+ years of experience in software quality engineering, test automation, data validation, and automation framework development. The ideal candidate will have strong hands-on experience building and maintaining automation frameworks using Python and PySpark, with a focus on testing large-scale data processing, ETL/ELT pipelines, data transformations, and data quality. Strong experience with Databricks and Azure Cloud is highly desirable. Candidates should have hands-on engineering experience and be comfortable developing reusable automation frameworks rather than simply executing existing test scripts.
Key Responsibilities
- Design, develop, and maintain Python/PySpark-based automation frameworks for data and ETL testing.
- Build reusable automation utilities, libraries, test components, and validation frameworks.
- Develop automated tests for PySpark transformations, joins, aggregations, data processing logic, and business rules.
- Validate large-scale data pipelines and perform source-to-target data reconciliation.
- Perform functional, integration, regression, data quality, and end-to-end testing.
- Develop automated validation checks for data accuracy, completeness, consistency, and integrity.
- Work with Data Engineers and Developers to identify and troubleshoot data and pipeline issues.
- Integrate automated testing frameworks into CI/CD pipelines.
- Develop framework-level logging, reporting, error handling, and failure diagnostics.
- Participate in code reviews and follow software engineering and test automation best practices.
- Continuously improve automation coverage, framework scalability, performance, and maintainability.
Required Skills
- 8+ years of experience in SDET, QA Automation, Software Engineering, or related roles.
- Strong hands-on experience with Python and PySpark.
- Proven experience building automation frameworks from the ground up.
- Strong understanding of PySpark DataFrames, transformations, actions, joins, aggregations, and partitioning.
- Experience testing ETL/ELT and large-scale data pipelines.
- Strong SQL skills for backend and data validation.
- Experience with functional, integration, regression, and data testing.
- Experience developing reusable automation libraries and test utilities.
- Experience with Git and CI/CD practices.
- Strong debugging and problem-solving skills.
Highly Desirable Skills
- Hands-on experience with Databricks.
- Experience with Azure Cloud and Azure data services.
- Experience with Azure Databricks.
- Experience with Delta Lake and Lakehouse architecture.
- Experience with Azure Data Factory (ADF), ADLS Gen2, or other Azure data services.
- Experience implementing and executing PySpark automation in cloud environments.
- Experience with Databricks notebooks, jobs, workflows, and clusters.
Additional Preferred Skills
- Experience with PyTest, Robot Framework, or similar Python automation frameworks.
- Knowledge of Jenkins, GitHub Actions, Azure DevOps, or similar CI/CD tools.
- Experience with REST API testing and automation.
- Experience with Kafka or other distributed data technologies.
- Knowledge of Docker and Kubernetes.
- Financial Services or Capital Markets experience is a plus.
Ideal Candidate
The ideal candidate is a hands-on SDET with strong Python and PySpark expertise who has successfully built automation frameworks for large-scale data pipelines. Candidates with Databricks, Azure Cloud, and Azure Databricks experience will be highly preferred.
#LI-HK1