Location: Los Angeles area (4 days a week)
W2 Candidates only
Long-term contract
Job Description
We are seeking a Data Engineer with 5+ years of experience in data engineering, ETL, data warehousing, and database design. The ideal candidate will have strong hands-on experience with AWS data services, PySpark/Spark, Python, Airflow, and Redshift, with a focus on designing, developing, optimizing, and supporting scalable data pipelines.
Key Responsibilities
Data Integration & Pipeline Development
- Design, develop, and maintain data integration workflows using AWS Glue, EMR, MWAA/Airflow, Lambda, and Redshift.
- Build scalable ETL/ELT pipelines for processing large datasets.
- Use Python, PySpark, and Apache Spark for data transformation and processing.
- Ensure accurate and efficient extraction, transformation, and loading of data into target systems.
- Design and support data pipelines throughout the development and production lifecycle.
Data Quality & Integrity
- Validate, cleanse, and transform data to maintain high levels of data quality.
- Implement monitoring, validation, error handling, and recovery mechanisms within data pipelines.
- Identify and resolve data quality and integration issues.
Performance & Cloud Optimization
- Optimize data workflows for performance, scalability, reliability, SLAs, and cost efficiency within AWS.
- Identify and resolve pipeline and processing bottlenecks.
- Tune SQL queries and optimize Amazon Redshift performance.
- Continuously review and improve existing data integration processes.
Business Intelligence & Analytics
- Translate business requirements into technical specifications and data pipelines.
- Ensure timely availability of integrated data for analytics and reporting.
- Collaborate with data analysts, business stakeholders, and technical teams to understand and deliver data requirements.
Documentation & Compliance
- Document data pipelines, workflows, architecture, technical specifications, and system processes.
- Follow data governance, security, compliance, and regulatory requirements.
Required Qualifications
- 5+ years of experience in data engineering, database design, ETL, and data warehousing.
- 3+ years of experience with AWS data services, including:
- AWS S3
- AWS EMR
- AWS Glue
- AWS Athena
- Amazon Redshift
- Amazon RDS
- Redshift Spectrum
- AWS MWAA / Airflow
- 2+ years of experience with CI/CD tools and practices.
- Strong knowledge of data storage, data lakes, databases, and distributed data-processing frameworks.
- Hands-on experience with Apache Spark, PySpark, or Hadoop.
- 3+ years of programming experience with Python, Java, or Scala.
- Experience designing, developing, and supporting production data pipelines.
Primary Skills
AWS EMR | AWS Glue | Airflow/MWAA | Apache Iceberg | Amazon Redshift | Amazon RDS | PySpark | Python | CI/CD
Nice to Have
- Informatica Cloud / IDMC experience.
- Agentic AI / Amazon Kiro experience.
- Experience with Apache Iceberg and modern data lake architectures.
#LI-JS2