Data / AI Native Engineer
Toronto, Ontario Canada (Remote)
Overall Experience: 7+ Years
Role Summary
We are seeking a hands-on Senior Data / AI Native Engineer with strong expertise in Databricks, AWS, PySpark, and modern Lakehouse architecture. The ideal candidate combines deep Data Engineering expertise with strong Software Engineering practices and has experience leading complex enterprise data initiatives.
This role requires daily hands-on use of GitHub Copilot and/or Claude Code CLI to accelerate software and data pipeline development. The candidate should be comfortable rapidly building POCs and MVPs and applying AI directly within data engineering workflows - not just using AI for application development.
Day to Day Job Duties:
Design, develop, and optimize high-volume enterprise data pipelines using PySpark, Spark, and Databricks
Build scalable Lakehouse solutions using Delta Lake and Medallion Architecture (Bronze/Silver/Gold)
Implement data governance, access controls, and cataloging using Databricks Unity Catalog
Design and develop cloud-native data solutions using AWS S3, EMR, Glue, Lambda, and Redshift
Lead complex Data Engineering initiatives and provide technical guidance to engineering teams
Develop batch and real-time data processing pipelines for large-scale enterprise datasets
Use GitHub Copilot and/or Claude Code CLI daily to accelerate coding, pipeline development, testing, troubleshooting, and documentation
Rapidly develop POCs and MVPs using AI-assisted engineering practices
Apply AI within data pipelines for schema inference, automated data-quality rule generation, anomaly detection, and PySpark transformation generation
Build AI-ready data platforms supporting RAG, embeddings, vector stores, and AI/ML applications
Develop streaming pipelines using Structured Streaming, Kafka, Kinesis, or Databricks Auto Loader
Implement modern data engineering patterns including CDC, SCD Type 2, schema evolution, idempotent processing, and data contracts
Implement data quality, validation, monitoring, lineage, and observability across data pipelines
Conduct code/design reviews and establish reusable Data Engineering patterns and standards
Collaborate with Data Architects, AI/ML Engineers, Software Engineers, and business stakeholders to deliver enterprise data products
Basic Qualifications Must Have:
7+ years of experience in Data Engineering and Software Engineering, building production-grade enterprise data solutions
4+ years of hands-on experience with Databricks, Delta Lake, Medallion Architecture, and PySpark/Spark
4+ years of experience with the AWS data ecosystem, including S3, EMR, Glue, Lambda, and/or Redshift
Proven experience leading complex Data Engineering initiatives or providing technical leadership to Data Engineering teams
Strong hands-on experience processing high-volume datasets using PySpark, rather than Scala-only Spark development
Demonstrated daily use of GitHub Copilot and/or Claude Code CLI, with the ability to explain specific examples of how these tools improve Data Engineering productivity
Proven experience rapidly developing POCs and MVPs using AI-assisted development practices
Technical Skills:
Data Platform: Databricks, Delta Lake, Unity Catalog
Data Processing: PySpark, Apache Spark
Cloud: AWS S3, EMR, Glue, Lambda, Redshift
Architecture: Lakehouse, Medallion Architecture, Data Lakes
Streaming: Spark Structured Streaming, Kafka, Kinesis, Auto Loader
eye