Project Description
The Lead Data Platform Engineer plays a crucial role in developing and enhancing a cloud-native Data Platform for a large-scale e-commerce marketplace. The primary objective is to improve developer experience and streamline data processes through automation and standardization, ultimately enabling data-driven decisions and supporting rapid product growth.
Responsibilities
- Design and develop tools that standardize, automate, and simplify Data Platform operations.
- Build and maintain internal CLI tools for common platform and Data Product management tasks.
- Develop and evolve Data Product templates based on modern data engineering practices.
- Design and implement visualizations and operational insights within Backstage.
- Collect and analyze platform usage metrics, audit data, and adoption statistics.
- Design and implement Agentic AI workflows and developer-assistance capabilities.
- Develop reusable AI skills and automation components that standardize work with data products and data assets.
- Create mechanisms for distribution, lifecycle management, and monitoring of AI skills.
- Maintain and enhance a shared Apache Airflow platform based on Cloud Composer.
- Build reusable libraries, operators, and common components for Airflow DAG development.
- Develop and maintain infrastructure-as-code assets and Terraform modules.
- Contribute to GitOps adoption initiatives using tools such as Backstage, GitHub, ArgoCD, and Crossplane.
- Collaborate with platform, data engineering, analytics, and machine learning teams to improve platform usability and engineering efficiency.
Requirements
- Strong commercial experience with Python development.
- Hands-on experience with Google Cloud Platform (GCP).
- Practical experience with BigQuery and cloud-based data platforms.
- Experience with Apache Airflow, preferably Cloud Composer.
- Experience building platform engineering, developer tooling, or internal self-service solutions.
- Knowledge of Infrastructure as Code practices and Terraform.
- Experience with GitOps principles and modern software delivery practices.
- Good understanding of data engineering concepts and Data Product lifecycle management.
- Experience with PySpark and/or Apache Spark ecosystems.
- Familiarity with FastAPI, Pydantic, and modern Python tooling.
- Knowledge of CI/CD processes and source control best practices.
- Experience working in Agile development environments.
- Strong problem-solving skills and ability to work independently.
- Effective communication skills and ability to collaborate with cross-functional teams.
- Professional proficiency in Polish and English.
Nice to Have
- Experience with Vertex AI or other GenAI platforms.
- Hands-on experience with GitHub Copilot, Copilot extensions, plugins, or AI-assisted development solutions.
- Experience developing Agentic AI workflows or AI automation capabilities.
- Knowledge of Backstage plugin development.
- Experience with Dataproc.
- Familiarity with dbt and Kedro.
- Experience with ArgoCD and Crossplane.
- Experience building observability, telemetry, or platform analytics solutions.
- Experience working in large-scale data environments supporting analytics and machine learning workloads.