Data Engineer - ETL/PySpark (Banking Domain)
- Posted
- 5d ago
Role Summary We are looking for a hands-on Data Engineer with strong ETL and PySpark expertise to design, build, and support data pipelines and data marts within a banking environment. The ideal candidate will own the full SDLC lifecycle — from build through UAT, production deployment, and post-production support — while working across structured, semi-structured, and unstructured data. Key Responsibilities Design, develop, and maintain ETL pipelines and data marts using PySpark and Python Write clean, maintainable, and production-grade Python code following software engineering best practices Own end-to-end SDLC activities: build, UAT support, UAT bug fixes, production deployment, and post-production support Perform data analysis and debugging using Oracle SQL and PySpark Work across structured, semi-structured, and unstructured data sources Build and maintain data warehousing solutions supporting banking/financial reporting needs Debug and optimize PySpark jobs for performance and reliability Collaborate with cross-functional teams (QA, DBAs, business analysts) through the release cycle Participate in CI/CD pipeline processes, including testing and validation of data pipelines Ensure data pipeline reliability, scalability, and adherence to banking data governance/compliance standards Required Skills & Experience 5+ years of commercial experience in a data-driven engineering role Hands-on experience building data marts and ETL pipelines Expert-level PySpark and Python for ETL scripting Strong command of Oracle SQL for data analysis and debugging Proven experience across the full SDLC — build, UAT, bug fixing, deployment, post-prod support Strong understanding of software engineering concepts and best practices for production pipelines Experience working with structured, semi-structured, and unstructured data Prior experience with banking clients or strong banking domain knowledge Strong data warehousing fundamentals Tech Stack (Daily Use) Languages: Python Big Data: Spark / PySpark, Hadoop, MapReduce, Hive Data Libraries: Pandas Databases: SQL and NoSQL DBMS Tools: Jupyter Practices: CI/CD, data testing & validation Nice to Have (optional — add if applicable) Cloud experience (AWS/Azure/GCP) — not mentioned in your input, confirm with client Airflow or other orchestration tools Experience with regulatory/compliance reporting in banking