
Pyspark databricks
About the role
Design, develop, and maintain ETL/ELT pipelines using Py Spark and Databricks.
Build scalable data processing solutions for large datasets.
Develop and optimize Spark jobs for performance and efficiency.
Implement data ingestion pipelines from various structured and unstructured data sources.
Work with Delta Lake, Data Lake, and cloud storage solutions.
Collaborate with business stakeholders, data analysts, and architects to understand data requirements.
Ensure data quality, governance, and security standards are followed.
Troubleshoot performance issues and optimize data workflows.
Participate in code reviews and follow DevOps best practices.
Required Skills:
Technical Skills:
Strong experience with Py Spark and Apache Spark.
Hands-on experience with Databricks.
Proficiency in Python programming.
Experience in building ETL pipelines and data transformations.
Strong SQL knowledge and query optimization skills.
Experience with Azure Data Factory (ADF) / AWS Glue or similar tools.
Knowledge of Delta Lake, Medallion Architecture, and Data Lakes.
Experience with Git, CI/CD pipelines, and Agile methodologies.
Education: MCA,MSc,MTech,BCA,BSc,BTech
- Preferred skills: Technology->Data Engineering->Databricks,Technology->Big Data
- Data Processing->Py Spark
About Infosys
BANGALORE
Headquarters