职位介绍
Role demands a highly skilled Data Engineer to design, build, and optimize scalable data pipelines and data platforms. The ideal candidate will have strong expertise in data modeling, cloud-based data architectures, and modern data engineering tools across Azure, Snowflake, and Databricks environments.
Key Responsibilities:
-
Data Engineering & Pipeline Development
-
Design, develop, and maintain robust ETL/ELT pipelines using Databricks, PySpark, and Azure Data Factory (ADF).
-
Build scalable and efficient data ingestion frameworks for structured and unstructured data.
-
Optimize pipeline performance through performance tuning and orchestration best practices.
-
Data Modeling & Management
-
Develop and maintain data models using modern tools (DBT preferred).
-
Implement Master Data Management (MDM) solutions to ensure data consistency and integrity.
-
Design scalable and efficient Snowflake schemas (star/snowflake schema, dimensional modeling).
-
Database & Query Optimization
-
Write and optimize advanced SQL queries across Snowflake, Azure SQL, and Synapse.
-
Develop and manage stored procedures and database objects.
-
Ensure efficient data retrieval through indexing, partitioning, and query optimization.
-
Cloud & Platform Integration
-
Work with Azure data services including:
-
Azure Data Factory (ADF)
- Azure Data Lake Storage (ADLS)
- Azure Synapse Analytics
- Azure SQL Database
-
Integrate and maintain Snowflake with Azure ecosystem.
-
Python Development
-
Develop data transformation and automation scripts using Python libraries:
-
pandas
- pyodbc
- SQLAlchemy
-
Build reusable components for data processing and validation.
-
Data Quality, Validation & Monitoring
-
Implement data validation rules, quality checks, and anomaly detection frameworks.
-
Perform root cause analysis for data inconsistencies.
-
Develop dashboards or tools for data quality monitoring.
-
Collaboration & DevOps
-
Use GitHub for version control, branching strategies, and code reviews.
-
Manage workload scheduling and dependency management for pipelines.
-
Collaborate with cross-functional teams including data analysts, data scientists, and business stakeholders.
-
Required Skills & Qualifications
-
Bachelor’s or Master’s degree in Computer Science, Information Systems, or related field.
-
Strong experience in data engineering and data platform development.
-
Technical Skills
-
Expertise in DBT (preferred) for data modeling.
-
Strong SQL skills with hands-on experience in:
-
Snowflake
- Azure SQL
- Stored procedures
-
Proficiency in Python for data engineering workflows.
-
Hands-on experience with:
-
Databricks & PySpark
- Azure Data Services (ADF, ADLS, Synapse)
-
Strong knowledge of Snowflake architecture and schema design.
-
Experience with data validation, quality frameworks, and analysis tools.
-
Familiarity with GitHub and CI/CD practices.
-
Preferred Qualifications
-
Experience implementing data governance and MDM solutions.
-
Knowledge of performance tuning in distributed processing systems.
-
Familiarity with workflow orchestration tools.
-
Experience in agile environments (SCRUM/Kanban).
Education: Master Of Engineering,Master Of Technology,Bachelor of Engineering,Bachelor Of Technology
- Preferred skills: Technology->Data Engineering->Databricks,Technology->Cloud Integration->Azure Data Factory (ADF),Technology->Big Data
- Data Processing->Py Spark,Technology->Oracle->PL/SQL,Technology->Data on Cloud-Data Store->Snowflake,Technology->Cloud Platform->Databases on Azure->Azure SQL Database,Technology->Cloud Platform->Azure Analytics Services->Azure Data Lake
关于Infosys
PUNE
总部位置
