Infosys
Infosys

Data Engineer - DaAI

职能数据工程
级别中级
地点Bangalore, India
方式现场办公
类型Senior Consultant
发布2个月前
立即申请

职位介绍

As a Data Engineer, you will help build the data foundation for our agentic AI platform. You will work with senior data architects, AI/ML engineers, and platform engineers to implement data ingestion, transformation, profiling, enrichment, validation, and preparation pipelines across structured and unstructured enterprise data sources.
This is a hands-on engineering role for someone who enjoys working with real-world enterprise data, building reliable pipelines, writing robust Python and SQL, and helping convert raw enterprise information into AI-ready data assets.

  • Build and maintain data ingestion pipelines for structured enterprise systems such as ERP, CRM, billing, finance, HR, OSS/BSS, Service Now, Salesforce, SAP, Oracle, databases, and APIs.
  • Build pipelines for unstructured and semi-structured data sources such as documents, emails, logs, transcripts, PDFs, spreadsheets, and media metadata.
  • Develop ETL/ELT workflows using Python, SQL, Py Spark, Apache Spark, Airflow, dbt, Dagster, cloud-native services, or equivalent technologies.
  • Support data profiling routines to identify missing values, duplicates, inconsistent formats, incomplete master data, schema changes, and conflicting records.
  • Implement data quality checks using frameworks such as Great Expectations, dbt tests, AWS Glue Data Brew, custom validation scripts, or equivalent tools.
  • Support data labelling, contextualization, harmonization, enrichment, and classification workflows required for AI agent configuration.
  • Prepare data outputs for downstream AI consumption, including embeddings, metadata, semantic tags, graph-ready datasets, and retrieval-ready document chunks.
  • Working knowledge of data pipeline development using Py Spark, Apache Spark, Airflow, dbt, Dagster, or equivalent technologies.
  • Experience working with structured data from databases, APIs, enterprise applications, data lakes, warehouses, or lakehouse platforms.
  • Exposure to cloud data platforms such as Databricks, Snowflake, Big Query, Azure Data Lake, AWS S3, Google Cloud Storage, or equivalent platforms.
  • Understanding of data modelling, schema design, joins, keys, relationships, data validation, and data quality concepts.
  • Practical experience with data profiling, cleansing, transformation, and reconciliation.
  • Familiarity with Git, CI/CD basics, unit testing, and production-grade engineering practices.

Education: Bachelor of Engineering

  • Preferred skills: Technology->Big Data
  • Data Processing->Py Spark,Technology->Big Data
  • Data Processing->Spark->Apache Storm

福利待遇

Learning Budget

必备技能

Python

SQL

PySpark

Apache Spark

Airflow

dbt

Data quality

Data modeling

关于Infosys

BANGALORE

总部位置