The Data & Integration Engineer will design and implement data flows, APIs, and pipelines to support GenAI initiatives. The role involves working with enterprise data platforms like Informatica and Cloudera to ensure reliable data movement and preparation for AI workflows.
We are looking for a
Data & Integration Engineer who operates effectively at the intersection of
system analysis, enterprise integration and data engineering
to
support GenAI initiatives
.
The successful candidate will:
Understand business and functional requirements
Translate them into data flows and integration designs
Work across upstream and downstream systems
Ensure reliable movement, transformation and availability of data for GenAI use cases
Develop scripts / programs to get data from different integration systems
Develop APIs to integrated with relevant systems
Analyse business/technical requirements and translate them into data flows and integration designs
Work with upstream and downstream teams to define data contracts and interfaces
Support data ingestion and preparation for GenAI use cases
Coordinate integrations across systems in the DataLake ecosystem (Informatica, Cloudera, etc.)
Design and implement data movement across systems using:
APIs
SFTP and file based transfers
Batch pipelines
Requirements
Key Requirements
Below are the key skillsets that will be required for all relevant tasks mentioned:
Good years of experience in system analysis, integration engineering, data engineering or technical delivery roles
Strong ability to translate requirements into system flows, data flows, interface specifications and implementation plans
Experience working with upstream and downstream teams to define and deliver enterprise integrations
Practical experience with REST APIs, SFTP, batch processing, file based integration and data pipeline orchestration
Good understanding of data mapping, transformation, aggregation, reconciliation and data quality controls
Good SQL skills and basic to moderate Python skills for data handling, scripting, automation and troubleshooting
Exposure to Java
Exposure to Informatica, Cloudera or similar enterprise data platforms
Working knowledge of Git, branching, pull requests, code reviews and controlled release practices
Familiarity with CI/CD, Jira, Confluence and enterprise deployment processes
Experience with Control M or equivalent scheduling tools
Familiarity with logging (OTEL) and monitoring tools such as Splunk Elastic Stack
Exposure to GenAI concepts such as document ingestion, RAG, embeddings and data preparation for AI workflows
Working experience with Informatica is preferable
Strong communication skills, with the ability to challenge weak designs and coordinate across business, application, data, infrastructure and security teams.
Data Engineering,
System Integrations,
Python, SQL, Informatica