Data Engineer (Expert)
Imizizi
Reference: JHB001567-KF-1
ESSENTIAL SKILLS
- Hands-on experience with Kafka and event streaming platforms for real-time data movement.
- Proven experience with API integration patterns, webhooks and event/webhook ingestion.
- Strong proficiency in Python for data engineering, ingestion pipelines and automation.
- Solid competence with enterprise databases and query languages, with practical experience with database performance tuning and query optimisation for OLTP/operational workloads.
- Experience in data modelling to design schemas and standardized data representations.
- Experience with schema registries and contract-first designs (Avro, Protobuf) to manage producer/consumer compatibility.
- Strong understanding and practice of data quality techniques and tooling to ensure trusted data.
- Knowledge of metadata management and cataloguing to support discoverability and lineage.
- Familiarity with ETL/ELT patterns and best practices for performant, reliable data pipelines.
- Ability to normalise and correlate data across multiple sources (CMDB, identity, security tooling).
- Design and operate idempotent pipelines with retry, replay and compensation mechanisms.
ADVANTAGEOUS SKILLS
- Experience with Java for stream processing or connector development.
- Awareness of frontend frameworks (e.g., Angular) to better understand downstream consumers.
- Familiarity with infrastructure automation and IaC (e.g., Terraform) to deploy integration components.
- Experience operating container platforms and orchestration (Kubernetes) for scalable stream processing.
- Experience with NoSQL/document stores such as MongoDB for operational data needs.
- Familiarity with enterprise systems like SAP and working with their integration interfaces.
- Experience with big data ecosystems (CDH/Hadoop) and distributed storage/processing.
- Working knowledge of Azure Synapse or similar data platform services for integrated data processing.
- Observability for streaming: experience with metrics, tracing and logging for pipelines (Prometheus, Grafana, OpenTelemetry).
- Experience with connector ecosystems and managed services (Kafka Connect, Confluent Cloud, managed Kafka).
- Knowledge of message delivery semantics, partitioning strategies and capacity planning for high-throughput pipelines.
ROLE & RESPONSIBILITIES
- Design, build and operate operational data integration solutions focused on making enterprise data available, trusted and consumable.
- Ingest and stream data from Zero Trust data sources (CMDB, identity systems, security tooling) into operational pipelines rather than analytics stores.
- Build and maintain Kafka topics, producers/consumers, connectors and streaming applications to enable realtime flows.
- Implement API-based integrations, webhook listeners and file ingestion solutions to capture operational events.
- Develop robust ETL/ELT and stream processing logic to normalise, correlate and enrich data across sources.
- Ensure pipelines are idempotent and resilient (retries, replay, backpressure and compensation strategies).
- Apply data modelling and metadata practices to ensure consistent schemas and discoverability.
- Implement and monitor data quality checks, anomaly detection and validation rules; drive remediation where needed.
- Maintain data lineage, access control and security practices to meet governance and compliance requirements.
- Collaborate with automation and orchestration teams to ensure streamed operational data feeds automation workflows and runbooks.
- Own pipeline alerting, runbooks, and participate in incident response/on-call rotations
- Troubleshoot production issues, perform root cause analysis and implement long-term reliability improvements.
- Build testable pipelines with unit/integration tests, contract tests and CI/CD deployment pipelines for streaming workloads
- Data Quality Gates: enforce automated quality checks early in pipelines (validation, schema checks, anomaly detection) and block or quarantine bad data.
- Performance & Cost Awareness: implement efficient partitioning, batching and resource usage patterns while tracking cost and throughput trade-offs.
- Security & Privacy by Design: apply least-privilege, secrets management, encryption in transit and at rest, and data masking/anonymisation where required.
- Drive architectural decisions and define integration best practices
QUALIFICATIONS/EXPERIENCE
- Extensive hands-on experience (typically 6+ years) in data engineering, integration or streaming roles with demonstrable production experience.
- Proven track record building and operating streaming platforms (Kafka) and API-based integrations, with strong Python/Java and enterprise databases and query languages skills.
- Strong analytical thinking, curiosity about data, attention to detail, structured problem solving and ownership - able to figure things out and drive topics to completion.
Submit your CV to: ***email_hidden*** and Subject line Role title