Title: Senior Platform Data Engineer
Lisbon, PT
At Chain IQ, your ideas move fast.
Chain IQ is a global AI-driven Procurement Service Partner, headquartered in Baar, Switzerland, with operations across main centers and 16 offices worldwide. We provide tailored, end-to-end procurement solutions that enable transformation, drive scalability, and deliver substantial reductions in our clients' indirect spend. Our culture is built on innovation, entrepreneurship, ownership, and impact. Here, your voice matters - bold thinking is encouraged, and action follows ambition.
We are building a modular, enterprise-grade data platform with Databricks at its core. The platform is designed for flexibility at every layer: ingestion (batch connectors, APIs, event-driven streams), strong governance and privacy controls by default, and a data product mindset that lets consumers access trusted data through whichever channel suits them - analytics, APIs, streaming, or knowledge graphs - all reconciled through a shared canonical data model.
We're looking for a Senior Platform Data Engineer who can operate across this full stack: someone equally comfortable designing a governance model, and shaping how a canonical model surfaces as a well-documented data product. This is a hands-on, high-ownership role for someone who wants to build platform capability, not just pipelines.
We're looking for a Senior Platform Data Engineer who can operate across this full stack: someone equally comfortable designing a governance model, and shaping how a canonical model surfaces as a well-documented data product. This is a hands-on, high-ownership role for someone who wants to build platform capability, not just pipelines.
Responsibilities
- Design and build modular ingestion patterns supporting batch connectors (JDBC, SaaS/ERP connectors, file drops) and API-based ingestion as the primary focus, with the ability to extend to event-driven/streaming sources (e.g., Kafka, CDC feeds) where a use case calls for it.
- Establish reusable ingestion frameworks (metadata-driven where possible) so new sources can be onboarded quickly and consistently.
- Build and optimize pipelines using Databricks (Spark, Delta Live Tables / Lakeflow) across the medallion (bronze/silver/gold) architecture.
- Implement and evolve data governance: access control, lineage, cataloguing, auditing, and data classification.
- Design and implement data anonymisation/pseudonymisation approaches (masking, tokenisation, differential privacy where appropriate) to meet regulatory and internal privacy requirements (e.g., GDPR).
- Define and enforce data quality frameworks (expectations, contracts, monitoring, alerting) across ingestion and transformation layers.
- Design, build, and own data products against a well-defined canonical/conformed data model, with clear ownership, SLAs, and documentation.
- Work with domain teams to map source data into the canonical model without losing domain-specific nuance.
- Champion a "data as a product" approach — discoverability, self-service, versioning, and consumer-facing contracts.
- Build and maintain varied output/serving layers from the platform: BI/analytics datasets and APIs (REST/GraphQL) for operational consumption as the primary channels, with lighter-touch support for occasional streaming outputs or knowledge graph representations where a specific use case justifies them.
- Ensure consistent semantics across all output channels by anchoring them to the canonical model.
- Set technical direction and standards for the data platform (CI/CD, IaC, testing, observability, cost management).
- Mentor mid-level and junior engineers; review designs and code; contribute to architecture decision records.
- Partner with data governance, security, architecture, and business stakeholders to balance flexibility with control.
Requirements
- 6+ years in data engineering, with 2+ years designing/operating Databricks-based Lakehouse platforms at production scale.
- Strong hands-on expertise with Apache Spark and Delta Lake; experience with Delta Live Tables/Lakeflow Declarative Pipelines is a strong plus.
- Practical experience with Unity Catalog (or equivalent) for access control, lineage, and cataloguing.
- Demonstrated experience implementing data anonymisation/pseudonymisation techniques in a regulated environment.
- Solid understanding of canonical/conformed data modelling (e.g., data vault, dimensional modelling, domain-driven design for data).
- Experience exposing data via APIs (REST/GraphQL) - designing schemas, contracts, and versioning for external/internal consumers.
- Proficient in Python and/or Scala, SQL, and infrastructure-as-code (Terraform preferred).
- Strong grasp of data governance frameworks, data privacy regulation (GDPR/CCPA), and data security practices.
- Experience with cloud platforms (Azure, AWS, or GCP) and their native data/event services.
Nice to Have
- Exposure to Structured Streaming or event-driven ingestion (Kafka, CDC tools such as Debezium) - not a core requirement, but useful as the platform matures.
- Familiarity with knowledge graph concepts (e.g., RDF/property graphs, Neo4j, graph layers on Spark) - a plus rather than an expectation.
- Experience with data mesh or federated data product architectures.
- Experience with metadata management/catalog tools beyond Unity Catalog (e.g., OpenMetadata, Collibra, Purview).
- Exposure to MLOps or feature store patterns on Databricks.
- Relevant certifications (Databricks Data Engineer Professional, cloud architecture certs).
Join a truly global team.
We offer a dynamic and international environment where high performance meets real purpose. We're proud to be Great Place to Work-certified and even prouder of the people who make that possible. Let’s shape the future of procurement - together.
Chain IQ – Create. Lead. Make an impact.
Information for agencies: Applications sent or uploaded by placement agencies or similar are not desired, will therefore not be considered and will be deleted.