Senior Data Engineer
Hack The Box
| Company | Hack The Box |
| Category | Engineering |
| Location | London |
| Remote | On-site (inferred) |
| Employment | Full-time |
| Level | Senior |
| Salary | Not stated by the employer |
| Posted | 7 May 2026 |
| Last verified | 30 Jul 2026 |
| Source | Employer career page (workable) |
Description
Welcome! Super excited you dropped by 🥳 Let's redefine cyber security expertise standards and connect business - community through highly engaging hacking experiences. (Find out more insights about Hack The Box culture in our career site). ✨The core mission of the Senior Data Engineer: You will own and evolve our data pipelines on GCP — building new ones, hardening existing ones, improving data quality, and making clean, trustworthy data available across the organisation. You'll work end-to-end on streaming and batch pipelines, from CDC and event ingestion through transformation, serving, and the feature layer that powers our ML and AI products. Your day-to-day will include designing ELT/ETL processes on BigQuery and ClickHouse, building real-time pipelines on Pub/Sub and Kafka with Dataflow (and where it fits, Flink/Spark), orchestrating workflows with Airflow, and ensuring data is properly cleaned, modelled, and served for analytics, ML training, and online inference. You'll partner with ML engineers on feature pipelines, monitoring data drift, and keeping models well-fed and retrained as needed. You'll consume and build REST APIs, integrate with third-party SaaS sources, and treat infrastructure as code. 🏢 Location & Work Mode: Europe/Greece Choose between fully remote work across Europe or relocating to our Athens Tech Hub with relocation support. Fully Remote / Hybrid (2 days in the office, 3 days remote, plus one month of work from anywhere) - If only in Greece: When hiring in Greece, we’re open to candidates from all locations. Those based within 55 km of our Athens office will follow a hybrid work model. For candidates located beyond that radius, a fully remote arrangement is available. ✈️ Relocation We offer a 10% relocation bonus on the annual salary to support your move. In addition, you would benefit from a 50% tax reduction on your income as part of the relocation. We'll support you with accommodation during the first few weeks 🍺 The fellowship you’ll be joining: You will be part of the Data, Analytics & AI team, collaborating closely with Infrastructure, Software Engineering, Product, and ML/AI engineers. We're in the middle of a GCP-native modernisation — migrating away from Snowflake toward BigQuery, Bigtable, Pub/Sub, and Dataflow — so we're looking for someone who's opinionated about clean architecture, allergic to over-engineering, and comfortable owning systems end-to-end. If retiring a legacy warehouse and standing up its replacement sounds like a good time, you'll fit right in. ⚔️ Technology tools & weapons you’ll be using: Cloud & warehouse: GCP, BigQuery, Bigtable, Cloud Storage Streaming & messaging: Pub/Sub, Kafka Processing: Dataflow (Apache Beam), with Flink/Spark where appropriate Orchestration: Airflow (Cloud Composer) Analytical store: ClickHouse Languages: Python, SQL Modelling & quality: dbt, data quality gates Containers & CI/CD: Docker, Kubernetes, GitHub Actions / equivalent Legacy (being retired): Snowflake 🚀 The adventures that await you after becoming Senior Data Engineer at Hack The Box: Design and build batch and streaming pipelines on Dataflow, Pub/Sub, and Kafka feeding BigQuery, Bigtable, and ClickHouse Help drive the migration off Snowflake onto our GCP-native stack — and retire shadow pipelines along the way Own the orchestration layer in Airflow, including SLAs, retries, and data quality gates Model data for analytics and for ML — including feature pipelines that serve both training and low-latency online inference Partner with ML engineers on feature stores, drift monitoring, and retraining workflows Capture requirements from stakeholders and translate them into pragmatic, well-scoped data products Continuously improve data quality, reliability, observability, and cost efficiency Identify new data sources worth acquiring and integrate them cleanly �