Data Engineer
Preply
| Company | Preply |
| Category | Engineering |
| Location | Barcelona |
| Remote | Hybrid |
| Employment | Not stated |
| Level | Not stated |
| Salary | Not stated by the employer |
| Posted | 29 Jul 2026 |
| Last verified | 30 Jul 2026 |
| Source | Employer career page (ashby) |
Description
WE POWER PEOPLE’S PROGRESS.
At Preply, we’re all about creating life-changing learning experiences. We help people discover the magic of the perfect tutor, craft a personalised learning journey, and stay motivated to keep growing. Our approach is human-led, tech-enabled - and it’s creating real impact.
We’ve just reached unicorn status with a $150M Series D, accelerating our vision to transform education through human-led, AI-enhanced learning. Today, 100,000+ tutors teach 90+ languages to learners in 180 countries - and we’re only getting started. As a category-defining company, we’re shaping what the future of learning looks like at global scale.
Every Preply lesson sparks change, fuels ambition, and drives progress that matters. Joining Preply means helping define the future of education at global scale, and building something that truly matters for millions of people, every day.
MEET THE TEAM!
At Preply, the Data Ingestion and Enrichment team provides a single, trusted, and scalable data foundation. The team ensures that all analytics, machine learning, and product features are built on unified, governed, and production-grade data assets in Preply's Lake House, including the extraction, normalization, and generation of structured data from Preply's unstructured assets, forming a durable data moat for AI-driven products.
As a Data Engineer in the Data Ingestion and Enrichment team, you will build and contribute to the data layer that powers both Preply's analytics, machine learning, and product. You will work closely with ML Platform, Applied/Data Scientists, Analytics Engineering, and Product squads to ensure that features, datasets, and pipelines are production-ready, observable, and reusable within the team.
WHAT YOU'LL BE DOING:
Contribute to trusted ingestion & enrichment foundations (Data Lake and Data as a Product):
Build and maintain components of Preply's data lake. Ensure every dataset has clear ownership, purpose, schemas, and quality expectations from first ingestion through downstream consumption by analytics, product, and ML teams. Treat trust, correctness, and predictability as first-class features of the platform.
Develop end-to-end ingestion pipelines (batch & streaming):
Build and operate reliable batch and streaming ingestion pipelines that support both real-time and analytical use cases. Contribute to defining clear raw → standardized → consumption layers with explicit responsibilities, lineage, and retention strategies. Balance performance, cost, and reliability as the platform scales.
Data quality, contracts & early validation:
Implement data contracts between producers and consumers, covering schema, freshness, volume, and quality guarantees. Embed validation, anomaly detection, and quality checks early in the ingestion lifecycle to catch issues before they propagate. Apply standardized quality metrics.
Enrichment, modeling & lifecycle management:
Build enrichment logic that joins, standardizes, and contextualizes data across domains using shared definitions and reusable patterns. Support historical tracking, point-in-time correctness, and dataset versioning so downstream users can confidently analyze changes and impacts over time.
Observability, reliability & operational excellence:
Instrument ingestion pipelines with strong observability: freshness, latency, data quality, and cost metrics. Contribute to SLOs, alerting, and incident response playbooks so data failures are visible, diagnosable, and recoverable. Help move the platform from reactive firefighting to proactive reliability management.
Governance & compliance by design:
Apply consistent access control, classification, and privacy protections at ingestion time. Ensure sensitive data is properly masked, minimized, or anonymized by default, and that all data flows you own are auditable and traceable.
Enable self-service & standardization:
Contribute to standardized ingestion templates, shared librarie