AI Data Engineer
A 20-week data engineering bootcamp. Turn messy raw data into trustworthy, reproducible assets a model can use directly — the real foundation under every AI system.
Highlights
Batch through streaming, warehouse through vector store — one continuous pipeline
Data quality and lineage run throughout: a pipeline that merely runs is not one you can trust
An AI-specific track: document parsing, chunking and embeddings feeding straight into RAG
Curriculum
1 · Foundations: batch vs streaming trade-offs, advanced SQL, data modelling
2 · Pipelines and orchestration: ETL/ELT, Airflow scheduling, dependencies, backfills, idempotency
3 · Warehouse and lakehouse: dimensional modelling, layering and partitioning, dbt transforms and tests
4 · Streaming: Kafka, incremental sync and CDC, late-arriving data and windows
5 · Processing at scale: Spark, partitioning and skew, query and cost optimisation
6 · Preparing data for AI: parsing and cleaning documents, chunking strategy, generating and storing embeddings
7 · Retrieval and feature serving: tuning vector search, feature stores, offline/online consistency
8 · Quality and governance: validation rules, lineage, monitoring, alerting and SLAs
9 · Cloud deployment and cost: containers, cloud warehouses, storage tiers, cost governance
10 · Capstone: build and hand over an end-to-end AI data pipeline
Core technologies
Languages
Orchestration & transforms
Storage & warehouse
Streaming & compute
AI data
Cloud & delivery
