AI Academy

AI Data Engineer

A 20-week data engineering bootcamp. Turn messy raw data into trustworthy, reproducible assets a model can use directly — the real foundation under every AI system.

4.6(165 ratings)
Highlights

Batch through streaming, warehouse through vector store — one continuous pipeline

Data quality and lineage run throughout: a pipeline that merely runs is not one you can trust

An AI-specific track: document parsing, chunking and embeddings feeding straight into RAG


Curriculum
1

1 · Foundations: batch vs streaming trade-offs, advanced SQL, data modelling

2

2 · Pipelines and orchestration: ETL/ELT, Airflow scheduling, dependencies, backfills, idempotency

3

3 · Warehouse and lakehouse: dimensional modelling, layering and partitioning, dbt transforms and tests

4

4 · Streaming: Kafka, incremental sync and CDC, late-arriving data and windows

5

5 · Processing at scale: Spark, partitioning and skew, query and cost optimisation

6

6 · Preparing data for AI: parsing and cleaning documents, chunking strategy, generating and storing embeddings

7

7 · Retrieval and feature serving: tuning vector search, feature stores, offline/online consistency

8

8 · Quality and governance: validation rules, lineage, monitoring, alerting and SLAs

9

9 · Cloud deployment and cost: containers, cloud warehouses, storage tiers, cost governance

10

10 · Capstone: build and hand over an end-to-end AI data pipeline


Core technologies

Languages

Python
SQL
Bash

Orchestration & transforms

Airflow
dbt
Dagster

Storage & warehouse

PostgreSQL
Snowflake
BigQuery
S3 / 对象存储

Streaming & compute

Kafka
Spark
CDC

AI data

Embeddings
pgvector
文档解析
RAG 数据链路

Cloud & delivery

Docker
AWS
CI/CD
监控告警
Who this is for

Back-end, ops or analytics backgrounds moving into data engineering

People already writing reports or ETL who need streaming, lakehouse and AI pipelines

Candidates targeting North American data engineering or AI platform teams