Head of Data Engineering
– Present
Leading Automattic's data engineering team — the engineers behind the central self-hosted data platform powering analytics and AI across WordPress.com, Tumblr, and WooCommerce.
- Lead a team of 5–7 data-platform engineers owning the company's central self-hosted data platform (Spark, Trino, Airflow, Kafka, Apache Iceberg, Superset, Looker, JupyterHub) — roadmap, planning, reviews, and production operations.
- Set the team's AI-enablement strategy and led its delivery across the team: a program of MCP servers (Trino, Superset, Looker, data catalog) giving LLM agents governed access to company data; defined the next phase — AI evaluation, observability, and domain-scoped data assistants.
- Designed data-access governance for AI agents: least-privilege, tag-based, read-only warehouse access enforced with policy-as-code (Open Policy Agent).
- Put LLM evaluation into CI: an LLM-as-judge harness verifying documentation produces correct AI-agent behavior ("TDD for docs").
- Coordinated the team's migration of the warehouse to Apache Iceberg — the platform's move to an open lakehouse table format — while personally contributing 50+ merged improvements on Iceberg reliability, snapshot retention, and storage optimization.
- Drove a data-contract integrity program on Apache Iceberg: automated contract-violation checks, fixes to billion-row production tables, and major storage and compute cost optimizations.
- Represent the data platform across the company: AI data-tooling and access-policy decisions, security collaboration (member of the security incident response team), and the editorial group of Automattic's public data blog.
- Still hands-on: one of the platform's top contributors with 2,000+ merged changes; led every major Airflow production upgrade, through Airflow 3.
Technologies:Trino, Spark, Airflow, Kafka, Apache Iceberg, Scala, Python, Superset, Looker, JupyterHub, Open Policy Agent, MCP