
What We Learned Building Our ML Platform
Six ideas about feature stores, AST-generated ETLs, content-addressed artifacts, and control planes — distilled from building a semantic layer for ML.
Read articleBali · Asia/Makassar · Available for fractional engagements

Six ideas about feature stores, AST-generated ETLs, content-addressed artifacts, and control planes — distilled from building a semantic layer for ML.
Read article
Using S3 DAG bundles, a sidecar sync pattern, and a bundle watcher to onboard pipelines without Helm upgrades or downtime.
Read article
Replacing bespoke PySpark scripts with a config-driven, hook-based framework inspired by Rust's composition model.
Read article
Fixing the GET request explosion, prefix throttling, and threading edge cases that emerge when S3 shuffle meets production scale.
Read article
How S3 shuffle supports Spark executors on 100% spot capacity with Karpenter while cutting compute costs 70-85%.
Read article
An ops-first pattern for production Spark on Kubernetes with one SparkApplication template, pre-baked images, and sub-10-second warm starts.
Read article
Optimizing a daily Iceberg upsert from an hour to under 4 minutes with storage partition joins and shuffle hash hints.
Read article
A modern lakehouse architecture using Apache Iceberg, Lakekeeper, Spark, Trino, and AWS.
Read article