<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"><channel><title>Manoj Babu Katragadda — Blog</title><description>Production notes on data platforms, Spark, Iceberg, MLOps, and Kubernetes.</description><link>https://tardunge.com/</link><language>en-us</language><item><title>What We Learned Building Our ML Platform</title><link>https://tardunge.com/blog/what-we-learned-building-our-ml-platform/</link><guid isPermaLink="true">https://tardunge.com/blog/what-we-learned-building-our-ml-platform/</guid><description>Six ideas about feature stores, AST-generated ETLs, content-addressed artifacts, and control planes — distilled from building a semantic layer for ML.</description><pubDate>Thu, 23 Apr 2026 00:00:00 GMT</pubDate><category>mlops</category><category>semantic-layer</category><category>data-engineering</category><category>platform-engineering</category><category>apache-iceberg</category><author>Manoj Babu Katragadda, Meghanath Macha</author></item><item><title>Airflow DAG Bundles: Managing DAGs Across Teams Without Helm Upgrades</title><link>https://tardunge.com/blog/airflow-dag-bundles-without-helm-upgrades/</link><guid isPermaLink="true">https://tardunge.com/blog/airflow-dag-bundles-without-helm-upgrades/</guid><description>Using S3 DAG bundles, a sidecar sync pattern, and a bundle watcher to onboard pipelines without Helm upgrades or downtime.</description><pubDate>Wed, 15 Apr 2026 00:00:00 GMT</pubDate><category>airflow</category><category>kubernetes</category><category>data-engineering</category><category>dag-bundles</category><author>Manoj Babu Katragadda, Meghanath Macha</author></item><item><title>Building a Composable ETL Framework for Spark</title><link>https://tardunge.com/blog/building-a-composable-etl-framework/</link><guid isPermaLink="true">https://tardunge.com/blog/building-a-composable-etl-framework/</guid><description>Replacing bespoke PySpark scripts with a config-driven, hook-based framework inspired by Rust&apos;s composition model.</description><pubDate>Thu, 02 Apr 2026 00:00:00 GMT</pubDate><category>spark</category><category>data-engineering</category><category>architecture</category><category>python</category><category>iceberg</category><author>Manoj Babu Katragadda, Meghanath Macha</author></item><item><title>Taming S3 Shuffle at Scale</title><link>https://tardunge.com/blog/taming-s3-shuffle-at-scale/</link><guid isPermaLink="true">https://tardunge.com/blog/taming-s3-shuffle-at-scale/</guid><description>Fixing the GET request explosion, prefix throttling, and threading edge cases that emerge when S3 shuffle meets production scale.</description><pubDate>Mon, 30 Mar 2026 00:00:00 GMT</pubDate><category>spark</category><category>s3-shuffle</category><category>performance</category><category>data-engineering</category><category>scale</category><author>Manoj Babu Katragadda, Meghanath Macha</author></item><item><title>Spark on Spot with S3 Shuffle</title><link>https://tardunge.com/blog/spark-on-spot-with-s3-shuffle/</link><guid isPermaLink="true">https://tardunge.com/blog/spark-on-spot-with-s3-shuffle/</guid><description>How S3 shuffle supports Spark executors on 100% spot capacity with Karpenter while cutting compute costs 70-85%.</description><pubDate>Wed, 25 Mar 2026 00:00:00 GMT</pubDate><category>spark</category><category>kubernetes</category><category>spot-instances</category><category>karpenter</category><category>s3-shuffle</category><category>cost-optimization</category><author>Manoj Babu Katragadda, Meghanath Macha</author></item><item><title>Running Spark on Kubernetes: An Ops-First Approach</title><link>https://tardunge.com/blog/running-spark-on-kubernetes/</link><guid isPermaLink="true">https://tardunge.com/blog/running-spark-on-kubernetes/</guid><description>An ops-first pattern for production Spark on Kubernetes with one SparkApplication template, pre-baked images, and sub-10-second warm starts.</description><pubDate>Thu, 19 Mar 2026 00:00:00 GMT</pubDate><category>spark</category><category>kubernetes</category><category>data-engineering</category><category>airflow</category><category>developer-experience</category><author>Manoj Babu Katragadda, Meghanath Macha</author></item><item><title>From 60 Minutes to 4: Optimizing Spark MERGE INTO on a 2 Billion Row Iceberg Table</title><link>https://tardunge.com/blog/spark-merge-optimization-from-60-to-4-minutes/</link><guid isPermaLink="true">https://tardunge.com/blog/spark-merge-optimization-from-60-to-4-minutes/</guid><description>Optimizing a daily Iceberg upsert from an hour to under 4 minutes with storage partition joins and shuffle hash hints.</description><pubDate>Thu, 19 Mar 2026 00:00:00 GMT</pubDate><category>spark</category><category>iceberg</category><category>performance</category><category>data-engineering</category><author>Manoj Babu Katragadda, Meghanath Macha</author></item><item><title>Building Our Lakehouse with Apache Iceberg</title><link>https://tardunge.com/blog/building-our-lakehouse-with-iceberg/</link><guid isPermaLink="true">https://tardunge.com/blog/building-our-lakehouse-with-iceberg/</guid><description>A modern lakehouse architecture using Apache Iceberg, Lakekeeper, Spark, Trino, and AWS.</description><pubDate>Sun, 15 Mar 2026 00:00:00 GMT</pubDate><category>data-engineering</category><category>iceberg</category><category>lakehouse</category><category>architecture</category><author>Manoj Babu Katragadda, Meghanath Macha</author></item></channel></rss>