Posts
All the articles I've posted.
- 16 MIN READ•May 24, 2026
Choosing the Right Iceberg Control Plane: Polaris vs. Unity Catalog vs. Cloud REST
Choosing an Apache Iceberg catalog? Compare open-source Apache Polaris, open Unity Catalog, and managed cloud REST control planes to unify your lakehouse.
Apache Iceberg CatalogIceberg Rest CatalogApache Polaris - 13 MIN READ•May 24, 2026
Clean Rooms for Privacy-Preserving Analytics
Data clean rooms enable secure multi-party analytics without sharing raw data. Learn how Databricks Clean Rooms, AWS Clean Rooms, and BigQuery differential privacy work.
Data Clean Rooms Privacy-Preserving AnalyticsDatabricks Clean RoomsAws Clean Rooms - 14 MIN READ•May 24, 2026
Building Composable Query Engines with Rust Runtimes
Apache DataFusion, Velox, and Substrait form the foundation of modern composable query engine stacks. Learn how these components fit together and when to use each.
Composable Query Engine DatafusionApache Datafusion RustVelox C++ Engine - 11 MIN READ•May 24, 2026
Data Mesh After the Hype: What Actually Works
Three years after Zhamak Dehghani's original papers, data mesh has proven valuable in specific organizational contexts and impractical in others. Here's what the practical implementations look like.
Data Mesh Practical ImplementationData Mesh Reality CheckData Product Thinking - 12 MIN READ•May 24, 2026
How dbt Fusion Reshapes Analytics Engineering
dbt Fusion entered public beta in May 2025 with a Rust-powered runtime that changes how analytics engineers develop, validate, and deploy SQL models. Here's what changed.
Dbt Fusion Analytics EngineeringDbt Fusion RustDbt State-Aware Orchestration - 13 MIN READ•May 24, 2026
Using DuckDB and Polars to Query Iceberg Tables
DuckDB 1.4 LTS and Polars streaming engine now both support reading and writing Apache Iceberg tables. Learn how to use them for lakehouse analytics in 2025.
Duckdb Polars IcebergDuckdb Iceberg WritePolars Iceberg Sink - 12 MIN READ•May 24, 2026
FinOps for Data Warehouses with Open Billing Data
The FOCUS 1.3 specification and native warehouse cost views make real-time cost attribution practical. Learn how to build a FinOps pipeline for Snowflake, BigQuery, and multi-cloud environments.
Warehouse Finops Focus SpecificationSnowflake Cost ManagementBigquery Jobs View - 13 MIN READ•May 24, 2026
Designing Governed RAG on Data Products
Enterprise RAG architecture that trusts its own data requires governance at the retrieval layer. Learn how to build governed RAG using data products, access policies, and semantic layer routing.
Governed Rag Enterprise Data ProductsEnterprise Rag ArchitectureGoverned Retrieval Augmented Generation - 12 MIN READ•May 24, 2026
What Iceberg V3 Advances Mean for CDC Pipelines
Apache Iceberg V3 brings deletion vectors and row lineage that reshape CDC pipeline design. Learn what these features mean for your streaming data architecture.
Iceberg Cdc PipelineIceberg Deletion VectorsIceberg Row Lineage - 13 MIN READ•May 24, 2026
Kafka 4.0 Changes Streaming Platform Operations
Kafka 4.0 removes ZooKeeper and ships KRaft and KIP-848 by default. Learn what those changes mean for platform operations, upgrades, and client configurations.
Kafka 4.0 UpgradeKafka KraftZookeeper Removal Kafka - 13 MIN READ•May 24, 2026
Lance and Iceberg for Multimodal AI Data
LanceDB and Apache Iceberg serve complementary roles in a multimodal AI lakehouse. Learn when to use Lance for embeddings and random access, and Iceberg for structured metadata and SQL analytics.
Lancedb Iceberg Multimodal Ai DataLancedb FormatLance Vs Iceberg - 13 MIN READ•May 24, 2026
Bringing MLflow and Data Pipelines Closer Together
MLflow 3 extends observability from classic ML experiments to GenAI tracing and data pipeline lineage. Learn how to connect data quality monitoring with model performance tracking.
Mlflow Data Pipeline ObservabilityMlflow 3 Genai TracingMlflow Data Quality Monitoring