Posts
All the articles I've posted.
- 14 MIN READ•May 24, 2026
Modern Feature Stores Beyond Batch Pipelines
Feature stores like Feast now support streaming feature views from Kafka and Kinesis alongside batch pipelines. Learn how to build real-time features that maintain training-serving consistency.
Feature Store Streaming Real-Time MlFeast Streaming FeaturesOnline Offline Feature Store - 12 MIN READ•May 24, 2026
OpenLineage as the Spine of Data Observability
OpenLineage provides a standard API for collecting pipeline lineage across Airflow, Spark, Flink, and dbt. Learn how it powers blast radius analysis and incident triage.
Openlineage Data ObservabilityOpenlineage AirflowOpenlineage Spark - 12 MIN READ•May 24, 2026
When Paimon Beats Iceberg for Mutable Streams
Apache Paimon uses LSM-Tree storage for native CDC upserts without restart. Learn when Paimon outperforms Iceberg for high-churn mutable streaming workloads.
Apache Paimon Mutable StreamsPaimon Vs IcebergCdc Streaming Lakehouse - 13 MIN READ•May 24, 2026
Policy as Code for Lakehouse Governance
OPA, ABAC, row filters, and column masks make lakehouse governance programmable and scalable. Learn how Databricks, Snowflake Horizon, and BigQuery implement policy-as-code.
Policy As Code Data Governance LakehouseOpa Rego LakehouseAbac Databricks - 11 MIN READ•May 24, 2026
Real-Time Lakehouse Patterns with Apache Flink and Iceberg
Learn how to build a real-time lakehouse with Apache Flink 2.1 and the Dynamic Iceberg Sink, covering schema evolution, exactly-once delivery, and compaction.
Real-Time Lakehouse FlinkFlink Iceberg SinkKafka To Iceberg - 14 MIN READ•May 24, 2026
Why Semantic Layers Make Enterprise Text-to-SQL Safer
Text-to-SQL accuracy jumps from 40% to 85-95% when grounded in a semantic layer. Learn how Dremio, Snowflake Cortex Analyst, and dbt Semantic Layer improve AI analytics reliability.
Semantic Layer Text-To-SqlDremio Semantic LayerSnowflake Cortex Analyst - 14 MIN READ•May 24, 2026
Choosing Vector Stores for Retrieval Workloads
pgvector, Milvus, Weaviate, and LanceDB each make different tradeoffs on index type, hybrid search, scale, and operational complexity. Learn which fits your retrieval workload.
Vector Store Comparison Retrieval WorkloadsPgvector HnswMilvus Hybrid Search - 24 MIN READ•May 23, 2026
Single-Node Data Engineering: DuckDB, DataFusion, Polars, and LakeSail
Optimize single-node data engineering with DuckDB, DataFusion, Polars, and LakeSail. Compare architectures and learn when to transition to Dremio MPP.
DuckDBApache ArrowDataFusion - 20 MIN READ•May 23, 2026
An In-Depth Overview of the Apache Iceberg 1.11.0 Release
Apache Iceberg 1.11.0 delivers manifest list encryption, the new pluggable File Format API, credential lifecycle refreshes, and Spark/Flink improvements.
Apache IcebergData LakehouseOpen Table Format - 22 MIN READ•May 22, 2026
Open Table Format Benchmarks: Why They Require Critical Evaluation
An in-depth analysis of open table format benchmarks comparing Apache Iceberg, Delta Lake, and Apache Hudi, detailing the pitfalls of standard benchmarks and how to choose a format.
open table formatsapache icebergdelta lake - 27 MIN READ•May 22, 2026
Apache Iceberg SCD Type 2 and CDC Patterns: Building Historical Lakehouse Tables
A deep dive into implementing Slowly Changing Dimension Type 2 (SCD Type 2) patterns and Change Data Capture (CDC) pipelines on Apache Iceberg, using PySpark and Dremio.
apache icebergcdcscd type 2 - 25 MIN READ•May 22, 2026
Setting Up an AWS-Native Open Lakehouse: Querying Apache Iceberg with AWS Athena and AWS Glue Catalog
A comprehensive guide to building an open, high-performance lakehouse on AWS using Apache Iceberg, AWS Glue Catalog, Amazon S3, and S3 Tables, with query acceleration via the Dremio engine.
Apache IcebergAWS AthenaAWS Glue Catalog