Notes on data & systems

Articles on Data Engineering, Big Data, Cloud, and AI/ML — practical deep dives from building large-scale data platforms.

Feb 20, 2026

The 60-Minute Protocol for Staying Sharp in the Age of AI

mental-modelsartificial-intelligenceneural-networks

Jan 11, 2026

Engineers in 2026 Won’t Be Hired for Syntax. They’ll Be Hired for Leverage

distributed-systemsai-agentllm

Nov 26, 2025

I Built an AI Code Reviewer in a Weekend — Here’s the Exact Prompt

code-reviewprompt-engineeringbig-data

Sep 19, 2025

Integrating LLMs and AI Agents into Data Engineering Workflows

aiai-agentllm

Sep 14, 2025

A Practical Guide to Spark Serialization and Deserialization

big-dataserializationspark

Aug 22, 2025

Zero-ETL & Cloud-Native Architectures: Building Real-Time Data Systems

streamingcloud-nativereal-time-analytics

Jul 11, 2025

Why Every Serious Data Engineer Should Understand Bloom Filters and HyperLogLog

data-structuresbig-databloom-filter

Jul 6, 2025

Embedding-Based Retrieval Is Making Search Smarter

vectorembeddingartificial-intelligence

Jul 2, 2025

MLOps and Data Engineering: Bridging the Gap for Machine Learning Pipelines

mlopsdata-engineeringfeature-engineering

Jun 17, 2025

Understanding Spark’s Catalyst Optimizer: Demystifying Query Optimization

sparkapache-sparkspark-optimization

May 27, 2025

Build Your First Baby Agent with OpenAI in 20 Minutes

ai-agentchatgptopenai

May 21, 2025

Say Goodbye to Dirty Data: Build Trustworthy Pipelines with These Pro Tips

Data EngineeringData QualityData Pipelines

Apr 20, 2025

No SQL? No Problem: Ask Your Database Questions in Plain English

Data EngineeringNLPMySQL

Apr 7, 2025

Catching Sneaky Data Drift Before It Wreaks Havoc

Data EngineeringMachine LearningData Quality

Mar 24, 2025

Your Spark Executors Are Wasting Memory — Here’s How to Fix It

sparkdistributed-systemsmemory-improvement

Mar 8, 2025

Building a Data Lakehouse with Iceberg, Spark, and AWS Glue

Data EngineeringApache IcebergApache Spark

Feb 11, 2025

From Data Lake to Lakehouse: A Migration Guide with Delta

Data EngineeringDelta LakeApache Spark

Feb 5, 2025

Mastering CDC in Delta Tables: A Use-case in Spark

Data EngineeringCDCDelta Lake

Jan 30, 2025

Indexing Strategies: B-Trees, Hash Indexes, Bitmaps & Beyond

indexingsqlbig-data

Jan 20, 2025

Handling Bottlenecks in Spark Streaming: Lessons Learned

Data EngineeringSpark StreamingPerformance Optimization

Jan 9, 2025

Demystifying Event-Driven Architecture with AWS

Data EngineeringEvent-Driven ArchitectureAWS

Dec 24, 2024

Hands-on Cloud: Build a Serverless To-Do List App on AWS

Cloud ComputingAWSServerless

Dec 16, 2024

Zstd vs Snappy vs Gzip: The Compression King for Parquet Has Arrived

parquetdata-engineeringspark

Dec 12, 2024

Building Real-Time ETL Pipelines with Flink? Here's How You Can Nail It!

Data EngineeringApache FlinkKafka

Nov 30, 2024

Building Real-Time Recommendations with Spark, ALS, and Kafka

Data EngineeringApache SparkKafka

Nov 24, 2024

Customer 360 in E-commerce: Real-Life Use Case with Delta Lake on Databricks

Data EngineeringDelta LakeDatabricks

Nov 18, 2024

Real-Time Use-case: Fraud Detection in Financial Transactions with Kafka and Spark Streaming

Data EngineeringKafkaSpark Streaming

Nov 12, 2024

Preventing Data Mix-ups: Understanding Database Isolation and Concurrency Management

Data EngineeringDatabaseConcurrency

Nov 9, 2024

Data Engineering for ML: Building a Customer Churn Prediction Pipeline with Airflow

Data EngineeringMachine LearningApache Airflow

Nov 3, 2024

Building End-to-End Customer Insights Pipeline by Integrating Multiple Data Sources in Spark with Airflow

Data EngineeringApache SparkApache Airflow
Showing 30 of 35 articles