A frustratingly simple way to create governed operational data from PostgreSQL for AI applications
Over the past couple of years I’ve been building infrastructure for data platforms and AI systems. My background is in analytics engineering, cloud data platforms, PostgreSQL, Snowflake, and AWS, but lately my focus has shifted toward building products rather than consulting.
My current project is pg-cdc, an open-source PostgreSQL Change Data Capture platform that continuously publishes changes to modern data lake formats like Parquet and Apache Iceberg. The motivation came from seeing teams spend weeks wiring together Debezium, Kafka, connectors, and custom pipelines when many simply want reliable, governed data flowing into their analytics or AI stack.
More broadly, I’m fascinated by how AI agents will consume operational data. I believe the next generation of software won’t just answer questions—it will perform work autonomously. That requires fresh, trustworthy, permission-aware business context, and I think data infrastructure is becoming a critical part of the AI stack.
I’m here to:
Learn from other founders building developer tools and AI infrastructure.
Share what I’ve learned about PostgreSQL, CDC, cloud data engineering, and analytics platforms.
Get honest feedback from builders who aren’t afraid to challenge assumptions.
Replies