From weeks to minutes: The new agentic era of data pipelines
Data pipelines are the backbone of the modern enterprise, yet a barrier to entry exists for orchestrating them, making this critical capability unavailable to many data professiona...
Search fresh public links, source activity, and ready-to-use post angles for Datapipeline.
Fresh curated links around datapipeline are collected here so marketers can spot useful updates and turn timely ideas into posts faster.
Recent items include:
Recent curated links from global sources. Generate one free draft from any story, then use SocialBu to schedule and refine your content calendar.
Data pipelines are the backbone of the modern enterprise, yet a barrier to entry exists for orchestrating them, making this critical capability unavailable to many data professiona...
the strange space in between
Most enterprise data problems are not caused by machine learning models or dashboard tools. They usually start much earlier in the pipeline. A reporting table misses records after...
The Agentic Data Operations Platform (ADOP) is a reference architecture on Amazon Bedrock that uses specialized AI agents to automate the full Bronze-to-Silver-to-Gold data pipelin...
AWS Glue makes it easy to get a PySpark pipeline running quickly. It is significantly harder to build one that stays maintainable as logic grows, performs reliably at scale, and do...
The goal is not custom AMIs for every workload. It is to make dependency management, patching and recovery explicit platform responsibilities rather than repeated job-level tasks.
With AWS Glue 6.0, you can build real-time, near-real-time, and batch data pipelines on a single platform. Using a financial market-risk example, learn how to flag high-risk trades...
Originally appeared on OmbuLabs Blog.A few months ago, we were tasked with building a platform that aggregates events across an entire city, concerts, gallery openings, museum exhi...
If you ask an AI coding agent to write a standalone Python script to parse a single JSON file, it will likely give you a perfect answer in seconds. But the same agent often breaks...
One of Canada’s largest railway networks spans roughly 20,000 route miles across...
Asynchronous processing pipelines are a cornerstone of modern distributed systems, but wiring them together reliably can be complex. A typical pipeline built with queues or message...
Introduction In the previous article, we learned about AWS CodeCommit, AWS's managed Git repository service. Although CodeCommit can store source code, simply storing code is no...
Every data team eventually asks the same question: what runs our pipelines, on what schedule, with what retry logic, and who gets paged when it fails. The answer used to default to...
Discover how to easily stream data into Snowflake-managed Apache Iceberg tables using Snowpipe Streaming.
Senior data engineers are trained to be skeptical of proprietary platforms. When I entered a Palantir Foundry training bootcamp, I expected to find a slow, expensive alternative to...
This project creates a daily digest for sellers in an enterprise system. Each seller handles accounts at a set of companies and needs to know when something happens at one of them:...
About a year ago, I was building a terminal-based workflow manager called Glyph.Flow. It was mostly a learning project. I wanted to understand Python better, experiment with Textu...
Data engineering teams running Apache Airflow across multiple AWS accounts have no built-in way to coordinate workflows between separate Amazon MWAA environments. With Airflow 3.0...
В исходном orders.csv было 11 строк. До BI-витрины дошло 7, а валовая сумма 4720.30 после применения бизнес-правил превратилась в 2200.30 выручки. Четыре строки не исчезли: каждая...
Session 2 of The Production RAG HandbookContinue reading on Medium »
The problem we couldn't ignoreGlaspoort builds and operates fiber infrastructure in the Netherlands...
Many Fabric migration estimates assume every ADF pipeline deserves a place in the target environment. This blog presents a business-led framework for retiring, consolidating, redes...
Learn how to build a dynamic streaming data lake on Amazon Managed Service for Apache Flink that adapts to new event types and schema changes without stopping the pipeline, using A...
Discover how to migrate Apache Spark pipelines to Snowflake (Snowpark Connect) effortlessly using the Snowflake CoCo spark-migration skill. Improve performance and reduce costs.
Use SocialBu to discover ideas, generate post drafts, and schedule them across your social channels.