Skip to content

UPDATED 14:30 EDT / SEPTEMBER 30 2026

BIG DATA

AWS embeds DuckDB in its PostgreSQL DBMS to speed queries of live and historical data

Amazon Web Services Inc. is adding the ability to query the popular Apache Iceberg data lake directly from its Aurora PostgreSQL database management system, allowing applications to combine live transactions with historical records without copying data or creating extract/transform/load pipelines.

The capability embeds the DuckDB analytical engine inside Aurora PostgreSQL to support data stored in Iceberg and Apache Parquet formats. Customers can use existing PostgreSQL applications, tools and endpoints to query operational records alongside data in Amazon S3, including S3 Tables.

AWS acquired DuckLabs B.V., developer of the popular open-source database, last month.

AWS is positioning the feature as a way to simplify application development and reduce the engineering work needed to maintain data pipelines. Potential uses include real-time dashboards, transactions enriched with historical information and artificial intelligence agents that need access to both current and archived records.

Previously, combining recent transactions in Aurora with historical records in S3 commonly required reverse ETL pipelines. That resulted in duplicated data, additional infrastructure costs and continuing work to keep records synchronized.

“This challenge only grows as you increasingly embed AI agents into your applications, where it is impractical to predict and pre-replicate every dataset an agent might need,” Esra Kayabali, a principal solutions architect at AWS, wrote in an announcement.

DuckDB processes analytical scans within Aurora, avoiding additional network hops for query processing. A single query can access data lake records alongside live operational data, including uncommitted writes. AWS said integration as an example of how it is incorporating the DuckDB engine widely into its services.

The feature also supports external catalogs compatible with the Iceberg Representational State Transfer Catalog specification through federation with the data catalog in the AWS Glue serverless data integration service. Customers register an external catalog with Glue and create foreign tables that reference its data. Applications can then join Aurora records with Iceberg tables registered across multiple catalogs.

To limit the amount of data read, Aurora filters records and selects relevant columns during query execution. It also caches frequently accessed data. Developers can inspect metrics including rows scanned, bytes read from S3 and cache hits.

In a financial example, Kayabali demonstrated a query combining seven days of customer transactions in Aurora with five years of historical transactions stored in a Parquet file in S3. Aurora inferred the historical table’s schema from file metadata, eliminating manual column definitions.

For workloads needing single-digit-millisecond latency, customers can copy selected data lake records into native Aurora tables using standard SQL commands. Read queries can run on the cluster’s writer or a read replica, offloading analytical scans from operational workloads. Commands that materialize data run on the writer.

Customers enable the capability through the aurora_analytics extension and an AWS Identity and Access Management role granting access to S3 and Glue. It supports Aurora PostgreSQL versions 17 and 18, starting with versions 17.11 and 18.6, respectively.

AWS said the feature is available in all commercial AWS regions without an additional feature charge. Customers pay for the incremental Aurora computing resources their queries consume and S3 requests used to read files.

Photo: AWS

A message from John Furrier, co-founder of SiliconANGLE:

Support our mission to keep content open and free by engaging with theCUBE community. Join theCUBE’s Alumni Trust Network, where technology leaders connect, share intelligence and create opportunities.

  • 15M+ viewers of theCUBE videos, powering conversations across AI, cloud, cybersecurity and more
  • 11.4k+ theCUBE alumni — Connect with more than 11,400 tech and business leaders shaping the future through a unique trusted-based network

Are you an AWS customer?  Support SiliconANGLE financially by buying your AWS services from our Marketplace portal page and links: https://siliconangle.com/aws-marketplace/

 

About SiliconANGLE Media
SiliconANGLE Media is a recognized leader in digital media innovation, uniting breakthrough technology, strategic insights and real-time audience engagement. As the parent company of SiliconANGLE, theCUBE Network, theCUBE Research, CUBE365, theCUBE AI and theCUBE SuperStudios — with flagship locations in Silicon Valley and the New York Stock Exchange — SiliconANGLE Media operates at the intersection of media, technology and AI.

Founded by tech visionaries John Furrier and Dave Vellante, SiliconANGLE Media has built a dynamic ecosystem of industry-leading digital media brands that reach 15+ million elite tech professionals. Our new proprietary theCUBE AI Video Cloud is breaking ground in audience interaction, leveraging theCUBEai.com neural network to help technology companies make data-driven decisions and stay at the forefront of industry conversations.

Send us a news tip

Send us a News Tip

  • This field is for validation purposes and should be left unchanged.
  • Max. file size: 244 MB.

Sign in

SIGN IN

Bio

Ethics statement

Extract the signal from the noise

Get SiliconANGLE updates and analysis.

Contact us

Partner with us

Contact us

Guest inquiry