AWS embeds DuckDB in its PostgreSQL DBMS to speed queries of live and historical data
Amazon Web Services Inc. is adding the ability to query the popular Apache Iceberg data lake directly from its Aurora PostgreSQL database management system, allowing applications to combine live transactions with historical records without copying data or creating extract/transform/load pipelines.
The capability embeds the DuckDB analytical engine inside Aurora PostgreSQL to support data stored in Iceberg and Apache Parquet formats. Customers can use existing PostgreSQL applications, tools and endpoints to query operational records alongside data in Amazon S3, including S3 Tables.
AWS acquired DuckLabs B.V., developer of the popular open-source database, last month.
AWS is positioning the feature as a way to simplify application development and reduce the engineering work needed to maintain data pipelines. Potential uses include real-time dashboards, transactions enriched with historical information and artificial intelligence agents that need access to both current and archived records.
Previously, combining recent transactions in Aurora with historical records in S3 commonly required reverse ETL pipelines. That resulted in duplicated data, additional infrastructure costs and continuing work to keep records synchronized.
“This challenge only grows as you increasingly embed AI agents into your applications, where it is impractical to predict and pre-replicate every dataset an agent might need,” Esra Kayabali, a principal solutions architect at AWS, wrote in an announcement.
DuckDB processes analytical scans within Aurora, avoiding additional network hops for query processing. A single query can access data lake records alongside live operational data, including uncommitted writes. AWS said integration as an example of how it is incorporating the DuckDB engine widely into its services.
The feature also supports external catalogs compatible with the Iceberg Representational State Transfer Catalog specification through federation with the data catalog in the AWS Glue serverless data integration service. Customers register an external catalog with Glue and create foreign tables that reference its data. Applications can then join Aurora records with Iceberg tables registered across multiple catalogs.
To limit the amount of data read, Aurora filters records and selects relevant columns during query execution. It also caches frequently accessed data. Developers can inspect metrics including rows scanned, bytes read from S3 and cache hits.
In a financial example, Kayabali demonstrated a query combining seven days of customer transactions in Aurora with five years of historical transactions stored in a Parquet file in S3. Aurora inferred the historical table’s schema from file metadata, eliminating manual column definitions.
For workloads needing single-digit-millisecond latency, customers can copy selected data lake records into native Aurora tables using standard SQL commands. Read queries can run on the cluster’s writer or a read replica, offloading analytical scans from operational workloads. Commands that materialize data run on the writer.
Customers enable the capability through the aurora_analytics extension and an AWS Identity and Access Management role granting access to S3 and Glue. It supports Aurora PostgreSQL versions 17 and 18, starting with versions 17.11 and 18.6, respectively.
AWS said the feature is available in all commercial AWS regions without an additional feature charge. Customers pay for the incremental Aurora computing resources their queries consume and S3 requests used to read files.
Photo: AWS
A message from John Furrier, co-founder of SiliconANGLE:
Support our mission to keep content open and free by engaging with theCUBE community. Join theCUBE’s Alumni Trust Network, where technology leaders connect, share intelligence and create opportunities.
- 15M+ viewers of theCUBE videos, powering conversations across AI, cloud, cybersecurity and more
- 11.4k+ theCUBE alumni — Connect with more than 11,400 tech and business leaders shaping the future through a unique trusted-based network
Are you an AWS customer? Support SiliconANGLE financially by buying your AWS services from our Marketplace portal page and links: https://siliconangle.com/aws-marketplace/
About SiliconANGLE Media
Founded by tech visionaries John Furrier and Dave Vellante, SiliconANGLE Media has built a dynamic ecosystem of industry-leading digital media brands that reach 15+ million elite tech professionals. Our new proprietary theCUBE AI Video Cloud is breaking ground in audience interaction, leveraging theCUBEai.com neural network to help technology companies make data-driven decisions and stay at the forefront of industry conversations.