Newsletters




Amazon Aurora PostgreSQL Now Delivers Direct Querying of Apache Iceberg and Parquet Data in the Data Lake


AWS is announcing that Amazon Aurora PostgreSQL now enables users to utilize direct query operational data together with data stored in the data lake in Apache Iceberg and Apache Parquet formats, using existing PostgreSQL applications and tools.

By eliminating the need to extract, transform, and load (ETL) structured data from data lakes into the operational database, users can reduce operational complexity and simplify application development, according to AWS.

Users can also utilize Aurora PostgreSQL to query data from data lakes managed in Iceberg REST Catalog (IRC)-compatible catalogs, enabling access to data across a breadth of analytics systems without moving or duplicating it.

“Whether you’re powering real-time dashboards, enriching transactions with historical context, or building AI agents that reason over both live and archived data, you can now do it all through a single, familiar interface,” AWS said.

DuckLabs, the team that maintains the DuckDB project, recently joined Amazon, and now DuckDB is now embedded directly within Aurora PostgreSQL. This means users can query live operational data (including uncommitted writes) alongside the data lake in a single query.

Query processing stays within Aurora, with no additional network hops and no ETL pipelines that duplicate data.

This capability is supported on two Aurora PostgreSQL major versions: 17 (starting with 17.11) and 18 (starting with 18.6).

Direct querying of Apache Iceberg and Parquet data from Amazon Aurora PostgreSQL is available now in all commercial AWS Regions and AWS GovCloud (U.S.) Regions, at no additional charge. Users pay only for the incremental Aurora compute the queries consume and Amazon S3 request costs for reading data lake files.

For more information about this news, visit https://aws.amazon.com.


Sponsors