v4.6.0 Release Notes
Release date: June, 2026
Version: v4.6.0
SynxDB v4.6.0 adds native Apache Iceberg table support, extends vectorized execution and GPORCA parallel query capabilities, and introduces type-specific column encodings for PAX tables.
Data federation and lakehouse integration: Manages Apache Iceberg tables as native tables with full
INSERT,UPDATE, andDELETEsupport, works with the catalog and storage system you already use (Hive Metastore, Apache Polaris, Hadoop, S3, or a built-in catalog), and automatically compacts data in the background so query performance stays stable over time.Query processing and optimization: Speeds up recurring multi-table analytical queries by reusing materialized views, accelerates common aggregations and top-N queries through vectorized execution improvements, and runs more join types in parallel for faster large-scale queries.
Storage: Introduces type-specific column encodings (
deltadelta,gorilla,bool) for PAX tables to compress time-series data.
New features
Database
Category |
Feature |
User documents |
|---|---|---|
Data federation and lakehouse integration |
Manages Apache Iceberg tables as native relations through new |
|
Data federation and lakehouse integration |
Supports multiple Iceberg catalog backends (builtin, Hive Metastore, Apache Polaris, Hadoop, and S3) and S3-compatible or HDFS storage volumes for Iceberg tables. |
|
Data federation and lakehouse integration |
Compacts Iceberg data files through |
|
Query processing and optimization |
Multi-table JOIN exact-match rewrite for AQUMV. |
|
Query processing and optimization |
Vectorized hash aggregation enhancements. |
|
Query processing and optimization |
Vectorized scan and sort optimizations for top-N queries on PAX tables. |
|
Query processing and optimization |
GPORCA parallel outer hash join and streaming hash aggregate enhancements. |
|
Query processing and optimization |
Parallel outer hash joins in the Postgres query optimizer. |
|
Storage |
Type-specific column encodings for PAX tables. |
New feature details
Data federation and lakehouse integration
Native Apache Iceberg tables: SynxDB manages Apache Iceberg tables as native relations instead of foreign tables, so
CREATE,SELECT,INSERT,UPDATE,DELETE, andVACUUMrun directly against Iceberg tables through the native executor. AFOREIGN CATALOGresolves Iceberg metadata, aFOREIGN VOLUMEaccesses the underlying storage, and anICEBERG TABLEbinds the two into a queryable relation.INSERT,UPDATE, andDELETEuse Iceberg v2 merge-on-read semantics, producing position-delete records instead of rewriting data files. As a result, you can run ACID reads and writes against your Iceberg data lake directly from SynxDB, instead of treating it as read-only or routing every write through Spark or Trino.Multiple Iceberg catalog and storage backends: A
FOREIGN CATALOGcan resolve Iceberg metadata through a builtin catalog, a Hive Metastore, an Apache Polaris REST catalog, or a Hadoop or S3 catalog, and aFOREIGN VOLUMEreads and writes data on S3-compatible storage or HDFS. Session-level defaults (iceberg_default_catalog,iceberg_default_volume) letCREATE ICEBERG TABLEomit these clauses. This lets you connect SynxDB to an Iceberg data lake you already share with Spark, Trino, or Flink, regardless of catalog or storage backend, without migrating metadata or data.See Create and Manage Apache Iceberg Tables - Configure a foreign catalog and Create and Manage Apache Iceberg Tables - Configure a foreign volume.
Iceberg VACUUM compaction and autovacuum: Running
VACUUMon an Iceberg table compacts small data files and reduces the position-delete accumulation from merge-on-read updates and deletes. Compaction thresholds are tunable at the session level, and a background autovacuum worker, enabled by default, checks tables and triggers compaction automatically. This keeps query performance on frequently updated Iceberg tables from degrading, without requiring you to script and schedule compaction jobs yourself.See Create and Manage Apache Iceberg Tables - Compact tables with VACUUM.
Query processing and optimization
Multi-table JOIN exact-match rewrite for AQUMV: Answer Query Using Materialized Views (AQUMV) now rewrites multi-table JOIN queries. When a query is structurally identical to a populated, up-to-date materialized view, the planner reads from the view instead of recomputing the join. This lets repeated multi-table analytical queries, such as a recurring dashboard query, reuse pre-computed results without requiring you to rewrite the query or maintain a custom rewrite rule.
See Use automatic materialized views for query optimization.
Vectorized hash aggregation enhancements: The vectorized executor improves hash aggregation in the following areas:
Sonic and Normal hash aggregate engines: Hash aggregation now runs on a faster default engine (reported as
SonicinEXPLAIN VERBOSE) for common analytic aggregations, falling back to the legacyNormalengine otherwise. Both return identical results, soGROUP BYaggregations run faster with no query changes required. Check theVec HashAgg Method:line inEXPLAIN VERBOSEto see which engine handled a query. See Hash aggregate execution engines.date_truncvectorization:date_trunc(unit, timestamp)now runs through the vectorized engine when the column istimestampandunitis a constant string literal, falling back to row-based execution otherwise with identical results. Time-bucketed rollups, such as aggregating by day or month for a dashboard, pick up this speedup automatically. See Features supported by vectorization.Limit+HashAgg early stop: When a
GROUP BYfeeds directly intoLIMITwithout anORDER BYin between, the executor stops aggregating after the first N groups instead of building the full hash table, delivering roughly a 3x speedup on top-N-by-group workloads such as TPC-H Q18. Control the threshold withvector.limit_hashagg_max_total. See Speed up top-N group queries with Limit+HashAgg early stop.Direct-send routing for redistribute Motion: For distributed queries that redistribute a hash aggregation’s result, the executor can forward output batches to the Motion operator without rehashing every row, reducing per-row overhead and latency on queries that span segments, such as a
GROUP BYover a large fact table. Enable it withvector.sonic_motion_direct_send. See Reduce cross-segment latency with direct-send routing.
Vectorized scan and sort optimizations for top-N queries: The vectorized executor speeds up
ORDER BY ... LIMIT Nqueries on PAX tables in the following areas:TopK operator for
ORDER BY ... LIMIT N: Instead of fully sorting the input, the executor maintains a bounded heap of the best N rows, so it never materializes a full sorted set. It applies whenLIMIT + OFFSETis withinvector.topk_bound_threshold(default2000); larger values, or a query withOFFSET, fall back to a full sort. This speeds up common “top N rows” queries, such as the 10 most recent orders, without sorting the entire table. Look forVec Sort Method: TopKinEXPLAINto confirm. See Speed up ORDER BY … LIMIT N with TopK.PAX row-group skipping for TopK queries: When
vector.topk_runtime_filteris on (the default), the executor pushes theTopKheap’s current worst value down to the PAX scan as a threshold, so the scan skips any row group whose min/max statistics cannot improve the top N result. This reduces I/O forORDER BY ... LIMIT Nqueries on large PAX tables. See Skip PAX row groups for TopK queries.PAX fast filter with two-phase column read: For simple predicates on wide PAX tables, the vectorized scan reads only the filter columns first, evaluates the predicate natively, and reads remaining output columns only for rows that pass, skipping a group’s output columns entirely when no row passes. The savings grow with table width and predicate selectivity, so this has the biggest impact on selective queries against wide tables. Controlled by
pax.enable_fast_filter(on by default). See Use fast filter for two-phase column read.
GPORCA parallel outer hash join and streaming hash aggregate enhancements: GPORCA improves parallel query plans in the following areas:
Parallel outer hash joins: GPORCA now generates parallel-aware hash plans for Right Outer and Full Outer joins, in addition to Inner and Left Outer joins, extending intra-segment parallelism to ORCA-planned outer-join workloads. For a
LEFT JOIN, GPORCA can also flip the join into aParallel Hash Right Joinwhen the inner side is smaller, building the hash table on fewer rows. See Parallel hash join.Streaming hash aggregate control: By default, GPORCA uses a streaming hash aggregate for local partial aggregation, avoiding disk spills; this helps queries that aggregate over large datasets or use
GROUP BYwith many distinct values. When the streaming plan doesn’t fit a workload, setoptimizer_use_streaming_hashaggtoofffor a non-streaming aggregate that spills to disk and fully deduplicates. Applies only whenoptimizerison; the Postgres optimizer usesgp_use_streaming_hashagginstead. See Parallel aggregation.
Parallel outer hash joins in the Postgres query optimizer: The Postgres query optimizer now parallelizes
FULL JOINandRIGHT JOINusingParallel Hash Full JoinandParallel Hash Right Joinnodes instead of falling back to serial execution. The join preserves the parallel locus, so aggregations and further joins built on top stay parallel instead of losing it partway through the plan. When the probe side is larger than the build side, the planner builds the hash table on the smaller table.
Storage
Type-specific column encodings for PAX tables: PAX columns can now use encodings tuned to their data type:
deltadeltafor integer, date, and timestamp columns,gorillafor floats, andboolfor booleans, compressing time-series data far more than general compression does.deltadeltasuits timestamps and monotonic counters,gorillasuits slowly changing metrics such as CPU usage or sensor readings, andboolsuits flag columns. Apply per column with theENCODINGclause; writing data in sorted order improves the compression ratio.
Product change information
GUC configuration parameters
Newly added GUCs
The following configuration parameters are added:
iceberg_default_catalog: default''(empty string). Sets the default foreign catalog used byCREATE ICEBERG TABLEwhen the statement omits theCATALOGclause. See Configuration parameters.iceberg_default_volume: default''(empty string). Sets the default foreign volume used byCREATE ICEBERG TABLEwhen the statement omits theVOLUMEclause. See Configuration parameters.datalake.iceberg_autovacuum: defaulton. Enables the background autovacuum worker for Iceberg tables. See Configuration parameters.datalake.iceberg_autovacuum_naptime: default600(s). Sets the interval at which the Iceberg autovacuum worker checks tables. See Configuration parameters.datalake.iceberg_vacuum_compact_min_input_files: default5. Sets the minimum number of small data files required beforeVACUUMcompacts them into a larger file. See Configuration parameters.datalake.iceberg_vacuum_rewrite_target_file_size_mb: default512(MB). Sets the target file size produced by IcebergVACUUMcompaction. See Configuration parameters.datalake.iceberg_postion_deletes_threshold: default100000, range[100000, 10000000]. Sets the maximum number of position-delete records accumulated per Iceberg data file before compaction is triggered. See Configuration parameters.datalake.iceberg_max_compactions_per_vacuum: default100. Sets the maximum number of compaction operations performed within a singleVACUUMon an Iceberg table. See Configuration parameters.datalake.iceberg_max_file_removals_per_vacuum: default100000. Sets the maximum number of orphan or expired files the background deletion queue removes perVACUUMinvocation. See Configuration parameters.datalake.iceberg_max_snapshot_age: default432000(s, 5 days). Sets the maximum retention age for Iceberg snapshots before they become eligible for cleanup by the background deletion queue. See Configuration parameters.datalake.iceberg_log_autovacuum_min_duration: default600000(ms). Logs autovacuum runs that exceed this duration;-1disables logging and0logs every run. See Configuration parameters.datalake.enable_iceberg_fragment_cache: defaulton. Controls whether Iceberg scan fragments (data file split metadata) are cached across queries to reduce metadata lookup overhead. See Configuration parameters.datalake.disable_filter_pushdown: defaultoff. Whenon, disables predicate pushdown fordatalake_fdwexternal tables and Iceberg tables, applying filters in upper plan nodes instead. See Configuration parameters.vector.limit_hashagg_max_total: default1000. Controls the Limit+HashAgg early-stop optimization for vectorized top-N group queries. See Vectorization query computing.vector.sonic_motion_direct_send: defaultoff. Enables direct-send routing between a vectorized hash aggregate and a redistribute Motion. See Vectorization query computing.vector.topk_bound_threshold: default2000. Controls the bounded-heapTopKoptimization for vectorizedORDER BY ... LIMIT Nqueries. See Vectorization query computing.vector.topk_runtime_filter: defaulton. Controls PAX row-group skipping driven by theTopKthreshold. See Vectorization query computing.pax.enable_fast_filter: defaulton. Controls the PAX fast filter with two-phase column read for simple predicates. See Use fast filter for two-phase column read.optimizer_use_streaming_hashagg: defaulton. Controls whether GPORCA uses a streaming hash aggregate for local partial aggregation. See Parallel aggregation.
Changed GUCs
The following configuration parameters are changed:
vector.winagg_spill_work_memis renamed tovector.winagg_spill_memory_mb, with the unit changed from kB to MB, the default changed from0to512, and0now disabling spill-to-disk instead of falling back towork_mem. See Set window aggregate spill memory budget.
Bug fixes and improvements
Data federation and lakehouse integration
Bumped
CATALOG_VERSION_NOfor laketable catalogs so upgraded clusters reject catalog files written by a previous version instead of silently reading a mismatched layout, preventing data corruption after upgrade.Fixed data loss when concurrent
INSERTstatements into an Iceberg table both retried a metadata commit, which could drop one writer’s files from the resulting snapshot.Stopped sharing the
IcebergMetadataFetcheracross concurrent requests indatalake_agent, eliminating metadata corruption and intermittent failures under concurrent Iceberg queries.Fixed a double-free in
datalake_fdwIceberg End hooks triggered when an error occurred during cleanup, eliminating a class of crashes on the Iceberg query error path.Fixed a backend crash at
INSERTplan time whenCREATE FOREIGN TABLEomitted required Iceberg options, replacing it with a clear error naming the missing option.Fixed the
dlproxywrite request to setContent-Type: application/json, restoring compatibility with catalog servers that strictly validate the header.Rejected all
ALTER TABLEsubcommands on Iceberg tables and fixed a crash that could leave the relation in a half-applied state, including a related fix soALTER TABLE DROP COLUMNon the source heap no longer breaks the Parquet writer.Fixed external-catalog Iceberg tables storing the metadata file URI as the table location, which broke later scans.
Fixed an orphaned
pg_dependrow left behind after dropping an Iceberg table underdatalake_fdw.Fixed a failure when an Iceberg table on an S3 catalog received two
UPDATEstatements in the same transaction.Fixed Iceberg
byteavalues being silently truncated at the first NUL byte during write.Fixed an incorrect type-mapping warning for
CHAR(N)columns in Iceberg tables.Fixed
INSERTfailures on Iceberg tables containingTIMESTAMPTZcolumns.Fixed Iceberg merge-on-read
UPDATEwriting empty position-delete entries that broke downstream readers such as Spark and Trino.Fixed silent data corruption for Iceberg columns declared
numeric(p,s)with scale 37 or 38.Fixed Iceberg
SeqScanrewrites droppingSequencenode children from the plan tree.Fixed
UPDATEandDELETEstatements on Iceberg tables that reference cross-table subqueries.Fixed a C++ handle leak in
datalake_fdwIceberg catalog End hooks that slowly leaked resources on long-lived sessions.Fixed a null
catalogHandleinfdwContextafter the context is deleted, closing a use-after-free in thedatalake_fdwerror-recovery path.Required the
datalake_fdwextension forCREATE TABLE ... USING icebergandCREATE ICEBERG TABLE, returning a clear error instead of leaving an unreadable, undroppable relation when the extension is missing.Routed Iceberg commits through
commitAppendfor non-builtin catalogs so the external catalog pointer advances atomically on each commit, ensuring external readers see the new snapshot.Extended the external-catalog commit path in
datalake_agentto coverUPDATE,DELETE, and rewrite operations, matching the coverage already available on the builtin catalog.Added
NULLguards and clearer missing-option errors in the shared Iceberg/Hudi option builder.
Query optimizer and executor
Fixed
ORDER BYbeing lost after the AQUMV join rewrite.Added
C.utf8andC.UTF-8to the vectorized sort collation whitelist.Fixed an attnum-index mismatch in column-specific
ANALYZEthat produced incorrect per-segment NDV statistics.Fixed a hang in vectorized execution when
ShareInputScancrosses slice boundaries.Fixed
CREATE TABLE ... LIKE ... INCLUDING INDEXESproducing duplicated distribution keys on the new table.Made ORCA fall back to the PostgreSQL planner for queries with KNN
ORDER BY, because ORCA lacks support for the operator.Fixed ORCA incorrectly decorrelating
GROUP BY () HAVING <outer_ref>queries, which produced wrong results.Initialized previously uninitialized
PlannedStmtfields in ORCA.Fixed ORCA misdetecting mixed storage in partitioned tables whose partitions include foreign tables.
Set the
FRAMEOPTION_BETWEENflag on ORCA-generated window frames so downstream consumers interpret them correctly.Fixed a use-after-free in
flatten_join_alias_var_optimizerthat could crash ORCA under specific query shapes.Fixed ORCA picking the wrong column type for
CREATE TABLE AS SELECTplans that containUNION ALL.Kept the numeric output path for
sum(bigint), fixing a crash in vectorizedSUMaggregates overBIGINTcolumns.Fixed vectorized Motion hashing on
TIMESTAMP,TIMESTAMPTZ, andTIMEcolumns, which previously produced inconsistent hashes between segments.Fixed incorrect row counts on queries whose
GROUP BYtarget list contains only constant expressions.Eliminated redundant derived
GROUP BYexpressions before execution, speeding up affected grouping queries.
Storage and access methods
Fixed
aoco_relation_size()readingpg_aocssegwith the wrong snapshot, which could return stale or inconsistent sizes.Fixed concurrent
palloc/pfreecalls in the PAX TopK runtime filter that could crash under parallel scan.Fixed a
SIGSEGVinfsm_extendwhen vacuuming a table stored in a non-default tablespace.Fixed a
SIGSEGVon segments when creating an in-place tablespace.Fixed typos and added MPP support for the
allow_inplace_tablespaceGUC.
Processes and concurrency
Fixed a session lock leak in
ALTER DATABASE ... SET TABLESPACEwhen run in utility mode.Fixed a
SIGSEGVingetCdbComponentInfo()when the standby coordinator is deployed on a dedicated host rather than co-located with a primary.
Security
Backported the upstream libpq fix to bail out immediately on SSL/GSS negotiation errors instead of retrying with a downgraded protocol, closing a downgrade-attack vector during connection negotiation.
Tools and utilities
Fixed duplicate counting of
num_executedingp_toolkit.gp_resgroup_status.Fixed a
SyntaxWarningunder Python 3.12 inorphaned_toast_tables_check.py.Fixed stale
errnohandling ininitdb’ssetup_cdb_schema().Fixed
pg_dump/restore failures when expression indexes reference nested SQL functions.Advanced old-cluster checkpoint counters to new-cluster values during
pg_upgradeso post-upgrade WAL replay starts from a consistent point.Forward-ported upstream
pg_upgradexid fixes to segments for more reliable upgrades on clusters with high xid usage.Froze coordinator data after relfilenode transfer during
pg_upgradeto prevent xid-wraparound issues post-upgrade.Invalidated BRIN indexes on AO/CO tables after
pg_upgradeso they are rebuilt instead of returning stale summaries.Preserved
gp_fastsequencevalues acrosspg_upgrade, preventing duplicate-row or sequence-skip issues on AO/CO tables after upgrade.Fixed a syntax error in a packaged bash script that handles
LD_LIBRARY_PATH.Fixed
COPY FROMdouble-counting encoding errors and enabled single-row error handling for transcoding errors.Improved
psqlSQL tab-completion for resource group commands.
Observability
Fixed places that displayed Oids incorrectly in log and error messages.
Widened
MotionLayerStatestat counters fromuint32touint64to prevent overflow on long-running queries with high motion volume.