Copy-on-write (COW) tables store data in Parquet files. Internal update operations need to be performed by rewriting the original Parquet files.
Merge-on-read (MOR) tables store data in a hybrid format combining columnar-based Parquet and row-based format Avro. Parquet files are used to store base data, and Avro files (also called log files) are used to store incremental data.
Trade-off | CopyOnWrite | MergeOnRead |
|---|---|---|
Data latency | High | Low |
Query latency | Low | High |
Update cost (I/O) | High (rewriting the entire Parquet file) | Low |
Parquet file size | Small (high update cost) | Large (low update cost) |
Write amplification | High | Low (depending on the compaction policy) |
Compute Engine | Version | Hudi Version |
|---|---|---|
Spark | 3.3.1 | 0.11.0 |
Flink | 1.15 | 0.11.0 |
HetuEngine | 2.1.0 | 0.11.0 |
To determine the compute engine versions supported by a queue, follow these steps: Log in to the DLI console and choose Resources > Queue Management in the navigation pane on the left. Locate the queue you want to query on the queue management page. Click the icon next to queue name to expand its details. Look for the supported versions in Supported Versions. For SQL queues, you cannot switch versions, so checking the default version will indicate the current compute engine version in use.