[GLUTEN-12474][CORE] Preserve V1 write ordering for dynamic partition writes#12514
Draft
wForget wants to merge 6 commits into
Draft
[GLUTEN-12474][CORE] Preserve V1 write ordering for dynamic partition writes#12514wForget wants to merge 6 commits into
wForget wants to merge 6 commits into
Conversation
|
Run Gluten Clickhouse CI on x86 |
Member
Author
|
Successfully reproduced this issue, https://github.com/apache/gluten/actions/runs/29396675906/job/87292215408?pr=12514 |
|
Run Gluten Clickhouse CI on x86 |
1 similar comment
|
Run Gluten Clickhouse CI on x86 |
|
Run Gluten Clickhouse CI on x86 |
1 similar comment
|
Run Gluten Clickhouse CI on x86 |
Member
Author
|
|
Run Gluten Clickhouse CI on x86 |
|
Run Gluten Clickhouse CI on x86 |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What changes are proposed in this pull request?
Fixes #12474.
Spark V1 writes prepare an ordering for partition and bucket columns before physical planning.
WriteFilesExecassumes this ordering is preserved and therefore does not expose it throughrequiredChildOrdering.Gluten may replace
SortAggregateExecwith a hash aggregate, which invalidates the ordering prepared by Spark. For dynamic partition writes using a single output writer, rows from the same partition may then become non-contiguous. This can cause the writer to reopen an existing output file and fail withFileAlreadyExistsException.This PR:
WriteFilesExecinEnsureLocalSortRequirements;V1WritesUtils.getSortOrderfor Spark 3.4 through 4.1, while retaining empty ordering for Spark 3.3;How was this patch tested?
Added a regression test to
GlutenV1WriteCommandSuitefor Spark 3.4, 3.5, 4.0, and 4.1.The test performs a dynamic partition overwrite after an aggregation with:
spark.sql.maxConcurrentOutputFileWriters=0spark.sql.sources.partitionOverwriteMode=DYNAMICIt verifies that the write completes successfully and that the target table contains the expected rows. The test was also confirmed to reproduce the original write failure before applying the fix.
Was this patch authored or co-authored using generative AI tooling?
Generated-by: OpenAI Codex (GPT-5)