[SPARK-48308][CORE][3.5] Unify getting data schema without partition columns in FileSourceStrategy #47483

vkorukanti · 2024-07-25T06:44:08Z

What changes were proposed in this pull request?

(Cherry-pick of 57948c8 to branch-3.5)

Compute the schema of the data without partition columns only once in FileSourceStrategy.

Why are the changes needed?

In FileSourceStrategy, the schema of the data excluding partition columns is computed 2 times in a slightly different way, using an AttributeSet (partitionSet) and using the attributes directly (partitionColumns) These don't have the exact same semantics, AttributeSet will only use expression ids for comparison while comparing with the actual attributes will use the name, type, nullability and metadata. We want to use the former here.

Does this PR introduce any user-facing change?

No

How was this patch tested?

Existing tests

Was this patch authored or co-authored using generative AI tooling?

No

Authored-by: Johan Lasperas [email protected]

…columns in FileSourceStrategy Compute the schema of the data without partition columns only once in FileSourceStrategy. In FileSourceStrategy, the schema of the data excluding partition columns is computed 2 times in a slightly different way, using an AttributeSet (`partitionSet`) and using the attributes directly (`partitionColumns`) These don't have the exact same semantics, AttributeSet will only use expression ids for comparison while comparing with the actual attributes will use the name, type, nullability and metadata. We want to use the former here. No Existing tests No Closes apache#46619 from johanl-db/reuse-schema-without-partition-columns. Authored-by: Johan Lasperas <[email protected]> Signed-off-by: Wenchen Fan <[email protected]>

cloud-fan · 2024-07-25T08:32:08Z

This fixes a regression caused by https://github.com/apache/spark/pull/46565/files#diff-fbc6da30b8372e4f9aeb35ccf0d39eb796715d192c7eaeab109376584de0790eR121 , and make Delta Lake pass all tests. I think we should include it in 3.5.2. cc @yaooqinn

yaooqinn · 2024-07-25T08:36:59Z

Do we need this for branch 3.4?

…columns in FileSourceStrategy ### What changes were proposed in this pull request? (Cherry-pick of 57948c8 to branch-3.5) Compute the schema of the data without partition columns only once in FileSourceStrategy. ### Why are the changes needed? In FileSourceStrategy, the schema of the data excluding partition columns is computed 2 times in a slightly different way, using an AttributeSet (`partitionSet`) and using the attributes directly (`partitionColumns`) These don't have the exact same semantics, AttributeSet will only use expression ids for comparison while comparing with the actual attributes will use the name, type, nullability and metadata. We want to use the former here. ### Does this PR introduce _any_ user-facing change? No ### How was this patch tested? Existing tests ### Was this patch authored or co-authored using generative AI tooling? No Authored-by: Johan Lasperas <johan.lasperasdatabricks.com> Closes #47483 from vkorukanti/partitionCols. Authored-by: Johan Lasperas <[email protected]> Signed-off-by: Kent Yao <[email protected]>

yaooqinn · 2024-07-25T08:52:43Z

Merged to branch-3.5, thank you all

cloud-fan · 2024-07-25T09:31:27Z

It's fine to skip 3.4 as #46565 was not merged to 3.4 either.

yaooqinn · 2024-07-25T09:38:03Z

Thank you @cloud-fan

dongjoon-hyun

It's great to find and fix this during RC2 vote period. Thank you, @vkorukanti , @cloud-fan , @yaooqinn .

…columns in FileSourceStrategy (apache#538) ### What changes were proposed in this pull request? (Cherry-pick of 57948c8 to branch-3.5) Compute the schema of the data without partition columns only once in FileSourceStrategy. ### Why are the changes needed? In FileSourceStrategy, the schema of the data excluding partition columns is computed 2 times in a slightly different way, using an AttributeSet (`partitionSet`) and using the attributes directly (`partitionColumns`) These don't have the exact same semantics, AttributeSet will only use expression ids for comparison while comparing with the actual attributes will use the name, type, nullability and metadata. We want to use the former here. ### Does this PR introduce _any_ user-facing change? No ### How was this patch tested? Existing tests ### Was this patch authored or co-authored using generative AI tooling? No Authored-by: Johan Lasperas <johan.lasperasdatabricks.com> Closes apache#47483 from vkorukanti/partitionCols. Authored-by: Johan Lasperas <[email protected]> Signed-off-by: Kent Yao <[email protected]> Co-authored-by: Johan Lasperas <[email protected]>

github-actions bot added the SQL label Jul 25, 2024

cloud-fan approved these changes Jul 25, 2024

View reviewed changes

yaooqinn approved these changes Jul 25, 2024

View reviewed changes

yaooqinn closed this Jul 25, 2024

dongjoon-hyun reviewed Jul 25, 2024

View reviewed changes

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Uh oh!

[SPARK-48308][CORE][3.5] Unify getting data schema without partition columns in FileSourceStrategy #47483

[SPARK-48308][CORE][3.5] Unify getting data schema without partition columns in FileSourceStrategy #47483

Uh oh!

vkorukanti commented Jul 25, 2024 •

edited

Loading

Uh oh!

cloud-fan commented Jul 25, 2024

Uh oh!

yaooqinn commented Jul 25, 2024

Uh oh!

yaooqinn commented Jul 25, 2024

Uh oh!

cloud-fan commented Jul 25, 2024

Uh oh!

yaooqinn commented Jul 25, 2024

Uh oh!

dongjoon-hyun left a comment

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

5 participants

[SPARK-48308][CORE][3.5] Unify getting data schema without partition columns in FileSourceStrategy #47483

[SPARK-48308][CORE][3.5] Unify getting data schema without partition columns in FileSourceStrategy #47483

Uh oh!

Conversation

vkorukanti commented Jul 25, 2024 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

What changes were proposed in this pull request?

Why are the changes needed?

Does this PR introduce any user-facing change?

How was this patch tested?

Was this patch authored or co-authored using generative AI tooling?

Uh oh!

cloud-fan commented Jul 25, 2024

Uh oh!

yaooqinn commented Jul 25, 2024

Uh oh!

yaooqinn commented Jul 25, 2024

Uh oh!

cloud-fan commented Jul 25, 2024

Uh oh!

yaooqinn commented Jul 25, 2024

Uh oh!

dongjoon-hyun left a comment

Choose a reason for hiding this comment

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

5 participants

vkorukanti commented Jul 25, 2024 •

edited

Loading