feat: support array_insert #1073

SemyonSinchenko · 2024-11-09T16:13:44Z

Which issue does this PR close?

Related to #1042

array_insert: SELECT array_insert(array(1, 2, 3, 4), 5, 5)

Rationale for this change

As described in #1042

What changes are included in this PR?

QueryPlanSerde.scala: I added an additional case for the array insert;
expr.proto: I added a new message for the ArrayInsert;
planner.rs: I added a case for the array_insert;
list.rs:
- I added a new ArrayInsert struct;
- I implemented PhysicalExpr, Display and PartialExpr for it;
- The main logic of insertion is in fn array_insert

How are these changes tested?

At the moment I added a simple tests for fn array_insert and a test for QueryPlanSerde.

SemyonSinchenko · 2024-11-09T17:47:44Z

As I was able to realize, array_insert does not supported in datafusion. Is the list.rs a good place to have an implementation of ArrayInsert and PhysicalExpr for it?

SemyonSinchenko · 2024-11-11T14:34:56Z

@andygrove Sorry for tagging but I have questions about the ticket (array_insert).

[RESOLVED] array_insert was added in spark 3.4, so all the 3.3.x tests are obviously failed. I checked and it looks like the EoL for 3.3 is about the end of 2024. Technically I think I can try to workaround tests in 3.3.x by reflection API, my question is mostly should I do it due to soon EoL of the 3.3.x?
array_insert is not supported in DataFusion. I made an implementation (and it looks like it works, except negative indices and corner cases). Is the list.rs a good place for it? Or should I move my code somewhere else?
Spark does not support anything except Int32 for position argument, is it OK if I will support only int32 too? In theory, other types can be supported too, but I'm still trying to realize how to achieve it and it may become complex...

Thanks in advance! That is my first serious attempt to contribute to the project, so sorry If I'm annoying.

+ fix tests for spark < 3.4

- added test for the negative index - added test for the legacy spark mode

spark/src/test/scala/org/apache/comet/CometExpressionSuite.scala

andygrove · 2024-11-13T22:00:06Z

Thanks, @SemyonSinchenko. I think it's fine to skip the test for Spark 3.3. I plan on reviewing this PR in more detail tomorrow, but it looks good from an initial read.

codecov-commenter · 2024-11-14T06:58:20Z

Codecov Report

Attention: Patch coverage is 59.09091% with 9 lines in your changes missing coverage. Please review.

Project coverage is 34.27%. Comparing base (845b654) to head (8b58d8d).
Report is 18 commits behind head on main.

Files with missing lines	Patch %	Lines
.../scala/org/apache/comet/serde/QueryPlanSerde.scala	59.09%	7 Missing and 2 partials ⚠️

Additional details and impacted files

@@             Coverage Diff              @@
##               main    #1073      +/-   ##
============================================
- Coverage     34.46%   34.27%   -0.20%     
  Complexity      888      888              
============================================
  Files           113      113              
  Lines         43580    43355     -225     
  Branches       9658     9488     -170     
============================================
- Hits          15021    14860     -161     
- Misses        25507    25596      +89     
+ Partials       3052     2899     -153

☔ View full report in Codecov by Sentry.
📢 Have feedback on the report? Share it here.

SemyonSinchenko · 2024-11-16T13:42:15Z

This PR is ready for the review. The only failed check failed due to the internal GHA error:

GitHub Actions has encountered an internal error when running your job.

native/spark-expr/src/list.rs

spark/src/test/scala/org/apache/comet/CometExpressionSuite.scala

Co-authored-by: Andy Grove <[email protected]>

NoeB · 2024-11-19T07:34:09Z

I am not sure if this should be done together with this PR but it would some add "free" tests. Spark introduced with 3.5 array_prepend which it implements with array_insert and starting 4.0 it also implements array_append with array_insert. If you want you can copy the array_append tests and replace array_append with array_prepend for Spark 3.5+ and enable the array_append tests for spark 4.0. You can also ignore this comment if you do not agree or if it leads to unrelated errors.

- fixes; - tests; - comments in the code;

In one case there is a zero in index and test fails due to spark error

SemyonSinchenko · 2024-11-20T18:01:58Z

Thanks for comments and suggestions!

This PR is ready for the review again.

What were changed from the last round of the review:

I added tests for the array_prepend (spark 3.5+) and enabled the test for array_append (spark 4.0+);
I added tests for the corner cases:
- Index is negative;
- Index is bigger than the array length;
- Index is negative and it's abs is greater than the array length;
- Test of the fallback to spark (udf as child);
- Value to insert is actually null;
I fixed the native part behavior for some of the corner cases;

These tests pointed my attention to the uncovered parts of the code that I fixed and I also added couple of additional tests to the native part. I revisited how spark do array_insert and it is more tricky than I realized at the first glance. I added few additional comments to the native part of the implementation and fixed the behavior.

At the moment this PR is tested in multiple ways:

Basic tests in native part that are useful because easy to debug and very fast to run;
Basic tests in native part for the so called "legacy mode" in spark;
Basic tests on the scala side;
Logical tests on small data for corner cases (negative, positive, long, short, null, etc.) on the scala side;
Tests in array_prepend and array_append that are calling array_insert under the hood (NULLS, different data types, etc.);
Test of the fallback to the Spark in case when one of children is not supported by Comet;

So, it looks to me, that all the possible cases are covered and the behavior is the same like in spark.

SemyonSinchenko · 2024-11-20T18:02:14Z

I am not sure if this should be done together with this PR but it would some add "free" tests. Spark introduced with 3.5 array_prepend which it implements with array_insert and starting 4.0 it also implements array_append with array_insert. If you want you can copy the array_append tests and replace array_append with array_prepend for Spark 3.5+ and enable the array_append tests for spark 4.0. You can also ignore this comment if you do not agree or if it leads to unrelated errors.

Done!

andygrove · 2024-11-20T22:45:58Z

native/spark-expr/src/list.rs

+        let src_element_type = match src_value.data_type() {
+            DataType::List(field) => field.data_type(),
+            DataType::LargeList(field) => field.data_type(),
+            data_type => {
+                return Err(DataFusionError::Internal(format!(
+                    "Unexpected src array type in ArrayInsert: {:?}",
+                    data_type
+                )))
+            }


minor nit: this logic for extracting a list type is repeated a few times and could be factored out into a function

@andygrove Thanks for the suggestion!
I moved a checking of the array type (and the exception logic) to the method:

pub fn array_type(&self, data_type: &DataType) -> DataFusionResult<DataType> { match data_type { DataType::List(field) => Ok(DataType::List(Arc::clone(field))), DataType::LargeList(field) => Ok(DataType::LargeList(Arc::clone(field))), data_type => { return Err(DataFusionError::Internal(format!( "Unexpected src array type in ArrayInsert: {:?}", data_type ))) } } }

It allows at least to avoid returning the same error multiple time. Is it what you suggested? Or should I move this method to a helper function and refactor also GerArrayStructField to use such a function?

P.S. Sorry for the stupid question... But can you please explain to me why we always check both List and LargeList, while Apache Spark only supports i32 indexes for arrays (max length is Integer.MAX_VALUE - 15), which is the case of List to my understanding? All the code in the list.rs might become a bit simpler if we make it non-generic (it also makes implementation of other missing methods like array_zip simpler).

That's a good question. I wonder if the existing code for LargeList is actually being tested. It would be interesting to try removing it and see if there are any regressions. It makes sense to only handle List if Spark only supports i32 indexes.

andygrove

LGTM. Thanks @SemyonSinchenko!

Part of the implementation of array_insert

f583e5c

SemyonSinchenko added 5 commits November 11, 2024 09:11

Missing methods

e870c21

Working version

ac7a2b3

Reformat code

9d9518e

Fix code-style

6e0d5f4

Add comments about spark's implementation.

e4b5e4c

SemyonSinchenko added 2 commits November 13, 2024 12:34

Implement negative indices

19230bf

+ fix tests for spark < 3.4

Fix code-style

58ecb82

SemyonSinchenko changed the title ~~[WIP][DO-NOT-MERGE] feat: support array_insert~~ [WIP] feat: support array_insert Nov 13, 2024

SemyonSinchenko changed the title ~~[WIP] feat: support array_insert~~ feat: support array_insert Nov 13, 2024

Fix scalastyle

a248567

SemyonSinchenko changed the title ~~feat: support array_insert~~ [WIP] feat: support array_insert Nov 13, 2024

SemyonSinchenko added 2 commits November 13, 2024 13:13

Fix tests for spark < 3.4

e4349f5

Fixes & tests

0d38ef0

- added test for the negative index - added test for the legacy spark mode

SemyonSinchenko changed the title ~~[WIP] feat: support array_insert~~ feat: support array_insert Nov 13, 2024

SemyonSinchenko marked this pull request as ready for review November 13, 2024 18:04

andygrove reviewed Nov 13, 2024

View reviewed changes

spark/src/test/scala/org/apache/comet/CometExpressionSuite.scala Show resolved Hide resolved

SemyonSinchenko added 2 commits November 14, 2024 06:55

Use assume(isSpark34Plus) in tests

c7f26f9

Merge remote-tracking branch 'refs/remotes/origin/main'

8b58d8d

Test else-branch & improve coverage

f832cf0

andygrove reviewed Nov 18, 2024

View reviewed changes

native/spark-expr/src/list.rs Outdated Show resolved Hide resolved

andygrove reviewed Nov 18, 2024

View reviewed changes

spark/src/test/scala/org/apache/comet/CometExpressionSuite.scala Outdated Show resolved Hide resolved

Update native/spark-expr/src/list.rs

6e41858

Co-authored-by: Andy Grove <[email protected]>

Merge main + add tests

659ab7a

- fixes; - tests; - comments in the code;

SemyonSinchenko added 2 commits November 19, 2024 19:33

Fix fallback test

4770fce

In one case there is a zero in index and test fails due to spark error

Adjust the behaviour for the NULL case to Spark

e9ef941

SemyonSinchenko closed this Nov 20, 2024

SemyonSinchenko reopened this Nov 20, 2024

SemyonSinchenko requested a review from andygrove November 20, 2024 18:02

andygrove reviewed Nov 20, 2024

View reviewed changes

andygrove approved these changes Nov 20, 2024

View reviewed changes

SemyonSinchenko added 2 commits November 21, 2024 10:25

Move the logic of type checking to the method

6431ad9

Fix code-style

e02d20f

andygrove merged commit 9990b34 into apache:main Nov 22, 2024
74 checks passed

SemyonSinchenko deleted the array-insert branch November 23, 2024 09:02

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

feat: support array_insert #1073

feat: support array_insert #1073

SemyonSinchenko commented Nov 9, 2024 •

edited

Loading

SemyonSinchenko commented Nov 9, 2024

SemyonSinchenko commented Nov 11, 2024 •

edited

Loading

andygrove commented Nov 13, 2024

codecov-commenter commented Nov 14, 2024 •

edited

Loading

SemyonSinchenko commented Nov 16, 2024

NoeB commented Nov 19, 2024

SemyonSinchenko commented Nov 20, 2024

SemyonSinchenko commented Nov 20, 2024

andygrove Nov 20, 2024

SemyonSinchenko Nov 21, 2024 •

edited

Loading

andygrove Nov 22, 2024

andygrove left a comment

feat: support array_insert #1073

feat: support array_insert #1073

Conversation

SemyonSinchenko commented Nov 9, 2024 • edited Loading

Which issue does this PR close?

Rationale for this change

What changes are included in this PR?

How are these changes tested?

SemyonSinchenko commented Nov 9, 2024

SemyonSinchenko commented Nov 11, 2024 • edited Loading

andygrove commented Nov 13, 2024

codecov-commenter commented Nov 14, 2024 • edited Loading

Codecov Report

SemyonSinchenko commented Nov 16, 2024

NoeB commented Nov 19, 2024

SemyonSinchenko commented Nov 20, 2024

SemyonSinchenko commented Nov 20, 2024

andygrove Nov 20, 2024

Choose a reason for hiding this comment

SemyonSinchenko Nov 21, 2024 • edited Loading

Choose a reason for hiding this comment

andygrove Nov 22, 2024

Choose a reason for hiding this comment

andygrove left a comment

Choose a reason for hiding this comment

SemyonSinchenko commented Nov 9, 2024 •

edited

Loading

SemyonSinchenko commented Nov 11, 2024 •

edited

Loading

codecov-commenter commented Nov 14, 2024 •

edited

Loading

SemyonSinchenko Nov 21, 2024 •

edited

Loading