Skip to content

[GLUTEN-12538][VL] Unblock TIMESTAMP_NTZ min/max in Delta statistics - #12967

Open
felipepessoto wants to merge 8 commits into
apache:mainfrom
felipepessoto:gluten-12538-timestamp-ntz-sort-aggregate
Open

[GLUTEN-12538][VL] Unblock TIMESTAMP_NTZ min/max in Delta statistics#12967
felipepessoto wants to merge 8 commits into
apache:mainfrom
felipepessoto:gluten-12538-timestamp-ntz-sort-aggregate

Conversation

@felipepessoto

@felipepessoto felipepessoto commented Sep 4, 2026

Copy link
Copy Markdown
Contributor

What changes are proposed in this pull request?

Enable native min/max aggregation for Spark TIMESTAMP_NTZ in the Velox backend, including the task-local SortAggregateExec used to collect Delta write statistics.

The implementation reuses Velox's existing Spark timestamp min/max kernels rather than introducing new aggregate kernels. TIMESTAMP_NTZ remains represented by the distinct Velox TIMESTAMP_UTC logical type, with microsecond precision (Velox implementation, Spark registration).

  • Accept TimestampNTZType in aggregate buffer/result and grouping type checks, and allow aggregate, shuffle, and the projection expressions needed by Delta statistics through the coarse NTZ fallback validator. Existing native function validation remains in place.
  • Use tsntz as the internal function-signature token. The previous ts_ntz token collides with _, which the native signature parser uses to separate argument types (parser).
  • Preserve TIMESTAMP_UTC as Substrait PrecisionTimestamp with precision 6 during type round-tripping.
  • Remove the 42 now-passing TIMESTAMP_NTZ data-skipping cases from the Delta known-failure baseline.

This addresses the NTZ aggregation failure encountered by the Delta statistics tracker in #12538. It does not implement a general fallback for arbitrary unsupported statistics plans or claim support for every NTZ expression. DATE/TIMESTAMP_NTZ cast support is handled separately in #12966.

How was this patch tested?

  • C++ signature, type, and min/max plan round-trip tests verify distinct NTZ type identity, Substrait precision 6, and preservation of microsecond values using Spark min/max.
  • GlutenTimestampNtzAggregateSuite on Spark 4.1 checks native SQL min/max execution and microsecond results, plus Spark fallback for NTZ JSON serialization in a non-UTC session timezone.
  • GlutenDeltaStatsSuite on Spark 3.5/Delta 3.3 and Spark 4.1/Delta 4.x checks writes and read-back of top-level and nested NTZ columns, with exact JSON-path assertions for their min/max statistics.
  • Existing DateFunctionsValidateSuite coverage exercises NTZ scalar functions after the signature-token change.
  • The Delta Spark UT workflow confirms the 42 data-skipping cases pass; unrelated type-widening failures remain in the baseline.

Was this patch authored or co-authored using generative AI tooling?

Generated-by: GitHub Copilot CLI 1.0.83

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@github-actions github-actions Bot added CORE works for Gluten Core VELOX labels Sep 4, 2026
@github-actions

github-actions Bot commented Sep 4, 2026

Copy link
Copy Markdown

Run Gluten Clickhouse CI on x86

@github-actions

github-actions Bot commented Sep 4, 2026

Copy link
Copy Markdown

🔄 Delta Spark UT started by @felipepessoto (~2.5 h). View run

@github-actions

github-actions Bot commented Sep 4, 2026

Copy link
Copy Markdown

Run Gluten Clickhouse CI on x86

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@github-actions

github-actions Bot commented Sep 4, 2026

Copy link
Copy Markdown

Run Gluten Clickhouse CI on x86

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@github-actions

github-actions Bot commented Sep 4, 2026

Copy link
Copy Markdown

Run Gluten Clickhouse CI on x86

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@github-actions github-actions Bot added the INFRA label Sep 4, 2026
@github-actions

github-actions Bot commented Sep 4, 2026

Copy link
Copy Markdown

Run Gluten Clickhouse CI on x86

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
@github-actions

github-actions Bot commented Sep 4, 2026

Copy link
Copy Markdown

Run Gluten Clickhouse CI on x86

@felipepessoto
felipepessoto marked this pull request as ready for review September 4, 2026 09:33
Copilot AI lite review requested due to automatic review settings September 4, 2026 09:33

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

A newly added Velox round-trip test uses millisecond-precision timestamps despite declaring microsecond precision, weakening coverage for the intended NTZ precision behavior.

Once you've addressed the issues Copilot identified, you can request another Copilot review.

Pull request overview

Enables native handling of TIMESTAMP_NTZ (Spark) / TIMESTAMP_UTC (Velox) in aggregation and related Delta statistics collection paths for the Velox backend, reducing Spark fallbacks and unblocking Delta stats plans previously rejected due to NTZ types.

Changes:

  • Extend the TimestampNTZ fallback validator to allow aggregates, shuffles, and direct NTZ projections needed by Delta stats plans.
  • Accept TimestampNTZType in aggregation buffer type checks and preserve a distinct native signature token (tsntz) plus Substrait PrecisionTimestamp encoding.
  • Add native regression coverage (Spark UT + Velox C++ tests) and remove now-fixed Delta data-skipping cases from the known-failure baseline.
File summaries
File Description
gluten-ut/spark41/src/test/scala/org/apache/spark/sql/GlutenTimestampNtzAggregateSuite.scala Adds Spark 4.1 regression coverage for NTZ min/max aggregation and a projection fallback case.
gluten-ut/spark41/src/test/scala/org/apache/gluten/utils/velox/VeloxTestSettings.scala Enables the new Spark 4.1 UT suite in Velox test settings.
gluten-substrait/src/main/scala/org/apache/gluten/extension/columnar/validator/Validators.scala Broadens NTZ fallback validation to permit additional plan nodes/expressions used by native stats aggregation.
gluten-substrait/src/main/scala/org/apache/gluten/expression/ConverterUtils.scala Maps Spark TimestampNTZType to Substrait timestamp NTZ type node and uses tsntz in signature naming.
gluten-substrait/src/main/scala/org/apache/gluten/execution/HashAggregateExecBaseTransformer.scala Allows TimestampNTZType in supported aggregation buffer type checks.
cpp/velox/tests/VeloxToSubstraitTypeTest.cc Adds a type-conversion test for TIMESTAMP_UTC -> Substrait PrecisionTimestamp.
cpp/velox/tests/VeloxSubstraitSignatureTest.cc Adds tsntz signature coverage for TIMESTAMP_UTC mapping in both directions.
cpp/velox/tests/VeloxSubstraitRoundTripTest.cc Adds a min/max aggregation round-trip test over TIMESTAMP_UTC.
cpp/velox/substrait/VeloxToSubstraitType.cc Encodes TIMESTAMP_UTC as Substrait precision_timestamp(6).
cpp/velox/substrait/VeloxSubstraitSignature.cc Maps TIMESTAMP_UTC to/from the tsntz signature token.
backends-velox/src-delta40/test/scala/org/apache/spark/sql/delta/GlutenDeltaStatsSuite.scala Adds Delta 4.0 stats regression coverage including nested NTZ columns.
backends-velox/src-delta33/test/scala/org/apache/spark/sql/delta/GlutenDeltaStatsSuite.scala Adds Delta 3.3 stats regression coverage including nested NTZ columns.
.github/workflows/util/delta-spark-ut/known-failures.txt Removes Delta data-skipping known failures resolved by native NTZ stats aggregation support.
Review details
  • Files reviewed: 13/13 changed files
  • Comments generated: 1
  • Review effort level: Lite

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread cpp/velox/tests/VeloxSubstraitRoundTripTest.cc Outdated
@felipepessoto

Copy link
Copy Markdown
Contributor Author

@Mariamalmesfer @rui-mo I think you two have worked on other timestamp_ntz PRs. Could you take a look, please?

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot AI review requested due to automatic review settings September 4, 2026 17:57

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Warning

Copilot couldn't run its full agentic review because it didn't start before the timeout. Make sure your repository has a runner available, or add a copilot-code-review.yml file specifying one with the runs-on attribute. See the docs for more details.

Pull request overview

Copilot reviewed 13 out of 13 changed files in this pull request and generated 6 comments.

Comment thread cpp/velox/tests/VeloxSubstraitRoundTripTest.cc Outdated
Comment thread cpp/velox/substrait/VeloxToSubstraitType.cc Outdated
@github-actions

github-actions Bot commented Sep 4, 2026

Copy link
Copy Markdown

Run Gluten Clickhouse CI on x86

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot AI review requested due to automatic review settings September 4, 2026 19:06
@github-actions

github-actions Bot commented Sep 4, 2026

Copy link
Copy Markdown

Run Gluten Clickhouse CI on x86

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Warning

Copilot couldn't run its full agentic review because it didn't start before the timeout. Make sure your repository has a runner available, or add a copilot-code-review.yml file specifying one with the runs-on attribute. See the docs for more details.

Pull request overview

Copilot reviewed 13 out of 13 changed files in this pull request and generated 2 comments.

Suppressed comments (1)

gluten-substrait/src/main/scala/org/apache/gluten/extension/columnar/validator/Validators.scala:1

  • The previous implementation used dataType.typeName == "timestamp_ntz", which is resilient across Spark versions/shims. Switching to a direct TimestampNTZType reference introduces a compile-time dependency that can break builds for Spark variants where TimestampNTZType is absent or shaded differently. If this module is cross-built across multiple Spark versions, consider keeping a version-tolerant check (e.g., match TimestampNTZType when available and fall back to typeName), ideally via a shim utility.
/*

Comment thread cpp/velox/substrait/VeloxSubstraitSignature.cc
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot AI review requested due to automatic review settings September 4, 2026 20:13
@github-actions

github-actions Bot commented Sep 4, 2026

Copy link
Copy Markdown

Run Gluten Clickhouse CI on x86

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟢 Approval recommended

The changes are cohesive and well-covered by targeted Spark + native regression tests, and the updated validator/signature/type mappings are consistent across JVM and C++ paths.

Review details
  • Files reviewed: 13/13 changed files
  • Comments generated: 0 new
  • Review effort level: Lite

@github-actions

github-actions Bot commented Sep 4, 2026

Copy link
Copy Markdown

Run Gluten Clickhouse CI on x86

@rui-mo rui-mo left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @felipepessoto. Could you please clarify whether this implementation is based on the aggregate implementation for the Timestamp type in Velox, and whether it can be fully reused for the TimestampNTZ type?

"1969-12-31T23:59:59.999",
"2024-01-01T00:00:00.123",
"2024-01-01T00:00:00.123"),
stats)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Could you please assert that the plan contains HashAggregateExecTransformer to ensure aggregate has been offloaded into native?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I agree that an explicit offload assertion would strengthen this regression. The statistics aggregate is constructed locally inside GlutenDeltaJobStatsTracker in each write task and submitted directly to NativePlanEvaluator, so it is not visible in the outer DataFrame's executed plan (code).

The approach I found would require changing production tracker code in both Delta variants to add a test observation hook. The test could then assert that the hook was invoked and that the captured plan contains HashAggregateExecTransformer. Would you be comfortable with that, or do you know a simpler approach, such as an existing way to observe this internal plan from tests without adding a new hook?

return "date";
}
if (type->equivalent(*TIMESTAMP_UTC())) {
return "tsntz";

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: use ts_ntz to be aligned with Spark.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ts_ntz would align visually with Spark, but it cannot represent one type in Gluten's current function-signature grammar. The native parser uses _ as the delimiter between argument types
(parser), so min:ts_ntz is parsed as two argument types, ts and ntz, and aggregate validation fails when it tries to resolve ntz.

tsntz is deliberately an internal, separator-free signature token (mapping). The Spark-facing type remains timestamp_ntz, and the Substrait representation remains PrecisionTimestamp. Using ts_ntz would require first changing the signature grammar to support escaping or structured type tokens.

@felipepessoto

felipepessoto commented Sep 8, 2026

Copy link
Copy Markdown
Contributor Author

Thanks @felipepessoto. Could you please clarify whether this implementation is based on the aggregate implementation for the Timestamp type in Velox, and whether it can be fully reused for the TimestampNTZ type?

@rui-mo, yes, for min and max, this PR reuses Velox's existing timestamp aggregate implementation; it does not add a separate TIMESTAMP_NTZ aggregate kernel (code).

Velox represents Spark TIMESTAMP_NTZ as TIMESTAMP_UTC. TimestampUtcType derives from TimestampType, so it has the same TypeKind::TIMESTAMP and native Timestamp value representation while retaining distinct logical type identity (Velox type). The shared min/max factory dispatches TypeKind::TIMESTAMP to the existing SimpleNumericMin/MaxAggregate<Timestamp> implementation and carries the supplied result type through (factory). Spark's registration configures that implementation with microsecond precision (registration).

This reuse is appropriate for min and max because their kernels compare the stored timestamp values directly, without session-timezone conversion (min comparison, max comparison). Gluten preserves TIMESTAMP_UTC in the argument and result types; it does not cast NTZ values to ordinary TIMESTAMP to reuse the implementation. This is specific to the min/max path used by Delta statistics, not that every timestamp operation can be reused unchanged.

@felipepessoto felipepessoto changed the title [GLUTEN-12538][VL] Support TIMESTAMP_NTZ aggregation [GLUTEN-12538][VL] Unblock TIMESTAMP_NTZ min/max in Delta statistics Sep 8, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CORE works for Gluten Core INFRA VELOX

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants