Skip to content

12728: Fix O(N) database updates during file upload on large datasets - #12729

Open
tuannx wants to merge 4 commits into
IQSS:developfrom
tuannx:perf/large-upload-modificationtime-fix
Open

tuannx wants to merge 4 commits into
IQSS:developfrom
tuannx:perf/large-upload-modificationtime-fix

Conversation

@tuannx

@tuannx tuannx commented Sep 21, 2026 •

Copy link
Copy Markdown

What this PR does / why we need it:
Fixes an $O(N)$ database write bottleneck in UpdateDatasetVersionCommand.java where adding a file or updating a dataset version unconditionally touched modificationTime on all existing files in the dataset.

In datasets with thousands of files (e.g. 14,000+ files), mutating every file's modification timestamp on each upload caused JPA dirty checking to execute an UPDATE DVOBJECT SET MODIFICATIONTIME = ... statement for each existing file on every flush ($\frac{N(N+1)}{2}$ cumulative updates). This caused per-file upload latency to degrade from 1s to over 30s.

The fix ensures UpdateDatasetVersionCommand only initializes timestamps on newly created files (dataFile.getCreateDate() == null), safely handles null timestamps on legacy files, or updates the target file when single-file variable metadata (fmVarMet) is explicitly passed.

Which issue(s) this PR closes:

Special notes for your reviewer:
Includes:

  1. src/main/java/edu/harvard/iq/dataverse/engine/command/impl/UpdateDatasetVersionCommand.java: Targeted fix to avoid dirtying existing files in JPA.
  2. src/test/java/edu/harvard/iq/dataverse/engine/command/impl/UpdateDatasetVersionCommandTest.java: Acceptance tests asserting that:
    • Adding a file does NOT modify existing files' modificationTime (reproduces and fails on unpatched code, passes with fix).
    • Updating variable metadata only touches the target file.
    • Null modification times are safely initialized.
  3. doc/release-notes/12728-perf-large-upload.md: Release notes following issue naming convention.

A standalone Docker reproduction benchmark environment and Python benchmark tool are provided in the companion repo:

Suggestions on how to test this:
Run the JUnit 5 acceptance test:

mvn test -Dtest=UpdateDatasetVersionCommandTest

Or run the full command regression suite:

mvn test -Dtest=UpdateDatasetVersionCommandTest,CreateDatasetVersionCommandTest,AbstractDatasetCommandTest,UpdateDatasetLicenseCommandTest,UpdateDatasetFieldsCommandTest

Does this PR introduce a user interface change? If mockups are available, please link/include them here:
No.

Is there a release notes update needed for this change?:
Yes, added doc/release-notes/12728-perf-large-upload.md.

Additional documentation:

- Prevent UpdateDatasetVersionCommand from modifying modificationTime on all existing files in the dataset.
- Only initialize timestamps on newly created files (createDate == null) or when single-file variable metadata (fmVarMet) is explicitly updated.
- Resolves quadratic O(N^2) cumulative UPDATE DVOBJECT queries causing upload times to degrade from 1s to 30s+ on datasets with 10k+ files.
- Adds comprehensive JUnit 5 acceptance tests asserting existing files are not dirtied.

Reference: https://groups.google.com/g/dataverse-community/c/pzSjF1YPaJw/m/bQWX3W_kBwAJ
…pload bottleneck

- Configured 2 CPU, 8GB RAM environment with Postgres query logging.
- Includes investigate.py benchmark tool measuring per-file latency and SQL UPDATE count.
- Includes technical report and replication guide.

Reference: https://groups.google.com/g/dataverse-community/c/pzSjF1YPaJw/m/bQWX3W_kBwAJ
…ion environment

- Renamed release note to doc/release-notes/12728-perf-large-upload.md matching IQSS#12728
- Moved Docker reproduction benchmark environment to dedicated repository: https://github.com/tuannx/dataverse-large-upload-repro
@pdurbin pdurbin moved this to Ready for Triage in IQSS Dataverse Project Sep 21, 2026
@pdurbin pdurbin moved this from Ready for Triage to Ready for Review ⏩ in IQSS Dataverse Project Sep 22, 2026
@pdurbin pdurbin added the Size: 20 A percentage of a sprint. 14 hours. label Sep 22, 2026
@pdurbin

pdurbin commented Sep 22, 2026

Copy link
Copy Markdown
Member

We discussed this at tech hours. Overall, we like the approach.

When it comes to how to test this, the following PR is related, and @landreev is actively using it for testing:

@tuannx also created this related PR:

@cmbz cmbz added FY27 Sprint 6 FY27 Sprint 6 (2026-09-09 - 2026-09-23) FY27 Sprint 7 FY27 Sprint 7 (2026-09-23 - 2026-10-07) labels Sep 23, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

FY27 Sprint 6 FY27 Sprint 6 (2026-09-09 - 2026-09-23) FY27 Sprint 7 FY27 Sprint 7 (2026-09-23 - 2026-10-07) Size: 20 A percentage of a sprint. 14 hours.

Projects

Status: Ready for Review ⏩

Development

Successfully merging this pull request may close these issues.

Performance: O(N) database updates on every file upload degrade large dataset ingestion (1s -> 30s+)

3 participants