feat(dedupe): injectable candidate scope and winner order, plus dedupe pools docs - #15819
Open
blakeaowens wants to merge 11 commits into
Open
feat(dedupe): injectable candidate scope and winner order, plus dedupe pools docs#15819blakeaowens wants to merge 11 commits into
blakeaowens wants to merge 11 commits into
Conversation
blakeaowens
marked this pull request as draft
August 28, 2026 11:54
blakeaowens
marked this pull request as ready for review
September 5, 2026 01:55
…order Deduplication derives its own candidate scope: a finding matches within its own product, narrowed to a single engagement by the engagement checkbox, and that is the only scope an installation can get. This makes both the scope and the preference order injectable, so a plugin can express a different one, while leaving every existing call site on exactly the behaviour it had. build_candidate_scope_queryset gains candidate_qs. When supplied it replaces the scope derivation; the loading strategy (defer, select_related, prefetch_related) is still applied here, so a caller decides which findings are candidates while the engine keeps deciding how to load them. Candidate confirmation walks locations, vulnerability ids and CWEs per candidate, so a scope handed in without those prefetches would silently turn one query into thousands. The four match generators gain ordering_key: a plain sort key over candidates that only changes which of several valid candidates is preferred. _is_candidate_older still runs afterwards, so an ordering key cannot make a newer finding win, and the global antisymmetry concurrent batches depend on is unaffected. uid_or_hash applies it to the merged candidate set, since merging two buckets loses their query order. Both kwargs thread through find_candidates_for_deduplication_*, match_batch_*, _dedupe_batch_* and their dispatchers. dedupe_batch_of_findings forwards them to a custom deduplication method only when set, so a plugin that predates them sees exactly the arguments it saw before. false_positive_history gains the same seam as scope_filter, since it builds its own queryset per algorithm from filter kwargs rather than filtering a supplied one. Every kwarg defaults to None and every default path is the previous code, so an open-source install is unaffected. The pairwise engagement guard is deliberately untouched: it stays correct for installs that rely on it, and a caller supplying a scope owns expressing its own isolation.
…guard to bite The candidate was created first, so it held the lower id and was a legal original after all. Creating the target first makes the candidate genuinely newer, which is the case the assertion is about.
Deduplication has always been scoped to one Asset, narrowable to an Engagement. Pools are the other direction: a named group of Assets whose Findings deduplicate against each other, per matching kind. The new page covers what a pool is and is not (it changes which Findings are eligible to be compared, never how two are compared), the per-kind membership rule, why reimport appears alongside the two kinds that scope and yet cannot widen scope, the preview-then-acknowledge contract on both retroactive actions, where originals collect and why there is no newest-wins, and the parent-edges- only subtree toggle. Three statements elsewhere became incomplete rather than wrong, so they are updated in the same batch: the scope paragraph and the Pro algorithm summary in About Deduplication, its troubleshooting table (which offered only instance-wide answers to a per-Asset scope problem), and the Enabling Deduplication intro.
The pools page said a reimport "does not widen what is compared" and left it
there. True of a reimport's own matching, and it reads as "reimports never
deduplicate across a pool", which is false. That reading is exactly how it was
caught.
Both halves are now stated. A reimport's own matching stays inside its Test,
because that matching decides whether a Finding is updated, created or closed
and is scoped to what the scan is authoritative over. The Findings it creates
are then deduplicated under same tool and cross tool, which are pool-scoped.
About Deduplication already said the second half in general terms ("Findings
that remain after Reimport Deduplication are still subject to Same-Tool
Deduplication"); it now names pools as a case of it, so the two pages agree.
The apply-now paragraph carried the same framing and is reworded: pooling for
reimport picks a formula rather than a scope, so there is no widened scope to
re-run, and Findings a reimport created are covered by the other two kinds.
The three tuner deduplication pages were replaced by a single Matching Configuration page, so this page was describing a UI that no longer exists: a menu path that is gone, a tool dropdown on one of three pages, and four screenshots of retired screens. Two claims were not merely stale but wrong in a way that matters. It promised that changing a tool's settings "automatically triggers a background re-hash of all existing Findings". Changing an algorithm does not re-hash anything, and cannot: the algorithm selects which already-stored value is compared, so there is nothing to recompute. A reader following the old text would wait for a backlog re-hash that is never coming. That section is replaced with what actually happens, and points at a pool's Apply Now for the case it was reaching for. It also described selecting hash fields, which is not possible in this release. Changing hash fields changes how every stored hash was computed, so it needs a new generation written behind it before matching moves across, and that is not shipped. The page says so and points at support rather than describing a control that is not there. The reference material that is still accurate is kept as is: the algorithms, Content Fingerprint, the set-based vulnerability-id and CWE matchers, and location drift tracking.
Follows the Pro change restoring hash-field editing to Matching Configuration. The previous revision said they were not editable and pointed at support, which was true for one commit and is not now. Splits the retroactive-re-hash guidance by axis rather than making one claim about both, since they behave oppositely and conflating them is what made the original page wrong: changing hash fields recomputes the tool's whole backlog in the background, and changing the algorithm recomputes nothing at all. Also records that hash fields are set on the instance default rather than per pool, and why: a finding stores one hash and every other view of that finding reads it.
…d the nav paths Four review findings on the dedupe docs batch. **PRO__location_drift_matching.md still walked users to a deleted page.** It said Settings > Finding Workflow > Reimport Deduplication and "Enable Track findings as locations change" — a page this batch removes — while the tuning page claimed the toggle was Support-set. One of those had to be wrong, and the drift page is the one a user follows step by step. The toggle is self-serve again on Matching Configuration (the Pro side of this batch adds it), so both pages now describe that, including the impact review, because turning it on or off changes which fields the reimport hash is built from and recomputes the tool's backlog. **Pooling narrows a global algorithm rather than leaving it alone.** The "Pools vs. the global algorithms" table read as two independent choices at different blast radii, but pro/dedupe/scope.py bounds Global Component and Global Locations to the pool once an Asset joins one for that kind. A Global Component user who creates a pool would silently narrow matching they believed was instance-wide. Called out under the table and on both PRO__global_* pages, which are the ones that promise "across all Assets". **The pages disagreed about where the feature lives.** Dedupe Pools said Settings > Deduplication Settings > Dedupe Pools; the tuning page said Settings > Matching Configuration. Both entries are in the same nav group, so every page now names the full path and says the two sit beside each other. Translations are left for the usual separate pass; only the English pages are updated here.
…dedupe pages Apply Now runs inside the request and refuses above 50,000 findings, and reports a reason rather than a silent zero when deduplication is off instance-wide. Neither was written down, so a user with a large pool met the refusal with no way to anticipate it. Also replaces the em dashes these pages introduced with colons, parentheses and periods, per the house copy style.
…ing on every finder build_candidate_scope_queryset took candidate_qs positionally while all five finders made it keyword-only. Every caller already passes it by keyword, so closing the asymmetry costs nothing and stops a positional service argument from ever landing in the wrong slot. The tests reached candidate_qs only through the hash path. Each of the five finders builds its own base queryset, so any of the others could have dropped the argument and stayed green. All five now assert the same property directly and in both directions: the derived scope does not reach the other product, and a supplied one does. The finders returning two maps (uid-or-hash, legacy) are checked on both, since forwarding could reach one and not the other.
…mport membership as they ship, upgrade and custom-role notes False-positive history now uses the same scope as deduplication, which narrows replication for instances with engagement-scoped deduplication: an isolated engagement is excluded from every other engagement's history and its own imports read only its own. About Deduplication says so; before, the search always covered the whole asset. The pools page and Deduplication Tuning both claimed a pool can give its members a different algorithm. No shipping path creates a pool-level configuration row, so both now say per-pool overrides are not yet available and every member uses the instance default. The pools page offered a reimport membership kind and described it as selecting the formula a reimport uses. The resolver never consults a pool for that kind, so the membership did nothing; the kind is withdrawn from the page and the text explains why, and that findings a reimport creates still deduplicate across the pool. Deduplication Tuning gains an Upgrading section: the cutover from the tuner pages is one-way and a database backup should precede the upgrade. The pools page gains a custom-roles note: the upgrade carries Edit Tuner to all four pool permissions and View Tuner to View Dedupe Pool, and only roles created afterwards need the grants made by hand.
…ng cross tool Adds the 3.3.0 section to the Pro changelog with the two behaviour changes the matching-configuration work introduces (false-positive history follows deduplication scope; Global Component and Global Locations matching bounded to the pool for a pooled Asset) and the two feature lines. The dedupe-pools page states the lowered Apply Now ceiling (10,000 Findings), and the tuning page says what setting a tool's cross-tool algorithm back to Disabled does.
blakeaowens
force-pushed
the
feature/dedupe-injectable-candidate-scope
branch
from
September 6, 2026 18:16
7e52490 to
fb6cb0b
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
Two related changes: a small seam in the deduplication engine, and the documentation for the Pro
feature built on top of it.
The seam.
build_candidate_scope_queryset()decides which existing findings a new finding iscompared against, and the winner order decides which one becomes the original. Both were computed
internally with no way for a caller to influence them. This adds two optional keyword arguments,
candidate_qsandordering_key, threaded through the finders, the candidate generators, thebatch matchers and both dispatchers, plus a
scope_filteron the false-positive-history path.Nothing changes when they are not supplied. Every call site forwards them only when set:
so the default path builds exactly the queryset and ordering it always did. The seam is
deliberately queryset-shaped rather than a callback or a settings flag: a caller narrows or widens
the candidate set by handing in a queryset, and the engine keeps ownership of how matching itself
works.
The documentation covers DefectDojo Pro's dedupe pools, which are the first consumer of the
seam: named groups of Assets whose findings deduplicate against each other, per matching kind.
The new page explains what a pool does and does not do (it changes which findings are eligible to
be compared, never how two are compared), why the reimport kind cannot widen scope even though it
appears alongside the two kinds that can, and the preview-then-acknowledge contract on the
retroactive actions.
Three existing statements became incomplete rather than wrong once pools exist, so they are
updated in the same change: the scope paragraph and the Pro algorithm summary in About
Deduplication, and that page's troubleshooting table, which previously offered only instance-wide
answers to a per-Asset scope problem. The Enabling Deduplication intro gains a pointer.
Test results
unittests/test_dedupe_injectable_scope.pyis new and covers the seam directly: that an injectedqueryset narrows and widens the candidate set, that an injected ordering key selects the original,
that omitting both reproduces the default scope and order exactly, and that the false-positive
history path honours the same filter.
Run locally against a Pro-flavoured checkout:
12 tests, all passing. The wider deduplication suites were also run green on the consuming side
(1060 passed, 1 skipped).
Release ordering
This has to land first, or in the same release train as the DefectDojo Pro change that consumes
it. That change calls the new signature in thirteen places, so against a released engine without
candidate_qsevery one of them raises aTypeError. Both are milestoned for the same release,which satisfies the constraint; the note is here so a later retarget of either one does not
quietly break the pairing.
The reverse direction is safe: both parameters default to
Noneand reproduce the previousqueryset and ordering exactly, so this engine works unchanged with no consumer at all.
Documentation
Included in this PR, under
docs/content/triage_findings/finding_deduplication/, plus a 3.3.0section in
docs/content/releases/pro/changelog.mdcarrying the two behaviour changes the Pro PRintroduces (false-positive history follows deduplication scope; Global Component and Global
Locations matching bounded to the pool for a pooled Asset) and the two feature lines. The
dedupe-pools page states the Apply Now ceiling (10,000 Findings) and the tuning page describes
what disabling Cross Tool Deduplication does.