Pass 1 deferred cross-file exact matches to Pass 2, but Pass 2's candidate
filter keeps only the first node per normalized label, so identical-label
cross-file concept pairs could never merge (while fuzzy pairs did). Pass 1
now unions the cross-file residue of each label group, gated to concept
nodes with provenance and above the entropy floor, so code/rationale/
document/image/empty-source and cross-repo guards are all preserved.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Two tweaks on top of @bchan84x's #1852:
1. The definer heuristic decided the survivor pairwise, so when several
same-file nodes co-defined one ID the surviving label still depended on
arrival order (the exact 3-node case #1851 reports). Choose the survivor
by a total order (definer first, then shorter/canonical label, then
lexically) via a min over _collision_rank, so it is order-independent.
Direction preserves #1504 (lexically-first source path wins).
2. _defines_id now also recognizes a bare file-level node whose id is
exactly the slugified path (nid == prefix), not only <path>_<entity>.
Adds an order-independence test over all 6 permutations of definer +
same-file relabel + cross-file reference, and a bare-file-node test.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Node IDs are <source-path>_<entity>, so a doc that merely references an entity
mints the entity's own ID and collides with the entity's node by construction.
The pre-dedup pass kept whichever arrived first, so chunk order decided whether
an entity kept its own attributes or a passing cross-reference's, and the
warning that fired (#1504) told the user a same-name-different-directory clash
had lost their data — while the one drop that really is lossy, two labels for
one ID from the same file, was silent because the warning was gated on
source_file differing.
The survivor is now the node whose source_file is the file its ID encodes (any
trailing slice of the path, so absolute, repo-relative, and pre-#1504
bare-stem IDs all resolve), falling back to first-seen. Reporting follows what
the drop actually costs: a cross-reference folding into the node it references
loses nothing (edges are keyed by ID and rewire to the survivor) and is silent;
a same-file relabel notes the discarded label; two files that both encode the
ID are distinct entities and still WARNING, now stating the ID-scheme cause
rather than a filename clash.
Refs #1851
When two LLM extraction chunks each process a file with the same name in
different directories, they independently generate the same node IDs and
deduplicate_entities() silently drops one node (first-writer-wins). The
data loss had no indication in any log, counter, or output.
Adds a stderr WARNING when a duplicate ID comes from a different
source_file, telling the user which files collided and recommending the
per-subfolder extract + merge-graphs workflow to avoid it.
Three pass-2 guards (mirrored in the --dedup-llm pair collection): block merges
when labels' embedded numbers differ as zero-padding-insensitive multisets;
block cross-file merges of file-anchored rationale/document nodes (same-file
still merges); and score cross-file long labels on plain Jaro instead of
Jaro-Winkler so the prefix bonus can't fabricate merges of shared-prefix but
token-divergent entities (jest-native vs react-native), while genuine cross-file
duplicates still clear Jaro and same-file near-duplicates keep Jaro-Winkler.
Co-Authored-By: van4oza <van4oza@users.noreply.github.com>
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Replaced O(n) linear scan `next((n for n in candidates if n["id"] == neighbor_id))` with O(1) dict lookup via pre-built `candidates_by_id`. Also pre-caches `_norm()` results in `norm_cache` to avoid recomputing per inner iteration.
For a 36k-file codebase (~100k high-entropy candidates) this reduces the pass-2 loop from O(n^2*B) (~30–100B iterations) to O(n*B), eliminating the multi-minute CPU hang after AST extraction completes.
Pass 2 selected the merge winner from the union of both normalized-label
groups, so never-compared same-label/cross-file nodes could be pulled
into the union, bypassing the #1046/#1178 guards. Pick the winner from
[node, neighbor] only; group members that belong together still merge
via pass 1 (same file) or their own verified comparison.
Co-authored-by: Cursor <cursoragent@cursor.com>
- graphify/_minhash.py: self-contained MinHash/MinHashLSH using pure numpy,
byte-identical hash math to datasketch (sha1_hash32, Mersenne-prime permutation).
Drops datasketch + scipy transitive dep — eliminates EDR hang on Windows where
numpy.testing platform.machine() subprocess spawn was intercepted at import time
- dedup.py: import from graphify._minhash instead of datasketch
- pyproject.toml: replace datasketch>=1.6 with numpy>=1.21
- detect.py: memoize _is_ignored/_eval results in a dict[Path,bool] cache per
detect() call; each unique ancestor dir evaluated once across all sibling files,
eliminating ~42M redundant fnmatch calls on large repos (~34% whole-run speedup)
- tests/test_minhash.py: 11 tests including import-isolation guard asserting scipy
and numpy.testing are not loaded after import graphify.dedup
- tests/test_detect.py: 2 cache tests — correctness (cached==uncached with negation
patterns) and hit-count (each dir evaluated exactly once across siblings)
Closes#1234, closes#1235
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- security.py: replace global socket.getaddrinfo monkey-patch with per-connection
_SSRFGuardedHTTPConnection/HTTPSConnection subclasses (thread-safe, closes TOCTOU)
- security.py: add GRAPHIFY_MAX_GRAPH_BYTES env var override for 512MB cap (MB/GB suffix
supported); improve cap error message to cite the env var
- llm.py: wrap untrusted source files in XML delimiters with sha256 fingerprint;
neutralise jailbreak sentinel tokens to mitigate prompt injection
- dedup.py: skip code nodes in label-based dedup passes; code symbols now deduplicated
by ID only, preventing distinct same-named symbols from merging
- extract.py: cross-file calls resolution now consults import evidence before bailing
on ambiguous callee names; emits EXTRACTED edges when named import is unambiguous
- analyze.py: extend _BUILTIN_NOISE_LABELS with stdlib types and modules
- __main__.py: CLAUDE.md template uses MANDATORY language for graphify-first rule;
PreToolUse hook message hardened to imperative; graphify export html auto-falls
back to community-aggregation view when graph.json exceeds size cap
- tests/test_pg_introspect.py: add importorskip guard for tree_sitter_sql
Closes#1211, #1210, #1205, #1219, #1227; resolves discussion #1019
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- analyze.py: pass length_bound=max_cycle_length to nx.simple_cycles() so
networkx prunes during enumeration instead of post-filtering; drops report
generation from never-returns to ~0.1s on dense graphs (#1196)
- llm.py: replace hardcoded min(40+16*n,4096) label_communities token budget
with _resolve_max_tokens(min(64+24*n,8192)) — 24 tok/community covers 5-word
JSON entries; 8192 cap fits 16k-context models; env var now honoured (#1200)
- dedup.py: add prefix-extension guard in Pass 2 and _llm_tiebreak — skip merge
when one normalised label is a strict prefix of the other (getActiveSession /
getActiveSessions, parseConfig / parseConfigFile). Option (a) rejected: dropping
the >=12 early-out from _short_label_blocked breaks test_typo_merged (#1201)
- tests/test_dedup.py: two new regression tests verifying prefix guard fires for
extension pairs and does not fire for same-length typo pairs
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Guards _norm, _norm_label, and _strip_diacritics against None node labels that cause TypeError in unicodedata.normalize(). Fixes#1194. Consistent with existing security.py:270 precedent.
Co-authored-by: freiit <freiit@users.noreply.github.com>
Three-part fix:
dedup.py: Pass 1 exact-merge now skips nodes with an empty source_file.
Previously all no-source_file nodes with the same label landed in one
bucket and were merged, destroying distinct symbols (third-party deps,
standalone functions) that happened to share a short name.
update.md (skillgen + all 13 host variants): the --update merge now
passes both deleted AND changed files to prune_sources, mirroring what
watch._rebuild_code already does correctly. Old nodes for re-extracted
files are pruned before fresh AST is inserted — no fuzzy reconciliation
needed, no cross-file collapse possible.
export.py: anti-shrink guard message now names fuzzy dedup as a
possible cause (not only "missing chunk files"), and advises a full
rebuild as the safe recovery path.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Revert .h -> extract_c (C++ grammar rejects C++ keywords used as identifiers
in Linux-kernel-style headers; .hpp/.hxx/.hh already route to extract_cpp)
- Fix field_declaration block: use children_by_field_name("declarator") instead
of iterating all children with wrong type guard; replace ensure_node (undefined)
with add_node
- Fix _import_c include resolution: use _make_id(str(resolved)) to match the
file_nid scheme _extract_generic uses, not _make_id(_file_stem(resolved))
- Fix exact_merges counter in dedup Pass 1 to count only within-file merges
actually performed, not the raw unpartitioned group sizes
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Route .h files through extract_cpp (was extract_c), fixing missing method nodes in C++ headers
- Extend _get_cpp_func_name to handle field_identifier, destructor_name, operator_name
- Add CPP-specific field_declaration branch in _extract_generic to emit class method/field nodes
- Partition dedup Pass 1 by source_file: only union same-label nodes within the same file;
cross-file matches fall through to Pass 2 fuzzy, preventing generic-label god nodes
- Add _resolve_c_include_path: resolve quoted #include paths to real files on disk so
target node IDs match what _extract_generic creates, fixing dangling include edges
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- extract/_make_id + build/_normalize_id: use NFKC normalization and casefold
so composed/decomposed Unicode forms produce the same ID; collapse consecutive
underscores; both functions are now byte-for-byte equivalent (#811)
- dedup: use explicit key-presence check instead of `or` for source/from
fallback; pop stale from/to keys so they don't leak into graph.json attrs (#803)
- skill --update: use build_merge() to avoid NetworkX round-trip direction flip;
fix dict merge ordering so explicit source/target win; pull hyperedges from
G.graph (merged) not new_extraction only (#801)
- skill subagents: inject absolute CHUNK_PATH so Write tool doesn't lose chunk
files to undefined cwd (#808)
- __main__: skip skill version check during hook-check (runs on every editor
tool use, must be silent); move warning to stderr
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Gemini is often the cheaper available quota for low-stakes semantic graph extraction, while OpenAI is a useful fallback. Extend the direct extraction backend registry, CLI validation, docs, and tests so headless extraction can use GEMINI_API_KEY, GOOGLE_API_KEY, or OPENAI_API_KEY without changing the existing Claude and Kimi paths.
Constraint: Gemini supports OpenAI-compatible chat completions at the Google generative-language endpoint
Rejected: Native google-genai integration | higher dependency and response-shape churn for the same chat-completions path
Confidence: medium
Scope-risk: moderate
Directive: Keep backend detection explicit and test every accepted API-key environment variable before adding new providers
Tested: uv run --directory vendor/graphify pytest tests/test_llm_backends.py tests/test_chunking.py -q
Not-tested: Live Gemini/OpenAI API calls; no GEMINI_API_KEY or OPENAI_API_KEY present in this environment