A single `!` rule in .graphifyignore set a blanket `has_negation` flag that
disabled directory-level pruning for EVERY ignored directory during the
os.walk in detect(). One unrelated `!docs/**` therefore made the walk descend
bin/, obj/, wwwroot/, generated/, … on large repos — a pathological slowdown.
Output stayed correct (the per-file `_is_ignored` filter still excluded those
files), but the walk visited the entire tree.
The bypass was unnecessary: `_is_ignored` already honours negations correctly —
last-match-wins lets `!dir/` un-ignore a directory (so it is not pruned), and
the gitignore parent-exclusion rule means a `!` cannot rescue a file beneath an
excluded directory, so descending an ignored dir to find a re-included file is
never needed. Prune purely on `_is_noise_dir` + `_is_ignored`.
Adds a regression test that tracks os.walk and asserts the ignored dir is never
descended while the negation still re-includes its target.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Closes#1264
Also fixes a narrow watch-mode path-filtering bug found while running the required local
full-suite gate for this PR. Hidden directories inside the watched corpus are still ignored;
hidden ancestors outside the watched root no longer suppress valid file events.
_read_tsconfig_aliases joined alias targets onto the tsconfig's own
directory and ignored compilerOptions.baseUrl. In the common monorepo /
NestJS layout (baseUrl "./src" with "@services/*": ["services/*"]), the
alias resolved to <dir>/services instead of <dir>/src/services, so every
aliased import failed to resolve and the import edge was silently dropped
— leaving cross-file caller graphs nearly empty on alias-heavy TS repos.
Resolve `paths` relative to `baseUrl` (TypeScript's actual semantics),
defaulting to "." so configs without baseUrl keep their current behavior.
Add a regression test covering a subdirectory baseUrl; the existing alias
test only exercised baseUrl ".", which is why this slipped through.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Two Windows fixes for the claude-cli backend:
1. _call_llm (the label/community-naming path) spawned a bare ["claude", ...],
which CreateProcess cannot resolve to the npm claude.cmd shim (PATHEXT does
not apply) — every labeling batch failed with WinError 2. Mirror the
extraction path's resolution: prefer shutil.which("claude.cmd") and pass
the resolved path. Regression test included.
2. Both claude -p spawn sites now pass CREATE_NO_WINDOW on Windows. Without
it, each per-batch claude.cmd spawn allocates a console — with Windows
Terminal as the default terminal, a labeling run pops one visible window
per batch on the user's desktop for the duration of the model call.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Callable nodes are labeled with a trailing "()" (e.g. classifyProperty()),
so a bare-name query like "classifyProperty" falls through the exact-label
pass and then ties with any prefix sibling (classifyPropertySafe()) in the
contains pass, returning "No unique node match" even though exactly one
callable with that name exists.
Add a bare-name pass between exact-label and exact-source matching that
compares query and label with the trailing "()" stripped from each. Bare
queries now resolve decorated labels and vice versa; genuine duplicates
still return None.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
llm.py: add explicit edge direction rule to extraction system prompt —
source = actor (caller/importer/subclass), target = acted-upon (callee/
imported/base). LLM was systematically emitting callee->caller for calls
edges because the schema never stated direction semantics.
build.py: extend ghost-node merge to catch LLM nodes that populate
source_location (bypassing the old None check). Now uses _origin=="ast"
as the canonical signal — AST nodes always win; any non-AST node sharing
(basename, label) with an AST node is collapsed into the AST canonical.
Fixes LLM bare-stem IDs (bpe_get_pairs) surviving alongside AST
parent-qualified IDs (mingpt_bpe_get_pairs) and carrying reversed edges.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
collect_files ran one full recursive rglob pass per supported extension
(~85 walks), descending into node_modules/.git/venv and filtering those
paths only after enumerating them, and re-evaluated gitignore patterns
per file without the shared cache.
Replace the loop with a single os.walk that prunes noise dirs (and, when
no negation patterns exist, ignored dirs) in place -- the same pattern
the follow_symlinks=True branch and detect.py's scan walk already use --
and pass the shared _is_ignored cache. Extension matching switches to
p.suffix in _EXTENSIONS, matching the follow_symlinks branch and
preserving the .f/.F Fortran case distinction.
Synthetic benchmark (200 source files + 5,000-file node_modules):
0.295s -> 0.007s, identical result set.
Tests: parity oracle against the old implementation on the fixtures and
on a synthetic tree (noise dirs, hidden dirs, gitignore negation), plus
a scandir-counting test asserting each directory is read at most once
and noise dirs are never entered.
Co-authored-by: Cursor <cursoragent@cursor.com>
_body_content used substring checks (startswith("---") / find("\n---")),
so `----` thematic breaks and `--- text` prose lines were mistaken for
frontmatter delimiters and everything above them was silently excluded
from the hash -- edits there never invalidated the cache.
Both delimiters must now be a whole line of exactly three dashes
(optional trailing whitespace). For well-formed frontmatter the stripped
body stays byte-identical to the previous implementation, so existing
semantic-cache hashes do not churn.
Co-authored-by: Cursor <cursoragent@cursor.com>
global_add deduplicates external-library nodes (no source_file) by label
against externals already in the global graph, but dropped every edge
incident to a skipped node. Each repo after the first lost its edges to
shared externals, so cross-repo "what uses library X" queries only saw
the first-added repo.
Build a remap from each deduplicated external to the surviving global
node and rewrite edge endpoints through it before add_edge, skipping
self-loops introduced by the remapping.
Co-authored-by: Cursor <cursoragent@cursor.com>
Pass 2 selected the merge winner from the union of both normalized-label
groups, so never-compared same-label/cross-file nodes could be pulled
into the union, bypassing the #1046/#1178 guards. Pick the winner from
[node, neighbor] only; group members that belong together still merge
via pass 1 (same file) or their own verified comparison.
Co-authored-by: Cursor <cursoragent@cursor.com>
graphify extract <path> --out R is documented to send all output to
R/graphify-out/, but the AST and semantic extraction caches were still
anchored at the scanned project (cache_root=target / root=target). That
re-created a graphify-out/ directory inside the project the user
explicitly asked to keep clean.
Anchor both caches at out_root instead. With --out unset, out_root
equals target, so existing in-project behavior is unchanged.
Adds test_extract_out_keeps_project_root_clean: runs extract from a
project root with --out pointing elsewhere and asserts the artifacts
land under --out while the project directory stays byte-identical.
`claude -p --output-format json` now emits a JSON array of streamed event
objects instead of a single envelope dict, so `_call_claude_cli` crashed
with `'list' object has no attribute 'get'` on every chunk. Add a shared
`_claude_cli_envelope` helper that normalizes both shapes (extracts the
terminal {"type":"result"} object from an array) and use it at both CLI
call sites.
extract intentionally stops at graph.json; GRAPH_REPORT.md requires cluster-only.
Use --no-label to skip LLM community naming (no API key needed in CI).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
graphify/skills/ contains 126 markdown files that trigger semantic extraction.
Add a temporary .graphifyignore entry during CI to keep the build pure AST.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- extract graphify/ (code only) instead of . to avoid LLM API key requirement;
the repo root contains docs/skills/.md files that trigger semantic extraction
- use --out . so output writes to ./graphify-out/ not ./graphify/graphify-out/
- remove --out from export html (flag does not exist; HTML auto-written next to graph)
- drop nonexistent --code-only flag from extract command
- add comments explaining each flag's behavior
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Runs graphify against its own source on every GitHub release (AST-only,
no API cost) and attaches graph.json + graph.html + GRAPH_REPORT.md as
graphify-self-graph.tar.gz to the release. Also supports manual runs via
workflow_dispatch, uploading the bundle as a 7-day workflow artifact.
Closes#1238
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- export.py: guard to_obsidian/to_canvas against dangling community member IDs
(KeyError crash when a node in communities dict is absent from graph, #1236)
- detect.py: NFC-normalize path before hashing Office sidecar filename to fix
macOS NFC/NFD mismatch causing --update to re-extract all Office files (#1226)
- extract.py: add _is_config_json() to skip data JSON files (only extract
package.json, tsconfig.json, eslint, deno, JSON Schema etc.) eliminating
561 orphan key-nodes on large repos (#1224)
- llm.py: add GRAPHIFY_LLM_TEMPERATURE env var + _resolve_temperature() helper;
auto-omit temperature for o1/o3/o4/gpt-5 reasoning models that reject temp=0;
mirrors GRAPHIFY_MAX_OUTPUT_TOKENS precedence pattern (#1191)
- tests: 20 new regression tests across obsidian, detect, extract, llm_backends
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- graphify/_minhash.py: self-contained MinHash/MinHashLSH using pure numpy,
byte-identical hash math to datasketch (sha1_hash32, Mersenne-prime permutation).
Drops datasketch + scipy transitive dep — eliminates EDR hang on Windows where
numpy.testing platform.machine() subprocess spawn was intercepted at import time
- dedup.py: import from graphify._minhash instead of datasketch
- pyproject.toml: replace datasketch>=1.6 with numpy>=1.21
- detect.py: memoize _is_ignored/_eval results in a dict[Path,bool] cache per
detect() call; each unique ancestor dir evaluated once across all sibling files,
eliminating ~42M redundant fnmatch calls on large repos (~34% whole-run speedup)
- tests/test_minhash.py: 11 tests including import-isolation guard asserting scipy
and numpy.testing are not loaded after import graphify.dedup
- tests/test_detect.py: 2 cache tests — correctness (cached==uncached with negation
patterns) and hit-count (each dir evaluated exactly once across siblings)
Closes#1234, closes#1235
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- security.py: replace global socket.getaddrinfo monkey-patch with per-connection
_SSRFGuardedHTTPConnection/HTTPSConnection subclasses (thread-safe, closes TOCTOU)
- security.py: add GRAPHIFY_MAX_GRAPH_BYTES env var override for 512MB cap (MB/GB suffix
supported); improve cap error message to cite the env var
- llm.py: wrap untrusted source files in XML delimiters with sha256 fingerprint;
neutralise jailbreak sentinel tokens to mitigate prompt injection
- dedup.py: skip code nodes in label-based dedup passes; code symbols now deduplicated
by ID only, preventing distinct same-named symbols from merging
- extract.py: cross-file calls resolution now consults import evidence before bailing
on ambiguous callee names; emits EXTRACTED edges when named import is unambiguous
- analyze.py: extend _BUILTIN_NOISE_LABELS with stdlib types and modules
- __main__.py: CLAUDE.md template uses MANDATORY language for graphify-first rule;
PreToolUse hook message hardened to imperative; graphify export html auto-falls
back to community-aggregation view when graph.json exceeds size cap
- tests/test_pg_introspect.py: add importorskip guard for tree_sitter_sql
Closes#1211, #1210, #1205, #1219, #1227; resolves discussion #1019
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
bandit, pip-audit, and safety are already declared in the dev dependency
group but nothing in CI invokes them, so a new HIGH-severity finding or
a newly-disclosed CVE in a pinned dep can land without anyone noticing
until the next manual audit.
Add a security-scan job that runs bandit (-ll, HIGH-severity only) and
pip-audit (--strict) on every push and PR. Marked continue-on-error so
this doesn't block PRs on pre-existing findings -- a follow-up should
do the cleanup pass and flip the flag.
safety intentionally omitted: it requires a free-tier API key for the
new commercial backend, which is a setup burden for forks. pip-audit
covers the same ground using the PyPI JSON advisory feed and OSV.
1. graphify merge-chunks dumped the entire node list to the terminal
instead of the node count. __main__.py concatenated merged['nodes']
(a list of dicts) into an f-string where it clearly meant
len(merged['nodes']) -- the other two values in the same line use
len() correctly.
2. global_graph._load_manifest silently returned a fresh empty manifest
on any JSON parse error. That is reachable through normal interrupted
writes (the manifest is rewritten in full on every global_add /
global_remove, with no fsync or atomic rename), and the failure mode
is total data loss: every tracked repo disappears from
~/.graphify/global-manifest.json on the next read.
Back the corrupt file up to <path>.corrupt.<unix_ts> and print to
stderr before returning the empty default. Users can then recover
manually and the failure is visible rather than silent.
All 27 tree-sitter-* deps were unversioned in pyproject.toml. Users
installing via 'pip install graphifyy' (the README's primary install
path) bypass uv.lock entirely and resolve whatever tree-sitter-*
versions PyPI happens to serve. A breaking minor bump in any grammar
package can land in user installs without notice.
Add explicit lower bounds (matching uv.lock) and upper bounds one
minor above (or one major above for 0.x packages with frequent breaks).
Ranges chosen to allow patch updates without re-pinning while blocking
incompatible major/minor jumps.
- analyze.py: pass length_bound=max_cycle_length to nx.simple_cycles() so
networkx prunes during enumeration instead of post-filtering; drops report
generation from never-returns to ~0.1s on dense graphs (#1196)
- llm.py: replace hardcoded min(40+16*n,4096) label_communities token budget
with _resolve_max_tokens(min(64+24*n,8192)) — 24 tok/community covers 5-word
JSON entries; 8192 cap fits 16k-context models; env var now honoured (#1200)
- dedup.py: add prefix-extension guard in Pass 2 and _llm_tiebreak — skip merge
when one normalised label is a strict prefix of the other (getActiveSession /
getActiveSessions, parseConfig / parseConfigFile). Option (a) rejected: dropping
the >=12 early-out from _short_label_blocked breaks test_typo_merged (#1201)
- tests/test_dedup.py: two new regression tests verifying prefix guard fires for
extension pairs and does not fire for same-length typo pairs
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Guards _norm, _norm_label, and _strip_diacritics against None node labels that cause TypeError in unicodedata.normalize(). Fixes#1194. Consistent with existing security.py:270 precedent.
Co-authored-by: freiit <freiit@users.noreply.github.com>
Adds extra_body parameter support for custom/OpenAI-compat providers so users can pass provider-specific params (e.g. thinking budget for Claude via Bedrock compat). Adds multi-batch label_communities for 16k-context models — batches multiple community descriptions into a single LLM call instead of one per community. Partial batch failures are handled gracefully.
Co-authored-by: EirikWolf <EirikWolf@users.noreply.github.com>
The --push/--user/--password export flags feed both the neo4j and falkordb
dispatch branches, so the neo4j_ prefix was misleading - a neo4j_password that
reads FALKORDB_PASSWORD made no sense. Renamed to push_uri/push_user/
push_password, and the password env lookup now reads the backend-specific var
(FALKORDB_PASSWORD for falkordb, NEO4J_PASSWORD otherwise) instead of OR-ing both.
The no-push 'graphify export falkordb' path advertised
'redis-cli -x GRAPH.QUERY graphify < cypher.txt', but FalkorDB rejects that
with 'query with more than one statement is not supported' - cypher.txt is a
multi-statement Neo4j script. The individual statements ARE valid OpenCypher
(verified by loading them one at a time), only bulk script import is unsupported.
Message + skill docs now say so and point to --push (the verified load path).
Makes the FalkorDB option a first-class sibling of Neo4j in the agent skill,
not just the export CLI:
- --falkordb / --falkordb-push shorthands documented in core.md + the shared
exports.md reference, so they render into all modular platform skills and
read exactly like --neo4j / --neo4j-push. (The aider/devin monoliths are
diff-frozen vs v8 by skillgen's roundtrip guard, so they are left untouched.)
- README command reference switched to the /graphify ./raw --falkordb-push form.
- Documented URI scheme is now falkordb://localhost:6379; the scheme is only
informational (host/port are parsed out), so redis:// or a bare host:port
remain equivalent. Regenerated skill artifacts + expected/ snapshots.
Adds support for the XML-based `.slnx` solution format (VS 2022 17.13+ replacement for `.sln`). Extracts project references as `contains` edges and build dependencies as `imports` edges. XXE-protected XML parsing with size cap. Wired into `_DISPATCH` and `CODE_EXTENSIONS`. 6 new tests passing.
Co-authored-by: bakgaard <bakgaard@users.noreply.github.com>
Adds `graphify-mcp` as a named console script pointing to `graphify.serve:_main`, making the MCP stdio server directly invocable as a first-class CLI command from uv tool / pipx installs. MCP client configs can now use `"command": "graphify-mcp"` instead of `python -m graphify.serve`.
Co-authored-by: jr2804 <jr2804@users.noreply.github.com>
The Agent Skills spec only defines name, description, license,
compatibility, metadata, and allowed-tools as valid frontmatter fields.
The trigger: /graphify line was non-spec, silently ignored by spec-
following hosts, and flagged by agentskills validate CI checks.
- gen.py: removed trigger emission from _render_frontmatter; added
_is_trigger_line() helper for roundtrip allow-list
- fragments/core/aider.md: removed hardcoded trigger: /graphify
- platforms.toml: removed trigger doc comment and trigger="" entries
- test_skillgen.py: replaced trigger-assertion tests with a single
test asserting no host has trigger: in frontmatter
- Regenerated all 125 skill artifacts
Routing intent is preserved: the description field already contains
"treated as a graphify query first" and "graphify-out/ exists".
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
graphify codebuddy install was writing CODEBUDDY.md and settings.json
but not copying the SKILL.md. Added _copy_skill_file("codebuddy") call
to match the --platform codebuddy path. README hook description updated
from "Glob and Grep" to "Bash search and file reads" to match actual
hook matchers (Bash + Read|Glob).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Three-part fix:
dedup.py: Pass 1 exact-merge now skips nodes with an empty source_file.
Previously all no-source_file nodes with the same label landed in one
bucket and were merged, destroying distinct symbols (third-party deps,
standalone functions) that happened to share a short name.
update.md (skillgen + all 13 host variants): the --update merge now
passes both deleted AND changed files to prune_sources, mirroring what
watch._rebuild_code already does correctly. Old nodes for re-extracted
files are pruned before fresh AST is inserted — no fuzzy reconciliation
needed, no cross-file collapse possible.
export.py: anti-shrink guard message now names fuzzy dedup as a
possible cause (not only "missing chunk files"), and advises a full
rebuild as the safe recovery path.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Adds FalkorDB as a sibling option to the existing Neo4j sink, selected via
`graphify export falkordb [--push redis://localhost:6379]`.
- New push_to_falkordb() in graphify/export.py mirrors push_to_neo4j; FalkorDB
is OpenCypher-compatible so the MERGE/SET upsert queries are identical.
- export falkordb subcommand wired in graphify/__main__.py (cypher.txt when no
--push, direct push otherwise). Auth is optional; target graph defaults to
"graphify".
- falkordb optional extra in pyproject.toml (and in the all extra).
- Tests: CLI cypher generation (CI-safe) + real-FalkorDB integration tests that
skip when no instance is reachable.
- README extras table + command reference and CHANGELOG updated.
#1174: affected.py load_graph now forces directed=True before
node_link_graph, matching the identical fix in serve.py and __main__.py.
Undirected graphs (directed:false in graph.json) were causing in_edges
to fall back to a direction-blind scan, missing true callers and
reporting false positives. Regression test added.
#1173: post-commit and post-checkout hook bodies now read
graphify-out/.graphify_root before calling _rebuild_code, falling back
to Path('.') if absent. A scoped build (graphify src/) no longer gets
silently expanded to the full repo on the next commit. Tests added.
#1172: Step 9 cleanup split into rm -f for fixed files and
find -maxdepth 1 -delete for the chunk glob. Under fish/zsh an
unmatched glob aborts the entire rm -f line, leaving temp files on disk.
Fixed in the three skillgen source fragments and regenerated.
#1163: detect_incremental type guard on stored mtime — if the manifest
contains a dict-valued mtime (schema drift from older versions), coerce
to None rather than propagating a non-numeric into comparisons.
Regression test added.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
#1170 — replace nohup with cross-platform Python detach in git hooks.
Git for Windows MSYS has no nohup so post-commit/post-checkout hooks
silently failed. Now uses subprocess.Popen with DETACHED_PROCESS |
CREATE_NEW_PROCESS_GROUP on Windows, start_new_session=True on POSIX.
Quoting-safe (argv list). Fixes#1161.
#1169 — fix _is_sensitive false positives on topic-mentioning filenames.
token-economics-of-recall.md and password-policy-discussion.md were
silently dropped as secrets. Generic keywords (token/secret/password)
now only fire when the keyword ends the filename stem or the stem is
≤2 words. Specific patterns (.env/.pem/id_rsa etc.) remain unconditional.
#1165 — fix multi-word endpoint resolution in _score_nodes.
graphify path "AuthService" "UserRepo" never fired the exact-match bonus
because per-token comparison never equalled the full label. Now joins
normalized tokens and compares against the full label and its tokenized
form. O(1) per node, affects query_graph and shortest_path uniformly.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
#1154: scope numpy>=2.0 constraint to python_version>='3.13' only.
numpy 1.26.4 ships no cp313 wheel so uv sync falls back to a source
build requiring a C compiler. The marker avoids forcing numpy 2.x on
3.10-3.12 users who have working 1.x environments.
#1160: codex platform skill now installs to .codex/skills/graphify/
instead of .agents/skills/graphify/. The hook already wrote to .codex/
so the skill destination was inconsistent. Propagates automatically
through install/uninstall (both read _PLATFORM_CONFIG dynamically).
Updated all codex-specific test assertions.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
#1118 — prune stale AST nodes on full re-extraction (#1116)
Stamps every AST-extracted node with _origin="ast" in extract(). On a
full rebuild _rebuild_code drops any AST-marked node absent from the
fresh output even when its source file survives, fixing stale symbols.
Backward-compat: marker-less nodes from pre-1118 graphs survive one
cycle then self-heal.
#1110 — stop reading images and PDFs as garbage in headless extract
Images route through per-backend vision payloads (base64/data-URI/bytes
for claude/openai/bedrock); non-vision backends get _strip_pixels for
graceful degradation. PDFs reuse pypdf. 5MB cap, 20-image chunk limit.
#1159 — Salesforce Apex extractor (.cls, .trigger)
Regex-based extractor: classes, interfaces, enums, methods, triggers,
SOQL/DML edges. No new dependency. Dispatched as .cls and .trigger.
#1107 — Azure OpenAI Service backend (--backend azure)
Uses AzureOpenAI SDK client (from existing openai package). Auto-detects
when AZURE_OPENAI_API_KEY + AZURE_OPENAI_ENDPOINT both set. Uses
max_completion_tokens (not deprecated max_tokens).
#1103 — live PostgreSQL introspection (--postgres DSN)
graphify extract --postgres "postgresql://..." introspects tables, views,
routines, and FK relations via information_schema (SERIALIZABLE READ ONLY).
Credentials sanitized on error. New graphify[postgres] extra (psycopg3).
Union-resolved llm.py conflict: Azure functions + bedrock images= param.
Fixed test_image_vision.py mock to accept timeout= kwarg (our #1112).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>