get_neighbors and get_community render every edge/member line unbounded — on a
god node or a large community that floods an MCP client's context window with
100KB+ of text in one tool result. query_graph already solves this with the
~3-chars/token cut in _subgraph_to_text; this applies the same budget rule to
the two line-list tools via a shared helper (_cut_lines_to_budget): cut at a
line boundary, report how many lines were dropped, and point at the narrowing
path (relation_filter / get_node). Default 2000 like query_graph; output under
budget is byte-identical to today.
The top-level `graphify --help` has listed --code-only since #1734, but the
`graphify extract` usage string and the README never mentioned it. A user
evaluating graphify on a "can this run with no network call" constraint sees
the extract usage / README first and can conclude the flag doesn't exist.
Add --code-only to the extract usage line, name it in the Privacy section and
the command reference, and add a test asserting the usage advertises it.
`uv run --with graphifyy python -m graphify` runs the system interpreter, so an
older `graphifyy` in system site-packages loads before the uv-managed one — with
no error. Env overrides like OPENAI_BASE_URL are then ignored and requests 401
against the default endpoint. The existing `skill is from graphify X, package is
Y` warning is the real fingerprint, but users dismiss it as a stale skill.
Add a troubleshooting entry next to the uvx-resolution one: name the symptom,
show how to confirm which module actually loaded, and point at
`uvx --from graphifyy` / uninstalling the stale copy.
Two correctness defects found in a head-to-head benchmark of 0.9.22.
Caller line numbers: explain/affected/get_neighbors/query printed the caller
node's def line for an incoming call, presented as a precise citation, so
click-through landed in the wrong place. The `calls` edge already carries the
true call-site line (engine.py sets it); every caller/relation listing now reads
the traversed edge's source_file:source_location, falling back to the node's own
line only when the edge lacks one.
Silent query truncation: rendered nodes were degree-ordered (a low-degree
definition node ranked last, cut first), the queried symbol wasn't guaranteed to
appear, and the truncation marker sat only at the end so silence read as absence.
Nodes are now ranked by hop distance from the seeds (deterministic), the seed the
question named is rendered first and never truncated, and a prominent TRUNCATED
notice at the top states shown/total counts and how to widen the budget. Also
rewires the seed-first ordering the renderer already supported — a branch merge
had silently dropped the `seeds=` argument, leaving it dead code.
Three issues found reviewing #2072:
- The import-edge repoint loop matched by target id regardless of the edge's
language, so a non-Python dotted import (C# `using Pkg.Mod;`, Java/Go) whose
dangling target coincided with a Python alias got repointed onto a Python file,
fabricating a cross-language import. Gate the rewrite on the edge being
Python-sourced.
- The resolver ancestor walk probed package dirs too, resolving an absolute
`from helpers import x` to a sibling in the current package (Python-2 implicit-
relative semantics) even when `helpers` is external. Only probe sys.path-root
candidates (ancestors without their own __init__.py).
- Bound the __init__.py package-root chain walk by the path depth so a
pathological `/__init__.py` can't loop.
Added a cross-language-guard regression test; fixed the tautological (or->and)
ambiguity assertion.
Python absolute imports were resolved only against the scan root, and file-node
ids are scan-root-relative, so a src-layout project (code under src/) lost most
of its imports/imports_from edges when scanned from the repo root — the dangling
edges were silently dropped, so the graph looked complete but wasn't. The chosen
scan root thus silently changed the graph.
Two fixes: (1) _resolve_python_module_path probes the scan root first, then walks
up from the importing file toward the root so a nested package root (src/pkg)
resolves (mirrors the Lua upward walk); (2) a Python post-pass detects each
file's package root via its __init__.py chain and repoints absolute-import edge
targets (dotted-module id -> real file-node id), guarded against shadowing an
existing id and against ambiguous aliases claimed by >1 file. Result: byte-
identical import edges whether scanned from the repo root or from src/.
build_from_json's #1145 ghost-duplicate merge keyed on (Path(source_file).name,
label), discarding the directory, so unrelated nodes from different files sharing
a common basename (index.md, README.md, ...) and a generic label were silently
merged onto one survivor with their edges rewired — corrupting multi-corpus doc
graphs. The AST/LLM ghost twins the merge legitimately targets always share the
same source_file, so keying on the full normalized source_file preserves #1145
while making cross-directory false merges impossible. This subsumes the
#1753/#1257 cross-file ambiguity guard (now removed as dead code). Independent of
the #2032 label pass. Updated the #1257 test to the now-correct precise merge.
A `cluster-only --no-label` run wrote "Community N" placeholders into
.graphify_labels.json plus a matching .sig, and the reuse path treated them as
fresh, so real labels were never regenerated on later runs. Two fixes: (a) don't
persist the labels sidecar (or its .sig) on a placeholder-only run, so a later
run generates real labels; (b) treat a stored "Community {cid}" as absent in the
reuse path so an already-polluted sidecar self-heals via the hub labeler while
genuine labels are still reused with no LLM call. The watch/update rebuild had
the same placeholder-perpetuation twin — fixed alongside.
`graphify path` (and the MCP shortest_path tool) ran shortest_path over
G.to_undirected(as_view=True), whose neighbor iteration is a hash-seeded set
union, so among equal-length paths BFS returned a route that varied per process.
Build a sorted, materialized undirected graph so the chosen path is canonical.
The hop label also printed a relation read from an arbitrarily-collapsed parallel
edge, so it could show `calls` on a pair that only carries `references`. Force
multigraph on the cli path reload so parallel links survive, and render the
ACTUAL stored relation(s) via edge_datas, falling back to an honest "related"
when the edge has none. Serve's shared graph is left untouched (its degree feeds
query-seed tie-breaks); the fix is applied locally in both path readers.
The uninstall strip used an unanchored regex `## graphify`, which matched inside
a user's `### graphify` heading and deleted hand-written content; the
`marker not in content` guard was a substring test that passed on the same
mention. Add a shared `_remove_marker_section` helper that matches the heading
only when a line is exactly the marker (mirroring the install-side #1688
hardening), running each section to the next same-level heading or EOF, and
returning None (leave the file untouched) when no exact heading exists. Replace
all six strip sites: GEMINI.md, copilot-instructions.md, AGENTS.md, CLAUDE.md
(_strip_graphify_md_section), CODEBUDDY.md, and the H1 skill-registration
(_remove_claude_skill_registration, which had the same bug with `# graphify`).
GitHub is the top organic source of waitlist signups, but the only path to the
site was the logo and a CTA at the very bottom of the README. Add a short,
plain-voice waitlist line right after the value proposition, and change the
Enterprise-section CTA from "Learn more" to an explicit "Join the waitlist".
The build_from_json label pass only covered the clustered/update paths; the raw
`extract --no-cluster` path writes the merged node list directly, so colliding
basenames stayed un-disambiguated there. Factor the logic into a shared
_file_label_reassignments core with a list-based variant
(disambiguate_file_labels_in_nodes) and apply it on the raw merged nodes.
Caught by the clean-venv edge-case battery.
In directory-per-entrypoint repos (Supabase Edge Functions, Next.js page.tsx,
Rust mod.rs, Python __init__.py) many files share a basename, so basename-only
file-node labels collided and `explain`/free-text discovery couldn't resolve
them — exactly the highest-value files. build_from_json now runs a final pass
that gives colliding-basename file nodes the shortest unique directory-qualified
label (`process-order/index.ts`); unique basenames stay bare, and node ids/edges
are never touched. The pass runs after the alias-competition (which still needs
bare basenames), is idempotent (labels derive from source_file), and the
downstream file-node predicates (analyze god-nodes, tree_html, serve lookup)
recognize the qualified form via a shared _is_file_node_label helper.
In _extract_generic's class branch the `contains` edge was hard-coded to source
from the file node, so a nested class/object/trait attached to the file instead
of its enclosing type across ~19 languages (only C# even flagged the node with
is_nested_type, and still emitted no edge to the parent). The edge now sources
from parent_class_nid when set, else the file node — keeping the containment
tree connected (file -> Outer -> Inner). A `!= class_nid` guard avoids a
self-loop when same-name nesting collides ids (class ids omit the enclosing
name). The C# is_nested_type flag is retained (load-bearing for cross-file
resolution). Methods were already parent-sourced and are unaffected.
Part 2: `god_nodes` was an analyzer, an MCP tool, and a README-advertised
capability, but `graphify god_nodes` errored with "unknown command". Add a
read-only `god-nodes`/`god_nodes` subcommand mirroring `affected` (--graph,
--top, --json), routing labels through sanitize_label.
Part 3: `--output DIR` on `extract` was silently dropped (output fell back to
the default dir). It is now an alias of `--out` (both space and =forms), matching
what `graphify tree` already documents. Help/usage text updated.
Part 1 (affected/reverse-dep import-id mismatch) is deferred — a build-time
id-resolution change, tracked separately.
#2058: `_is_noise_dir` treated any directory named `env`/`.env`/`*_env` as a
Python virtualenv and pruned it during the walk — before `.graphifyignore`
negation, with zero trace in any returned bucket. Real source dirs with those
names (common in UVM/ASIC verification trees) were silently lost. The venv
heuristic for those names is now gated on actual markers (`pyvenv.cfg`,
`bin`/`Scripts/activate`, `lib/python*`, `conda-meta/`); `venv`/`.venv`/`*_venv`
stay name-only. Pruned-as-noise dirs are recorded in a new `pruned_noise_dirs`
bucket for traceability, and extract.py's walk call sites pass the parent so
genuine venvs are still marker-checked and pruned.
#2059: Office and Google-Workspace sidecars were named with a hash of the
resolved ABSOLUTE source path, so the same tracked file in two clones/worktrees
produced two differently-named byte-identical sidecars — unbounded duplicates
when graphify-out/ is committed, each ingested as a distinct source doc. The
hash is now over the scan-root-relative (NFC-normalized) path, stable across
checkouts while still disambiguating same-stem files; out-of-root sources fall
back to the old absolute form. Also fixes the same bug in google_workspace's
`_sidecar_path` (which additionally never had the #1226 NFC fix).
Per review on #2017 (thanks @HerenderKumar): add a guard asserting a
non-decode ValueError (e.g. a non-.json path) still prints the plain
"must be a .json file" message, not the corrupted-graph hint. Locks in
the except-clause order from the parent fix so a future refactor can't
silently collapse the two branches back together.
json.JSONDecodeError subclasses ValueError, so the broader
`except (ValueError, FileNotFoundError)` clause always matched first,
making the intended "graph.json is corrupted (...). Re-run /graphify
to rebuild." recovery hint dead code — users with a truncated graph.json
got the bare json.JSONDecodeError message instead, contradicting the
behavior SECURITY.md documents for this exact threat.
Move the json.JSONDecodeError clause first so it actually catches.
Audited the rest of serve.py's exception handlers for the same
narrower-after-broader ordering bug; found no other instance.
Fixes#2005.
The #2051 disk-absence sweep guarded remote/virtual sources with a literal
`"://"` check. But the write-side path normalization (Path.as_posix) collapses
the double slash, so a stored `gdoc://x` reads back as `gdoc:/x` on the next
update; the literal guard then missed it and the node fell into the disk-absence
branch (`Path('gdoc:/x').exists()` is False), evicting it on the second
`graphify update`. Match the scheme with a regex tolerant of the collapse, with
a 2+ char scheme so a Windows drive letter (C:/) is not misread as remote.
Regression test runs three consecutive updates and asserts the remote node
survives every one.
The --update runbook's Step 9 stamped the entire detected corpus into the
manifest, so a semantic file (doc/paper/image) whose chunk failed or was omitted
was recorded as done and never re-queued on the next update — losing its content
permanently. The runbook now builds the manifest the way the library extract
path does: cli._stamped_manifest_files stamps only files that actually produced
nodes/edges/hyperedges, dispatched-but-empty files have their stale hash cleared,
and scan_corpus drops newly-excluded in-root rows.
Applied to the Claude, Aider, and Devin skill bodies and the shared update
reference; all 134 artifacts regenerated. gen.py gains a sanctioned monolith-
diff predicate for the new stamping lines.
build_merge now prunes a deleted file's nodes, edges, and hyperedges regardless
of whether their stored source_file is absolute or relative. When the caller
passed no root (the --update runbook), a node that kept an absolute path slipped
past the relative prune set and the deleted file's graph survived silently.
Matching is now form-insensitive (raw, normalized-relative, then an absolute-
identity fallback); a re-extracted file is still never pruned (#1796 preserved).
`graphify extract` also writes the .graphify_root marker after every graph write
so a later build_merge relativizes deleted-file paths correctly even under a
custom --out (its grandparent-of-graph.json fallback pointed at the wrong dir).
Regression tests in tests/test_build_merge_hyperedges_and_prune.py.
Three silent-data-loss fixes in the update/reconcile path:
- #2051: a full `graphify update` now evicts semantic nodes whose non-AST
source (a .txt/.pdf/.png with no code extractor) was deleted from disk.
The corpus sweep only checked re-extractable files, so deleted docs'/images'
LLM nodes survived as authoritative forever. Disk absence is now the deletion
signal; remote/virtual sources (`://`) are left untouched.
- #2056: an incremental rebuild whose change set names a present-but-
unextractable file no longer treats it as a deletion (which evicted its
semantic nodes and disabled the shrink guard). The guard now falls through to
per-source accounting instead of a wholesale bypass on any deletion.
- #2014: code-typed nodes the semantic pass surfaces from within a document now
count as that doc's semantic layer, so a rebuild doesn't re-scan and drop them.
Regression tests for each in tests/test_watch.py.
The barrel-chain resolver keyed (source, symbol) -> target as last-write-wins, so
a barrel re-exporting the same local name from two modules collapsed an importer's
edge onto whichever was learned last — a fabricated wrong edge. Learn into a set
per key and refuse to resolve when a name maps to more than one target; the edge
falls to the dangling-canonical fallback (dropped at build) instead. Adds a
regression test.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The PR resolved OLLAMA_HOST for the client base_url but not in detect_backend(),
so the headline #1940 case (OLLAMA_HOST set, no --backend) still errored with
"no LLM API key found". detect_backend() now uses _resolve_ollama_base_url (empty
default stays falsy, so ollama remains opt-in and never shadows a paid key). Also
default the port to 11434 when OLLAMA_HOST omits it (a bare host would otherwise
resolve to port 80), and handle bare-port / :port forms.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Ollama's own server has no OLLAMA_BASE_URL concept — it reads
OLLAMA_HOST (https://docs.ollama.com/faq#how-do-i-configure-ollama-server).
graphify only ever read OLLAMA_BASE_URL, so anyone who configured Ollama
the way its own docs describe (OLLAMA_HOST) had graphify silently ignore
it and fall back to the localhost:11434 default.
Add _resolve_ollama_base_url(default): OLLAMA_BASE_URL still wins when
set (unchanged behavior, graphify's existing convention for OpenAI-
compatible base_url overrides across backends). Otherwise falls back to
OLLAMA_HOST, normalized into an OpenAI-compatible URL (adds a scheme if
missing, ensures a trailing /v1). Wired into the BACKENDS dict's ollama
entry and the two ollama_url call sites that read the env var directly.
Fixes#1940.
Publishes graphifyy to PyPI via GitHub OIDC on a published Release — no API
token. Builds sdist+wheel with uv, guards that the package version matches the
release tag, twine-checks, then uploads via pypa/gh-action-pypi-publish with
id-token: write and the `pypi` environment. Requires a matching GitHub trusted
publisher configured on PyPI (Owner Graphify-Labs, repo graphify, workflow
publish.yml, environment pypi).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Homepage/Repository/Issues were stale from the org move (safishamsi/graphify).
Aligning them to Graphify-Labs/graphify fixes the PyPI "Issues" link and is a
prerequisite for PyPI to mark the Repository link Verified once releases publish
via a GitHub trusted publisher.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
With `graphify extract --out <dir>`, the semantic cache write and read
sides disagreed on both location and key anchoring, breaking the cache
round-trip in two ways:
- Checkpoints (#1990): `_checkpoint_chunk` called `save_semantic_cache`
with only `root=target`, so per-chunk recovery checkpoints were written
under `<corpus>/graphify-out/` while the reader consulted
`<out>/graphify-out/` — creating an unwanted graphify-out/ inside the
analyzed source tree and making every interrupted run re-extract (and
re-bill) completed chunks.
- Final save (#1991): cli.py passed `root=out_root`, so corpus-relative
`source_file` paths resolved against the --out directory, failed
`p.is_file()`, and every result group was silently skipped — the cache
the reader would consult was never populated at all, with no warning.
Fix, following the split the AST cache already uses (#1774):
- `save_semantic_cache` and `check_semantic_cache` gain a `cache_root`
parameter mirroring `load_cached`/`save_cached`: `root` stays the
source-key anchor (content-hash keys, source_file resolution and
relativization), `cache_root` selects where cache files live. Omitting
it keeps `root` for both, so existing callers are unchanged.
- `extract_corpus_parallel` plumbs `cache_root` into `_checkpoint_chunk`.
- cli.py extract passes `root=target, cache_root=out_root` at the cache
read, the checkpoint path, and the final save, and re-anchors the
prune sweep's live hashes to `target` (keys anchored to out_root would
mismatch every entry and sweep the fresh cache as orphaned).
- `save_semantic_cache` now warns loudly when every result group is
dropped because its source_file does not resolve to a real file — the
silent-0-writes failure mode #1991 asked to surface.
Regression tests cover: checkpoint written under cache_root (not the
corpus, no corpus graphify-out/ created), recovery read finds the
checkpoint via the same root/cache_root split, the final-save call shape
writes entries where the reader looks, the all-groups-dropped warning,
and backward compatibility when cache_root is omitted.
Fixes#1990Fixes#1991
ee1df22 narrowed the Claude Code search-guard matcher from "Glob|Grep" to
"Bash" on the premise that dedicated search tools were removed and searches
go through Bash. Current Claude Code routes content search through its
first-class Grep tool (its Bash tool description actively steers away from
shell grep), so the graphify-first nudge never fired on the agent's primary
exploration path and the graph was silently bypassed.
Three-part fix, per the issue's analysis:
- Matcher: "Bash" -> "Bash|Grep" in _claude_pretooluse_hooks. Glob already
fires the read nudge via "Read|Glob", so Grep was the only orphaned tool.
- Guard body: the hook-guard search branch only inspected tool_input.command,
which a Grep call doesn't carry (it has pattern/path/glob). A Grep-shaped
input (pattern present, no command) is now treated as a search — it IS one
by definition — and nudges whenever a fresh graph exists. The Bash
token-matching path is unchanged, and a command-carrying input never
triggers the Grep shape, so non-search Bash calls stay silent.
- Idempotency: "Bash|Grep" added to the four install/uninstall dedup filters
(claude + codebuddy), so upgrading replaces the stale "Bash" hook in place
instead of appending a duplicate — verified against a pre-fix settings.json.
Tests: new regression tests feed Grep-shaped tool_input through
hook-guard search and assert the nudge (with graph), silence (without),
valid PreToolUse JSON, and no blocking; plus a guard that a non-search Bash
command with a stray pattern key does not nudge. Existing matcher assertions
updated across test_search_hook/test_install/test_claude_md/test_codebuddy/
test_hook_strict. Hook+install suites: 397 passed. Full suite: 3224 passed;
the 13 failures are pre-existing on clean v8 in this environment.
Fixes#1986
Claude Code runs command-type PreToolUse hooks through Git Bash by default on
Windows. The resolved exe was emitted as a raw backslash path and quoted only
when it contained a space, so a space-free path like C:\Users\me\graphify.EXE
reached settings.json unquoted. Git Bash treats the unquoted backslashes as
escapes and strips them, producing 'C:Usersmegraphify.EXE: command not found',
so every graph guard silently fails. Normalize the path to forward slashes in
_resolve_graphify_exe (a no-op on POSIX), fixing the Claude, Gemini, and Codex
hook emitters at once.
Rename the "Penpax" section to "graphify Enterprise" (linking graphify.com) and
add a Website badge to the Community and links footer.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Stamp the 0.9.19 release date, add a README note for the strict PreToolUse hook
(graphify install --project --strict), and credit both the #1814 reporter
(@Greg-Moskalenko) and the fix author (@alphanury).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The PR always wrote gitignore=not no_gitignore into .graphify_build.json, so a
flag-less `graphify extract` after `--no-gitignore` reset it to True and the
git-ignored code silently disappeared again — the exact #1971 complaint. Write
False only when the flag is set (None = leave as-is, mirroring #1886 excludes),
and honor the persisted value for the run when the flag is absent.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>