GitHub is the top organic source of waitlist signups, but the only path to the
site was the logo and a CTA at the very bottom of the README. Add a short,
plain-voice waitlist line right after the value proposition, and change the
Enterprise-section CTA from "Learn more" to an explicit "Join the waitlist".
The build_from_json label pass only covered the clustered/update paths; the raw
`extract --no-cluster` path writes the merged node list directly, so colliding
basenames stayed un-disambiguated there. Factor the logic into a shared
_file_label_reassignments core with a list-based variant
(disambiguate_file_labels_in_nodes) and apply it on the raw merged nodes.
Caught by the clean-venv edge-case battery.
In directory-per-entrypoint repos (Supabase Edge Functions, Next.js page.tsx,
Rust mod.rs, Python __init__.py) many files share a basename, so basename-only
file-node labels collided and `explain`/free-text discovery couldn't resolve
them — exactly the highest-value files. build_from_json now runs a final pass
that gives colliding-basename file nodes the shortest unique directory-qualified
label (`process-order/index.ts`); unique basenames stay bare, and node ids/edges
are never touched. The pass runs after the alias-competition (which still needs
bare basenames), is idempotent (labels derive from source_file), and the
downstream file-node predicates (analyze god-nodes, tree_html, serve lookup)
recognize the qualified form via a shared _is_file_node_label helper.
In _extract_generic's class branch the `contains` edge was hard-coded to source
from the file node, so a nested class/object/trait attached to the file instead
of its enclosing type across ~19 languages (only C# even flagged the node with
is_nested_type, and still emitted no edge to the parent). The edge now sources
from parent_class_nid when set, else the file node — keeping the containment
tree connected (file -> Outer -> Inner). A `!= class_nid` guard avoids a
self-loop when same-name nesting collides ids (class ids omit the enclosing
name). The C# is_nested_type flag is retained (load-bearing for cross-file
resolution). Methods were already parent-sourced and are unaffected.
Part 2: `god_nodes` was an analyzer, an MCP tool, and a README-advertised
capability, but `graphify god_nodes` errored with "unknown command". Add a
read-only `god-nodes`/`god_nodes` subcommand mirroring `affected` (--graph,
--top, --json), routing labels through sanitize_label.
Part 3: `--output DIR` on `extract` was silently dropped (output fell back to
the default dir). It is now an alias of `--out` (both space and =forms), matching
what `graphify tree` already documents. Help/usage text updated.
Part 1 (affected/reverse-dep import-id mismatch) is deferred — a build-time
id-resolution change, tracked separately.
#2058: `_is_noise_dir` treated any directory named `env`/`.env`/`*_env` as a
Python virtualenv and pruned it during the walk — before `.graphifyignore`
negation, with zero trace in any returned bucket. Real source dirs with those
names (common in UVM/ASIC verification trees) were silently lost. The venv
heuristic for those names is now gated on actual markers (`pyvenv.cfg`,
`bin`/`Scripts/activate`, `lib/python*`, `conda-meta/`); `venv`/`.venv`/`*_venv`
stay name-only. Pruned-as-noise dirs are recorded in a new `pruned_noise_dirs`
bucket for traceability, and extract.py's walk call sites pass the parent so
genuine venvs are still marker-checked and pruned.
#2059: Office and Google-Workspace sidecars were named with a hash of the
resolved ABSOLUTE source path, so the same tracked file in two clones/worktrees
produced two differently-named byte-identical sidecars — unbounded duplicates
when graphify-out/ is committed, each ingested as a distinct source doc. The
hash is now over the scan-root-relative (NFC-normalized) path, stable across
checkouts while still disambiguating same-stem files; out-of-root sources fall
back to the old absolute form. Also fixes the same bug in google_workspace's
`_sidecar_path` (which additionally never had the #1226 NFC fix).
Per review on #2017 (thanks @HerenderKumar): add a guard asserting a
non-decode ValueError (e.g. a non-.json path) still prints the plain
"must be a .json file" message, not the corrupted-graph hint. Locks in
the except-clause order from the parent fix so a future refactor can't
silently collapse the two branches back together.
json.JSONDecodeError subclasses ValueError, so the broader
`except (ValueError, FileNotFoundError)` clause always matched first,
making the intended "graph.json is corrupted (...). Re-run /graphify
to rebuild." recovery hint dead code — users with a truncated graph.json
got the bare json.JSONDecodeError message instead, contradicting the
behavior SECURITY.md documents for this exact threat.
Move the json.JSONDecodeError clause first so it actually catches.
Audited the rest of serve.py's exception handlers for the same
narrower-after-broader ordering bug; found no other instance.
Fixes#2005.
The #2051 disk-absence sweep guarded remote/virtual sources with a literal
`"://"` check. But the write-side path normalization (Path.as_posix) collapses
the double slash, so a stored `gdoc://x` reads back as `gdoc:/x` on the next
update; the literal guard then missed it and the node fell into the disk-absence
branch (`Path('gdoc:/x').exists()` is False), evicting it on the second
`graphify update`. Match the scheme with a regex tolerant of the collapse, with
a 2+ char scheme so a Windows drive letter (C:/) is not misread as remote.
Regression test runs three consecutive updates and asserts the remote node
survives every one.
The --update runbook's Step 9 stamped the entire detected corpus into the
manifest, so a semantic file (doc/paper/image) whose chunk failed or was omitted
was recorded as done and never re-queued on the next update — losing its content
permanently. The runbook now builds the manifest the way the library extract
path does: cli._stamped_manifest_files stamps only files that actually produced
nodes/edges/hyperedges, dispatched-but-empty files have their stale hash cleared,
and scan_corpus drops newly-excluded in-root rows.
Applied to the Claude, Aider, and Devin skill bodies and the shared update
reference; all 134 artifacts regenerated. gen.py gains a sanctioned monolith-
diff predicate for the new stamping lines.
build_merge now prunes a deleted file's nodes, edges, and hyperedges regardless
of whether their stored source_file is absolute or relative. When the caller
passed no root (the --update runbook), a node that kept an absolute path slipped
past the relative prune set and the deleted file's graph survived silently.
Matching is now form-insensitive (raw, normalized-relative, then an absolute-
identity fallback); a re-extracted file is still never pruned (#1796 preserved).
`graphify extract` also writes the .graphify_root marker after every graph write
so a later build_merge relativizes deleted-file paths correctly even under a
custom --out (its grandparent-of-graph.json fallback pointed at the wrong dir).
Regression tests in tests/test_build_merge_hyperedges_and_prune.py.
Three silent-data-loss fixes in the update/reconcile path:
- #2051: a full `graphify update` now evicts semantic nodes whose non-AST
source (a .txt/.pdf/.png with no code extractor) was deleted from disk.
The corpus sweep only checked re-extractable files, so deleted docs'/images'
LLM nodes survived as authoritative forever. Disk absence is now the deletion
signal; remote/virtual sources (`://`) are left untouched.
- #2056: an incremental rebuild whose change set names a present-but-
unextractable file no longer treats it as a deletion (which evicted its
semantic nodes and disabled the shrink guard). The guard now falls through to
per-source accounting instead of a wholesale bypass on any deletion.
- #2014: code-typed nodes the semantic pass surfaces from within a document now
count as that doc's semantic layer, so a rebuild doesn't re-scan and drop them.
Regression tests for each in tests/test_watch.py.
The barrel-chain resolver keyed (source, symbol) -> target as last-write-wins, so
a barrel re-exporting the same local name from two modules collapsed an importer's
edge onto whichever was learned last — a fabricated wrong edge. Learn into a set
per key and refuse to resolve when a name maps to more than one target; the edge
falls to the dangling-canonical fallback (dropped at build) instead. Adds a
regression test.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The PR resolved OLLAMA_HOST for the client base_url but not in detect_backend(),
so the headline #1940 case (OLLAMA_HOST set, no --backend) still errored with
"no LLM API key found". detect_backend() now uses _resolve_ollama_base_url (empty
default stays falsy, so ollama remains opt-in and never shadows a paid key). Also
default the port to 11434 when OLLAMA_HOST omits it (a bare host would otherwise
resolve to port 80), and handle bare-port / :port forms.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Ollama's own server has no OLLAMA_BASE_URL concept — it reads
OLLAMA_HOST (https://docs.ollama.com/faq#how-do-i-configure-ollama-server).
graphify only ever read OLLAMA_BASE_URL, so anyone who configured Ollama
the way its own docs describe (OLLAMA_HOST) had graphify silently ignore
it and fall back to the localhost:11434 default.
Add _resolve_ollama_base_url(default): OLLAMA_BASE_URL still wins when
set (unchanged behavior, graphify's existing convention for OpenAI-
compatible base_url overrides across backends). Otherwise falls back to
OLLAMA_HOST, normalized into an OpenAI-compatible URL (adds a scheme if
missing, ensures a trailing /v1). Wired into the BACKENDS dict's ollama
entry and the two ollama_url call sites that read the env var directly.
Fixes#1940.
Publishes graphifyy to PyPI via GitHub OIDC on a published Release — no API
token. Builds sdist+wheel with uv, guards that the package version matches the
release tag, twine-checks, then uploads via pypa/gh-action-pypi-publish with
id-token: write and the `pypi` environment. Requires a matching GitHub trusted
publisher configured on PyPI (Owner Graphify-Labs, repo graphify, workflow
publish.yml, environment pypi).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Homepage/Repository/Issues were stale from the org move (safishamsi/graphify).
Aligning them to Graphify-Labs/graphify fixes the PyPI "Issues" link and is a
prerequisite for PyPI to mark the Repository link Verified once releases publish
via a GitHub trusted publisher.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
With `graphify extract --out <dir>`, the semantic cache write and read
sides disagreed on both location and key anchoring, breaking the cache
round-trip in two ways:
- Checkpoints (#1990): `_checkpoint_chunk` called `save_semantic_cache`
with only `root=target`, so per-chunk recovery checkpoints were written
under `<corpus>/graphify-out/` while the reader consulted
`<out>/graphify-out/` — creating an unwanted graphify-out/ inside the
analyzed source tree and making every interrupted run re-extract (and
re-bill) completed chunks.
- Final save (#1991): cli.py passed `root=out_root`, so corpus-relative
`source_file` paths resolved against the --out directory, failed
`p.is_file()`, and every result group was silently skipped — the cache
the reader would consult was never populated at all, with no warning.
Fix, following the split the AST cache already uses (#1774):
- `save_semantic_cache` and `check_semantic_cache` gain a `cache_root`
parameter mirroring `load_cached`/`save_cached`: `root` stays the
source-key anchor (content-hash keys, source_file resolution and
relativization), `cache_root` selects where cache files live. Omitting
it keeps `root` for both, so existing callers are unchanged.
- `extract_corpus_parallel` plumbs `cache_root` into `_checkpoint_chunk`.
- cli.py extract passes `root=target, cache_root=out_root` at the cache
read, the checkpoint path, and the final save, and re-anchors the
prune sweep's live hashes to `target` (keys anchored to out_root would
mismatch every entry and sweep the fresh cache as orphaned).
- `save_semantic_cache` now warns loudly when every result group is
dropped because its source_file does not resolve to a real file — the
silent-0-writes failure mode #1991 asked to surface.
Regression tests cover: checkpoint written under cache_root (not the
corpus, no corpus graphify-out/ created), recovery read finds the
checkpoint via the same root/cache_root split, the final-save call shape
writes entries where the reader looks, the all-groups-dropped warning,
and backward compatibility when cache_root is omitted.
Fixes#1990Fixes#1991
ee1df22 narrowed the Claude Code search-guard matcher from "Glob|Grep" to
"Bash" on the premise that dedicated search tools were removed and searches
go through Bash. Current Claude Code routes content search through its
first-class Grep tool (its Bash tool description actively steers away from
shell grep), so the graphify-first nudge never fired on the agent's primary
exploration path and the graph was silently bypassed.
Three-part fix, per the issue's analysis:
- Matcher: "Bash" -> "Bash|Grep" in _claude_pretooluse_hooks. Glob already
fires the read nudge via "Read|Glob", so Grep was the only orphaned tool.
- Guard body: the hook-guard search branch only inspected tool_input.command,
which a Grep call doesn't carry (it has pattern/path/glob). A Grep-shaped
input (pattern present, no command) is now treated as a search — it IS one
by definition — and nudges whenever a fresh graph exists. The Bash
token-matching path is unchanged, and a command-carrying input never
triggers the Grep shape, so non-search Bash calls stay silent.
- Idempotency: "Bash|Grep" added to the four install/uninstall dedup filters
(claude + codebuddy), so upgrading replaces the stale "Bash" hook in place
instead of appending a duplicate — verified against a pre-fix settings.json.
Tests: new regression tests feed Grep-shaped tool_input through
hook-guard search and assert the nudge (with graph), silence (without),
valid PreToolUse JSON, and no blocking; plus a guard that a non-search Bash
command with a stray pattern key does not nudge. Existing matcher assertions
updated across test_search_hook/test_install/test_claude_md/test_codebuddy/
test_hook_strict. Hook+install suites: 397 passed. Full suite: 3224 passed;
the 13 failures are pre-existing on clean v8 in this environment.
Fixes#1986
Claude Code runs command-type PreToolUse hooks through Git Bash by default on
Windows. The resolved exe was emitted as a raw backslash path and quoted only
when it contained a space, so a space-free path like C:\Users\me\graphify.EXE
reached settings.json unquoted. Git Bash treats the unquoted backslashes as
escapes and strips them, producing 'C:Usersmegraphify.EXE: command not found',
so every graph guard silently fails. Normalize the path to forward slashes in
_resolve_graphify_exe (a no-op on POSIX), fixing the Claude, Gemini, and Codex
hook emitters at once.
Rename the "Penpax" section to "graphify Enterprise" (linking graphify.com) and
add a Website badge to the Community and links footer.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Stamp the 0.9.19 release date, add a README note for the strict PreToolUse hook
(graphify install --project --strict), and credit both the #1814 reporter
(@Greg-Moskalenko) and the fix author (@alphanury).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The PR always wrote gitignore=not no_gitignore into .graphify_build.json, so a
flag-less `graphify extract` after `--no-gitignore` reset it to True and the
git-ignored code silently disappeared again — the exact #1971 complaint. Write
False only when the flag is set (None = leave as-is, mirroring #1886 excludes),
and honor the persisted value for the run when the flag is absent.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
`graphify prs` reads gh, git, and claude output via subprocess.run(text=True)
with no explicit encoding=, so on Windows stdout is decoded with the cp1252
locale codec. gh emits UTF-8 JSON whose PR titles and logins routinely carry
non-Latin1 bytes (emoji, or the Persian title in PR #1281), and cp1252 has no
mapping for bytes such as 0x81. The capture reader thread raises
UnicodeDecodeError, subprocess.run then returns stdout=None, and the caller dies
one line later on json.loads(None) (TypeError) or None.splitlines()
(AttributeError). Neither is caught by _gh's except clause, so the whole `prs`
command aborts, with a spurious reader-thread traceback on stderr, against any
repo that has non-ASCII PR metadata.
Pass encoding="utf-8", errors="replace" to all five subprocess reads in prs.py
so decoding matches Linux and macOS. This is the decode-side sibling of #1505,
which applied the same change to the llm.py claude-cli subprocess. CI runs Linux
only (a UTF-8 locale), so this path is never exercised there.
Reproduced on Windows 11 / CPython 3.13 with a cp1252 locale:
prs._gh("pr", "view", "1281", ...) crashed with TypeError before and returns the
decoded title after. The added regression tests assert encoding="utf-8" on each
call, mirroring tests/test_charmap_encoding.py, and are red without the fix.
`extract_dm` normalized include paths with `raw.lstrip("./")`. `str.lstrip`
treats its argument as a *set of characters*, not a prefix, so it strips
every leading `.` and `/` — destroying parent-relative includes:
"../shared/base.dm" -> "shared/base.dm"
"../../a.dm" -> "a.dm"
".hidden/x.dm" -> "hidden/x.dm"
`(path.parent / norm).resolve()` then points at the wrong directory,
`.exists()` returns False, and a correct internal `imports_from` edge to the
real included file becomes a bogus external `imports` edge to a phantom node,
silently losing the cross-file link.
The sibling local-include resolver in `extractors/bash.py` resolves
`(path.parent / raw)` from the raw path without stripping `..`, confirming
the intent. Strip only a literal leading `./` (re is already imported) so
`../` includes resolve correctly; the `./`-prefix and no-prefix cases are
unchanged.
The digest salts content with the path relative to root (portability, #1774),
but the stat-index memo was keyed by absolute path only — so the same file
hashed under two roots (which happens within one `--out` run) served whichever
digest was computed first, making file_hash order-dependent and poisoning the
persisted stat-index across runs. Store one digest per salt under a "hashes"
map; legacy un-salted "hash" entries are never trusted (recompute once). Digest
computation is byte-identical, so existing cache entries still hit.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Agents routinely ignore the advisory "run graphify query first" nudge and read
raw files anyway. `graphify install --project --strict` (or `graphify claude
install --strict`) now installs a hook that BLOCKS the first raw source read of a
session via permissionDecision:"deny" with a redirect to graphify query, then
downgrades to the soft nudge — it fires at most once per session (atomic
per-session marker) so it can never strand the agent, and a recent
query/explain/path refreshes a stamp that suppresses it. Claude Code only;
Bash-grep and Glob stay nudge-only; Gemini/Codex/OpenCode are unchanged.
GRAPHIFY_HOOK_STRICT=1/0 toggles at runtime without a reinstall; default installs
are byte-identical (soft nudge).
Also fixes#1840 for the default soft nudge: the guard no longer fires for reads
of out-of-project files, and softens to a non-mandatory nudge when the graph is
stale for the target file. Gating is ~3 stat calls (no corpus walk) and fails
open. Begins the 0.9.19 cycle.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Stamp the 0.9.18 release date and add a README troubleshooting entry for the
new #1951 behavior: an incomplete extraction refuses to overwrite a larger
existing graph, with --allow-partial as the escape hatch.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
`graphify extract <root> --out <dir>` reduced every node's `source_file` to a
bare filename, so a graph.json could no longer be resolved back to files on
disk by joining source_file onto the scan root.
--out passes the output dir as cache_root to relocate the cache, but that value
also anchored relativization. Every scanned file then failed relative_to(root),
fell through to the #1899 out-of-root fallback, tripped its `updepth > 3`
walk-up guard -- written for a stray ProjectReference, not a whole corpus -- and
collapsed to a basename. On Windows an --out on another drive hit the
cross-drive branch and basenamed unconditionally, which is what the reporter
saw: 0 of ~120k source_files kept a separator. The directory survived only in
the node id slug, lossily (`.`, `/`, `\`, `-`, spaces all map to `_`), and no
other field carried it -- origin_file is stripped (#1516) and the export has no
file table -- so 0% of nodes resolved.
extract() now takes an explicit `root` anchor for source_file/ids/symbol
resolution, which the CLI pins to the scan root independent of where the cache
lives. This completes the cache/anchor decoupling #1774 started and matches
build(root=target), which already anchored on the scan root -- extract was the
lone component keying off --out.
cache_root keeps its fallback-anchor role, so callers that pass the scan root as
cache_root (watch, the no---out CLI path, tests) are unchanged, as are cache
location and out-of-root portability (#1899).
Follow-up to #1956: a doc stamped complete on a prior run that truncates
(partial) this run must have its stale semantic_hash cleared so it is re-queued.
Partial files are dropped by _stamped_manifest_files and therefore land in
clear_semantic — pin that composed behavior.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The #1943 carve-out rescues genuine source under secrets/ / credentials/, but
.tfvars sits in CODE_EXTENSIONS while being Terraform's canonical values store
(routinely real secrets), so it would now be indexed. Add .tfvars to
_SECRET_PRONE_DATA_EXTS so the shared graphable-source predicate drops it in
both Stage 1 and Stage 3; .tf/.hcl stay graphable as genuine infra source.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>