Commit Graph
1144 Commits
Author SHA1 Message Date
tpateeqandsafishamsi f38e98012d fix(io): write graph.json and manifest.json atomically
graph.json (the clustered `to_json` write and the `--no-cluster`/merge raw
dumps) and manifest.json were written with a direct `open()`/`write_text`, so a
crash, kill, or disk-full mid-write left a truncated, unparseable file that the
next load or `detect_incremental` then failed on.

Add `write_text_atomic`/`write_json_atomic` in graphify.paths (temp file in the
same directory + `os.replace`; JSON is streamed into the temp, not materialized
as one string) and route the graph.json writers (export.to_json,
cli._prune_graph_json_sources, the merge driver) plus detect.save_manifest
through them. The helper preserves the destination's mode (an atomic replace
never tightens 0644 to mkstemp's 0600), writes through a symlinked destination
(shared-output setups), and falls back to copy-then-delete on a Windows
os.replace lock — matching graphify.cache's existing atomic writer. On failure
the previous file is left intact and the temp removed. Not a power-loss
durability guarantee (no fsync, consistent with the rest of the codebase).
2026-07-16 23:42:47 +01:00
safishamsiandClaude Opus 4.8 7d53a02704 fix(extract): fail closed on malformed existing graph + honor walk_errors (follow-up to #1951)
Two gaps the review found in the incomplete-build shrink guard:
- existing_graph_node_count() returned None ("proceed") on a present-but-
  unparseable graph.json, so the --no-cluster path could clobber a complete
  graph whose file was corrupt/mid-write. It now returns a MALFORMED_GRAPH
  sentinel and the caller fails closed, matching to_json's #479 handling.
- A walk that couldn't fully enumerate the corpus (permission-denied subtree,
  I/O error) is now treated as an incomplete extraction: detect()/
  detect_incremental() already record walk_errors; the extract path consumes
  them so a walk-truncated graph can't force-overwrite a complete one.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-16 23:42:47 +01:00
tpateeqandsafishamsi 0f936850ba fix(extract): extend the incomplete-build shrink guard to the --no-cluster path
The clustered write is guarded by to_json's #479 shrink check, but the
`--no-cluster` raw-dump path writes graph.json directly and had no guard, so an
incomplete `--no-cluster` build could still overwrite a larger complete graph
with a partial one — the residual gap noted in the original change.

Add `existing_graph_node_count` in graphify.export (mirrors to_json's guard and
respects the graph-size cap) and, on the raw path, refuse the write with exit 1
before the manifest when the build was incomplete and the new graph has fewer
nodes than the existing one — unless --allow-partial is passed. Both write paths
now enforce the same guarantee.
2026-07-16 23:38:09 +01:00
tpateeqandsafishamsi fcc01dbcdf fix(extract): don't force-write a partial graph over a complete one
A full `graphify extract` writes the final graph with `to_json(..., force=True)`,
which bypasses the #479 shrink guard. That is correct for a clean build that
legitimately shrinks (dedup collapse, deleted code), but when this run's
extraction was incomplete — an AST pass crashed, or some semantic chunks failed —
forcing the write lets a partial graph silently overwrite a good complete one.

The build now tracks incompleteness (AST-pass failure, semantic-pass crash, or
succeeded < total chunks) and falls back to the shrink guard (force=False) on an
incomplete run, so a smaller partial graph is refused rather than written. It
exits non-zero before the manifest is written, so the manifest is never stamped
for a graph we declined to write and the next run re-attempts. `--allow-partial`
restores force=True to override intentionally.

Note: the `--no-cluster` raw-dump path writes graph.json directly and has no
shrink guard; this change covers the normal clustered build path only.
2026-07-16 23:38:09 +01:00
safishamsiandClaude Opus 4.8 3305ef1059 fix(merge-chunks): coerce non-numeric chunk token counts (follow-up to #1953)
An untrusted chunk with a non-numeric input_tokens/output_tokens would abort
the whole merge with a TypeError after other chunks had already merged. Coerce
to 0 so a bad token field can't defeat the per-chunk validation guard.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-16 23:38:09 +01:00
tpateeqandsafishamsi ee74d80b45 fix(merge-chunks): validate untrusted subagent chunk JSON before merging
`graphify merge-chunks` concatenates agent-written `.graphify_chunk_*.json`
files with only a JSON-decode guard, so an oversized payload or a crafted
node/edge id (e.g. `../../etc/passwd`) flowed straight into the merged graph.

Route each chunk through `load_validated_semantic_fragment`, which stats the
file size BEFORE reading it (a multi-GB chunk can't blow up memory), parses the
JSON, and validates the byte/count caps + the node/edge id charset that blocks
path traversal (#825). An invalid chunk is skipped with a warning (filter
semantics; never abort). Left OUT of build_from_json/load_graph_json on purpose:
those must keep loading valid pre-existing graphs.

Also relax two over-strict checks in the shared validator that would otherwise
silently drop whole legitimate chunks (a relaxation — never a new rejection):
- file_type is no longer gated: build coerces any value via _FILE_TYPE_SYNONYMS
  (unknown -> "concept", #840), so synonyms like "markdown"/"tool" the loader
  maps must not fail validation.
- the id charset now allows Unicode word chars (build's normalize_id preserves
  CJK/Cyrillic/accented-Latin ids); the explicit path-separator/".." check still
  blocks directory escape.
Corrects the stale module docstring (the validator serves the devin skill path
and merge-chunks, not skill-opencode/codex).
2026-07-16 23:36:49 +01:00
safishamsiandClaude Opus 4.8 84638cb69c chore: begin 0.9.18 development cycle
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-16 16:42:18 +01:00
SinghAman21andsafishamsi 0d018b4c6d fix(cache): key semantic cache on the extraction prompt (#1939)
The semantic cache keyed entries on sha256(file content + path) alone, with
no component for the extraction prompt that produced them. After an upgrade
that changed the prompt, every unchanged file was a cache hit and replayed
the older prompt's extraction: the run exited 0, cost.json looked cheap, and
the graph silently carried two prompt generations side by side. The reporter
saw 506 of 512 docs replay an older vintage on a rebuild they expected to be
cold; deleting the whole cache was the only workaround.

output, and invalidating them on every release would re-bill extraction for
unchanged files. Fingerprinting the prompt itself keeps both properties:
entries survive releases that don't touch the prompt, and invalidate only
when it actually changed.

Semantic entries now live under cache/semantic/p{fingerprint}/, mirroring the
AST cache's v{version}/ layout. Both extraction paths pass their prompt: the
Python/CLI path from llm.py's _EXTRACTION_SYSTEM (shared by every backend),
and the skill path via a new prompt_file argument in Step B0/B3 naming the
references/extraction-spec.md the subagents were handed. The fingerprint
normalizes line endings so a CRLF checkout isn't mistaken for a new prompt.

Pre-existing entries predate fingerprinting and have unknowable vintage, so
they are still served rather than re-billing a whole corpus on upgrade — but
check_semantic_cache now warns with the count, turning "no signal at all"
into a visible one. merge_existing refuses to fuse such an entry into a
current-vintage write, which would mix two prompts inside one entry and then
attest the result to a prompt that produced half of it.

Old-fingerprint entries are pruned by liveness only, never swept wholesale
the way stale AST versions are: two hosts with different prompts (verbose vs
compact extraction-spec) can share one graphify-out/, and a wholesale sweep
would have each run delete the other's entries and re-bill on every
alternation. prune/clear/cached_files glob recursively so fingerprinted
entries can't become unprunable orphans (the #1527 failure mode).

The two monolith skills (aider, devin) inline their prompt instead of
shipping a spec sidecar and stay on the unfingerprinted path for now.
2026-07-16 16:38:12 +01:00
5840d0b44d fix(pg_introspect): emit FK DDL before function stubs so unparseable routines can't drop references edges (#1854)
A C-language routine's body from information_schema.routines is just the C
symbol name, so its reconstructed stub (CREATE FUNCTION ... AS $gfx$ name
$gfx$ LANGUAGE c) is unparseable by tree-sitter-sql, and the parser's error
recovery consumes the statements that follow. With FK ALTER TABLEs emitted
last, every references edge was silently lost on any DB with a common
extension installed (uuid-ossp, pgcrypto, pg_trgm, fuzzystrmatch, ...).

Emit the FK block before the routine DDL. FK statements only reference
tables, which are emitted first, so the order is always safe. On a real DB
with 35 tables / 41 FKs / 41 extension routines: 41/41 references edges vs
1/41 before.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-16 16:36:49 +01:00
safishamsiandClaude Opus 4.8 ecf1416a7e docs: cut 0.9.17 changelog and credit @zuwasi for Amp platform support (#948)
Stamp the 0.9.17 release date and add a belated acknowledgement that Amp
(ampcode.com) platform support, which shipped earlier in v8, was contributed
by @zuwasi in #948.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
v0.9.17
2026-07-16 11:47:08 +01:00
safishamsiandClaude Opus 4.8 cb96bdaa0c fix: preserve semantic layer, stamp hyperedges, PHP namespaces, ignore diagnostic (#1925 #1920 #1923 #1922)
- #1925: a missing manifest.json no longer degrades `extract --code-only`
  into a full scan that discards the committed semantic layer. An existing
  graph.json is a sufficient incremental baseline (detect_incremental treats
  an absent manifest as "all new / none deleted"), so out-of-scope doc/paper/
  image nodes are preserved while genuinely deleted sources still evict.
- #1920: _stamped_manifest_files now counts hyperedge output, so a doc whose
  only chunk output is a hyperedge is stamped instead of re-extracted forever.
- #1923: new namespace/use-aware PHP resolver (mirrors the Java resolver, runs
  before the unique-name rewire) so App\Models\Page and an imported
  Filament\Pages\Page stay distinct — no more false inherits/imports edge.
- #1922: detect() records ignored files/dirs in a new `ignored` diagnostic
  field (the nested-ignore scoping bug itself shipped in 0.9.16 / #1873).

Regression tests added for each; full suite 3325 passed, 3 skipped.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-15 19:51:53 +01:00
safishamsiandClaude Opus 4.8 19c496dd63 fix(build): make _semantic_id_remap idempotent to stop id accretion (#1917)
An id whose canonical stem contains its legacy stem as a prefix (parent
dir name == file stem, e.g. .claude/CLAUDE.md) re-matched the legacy
branch every build and grew another stem segment, defeating the
same_topology/no_change short-circuits and churning graph.json +
clustering on every zero-delta update. Skip an id that already carries
its canonical stem (mirrors graph_has_legacy_ids); genuine legacy
migration still applies. Also records the #1889/#1918 query-scoring perf
work in the changelog.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-15 15:14:34 +01:00
safishamsiandClaude Opus 4.8 c0875ccae9 test(serve): rebase #1900 test onto #1918 _score_query API; bench newline
PR #1918 replaced _pick_seeds(terms=) with best_seed_by_term= from
_score_query. Update the #1900 German-stopword seed test to the new
single-traversal API and add the missing trailing newline in the
(manual, non-collected) query-scoring benchmark.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-15 15:09:17 +01:00
Sirhan1andsafishamsi ff999a909a perf(serve): collapse T+1 per-term scoring passes into one traversal
Add _score_query() producing combined ranking + per-term singleton winners
in a single graph traversal. _pick_seeds consumes precomputed winners
instead of rescoring each term via _score_nodes.

Behavior byte-identical to the legacy T+1 path. Supersedes the rescoring
loop from #1596 while preserving its per-term coverage guarantee; folds
in coverage scaling from #1724 and multiplicity penalty from #1832.
Orthogonal to the trigram prefilter from #1431.

Benchmark (full sweep in #1889):
  1k-100k nodes, 1-10 terms: 1.43x-2.43x median speedup
  Single-term: no regression (1.70x-1.89x improvement)
  Traversals: T+1 -> 1 regardless of term count

Tests: 3221 passed, 3 skipped. Ruff clean. Graphify graph updated.

Refs #1889
2026-07-15 15:07:25 +01:00
safishamsiandClaude Opus 4.8 1c280e4576 docs(changelog): record #1894(PR-1) #1916 #1915 #1912 in unreleased 0.9.17
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-15 14:16:14 +01:00
safishamsiandClaude Opus 4.8 b7ddee3c28 fix(extract): make --mode deep effective over a warm cache; add --force (#1894)
`graphify extract --mode deep` over a warm tree was a silent no-op, for
three stacked reasons:

1. The semantic cache ignored mode: deep runs were served standard-mode
   entries (and vice versa). check_semantic_cache/save_semantic_cache now
   take `mode` (default None, byte-identical when omitted so older
   installed callers keep working) and map it to a namespaced kind —
   cache/semantic/ for None, cache/semantic-{mode}/ otherwise. The
   per-chunk checkpoint in llm.extract_corpus_parallel and the extract /
   cache-check consumers thread the run's mode through (cache-check grows
   --mode/--deep). cached_files, clear_cache, and prune_semantic_cache
   sweep BOTH namespaces; prune uses the same live-hash set for both
   (liveness is content-based, mode-independent) so semantic-deep/ can't
   regrow the #1527 unbounded-orphan problem and inherits the
   files_by_type-derived exclusion gating for free.

2. extract had no --force — the flag was silently swallowed by the
   parser's unknown-arg fallthrough. It is now real (plus GRAPHIFY_FORCE
   env parity with `update`): force disables the incremental gate so
   detection is a full scan and skips the semantic cache READ so every
   semantic file re-dispatches, while the post-run save and manifest
   stamping still happen.

3. The incremental gate dispatched zero files on a warm unchanged tree
   before the cache was ever consulted, so namespacing alone couldn't fix
   the repro. In deep+incremental runs the semantic pass now widens to the
   full live doc/paper/image set from detect_incremental's files_by_type
   (already exclusion-filtered, #1908/#1909) and lets the mode-namespaced
   cache decide hits/misses, with a loud count line so the first deep
   run's full re-dispatch is visible.

Skill-side threading of mode is deliberately deferred to PR-2; mode
defaults keep generated skills byte-compatible.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-15 14:14:13 +01:00
safishamsiandClaude Opus 4.8 d916c61559 fix(cache): stop persisting dangling edges/hyperedges in the semantic cache (#1916)
save_semantic_cache groups nodes/edges/hyperedges by their own source_file
and the write loop skips a group whose path is a ghost (not .is_file(),
silently) or out-of-scope per the #1757 guard — but an edge or hyperedge in
an ALLOWED group that references a node id from a skipped group was still
written verbatim, so on replay (check_semantic_cache) it dangled forever.
The #1895 filter does not cover this: it cleans the in-memory merged result,
while the checkpoint writes the cache BEFORE it runs and replay bypasses it
entirely.

Fix at the cache layer (the authoritative write path): before the write
loop, compute the node ids belonging to groups that will be skipped —
mirroring BOTH skip branches — minus ids also defined in a group that will
be written (duplicate-attribution nodes must not be over-pruned). Each
written group then drops edges whose source/target is a skipped id and
hyperedges whose member list intersects them (whole-hyperedge drop,
mirroring #1895). Pruning runs on the incoming result only, so with
merge_existing=True (the llm.py checkpoint path) the prior cached entry's
valid edges survive the union untouched. Everything is gated on
allowed_source_files being provided, so unscoped callers stay
byte-identical.

Complementary hardening in build_from_json: hyperedges were copied into
G.graph["hyperedges"] verbatim without member validation, so a dangling
hyperedge reached graph.json even from a live (non-cache) extraction.
Members are now remapped via the same normalization the pairwise-edge loop
uses and pruned when they still don't resolve; a hyperedge with no
surviving member is dropped whole with a stderr warning. (Single-member
hyperedges are legal in this codebase — per-file flows in the #1574 tests —
so the drop threshold is zero survivors, not two.)

Tests: scoped save with an edge to an out-of-scope real file, the same with
a ghost source_file, whole-hyperedge drop, unscoped save preserved
byte-identically (raw cache entry compared), merge_existing keeping the
prior slice's valid edges, and build_from_json pruning a dangling hyperedge
member / dropping an all-dangling hyperedge.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-15 13:30:25 +01:00
safishamsi 76f5bc98fe fix(watch): stop AST-quick-scanning docs that already have semantic nodes (#1915)
_rebuild_code fed Markdown/doc files (.md/.mdx/.qmd/.skill) into the AST
quick-scan while _reconcile_existing_graph preserved the same docs'
semantic (LLM) nodes, so every semantic-backed doc was represented twice
(heading nodes + concept nodes) and the graph bloated ~4x vs the CLI
`graphify . --update` path.

Semantic now supersedes AST per doc source: before choosing
extract_targets, read the existing graph's source identities that carry
semantic doc nodes (non-_origin=="ast", gated on file_type=="document"
so pre-#1865 marker-less graphs aren't misread) and exclude those docs
from the quick-scan in both the full-rebuild and changed_paths branches.
They stay in code_files, so the #1795 fail-closed corpus check and the
shrink accounting still cover them — a previously-bloated graph
self-heals on the next full rebuild (the AST ownership rule drops the
stale heading nodes) without the shrink guard refusing the smaller
write. On incremental rebuilds the excluded docs never enter
rebuilt/node-evicted identities, mirroring #1865's tier-scoped edge rule
at the node level, so a doc's semantic nodes and edges survive. Docs
with no semantic layer (and brand-new docs) keep the no-LLM quick-scan
structure from #09b33b7.
2026-07-15 13:30:25 +01:00
2851f31a07 fix: treat .cjs (explicit CommonJS) as a code extension
.cjs was half-registered: the language maps in build.py and
extract.py already routed it to the JS grammar, but it was missing
from CODE_EXTENSIONS, the extractor _DISPATCH, _LANG_FAMILY, and the
JS resolution/cache suffix sets — so .cjs sources (Electron
main/preload scripts, CommonJS escape hatches in "type": "module"
packages) were silently skipped during a build: classified as
non-code and never handed to the JS extractor.

Add .cjs alongside .mjs in the six lists that gate/route JS files
(detect.py, extract.py _DISPATCH, analyze.py _LANG_FAMILY,
extractors/models.py cache-bypass, extractors/resolution.py resolve
exts, cli.py _HOOK_SOURCE_EXTS), with regression locks mirroring the
.mts/.cts fix (#1607).

Real-world impact: an Electron app whose ~2000-line main.cjs backend
was invisible went from 201 to 294 nodes and 256 to 441 edges on
rebuild.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-15 13:13:30 +01:00
safishamsiandClaude Opus 4.8 0d783915a6 docs(changelog): record #1908/#1909 and #1910 in unreleased 0.9.17
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-15 12:18:55 +01:00
safishamsiandClaude Opus 4.8 49df466385 fix(sql): recover CREATE FUNCTION/PROCEDURE from tree-sitter ERROR nodes (#1910)
tree-sitter-sql parses PL/pgSQL CREATE FUNCTION statements (OUT/INOUT
params, tagged dollar quotes, PERFORM/:= body statements) as ERROR
nodes, and the dispatch loop had no branch for them, so the functions
were silently dropped from the graph.

Handle ERROR nodes inside walk() (they can nest inside a merged
create_function during multi-statement error recovery) and dispatch
top-level ERROR statements to it. The branch regex-scans the raw node
text for every CREATE [OR REPLACE] FUNCTION/PROCEDURE, mirroring the
existing fb_proc_or_trigger and has_error fallbacks, with a name class
that keeps schema-qualified names (exposed.important_function) whole.
The PL/pgSQL body is not scanned for FROM/JOIN references to avoid
junk reads_from targets, and node ids match the clean create_function
branch so seen_ids dedups consistently.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-15 10:58:57 +01:00
safishamsi 75cf56bfbc fix(detect,cli): excluded files are pruned from graph and manifest, not misreported as deleted (#1908 #1909)
Two coupled excluded-vs-deleted fixes:

#1909 — incremental extract's prune set was derived from the manifest
alone (manifest - corpus), so a file that became excluded without ever
being manifest-listed (every pre-#1897 graph) kept its stale nodes in
graph.json forever. The prune set is now also derived from the existing
graph's own node source_files reconciled against the post-exclude detect
corpus (_stale_graph_sources), restricted to in-root paths; out-of-root
--include/symlinked entries and remote (://) sources are never pruned.
Relative source_files are anchored against both the scan root and the
--out root (the #555/#1899 relativized form). The --no-cluster
incremental early exit never runs build_merge, so an exclusion-only
change now prunes the raw graph.json in place instead.

#1908 — save_manifest retained any prior row whose file still existed
on disk, so an excluded-but-alive file survived as a permanent phantom
that detect_incremental reported as deleted on every run. Full-scan
callers (extract's saves, watch._rebuild_code's saves) now pass the RAW
detect corpus via a new scan_corpus parameter and in-root rows outside
it are dropped; the corpus is deliberately not the #933 stamp-filtered
files dict, so failed-chunk/omitted-doc rows and --code-only doc rows
survive. Subset saves (changed_paths hooks, #917) keep the seeding
default. detect_incremental now splits manifest rows that left the scan
into deleted_files (gone from disk) and excluded_files (alive but out of
scan), mirroring the watch-side #1795 distinction, and the extract
summaries report the two separately.

Ordering matters: extract's cleanup of newly-excluded nodes previously
worked only through the #1908 conflation, so the graph-source prune
lands together with the manifest split to avoid regressing #1909.
2026-07-15 10:58:57 +01:00
safishamsiandClaude Opus 4.8 43b2affd56 chore: open 0.9.17 with 8 batch fixes (#1895 #1897 #1902 #1907 #1896 #1900 #1901 #1906)
Bumps to 0.9.17 (unreleased) and records the batch implemented via Fable
subagents in isolated worktrees, integrated onto v8: out-of-scope node
drop, manifest stamping, merge-driver registration, hooksPath configparser
fix, obsidian prune, multilingual query stopwords, .skill classification,
and the postgres package-name hint.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-15 00:43:46 +01:00
safishamsiandClaude Opus 4.8 b1e313cc5f fix(cli): stamp freshly-extracted semantic docs in the manifest (#1897)
The #933 manifest filter built its extracted-set from node/edge
source_file values, which are root-relative on a fresh extraction, and
compared them against files_by_type entries, which are absolute (from
detect()). The raw string membership test therefore never matched, so
every freshly-extracted semantic doc was dropped from the manifest and
re-queued as changed on the next run; only code files and cache-replayed
docs got stamped. The filter now lives in _stamped_manifest_files(),
which resolves BOTH sides against the scan root before the membership
test — the same Path/is_absolute/resolve normalization the #1890
reconciliation uses in graphify.llm. Genuinely omitted zero-node docs
still have no source_file entry and stay unstamped, preserving the
intentional #933 re-queue behavior.

Regression tests: a CLI extract with a mocked corpus extractor returning
root-relative source_files lands the doc in manifest.json with a
non-empty semantic_hash while a zero-node doc stays out; a unit test
covers relative (fresh) and absolute (cache-hit) source_file shapes.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-15 00:41:37 +01:00
safishamsiandClaude Opus 4.8 0cacd708e9 fix(llm): drop out-of-scope nodes from the merged extraction result (#1895)
The #1757 cache guard refuses to write a cache entry for a node whose
source_file resolves to a real corpus file that was not dispatched, but
the node itself still flowed into merged["nodes"] and landed in
graph.json (plus any edges/hyperedges built on it). extract_corpus_
parallel now filters the merged result right before the #1890
dispatched-vs-returned reconciliation: a node is dropped when its
source_file resolves (against root, same normalization as #1890) to an
existing file (.is_file(), mirroring the #1757 condition) outside the
dispatched set. Non-file source_files (concepts, model-invented anchors)
pass through untouched. Edges whose endpoint and hyperedges whose member
is a dropped node id go with it, plus any edge/hyperedge itself
attributed to an undispatched real file. One summary warning names the
offending files and merged["out_of_scope_dropped"] records the count.
Running before the reconciliation keeps covered/uncovered reflecting the
post-filter graph.

Regression tests: a chunk over A.md+C.md returning a stray B.py node
loses the stray (and its edge/hyperedge) while sibling and concept
attributions survive; a clean run records a zero count and no warning.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-15 00:41:36 +01:00
safishamsiandClaude Opus 4.8 f2a6139667 Fix spurious hooksPath warning and register graph.json merge driver
Fix #1907: _hooks_dir parsed .git/config with a strict configparser,
which rejects git-legal duplicate keys/sections (VS Code writes such
configs) and printed a spurious "could not read core.hooksPath" warning
on every hook command. Ask git itself (rev-parse --git-path hooks) as
the primary resolution instead; it handles core.hooksPath, includeIf
and linked worktrees, keeps the Windows-path and multiline/NUL guards,
and surfaces git's own stderr when the config is genuinely corrupt.
The .git/hooks final fallback is unchanged.

Fix #1902: hook install never registered the graph.json union merge
driver that README and CHANGELOG 0.7.0 document. install() now sets
merge.graphify.name/driver via git config (pinning sys.executable with
the same shell-safe allowlist the hook scripts use, falling back to the
graphify launcher) and idempotently appends the graph.json
merge=graphify line to .gitattributes without clobbering other entries.
uninstall() mirrors the removal, and status()/install() report the
merge-driver state.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-15 00:41:36 +01:00
safishamsiandClaude Opus 4.8 876c043c2f fix(export): prune orphaned Obsidian notes on re-export (#1896)
Re-exporting into an existing vault left notes for nodes that dropped
out of the graph, and rewriting the manifest to only this run's files
disowned those orphans - so a returning node's stale note became
permanently unwritable.

Before rewriting the manifest, delete stale = owned - written - skipped.
This only ever touches files graphify itself wrote (foreign files go to
_skipped, never the manifest), and each path is containment-checked
against the vault dir to defuse a corrupt/hostile manifest with ../
entries. The manifest rewrite then correctly drops the pruned files.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-15 00:41:36 +01:00
safishamsiandClaude Opus 4.8 107cea044c fix(serve): drop German and Romance question stopwords from query terms (#1900)
Full-sentence non-English queries picked wrong BFS seeds because
_QUERY_STOPWORDS was English-only: in a mostly-English code corpus,
fillers like wie/die/das/funktioniert are rare, get high IDF weight,
and out-seed the real keyword (~175x score collapse).

Extend the frozenset with a curated German set plus trimmed French/
Spanish/Portuguese/Italian question/filler words, diacritics intact.
English-collision words (war, bald, comment, come, son, sin, con,
pour, des) are deliberately omitted; die/hat are kept since the
all-stopword fallback and unfiltered find_node protect English use.
Filtering stays in _query_terms, so the per-term seed guarantee in
_pick_seeds composes unchanged.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-15 00:41:36 +01:00
safishamsiandClaude Opus 4.8 ef8771e8c1 fix: correct postgres install hint to graphifyy (#1906); classify .skill files as documents (#1901)
- pg_introspect: the psycopg-missing ImportError suggested pip install
  'graphify[postgres]', but the PyPI package is graphifyy (double-y)
- detect: add .skill (Markdown with YAML frontmatter agent files) to
  DOC_EXTENSIONS so they are no longer dropped as unclassified
- extract: route .skill through extract_markdown for the structural
  quick-scan

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-15 00:41:36 +01:00
safishamsiandClaude Opus 4.8 75b7279d92 docs(readme): point star-history at Graphify-Labs org; X handle to @graphify
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-15 00:00:56 +01:00
safishamsiandClaude Opus 4.8 a0e4a1c6bd docs(changelog): date 0.9.16 for release
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
v0.9.16
2026-07-14 23:57:12 +01:00
safishamsiandClaude Opus 4.8 a4ab6ed3f6 fix(llm): reconcile dispatched vs returned files in semantic extract (#1890)
A semantic chunk can return a clean, non-empty response that omits some
of the documents it was given. Those docs vanished from the graph with
no node, no warning, and no cache/manifest stamp, so they were silently
re-dispatched (and typically re-omitted) on every run. extract_corpus_
parallel now diffs the dispatched file set against the source_files that
returned, records the gap in merged["uncovered_files"], and prints a loud
warning naming the omitted files. Smallest visibility guard; routing docs
through the deterministic extractor for a guaranteed file node is a
separate, larger change.

Adds a regression test: a chunk that omits odd-numbered docs is caught
and warned, not silently dropped.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-14 23:45:46 +01:00
safishamsiandClaude Opus 4.8 b3dc15b839 fix(extract): close residual absolute-path/username leaks (#1899)
Completes #1789. Two residual leaks into committed graph.json:

(a) Out-of-root reference targets (a .csproj ProjectReference, .sln
    project, or bash `source` pointing outside the scan root) kept an
    absolute source_file and an absolute-derived id, because the
    relativization post-passes only handled paths under root and silently
    left out-of-root ones absolute. They now get a portable walk-up
    relative source_file and an `ext_`-namespaced id (basename fallback
    for far-outside or cross-drive targets), with edge endpoints remapped.

(b) A symbol whose name normalizes to nothing (minified `$`, a JSONC
    `"//"` key) collapsed _make_id(stem, name) to the bare absolute file
    stem, leaking the path and colliding with the file node. Guard at the
    three mint sites (json_config keys, engine function names) to skip
    these no-signal symbols.

Adds regression tests: an out-of-root ProjectReference stays portable
(no scan-path in any id/source_file/edge), and a minified `$` is dropped
while the real function survives.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-14 23:42:06 +01:00
safishamsiandClaude Opus 4.8 e1b8a17dc8 docs(changelog): record #1881, #1876, #1893 in the unreleased 0.9.16 section
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-14 23:32:28 +01:00
kebwlmbheeandsafishamsi 5c0a04c991 fix(extract): filter Kotlin builtin/stdlib types from references graph
_kotlin_collect_type_refs had no equivalent of _JAVA_BUILTIN_TYPES /
_PYTHON_ANNOTATION_NOISE / _GO_PREDECLARED_TYPES, so every Boolean/String/
Int/List/... type annotation in a .kt file became a real graph node/edge.
Unrelated files sharing only these language-level types were getting
merged into the same community by cluster-only.

Add _KOTLIN_BUILTIN_TYPES (kotlin.* scalars, collections, throwables -
deliberately not Android/JVM framework types like Context/View/Bundle,
matching how _JAVA_BUILTIN_TYPES doesn't filter framework types either)
and check it plus _JAVA_BUILTIN_TYPES (Kotlin freely references
java.util.Date/UUID/etc) at the three call sites that previously appended
unconditionally.

Deliberately excludes "Result" from the filter list even though
kotlin.Result exists - it collides with the very common user-defined
sealed-class name (confirmed against this repo's own sample.kt fixture,
which defines its own Result<T>).

Verified on a ~400-file Android/Kotlin project: 3693->3239 nodes (-12%),
6481->5221 edges (-19%), with a previously merged 5-file grab-bag
community cleanly splitting out an unrelated screen-dimension utility.
2026-07-14 23:32:03 +01:00
xkam7arandsafishamsi 8361a81f73 fix(extract): recognize uppercase TypeScript 2026-07-14 23:32:02 +01:00
mzt006andsafishamsi 4f6106a2f1 fix(cli): send missing-skill warning to stderr 2026-07-14 23:31:59 +01:00
safishamsiandClaude Opus 4.8 f80e4fac35 test(watch): pin update discovers newly-added files/dirs (#1837)
#1837 (update silently fails to discover new files) was the same
nested-gitignore scoping bug (#1873/#1880), already fixed. The existing
test only does a single build; add the build -> add new file+dir ->
update -> discovered sequence the report describes, with a nested `*`
scratch dir as a scoping guard. No production change.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-14 23:20:36 +01:00
safishamsiandClaude Opus 4.8 060dd6317f fix(extract): resolve Python module-qualified calls to calls edges (#1883)
`module.function()` where `module` is imported produced no calls edge,
while bare-name calls did. The member call was captured, but the shared
cross-file pass skips all member calls and _resolve_python_member_calls
only handled capitalized class receivers, so a lowercase module receiver
fell through. Add a module arm: a lowercase receiver naming a module
imported into the caller's file resolves to the single callable that
module contains. Keyed on stable node ids (imports edge source is the
caller file node; contains maps file -> children), not source_file
strings, so it holds under the cache_root id-remap relativization. Also
drop the no-classes early return that skipped the whole resolver when a
corpus had functions but no class methods.

Guards: only modules imported into the caller's own file match (so
self/obj/local instances never match), and a single-definition god-node
guard on the callable. Adds a positive test and a false-edge negative.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-14 23:19:06 +01:00
safishamsiandClaude Opus 4.8 2795a18bda fix(exclude): persist --exclude so rebuilds re-apply it (#1886)
--exclude was honored only by the initial extract scan; update/watch/hook
rebuilds re-ran detect() without it and silently re-indexed excluded
paths. extract now writes the patterns to a graphify-out/.graphify_build
.json sidecar, and _rebuild_code reads and re-applies them on every
rebuild (layered after the ignore files, so they still win). Pre-existing
graphs without the sidecar behave as before. Adds a regression test.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-14 14:03:20 +01:00
safishamsiandClaude Opus 4.8 fd54f0ef0b fix(dedup): deterministic collision survivor + bare-file definer (#1851 followup)
Two tweaks on top of @bchan84x's #1852:

1. The definer heuristic decided the survivor pairwise, so when several
   same-file nodes co-defined one ID the surviving label still depended on
   arrival order (the exact 3-node case #1851 reports). Choose the survivor
   by a total order (definer first, then shorter/canonical label, then
   lexically) via a min over _collision_rank, so it is order-independent.
   Direction preserves #1504 (lexically-first source path wins).

2. _defines_id now also recognizes a bare file-level node whose id is
   exactly the slugified path (nid == prefix), not only <path>_<entity>.

Adds an order-independence test over all 6 permutations of definer +
same-file relabel + cross-file reference, and a bare-file-node test.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-14 13:51:42 +01:00
bchan84andsafishamsi 70bc9ca7b2 fix(dedup): let the defining file win an ID collision, and warn only about real loss
Node IDs are <source-path>_<entity>, so a doc that merely references an entity
mints the entity's own ID and collides with the entity's node by construction.
The pre-dedup pass kept whichever arrived first, so chunk order decided whether
an entity kept its own attributes or a passing cross-reference's, and the
warning that fired (#1504) told the user a same-name-different-directory clash
had lost their data — while the one drop that really is lossy, two labels for
one ID from the same file, was silent because the warning was gated on
source_file differing.

The survivor is now the node whose source_file is the file its ID encodes (any
trailing slice of the path, so absolute, repo-relative, and pre-#1504
bare-stem IDs all resolve), falling back to first-seen. Reporting follows what
the drop actually costs: a cross-reference folding into the node it references
loses nothing (edges are keyed by ID and rewire to the survivor) and is silent;
a same-file relabel notes the discarded label; two files that both encode the
ID are distinct entities and still WARNING, now stating the ID-scheme cause
rather than a filename clash.

Refs #1851
2026-07-14 13:38:57 +01:00
safishamsiandClaude Opus 4.8 2210d69bc4 docs(changelog): record #1865 and #1858 in the unreleased 0.9.16 section
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-14 13:21:29 +01:00
xor_xeandsafishamsi 4ac5f07e88 fix(watch): preserve semantic edges of re-extracted sources (#1865)
A full graphify update evicted every edge whose source_file was a
re-extracted source, including LLM semantic edges (e.g.
semantically_similar_to) from Markdown docs that also have an AST
extractor. The AST pass cannot regenerate those edges, so they were
silently lost while their concept nodes survived — node eviction is
provenance-aware via _origin (#1116), edge eviction was not.

Tag AST-extracted edges with the same _origin=ast marker nodes already
carry, and scope rebuild-driven edge eviction to that tier: re-extraction
replaces a source's AST edges, while its semantic edges survive until a
semantic re-extraction supersedes them. Deletion-driven eviction stays
provenance-blind, so edges of deleted or excluded sources are still
purged regardless of tier.

Edges from graphs built before this change lack the marker; a stale
AST edge from a file changed exactly once between the old and new
version can linger until its source is deleted — the same migration
trade-off the #1116 node marker made.
2026-07-14 13:20:40 +01:00
thejeshandsafishamsi 1aca158ada fix(cargo): honor package = "..." rename when resolving internal deps (#1858) 2026-07-14 13:19:16 +01:00
safishamsiandClaude Opus 4.8 40eae3cea2 test(watch): pin update-with-nested-gitignore (#1880); changelog for #1873/#1880
#1880 is the update-layer symptom of the #1873/#1887 nested-ignore
subtree-scoping regression (fixed in fb4d452, @Alwyn93): a nested broad
`.gitignore` zeroed update's re-scan, so it built 0 nodes and the
shrink-guard refused to overwrite. Adds a _rebuild_code regression test
asserting update sees the real files (and still scopes the nested ignore
correctly) so this can't recur at the update layer. Records both
regressions in the unreleased 0.9.16 changelog.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-14 13:18:28 +01:00
fb4d45226e fix(detect): scope nested ignore patterns to their own subtree (#1873)
Non-anchored patterns from a nested .gitignore/.graphifyignore were
matched against root-relative paths first, so a nested ignore file's
patterns leaked outside its directory. In the wild, .hypothesis/.gitignore
(a bare "*" auto-written by the hypothesis library) ignored the entire
repository and detect() returned 0 files.

Per gitignore semantics, patterns from A/.gitignore apply only to paths
under A. Match every pattern against the path relative to its own anchor
(the anchor dir itself exempt — an ignore file governs its directory's
contents, not the directory), and skip patterns whose anchor does not
contain the target.

Regression introduced with nested-ignore support in 8a5287a (#1206).
tests/test_detect.py: 143 passed, including the existing nested-ignore
and nested-negation tests plus two new regressions.

Fixes #1873

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 13:14:24 +01:00
safishamsiandClaude Opus 4.8 368407874f docs(changelog): record #1862 and #1860 in the unreleased 0.9.16 section
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-14 00:26:13 +01:00
thejeshandsafishamsi e32be08c2a fix(detect): legacy-manifest branch must !=-compare mtime to catch reverse jumps (#1859) 2026-07-14 00:25:36 +01:00
thejeshandsafishamsi 20405a84bf fix(dedup): report fuzzy count in summary when exact_merges is 0 (#1857) 2026-07-14 00:25:35 +01:00