Java field declarations produced no `references` edge for their type, so a class's
data dependencies (its field types) were missing from the graph even though
parameter and return types were already captured. The field handler now collects
the declared type via the same `_java_collect_type_refs` helper used elsewhere,
preserving the `field` and `generic_arg` contexts and skipping primitives
(int/boolean/etc.), matching the existing C#/PHP/Kotlin field handlers.
Ported from PR #1485 by @oleksii-tumanov.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
A detailed report (#1475) showed the ObjC extractor silently dropping ~60% of
code-level relationships. Five fixes:
1. Dispatch: ObjC `.h` headers were parsed by extract_c (1 node, 0 edges, losing
every @interface/@protocol/@property/method). _get_extractor now routes a `.h`
to extract_objc when it contains an ObjC-only directive
(@interface/@protocol/@implementation/@import) — these are illegal in C/C++, so
the sniff never hijacks a genuine C/C++ header (verified: a plain C/C++ `.h`
stays on its existing extractor).
2. Calls: the method-body pass produced zero `calls` edges because it matched
child types `selector`/`keyword_argument_list`, but tree-sitter-objc tags
selector parts with the field name `method` (type `identifier`). The selector
is now reconstructed from every `method`-field child, explicitly skipping the
`receiver` field — so self/super/ClassName receivers are never mistaken for a
selector, and compound sends ([self a:x b:y]) resolve too. (Avoids the report's
suggested `"identifier"` fix, which would have matched receivers as selectors.)
3. Generic property types: NSArray<Product *> * wraps the type in a
generic_specifier, so the old direct-type_identifier scan saw nothing. The
element (Product) and container (NSArray) are now both referenced.
4. Class methods: `+ (…)shared` was labeled -shared; the +/- sigil is now read
from the method node's first child.
5. @import: `@import Foundation;` (a module_import node) now emits an imports edge.
Dot-syntax property `accesses` (Bug 5) and @selector(...) target-action edges
(Bug 6b) need type/name resolution policy and are left as follow-ups. Added six
regression tests; the reporter's claim that compound messages lack a
message_expression wrapper (Bug 6a) was checked against the real AST and refuted —
they share Bug 2's root cause, fixed by the same field-name change.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Builds on the initial XAML support (#1460). Resolves a view to its ViewModel from
an explicit <Window.DataContext><vm:MainViewModel/>, a design-time
d:DataContext="{d:DesignInstance Type=...}", the View->ViewModel naming
convention, or Prism ViewModelLocator.AutoWireViewModel="True". Resolution is
always against an actually-extracted C# class node, so a name matching no class
(or an ambiguous 2+) emits no edge -- explicit DataContext is EXTRACTED,
convention/Prism are INFERRED. Also extracts binding paths ({Binding User.Name},
Path=Order.Total), commands (Command="{Binding SaveCommand}"), converters, and
CommunityToolkit [ObservableProperty]/[RelayCommand] generated members.
The #1460 event-handler hardening is preserved unchanged: events still resolve
only to methods with a .NET (object sender, ...EventArgs e) signature, and the
free-form-attribute denylist still prevents values like Content="Save" from
fabricating event edges (both regression tests still pass). ViewModel discovery is
bounded to the active extraction root.
Ported from PR #1473 by @MikeKatsoulakis (clean 3-way merge onto current v8).
Maintainer fix on top: the CommunityToolkit member reader now reads the
code-behind with errors="replace", so a non-UTF8 ViewModel .cs can't raise
UnicodeDecodeError and abort extract_xaml (matches every other reader in the
module). Added a regression test for that case.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
.vue files were dispatched to extract_js, which picks a tree-sitter grammar by
suffix. .vue is neither .ts nor .tsx, so the whole SFC -- <template> markup,
<script>, and <style> -- was fed to the JavaScript grammar, producing a top-level
ERROR node and recovering no imports, symbols, or type references.
A dedicated extract_vue masks everything outside <script> (replacing it with
spaces so symbol line numbers stay accurate) and parses just the script with the
grammar named by `lang` (ts default; tsx/js/jsx honored). .vue also joins the
cross-file symbol-resolution pass now that it parses cleanly.
Ported from PR #1468 by @papinto. Maintainer fix on top: the <script> open-tag
scan now skips over quoted attribute values, so a `>` inside one (Vue 3.3+ generic
components, e.g. generic="T extends Record<string, unknown>") no longer ends the
tag early and swallow the body. Added a regression test for that case.
(CHANGELOG also records #1470, committed just prior.)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
graphify reflect renders LESSONS.md from the memory docs and the
.graphify_analysis.json / .graphify_labels.json sidecars, but reflect --if-stale
only stat'd the memory docs and graph.json. So after community analysis or labels
changed (without the graph changing), --if-stale wrongly reported lessons fresh and
skipped the regen, leaving LESSONS.md stale. The freshness check now also stats the
analysis and labels sidecars, using the same custom --analysis / --labels paths the
run itself uses (so the check and the run can't disagree about which files matter),
and treats a missing sidecar as not-an-input. This makes the documented "no-op when
LESSONS.md is newer than every input" contract actually true.
Ported from PR #1470 by @oleksii-tumanov.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The Read|Glob PreToolUse hook (the "run graphify first" nudge, shared by the
Claude Code and CodeBuddy installers via _READ_SETTINGS_HOOK) decided whether to
nudge by substring-scanning the joined file_path/pattern/path for known
extensions. That had two opposite failures: '.js' is a substring of '.json' so
package.json / tsconfig.json spuriously fired, and .astro/.vue/.svelte weren't in
the set so Astro/Vue/Svelte projects never nudged on their primary source type.
The hook now compares each value's real trailing extension (segment after the
last '/', then after the last '.') against the set, and adds .astro/.vue/.svelte.
package.json -> tail .json (silent); **/*.astro -> tail .astro (fires); an
extension on a directory component (my.ts/file) correctly stays silent. The
graphify-out/ suppression and fail-open behavior are unchanged.
Ported from PR #1464 by @marketechniks onto current v8. Added three regression
tests on top of the PR's (multi-dot a.test.tsx / foo.min.js, a Windows backslash
path, and the directory-extension trap) to pin the trickiest parts of the new
segment-split logic.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Hermes and the other AGENTS.md hosts (Codex, Aider, OpenClaw, Droid, Trae, ...)
run the graphify CLI directly and do not dispatch subagents. The Step 3 extraction
guidance only described the no-API-key path as "fall straight through to subagent
dispatch", so on `/graphify .` those agents had no path they could take, fixated on
the GEMINI/ANTHROPIC key language, and looped for minutes insisting on a missing key
before eventually proceeding (a pure-code corpus is AST-only and needs no key at all).
Step 3 now opens with a hoisted, host-agnostic statement -- graphify needs no API
key, never prompt for one, never block on one; code is AST-only; a code-only corpus
skips semantic extraction entirely -- and the no-key fallback now spells out a
terminal-only path (write the empty semantic file and continue, or extract content
inline) instead of assuming subagent dispatch.
Applied in the shared core fragment and in the aider/devin core variants, so all 16
skill bodies carry it; the aider/devin change is registered as a sanctioned
monolith-roundtrip diff. Regression test (test_extraction_states_no_api_key_required_
for_every_host) renders every host and pins the wording, ordering, and the
non-subagent fallback so this can't silently regress.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Makes .xaml a first-class code input. extract_xaml() uses stdlib XML (no new
parser dependency) behind the same DOCTYPE/ENTITY and size guards as the .csproj
extractor, and captures: the root element, named controls (x:Name/Name) and their
control types, {Binding ...} references, x:Class, and -- the useful part -- a
bridge from the view markup to its .xaml.cs code-behind by resolving event-handler
attributes to the matching methods on the partial class.
Ported from PR #1460 by @MikeKatsoulakis onto current v8.
Maintainer hardening on top of the original PR: event resolution is now gated so
it can't fabricate edges. The original matched any attribute value against
code-behind method names, so Content="Save" next to a business method Save(), or
Tag="<a-handler-name>", produced spurious "event" edges. Resolution now requires
(a) the attribute is not a known free-form/identity property (Content, Text, Tag,
Title, ToolTip, Header, ...), (b) the value is a bare identifier, and (c) the
matched method actually has the .NET event-handler signature
(object sender, <T>EventArgs e) -- read from the code-behind source since the C#
extractor does not record parameter lists on method nodes. Added regression tests
for both false-positive cases.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The kimi, gemini, and deepseek backends hardcoded their base_url, so users
behind an OpenAI-compatible proxy/gateway or running a self-hosted relay had no
way to redirect them (unlike ollama/openai, which already read *_BASE_URL). Each
backend now reads KIMI_BASE_URL / GEMINI_BASE_URL / DEEPSEEK_BASE_URL and falls
back to its official default when unset, so behavior is unchanged for anyone who
doesn't set the variable.
Ported from PR #1458 by @jc2shile onto current v8. The PR branch carried 624
unrelated files from a stale base; this lands just the clean 16-line llm.py
change. Added subprocess-based tests covering both the override and the default
for all three backends (BACKENDS reads the env at import time, so each case runs
in a fresh interpreter).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
to_canvas sized each community group box for a ceil(sqrt(n))-column grid but the
placement loop hardcoded 3 columns, so any community bigger than ~9 members
rendered as a cramped 3-wide strip in an over-wide, mostly-empty box (and the box
width/height didn't even agree — w used sqrt(n), h used /3). The column count is
now computed once per community (inner_cols) and reused for box width, box height,
and card placement, so the cards fill the box. Cosmetic, no data change.
Ported from PR #1459 by @TPAteeq onto current v8 (clean: only the grid math
changed, the #1457 dedup helper is untouched). Verified the geometry on a real
canvas: n=25 -> 5x5 grid with every card inside its box; n=10 -> 4 columns.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
to_obsidian / to_canvas / to_wiki keyed filename dedup on the exact-case name,
so two labels differing only by case (e.g. `References` vs `references`) counted
as non-colliding and the second write clobbered the first on case-insensitive
filesystems (macOS/APFS, Windows/NTFS) — silently, no suffix, no warning.
Dedup now folds case (keyed on the lowercased name) while emitting the
original-case filename, so any pair that would collide on disk gets a numeric
suffix. The obsidian/canvas dedup is one shared helper (`_dedup_node_filenames`)
so they can't drift; wiki's slug dedup gets the matching fix; the `_COMMUNITY_*`
overview notes (which had no dedup at all) are covered; and a generated `base_1`
is re-checked so it can't overwrite a node literally labelled `base_1`.
Ported from PR #1457 by @TPAteeq onto current v8. Verified with a rigorous
edge-case battery (case-only collision, base_1 literal re-check -> base_1_1,
community-label case fold, determinism) plus the PR's tests; full suite 2404
passed, ruff + skillgen clean.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
0.8.48 added graphify/extractors/ (the per-language split, #1212) but
[tool.setuptools] packages listed only "graphify", so the subpackage was
omitted from the wheel and `graphify extract` raised ModuleNotFoundError on a
fresh install. Editable/dev installs and the test suite import from the source
tree, so it passed CI and local smoke but broke the published package. Add
graphify.extractors to packages; verified the rebuilt wheel imports and extracts
in a clean venv. 0.8.48 is yanked.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
get_community shows the community name (#1448); starlette floored >=1.3.1 for
CVE-2026-48818 / CVE-2026-54283 (#1391, #1396); begin per-language extractor
split into graphify/extractors/ (#1212); parallel community labeling via
--max-concurrency / --batch-size (#1390); reflect dedups dead-ends/corrections;
and the work-memory loop no longer depends on the git hook (reflect --if-stale).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
get_community was the only graph tool still returning a bare numeric id, while
get_node and the query-traversal output already render the community_name
attribute to_json writes onto every node. Read the name from the community's
member nodes and put it in the header ("Community 12 — Auth & Sessions"),
sanitised like every other LLM-derived field.
Ported from PR #1448 by @rmart1308 onto current v8, with two additions: the name
is skipped when it is just the "Community N" placeholder (written for unnamed
communities) so the header never doubles to "Community 12 — Community 12", and
the formatting is extracted to a module-level _community_header() with focused
tests (named / placeholder / empty / sanitised). Full suite 2397 passed.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
starlette underpins the HTTP MCP transport (serve_http); graphify/serve.py
imports it directly (Starlette/Middleware/Route) but it was only present
transitively via mcp, with no version floor — so end users installing
graphifyy[mcp] could resolve a vulnerable starlette even after a lockfile bump.
Declare it in the mcp (and all) extras and floor at >=1.3.1, which carries the
fixes for both CVEs (1.3.1 >= the 1.1.0 fix for CVE-2026-48818 and is the fix
for CVE-2026-54283). Lock regenerated 1.0.0 -> 1.3.1.
Supersedes the lock-only bumps in #1391 and #1396 (both targeted 1.3.1 from a
stale base) with a pyproject floor that also protects end users. stdio MCP and
CLI are unaffected. serve/MCP/HTTP tests pass on 1.3.1; full suite 2393 passed.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The extractor-split port (b3ab221) moved `_make_id` (the only user of
`make_id`) into graphify/extractors/base.py, leaving the top-level
`from .ids import make_id` in extract.py unused. Remove it. Surfaced by
`ruff --select F` (F401); harmless at runtime and not in the enforced lint
set, but it is dead code from the move. extract.py F-issue count back to the
pre-port baseline; module imports and the zig/elixir/razor/blade tests pass.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Move the blade/elixir/razor/zig extractors and the shared primitives
(_make_id, _file_stem, _read_text, _LANGUAGE_BUILTIN_GLOBALS) out of the
13k-line extract.py into a graphify/extractors/ package: base.py holds the
shared pieces, one module per language, and __init__ seeds a
LANGUAGE_EXTRACTORS registry for future dispatch. Import direction is strictly
extract.py -> extractors/ (extractors never import extract), so there is no
cycle. extract.py re-exports every moved name, leaving all callers and the
dispatch table unchanged.
Ported from PR #1291 by @TheFedaikin onto current v8 as a thin, behavior-neutral
slice (the PR itself was branched 31 commits behind and entangled with unrelated
files). Verified the moved code is byte-identical to current v8 before porting;
full suite 2393 passed, the zig/elixir/razor/blade extractor tests pass, ruff and
skillgen --check clean. Also gitignores .DS_Store.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
label_communities ran batches one LLM call at a time, so a large graph needed
hundreds of sequential calls even on backends that allow heavy concurrency. It
now fans batches out across a thread pool, mirroring extract_corpus_parallel:
results are returned per batch and merged on the main thread (labels dict is
never mutated concurrently, no lock), and workers==1 keeps the original
sequential path verbatim. ollama and claude-cli are forced serial unless the
matching GRAPHIFY_*_PARALLEL env opt-in is set (same guard as extract).
generate_community_labels threads max_concurrency + batch_size through, and the
cluster-only/label CLI parses --max-concurrency and --batch-size (both `--flag N`
and `--flag=N` forms; the space form is parsed explicitly so the value is not
mistaken for the positional scan path by the arg-walk's catch-all).
Output is deterministic regardless of concurrency (keyed by community id). Tests:
parallel == sequential result, batch-size controls batch count, batches actually
run concurrently, ollama forced serial, and the CLI parses both new flags. Full
suite 2393 passed; ruff clean.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Saving the same Q&A more than once duplicated lines in the "known dead ends" and
"corrections" sections: both lists were appended per memory doc with no key, while
node scoring already dedups by node. They now collapse by question, keeping the
most recent entry (docs are processed oldest-first, so a re-corrected question
shows its latest correction). Output stays deterministic, ordered by (date,
question). Applied to both the flat lists and the per-community buckets.
Found by a user running it on a 104-file Go codebase. Added a regression test
covering the dedupe and the recency-wins correction. Full suite 2389 passed.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Both the agent (session start) and the post-commit hook can run reflect; the runs
are deterministic and idempotent, but back-to-back ones are wasted work. Add
`graphify reflect --if-stale`, which no-ops when LESSONS.md is already at least as
new as every input (the memory docs and the graph). The skill's session-start
guidance now uses `--if-stale`, so when the hook just refreshed the file the
agent's run costs almost nothing, while a skill-only install still refreshes
on demand.
New lessons_fresh() helper + 5 tests (mtime freshness in each direction, and the
CLI skip/run behavior). Regenerated per-host references + re-blessed expected/;
all five skillgen guards pass; full suite 2388 passed.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
A skill-only install (no `graphify hook install`) recorded outcomes via
save-result but never ran reflect, so LESSONS.md was never generated or
refreshed and the lessons never surfaced. The query reference now instructs the
agent to run `graphify reflect` itself at the start of graph work (cheap,
deterministic, no-op with no saved outcomes) before reading LESSONS.md. The
post-commit hook stays as a between-session freshness optimization, not a
requirement. Regenerated per-host references/query.md + re-blessed expected/;
all five skillgen guards pass; full suite 2383 passed.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Work-memory (#1441): save-result --outcome + graphify reflect, with recency-decayed
scoring, corroboration threshold, label-matched node-existence gate, contested
handling, and zero-config adoption (skill reads LESSONS.md + records outcomes; git
hooks auto-reflect). Fixes: Python ClassName.method() qualified-call edges (#1446);
validate_extraction/build crash on non-hashable id (#1447); graphify update now
prunes a symbol removed from a surviving file without --force.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Make the self-improving loop "just work" once graphify is installed:
- Skill: the query reference now tells the agent to read
graphify-out/reflections/LESSONS.md at the start of graph work (start from
preferred sources, skip known dead ends, see prior corrections) and to record
--outcome useful|dead_end|corrected (+ --correction) on save-result.
- Hooks: the post-commit and post-checkout rebuild bodies now auto-run reflect
after _rebuild_code — best-effort, only when graphify-out/memory/ holds saved
outcomes, and never fails the hook — so LESSONS.md refreshes on every rebuild
without a manual `graphify reflect`.
Regenerated the per-host references/query.md and re-blessed expected/; all five
skillgen guards pass (check, audit-coverage, schema-singleton, monolith-roundtrip,
always-on-roundtrip). Verified end-to-end: install hook, save-result --outcome,
commit a code change -> hook rebuilds and writes LESSONS.md with the outcome.
Full suite 2383 passed; ruff clean.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Adds the deterministic work-memory loop: `save-result --outcome
useful|dead_end|corrected [--correction]` records how a saved Q&A turned out, and
`graphify reflect` aggregates graphify-out/memory/ into a deterministic
reflections/LESSONS.md an agent loads next session.
Source nodes are scored, not counted: signed, recency-decayed (useful +,
dead_end/corrected -, configurable --half-life-days, default 30), so a fresh dead
end outweighs a stale useful. A node is "preferred" only once corroborated by
>=--min-corroboration distinct results (default 2); others are "tentative", and
mixed-signal nodes render once as "contested" (recency-wins). Source nodes are
matched to the graph by label OR id, and citations whose node no longer exists are
dropped, so a plain `graphify update` after deleting code clears stale lessons.
Deterministic, no LLM; bare save-result and existing behavior unchanged.
Rigorously verified end-to-end on real data: corroboration boundary, recency flip,
contested verdict, foreign/malformed memory docs, cold start, 300-doc scale +
byte-stable output, and the node-gate dropping deleted-code lessons after update.
Full suite 2383 passed; skillgen --check clean; ruff clean.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
`graphify update` after deleting a function left the stale node in graph.json.
The build correctly dropped it (#1116), but _check_shrink then refused to write
the smaller graph ("Refusing to overwrite — you may be missing chunk files"),
so the deletion never persisted without --force. That also starved the
work-memory node-existence gate, which relies on graph.json reflecting deletions.
The shrink-guard now takes the set of source files re-extracted this run
(rebuilt_sources). A net shrink is allowed when every lost node belongs to a
rebuilt source (a symbol genuinely removed) or a deleted file; it is still
refused when a node vanishes from a file we did NOT touch — the silent
failed/partial-extraction case the guard exists to catch.
The #1116 e2e test now asserts the prune happens with force=False (was force=True);
added two _check_shrink unit tests (allowed within rebuilt sources, refused
outside). Full suite 2339 passed; skillgen --check clean; ruff clean.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The #1447 cherry-pick only touched build.py/validate.py + tests, so it
shipped without a CHANGELOG note. Add it under Unreleased.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Cross-class qualified static calls like `CustomerTaskActions.approve(...)` did
not produce an EXTRACTED `calls` edge. Two compounding causes:
1. The shared cross-file pass skips all member calls (the #543/#1219 god-node
guard against bare `obj.method()` name collisions), and there was no Python
receiver-based resolver to recover the qualified ones.
2. When the called method shared its name with an in-file node — e.g. a viewset
action `approve()` delegating to a service `Service.approve()` — the in-file
bare-name lookup matched the caller's own node (tgt == caller), so the call
was silently dropped before any raw_call was recorded.
Fix: capture a simple-identifier receiver in the call walk (new
`call_accessor_object_field`, set to `object` for Python), defer capitalized-
receiver member calls to a new `_resolve_python_member_calls` pass (mirroring the
Swift resolver), and emit an EXTRACTED edge only when the receiver resolves to
exactly one class that owns the method (single-definition god-node guard).
Instance/module calls (`self.x()`, `obj.x()`, lowercase receivers) are unaffected.
Tests: cross-class resolution, the same-method-name collision shape from the
issue, instance-call non-over-connection, and the ambiguous-class guard.
Full suite 2337 passed; skillgen --check clean; ruff clean.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
validate_extraction() is documented to return a list of error strings ("empty
list means valid"), but raised TypeError: unhashable type: 'list' when a node
id -- or an edge source/target -- was a non-hashable value such as a list. This
occurs in practice when an LLM extraction subagent emits malformed JSON like
{"id": ["foo", "bar"], ...}. Crash sites: the node_ids set comprehension and
the `edge[...] not in node_ids` membership tests.
Because build_from_json() validates at its start, a single malformed node
aborted the entire build, losing an otherwise-complete extraction of a large
corpus. build_from_json() itself would also raise (G.add_node(<list>) and the
`not in node_set` test) if the validator were bypassed.
- validate.py: build node_ids during the node pass, adding only hashable ids;
report a non-hashable id/endpoint as an error string instead of crashing.
All existing messages and the dangling-edge checks are preserved.
- build.py: skip dict nodes with a missing/non-hashable id and edges with
non-hashable endpoints (stderr warning). Non-dict nodes are deliberately
left to raise so the multigraph diagnostic still observes shape errors.
- tests: 3 cases in test_validate.py and 2 in test_build.py.
`graphify install --platform agents` installs the skill to the generic
Agent-Skills locations: the spec's user-global ~/.agents/skills (global) and
./.agents/skills (--project) — the directories `npx skills` and spec-compliant
frameworks read. `--platform skills` is an alias. Previously that user-global
location was only reachable as an accidental side effect of the gemini-on-Windows
branch. Bare `graphify install` is unchanged (still single-platform claude/windows).
The platform is registered in tools/skillgen/platforms.toml (split, mirroring
amp's agents-md body) and rendered through the skillgen drift/coverage guards.
Since it is a post-v8 platform with no own v8 body, its --audit-coverage baseline
is amp's v8 body (the body it re-homes). The rendered skill body is byte-identical
to amp's; only the on-demand hooks reference differs (its own `graphify agents
install` wording).
The `graphify agents install` / `graphify skills install` subcommand is the
amp-twin: it also wires an AGENTS.md always-on section, keeping it honest with the
hooks reference it points at. The `--platform agents` path stays skill-only,
exactly as amp's `--platform amp` does.
Also: `skill-agents.md` added to package-data, and the wheel-packaging guard now
covers every platform's skill body (not just references/always-on), so a missing
skill body fails CI instead of only breaking install for real users.
Closes#1405. Implements #1432.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Two integrity improvements to the core skill runbook (rendered via skillgen):
1. Graph-health gate (new Step 4.5): runs diagnostics.diagnose_extraction
read-only after build, before labeling, and surfaces dangling/missing/
self-loop and same-endpoint-collapsed edges — the silent-corruption modes of
incremental updates and AST/LLM id mismatches. Never aborts.
2. Anchor the AST and semantic caches on the scan root (root/cache_root=
'INPUT_PATH') instead of the cwd, matching the CLI extract path
(ast cache_root=out_root, check/save_semantic_cache(root=out_root)) and
build_from_json/save_manifest (#1361 parity). Without this, running from a
cwd != scan root cold-misses the cache and can split AST vs semantic caches.
Source edited in tools/skillgen/fragments/core/core.md; artifacts + expected/
regenerated via 'python -m tools.skillgen' (+ --bless). Full suite green
(2229 passed); skillgen + cache suites green (80 passed).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
_score_nodes and _find_node scan every node per query (O(nodes x terms)),
so query latency scales with total graph size regardless of where the
answer lives. Add a lazily-built character-trigram index (cached on the
graph object, auto-invalidated on hot-reload like _idf_cache) that narrows
each query to a small candidate superset before the unchanged scoring loop.
Results are byte-identical: the index is a pure candidate generator over
the exact fields the scorer reads (norm_label, label_tokens, nid,
source_file); a non-candidate node always scores 0, and IDF stays a
whole-graph statistic. _find_node candidates are returned in graph
iteration order so its exact/prefix/substring ordering, and matches[0],
stay unchanged.
A selectivity guard falls back to the full scan when a query term is too
short to trigram or its rarest trigram is still common (broad terms like
model/client), preserving a never-worse contract.
The index builds eagerly at load and before a reloaded graph is swapped
in, so neither the first query nor the first post-reload query pays the
one-time build cost. Storage format is unchanged (in-memory index only).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
graphify install --platform hermes always wrote the skill to ~/.hermes/skills,
the POSIX path. On Windows, Hermes scans %LOCALAPPDATA%\hermes\skills, so the
installed skill was never discovered. _platform_skill_destination now has a
hermes branch: Windows -> %LOCALAPPDATA%\hermes\skills, other OSes unchanged
(~/.hermes/skills). Pure path logic — no skillgen regeneration.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
A class defined once but referenced via type annotations in N other files appeared
as 1+N nodes — the extras carrying the referencing file's path (with extension)
baked into the id (e.g. pkg_a_py_thing). ensure_named_node's cross-file fallback
called add_node, which stamps the referencing file as source_file; that sourced
stub then collided in _disambiguate_colliding_node_ids (baking the .py path into
the id) and _rewire_unique_stub_nodes skipped it (a node with a source_file is
treated as a real definition, not a stub).
The fallback now emits a SOURCELESS stub (mirroring the inheritance-base path), so
disambiguation ignores it and the rewire collapses it onto the canonical
definition. The helper is duplicated across all six language extractors, so the
fix is applied to all six. Genuinely-defined duplicates (same name, different
files) still stay separate — only cross-file references collapse.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Cover the #1410 fix: to_obsidian and to_canvas must never emit a punctuation-only
filename (e.g. `@.md` from a `@/*` tsconfig paths key) — valid on disk but empty
once a downstream tool re-slugs on word chars (crashes `qmd update`). Both tests
exercise the public exporters and fail against the pre-fix safe_name().
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
An all-punctuation node label (e.g. `@/*` from a tsconfig paths entry) survived
the unsafe-char strip in `to_obsidian`'s `safe_name()` as a bare `@`, producing
`@.md`. That filename is valid on disk but empty once a downstream tool re-slugs
on word chars — qmd's handelize() reduces "@" -> "" and raises, aborting the
entire `qmd update` (every collection on the machine stops reindexing).
Require at least one word char in the stem; otherwise fall back to "unnamed"
(the existing dedup handles collisions). Applied to both safe_name occurrences.
Fixes#1409
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
CUDA is a C++ superset, so .cu/.cuh files parse cleanly with
tree-sitter-cpp (already a dependency). Two registrations wire it up:
- detect.py: add .cu/.cuh to CODE_EXTENSIONS so they're detected and
watched (watch.py's _WATCHED_EXTENSIONS derives from CODE_EXTENSIONS).
- extract.py: route .cu/.cuh through extract_cpp in _DISPATCH, which
also makes collect_files() pick them up (_EXTENSIONS = _DISPATCH.keys()).
Adds tests/fixtures/sample.cu (kernel + __device__/host functions +
struct + includes) and CUDA cases in test_languages.py covering kernel/
device function extraction, structs, includes, and host call edges.
Documents the new extensions in the README extension table and CHANGELOG.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The opencode plugin template embedded the reminder string with backticks
around `graphify query "<question>"`. Because the plugin prepends
`echo "<reminder>" && <cmd>` to the user's bash command, those backticks
triggered bash command substitution: every grep/rg/find invocation silently
ran `graphify query "<question>"` and substituted its output
("No matching nodes found.") into the reminder text shown to the agent.
This both corrupted tool output with graphify noise, and actually spawned
a graphify process + loaded graph.json + ran a BFS traversal with the
literal token <question> on every search.
Fix: remove the backticks. Adds a guard comment in the template so future
editors don't reintroduce the bug, and a regression test that asserts the
reminder string contains no backticks and no $() constructs.
`graphify extract --backend <gemini|claude|claude-cli|openai|kimi|...>` produced
zero hyperedges for any corpus: llm._EXTRACTION_SYSTEM only showed
"hyperedges":[] in its output schema and never described what a hyperedge is, so
every model returned the empty array. Meanwhile the agent/skill path, whose
references/extraction-spec.md fully documents hyperedges ("3 or more nodes
participate together..."), produced them — the two prompts had drifted.
Bring the native prompt in line with the skill spec: add the hyperedge
instruction and a populated schema example. The parse/merge side already handled
hyperedges, so this is prompt-only. Verified with a real claude-cli run — a doc
that previously yielded 0 hyperedges now yields one, correctly relativized (#1418).
Adds two guard tests: the native prompt must request hyperedges with a populated
example, and it must share the skill spec's hyperedge wording so they can't drift
apart again.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The GRAPHIFY_OUT override (custom output-dir name / absolute path, #686) was only
respected by some readers. `graphify extract` and several commands hardcoded the
literal "graphify-out", so `GRAPHIFY_OUT=custom-out graphify extract` still wrote
to graphify-out/ and downstream query/serve/update looked in the wrong place.
Resolve the output-dir name through graphify.paths everywhere it matters:
- new graphify.paths.out_path()/default_graph_json() helpers
- __main__: extract write dir, cluster-only/label, query/affected/benchmark
defaults, save-result --memory-dir, uninstall --purge, cache-check
- detect: _MANIFEST_PATH, memory/ + converted/ dirs, and the scan-exclude (a
renamed output dir is no longer re-ingested as source input)
- transcribe._TRANSCRIPTS_DIR; build_merge/serve/benchmark/prs graph-path defaults
Default behaviour is unchanged: with no env var everything still uses graphify-out/.
Verified end-to-end (extract -> cluster-only -> query under GRAPHIFY_OUT=custom-out
writes/reads custom-out/, no stray graphify-out/) and added a CLI regression test.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The split-skill runbook passed '.' as report.generate's root argument in
Steps 4 and 5, so `/graphify /some/path` produced a report titled
"# Graph Report - ." regardless of the scanned directory. It now passes
'INPUT_PATH' (the Aider/Devin monoliths were already correct). Display-only:
no path written to graph.json or manifest.json was affected. Regenerated +
blessed the split-skill artifacts.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The skill runbooks called save_manifest(...) with no root=, so manifest keys
were stored as absolute paths (e.g. /Users/.../main.go). Cloning or moving the
repo then broke `graphify --update`: detect_incremental matched none of the
cached keys, so the entire corpus re-extracted (and the report showed ghost
nodes). The native `graphify extract`/`update` CLI already passed root=target;
the agent-executed runbooks did not.
All four runbook call sites now pass root='INPUT_PATH', relativizing manifest
keys to the scan root (portable forward-slash form, per save_manifest's #777
support): the lean-core skill.md Step 9, the shared --update reference, and the
Aider/Devin monoliths.
The monolith edit is registered as a new sanctioned change-class
(_is_manifest_root_fix_line) in the round-trip guard, mirroring how the #1392
runbook fixes were sanctioned. Regenerated + blessed all artifacts; added a
regression test asserting every shipped runbook threads root= into save_manifest.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
#1418: build_from_json relativized source_file on nodes and edges but stored
graph.hyperedges[] verbatim, so a semantic subagent's absolute path leaked into
graph.json. Relativize hyperedges in build_from_json (to_json has no root to
relativize against), mirroring the existing node/edge handling.
#1423: consolidate the GRAPHIFY_OUT output-dir name into a single graphify.paths
module (was duplicated in __main__, cache, watch) and route the path guards
through it — security.validate_graph_path's base=None discovery + fallback,
callflow_html's project-root resolution, and the post-commit/post-checkout hook
bodies (which now read the env var at hook-run time). A renamed output dir is no
longer validated against the wrong base or missed by the hook.
Tests: hyperedge relativization (test_hypergraph), GRAPHIFY_OUT discovery
(test_security), updated the hook-body contract assertion (test_hooks).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The monoliths are hand-maintained single files frozen against a pinned pristine-v8 blob by the round-trip guard, so they were excluded from the 0.8.44 #1392 batch. Evolve the guard from a positional zip (line-count-exact, single-line-class allowlist) to a multiset diff that classifies every added/removed line against documented sanctioned change-classes, so the multi-line fixes can land while any unsanctioned drift still fails. Add predicates for the four fix classes and broaden the enum/chunk-cleanup predicates to match both the v8 and rewritten forms.
Both monoliths now: thread directed=IS_DIRECTED through every build_from_json call (a --directed run no longer collapses reciprocal edges), scope semantic extraction to document/paper/image, unlink a stale .graphify_cached.json on a cache miss, and run Step 4's zero-node guard before any write with the report/analysis gated on to_json persisting the graph (#1392).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>