All-platform progressive-disclosure skill split + generator (addresses #1106).
Splits each platform's skill into a lean core (~615 lines, full default pipeline inline) + on-demand references/, generated from a single source via tools/skillgen with a CI/pre-commit drift gate. 13 hosts split, aider/devin stay monoliths. Also fixes the stale bare-path bugs across the previously hand-maintained variants and moves the always-on blocks into packaged markdown.
Verified: all 5 generator guards pass, byte-verbatim load-bearing slices, lean cores self-sufficient on the default path across all 13 split hosts, references gated to non-default branches, description preserves the graphify-out-query-first clause. Supersedes #1119 (Claude-first subset).
Known follow-up applied on top: harden _always_on() against a missing packaged file so a partial install can't brick the CLI.
Follow-up to the file-ordering fix. The from-scratch build writes each node's
community field straight from cluster()'s enumerate() after a STABLE size-sort,
so the hundreds of equal-sized small communities in a sparse graph were ordered
by the partitioner's (not seed-stable) enumeration order. Their integer IDs
permuted run-to-run, which reads as 77-88% "community churn" in a per-node cid
diff even though the actual grouping is reproducible.
Add a tuple(sorted(nodes)) tiebreak to make the sort a total order, so an
identical grouping always yields identical community IDs. Verified: with the
partition returned in shuffled order across 5 runs, the node->cid map is now
identical. (A separate ~0.06% community-count drift remains - likely
non-canonical edge weights upstream - tracked separately.)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
tree-sitter-dm (BYOND DreamMaker) publishes only a Windows wheel, so on
Linux/macOS uv/pip compiled it from source and aborted the entire
`uv tool install graphifyy` when a C toolchain or python3-dev was missing.
It is the only core grammar lacking Linux/Mac wheels. Move it from core
dependencies to an optional `dm` extra (also in `all`); the default install
now needs no compiler. extract_dm already imports the grammar lazily with a
graceful fallback, and .dmi/.dmm/.dmf use no tree-sitter, so only .dm/.dme
AST parsing is gated behind the extra.
Also: guard the .dm grammar tests with pytest.importorskip so a plain
`uv sync` (no extras) doesn't hard-fail; bump to 0.8.28 with a changelog
upgrade note; switch the README extras table + type notes from pip to
`uv tool install` (uv is the recommended installer, and pip-into-a-uv-tool-env
is the exact ModuleNotFoundError the README already warns about). 0.8.28 also
ships Kilo Code (#512) and the Dart parser modernization (#1098).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Modernizes the Dart extractor: comment stripping, `part of` redirection, nested-generic-aware extends/with/implements parsing, generic type-argument mapping, and generic call detection.
Verified against current v8: merges cleanly, full suite 1530 passed, and inheritance edges connect to same-file class nodes with no ghost-node splits or dangling edges (consistent with the #1033/#1096 node-ID invariants).
Adds Kilo Code as a supported platform: native skill + /graphify command, install/uninstall, and a .kilo tool.execute.before plugin (mirroring the OpenCode integration).
JSONC config is handled non-destructively - existing .kilo/kilo.jsonc is read but never rewritten; automated plugin registration goes to kilo.json, preserving user comments (addresses the Qodo review flag).
Verified locally: full suite 1525 passed, install/uninstall smoke test works.
The auto-labeling branch wrote .graphify_labels.json, but the unconditional
write at the end of cluster-only already persists the final labels dict
(LLM or placeholder). Removed the duplicate; no behavior change.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
No behavior change. label_communities / generate_community_labels live next
to the backend dispatch (_call_llm, detect_backend) they use, so the lazy
import and separate file are no longer needed. Tests unchanged except the
import path; they still monkeypatch graphify.llm._call_llm.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Community labeling was an agent-only step (skill.md Step 5): inside Claude Code
or Gemini CLI the agent names communities itself. Run as a bare CLI
(python -m graphify extract . --backend X), no agent does Step 5, so labels
stayed Community 0/1/2 for every backend.
Add graphify/labeling.py: one batched _call_llm asking the backend for a
{cid: name} map from each community's top node labels (god nodes first),
with per-community placeholder fallback and graceful degradation (no backend,
API error, or malformed reply -> Community N, never crashes).
Wire it in:
- cluster-only auto-labels when no .graphify_labels.json exists (the reported
repro path); --no-label opts out, --backend=<name> overrides auto-detect
- new `graphify label <path>` subcommand force-regenerates names
- extract prints a hint pointing standalone users at cluster-only
Tests mock the backend (no network): happy path, partial/fenced/malformed
replies, no-backend and error fallback, god_nodes dict shape.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
The #1033 remap relativized file node IDs but symbol IDs still embedded the
absolute parent-dir stem, so a root-level file's symbols became
<rootdir>_main_run while the file node was correctly main and the skill.md
spec (and semantic subagent) wants main_run -- splitting every top-level
file's symbols into AST/semantic ghost pairs.
Extend the remap chokepoint to also canonicalize the symbol stem prefix
(old _file_node_id(path) -> new _file_node_id(rel)), gated by source_file so
two files sharing a prefix cannot cross-contaminate. raw_calls.caller_nid is
rewritten too, since the cross-file call pass consumes it after the remap and
would otherwise dangle. No-op for nested files (immediate parent identical in
absolute and relative form).
The originally-filed root cause (facts-collector path form) was a stale-cache
red herring; the real bug is the symbol-level continuation of #1033.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Two gaps in TypeScript inheritance:
1. interface heritage is an extends_type_clause node, not class_heritage,
so the walker never saw it and interface extends produced no inherits
edge. Add an extends_type_clause branch reusing _ts_heritage_clause_entries
(handles multiple extends).
2. a same-file superclass has no import alias, so the use-fact resolver
(which only consulted the import table) dropped it; only imported bases
resolved. Add a same-file fallback against the file's own symbol_nodes,
scoped to inherits/implements so it does not duplicate same-file calls
that already resolve via the call-graph pass. Import resolution still
takes precedence.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
to_obsidian and to_canvas built note filenames from node labels with no
length cap, so a label >=255 bytes crashed write_text with OSError. Add a
shared _cap_filename helper that caps on UTF-8 bytes (not chars, so CJK
labels don't slip past) and appends an 8-char hash of the full label when
truncating, so two distinct labels sharing a long prefix stay distinct.
Both safe_name builders route node, community and canvas filenames through
it; wikilinks stay consistent because they read the same filename dict.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
AST file nodes were ID'd from the full relative path plus extension
(match_script_pipeline_step_py) while semantic subagents follow the
{parent_dir}_{stem} spec (script_pipeline_step), so every file split
into two disconnected ghost nodes.
Fix at the single remap chokepoint in extract(): file node IDs and all
edge endpoints already funnel through the #502 relative-path remap, so
changing that remap to emit _file_node_id (one parent dir, no extension)
converts the node and every referencing edge together - Python, TS, Lua,
C and bash import edges all stay connected. symbol_resolution pre-computes
the canonical form directly (bypassing the remap) so it is synced too.
Per-site conversion (as attempted in #1038/#1065) orphans edges because
it moves the node without the edge targets; the chokepoint approach
avoids that entirely.
Backward compat for existing graphs: graphify extract --force, as the
skill.md spec already documents.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
os.walk order is non-deterministic across runs (filesystem b-tree order
shifts with cache/mount state), causing first-writer-wins decisions in
cross-file resolution to produce different node IDs each run; sorting
all_files and per-FileType lists stabilises graph.json without changing
any parent-child or connectivity semantics (closes#1090)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
pytest derives tmp_path from the test name, so the path contained
"incremental" and the bare-word assertion always failed; check for
the actual incremental-mode phrases instead
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
reconfigure stdout/stderr to UTF-8 at startup; replace → and — in all
print statements with ASCII equivalents as belt-and-suspenders fallback
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* feat: detect circular import dependencies at file level
- Add find_import_cycles() to analyze.py
- Collapses symbol nodes to parent files, builds directed file graph
- Uses nx.simple_cycles() bounded by max_cycle_length (default 5)
- Deduplicates rotations, returns shortest cycles first
- Considers both imports_from and re_exports edges
Tested on a 976-file Next.js codebase: found 4 cycles including
a known utils↔barrel circular dependency and a 4-file API cycle.
* fix: resolve import-cycle merge blockers
- use source_file-only endpoint resolution (no label fallback)
- support Graph/DiGraph orientation via edge source_file
- return structured cycle records and include self-loops
- integrate Import Cycles section into GRAPH_REPORT.md
- expand cycle tests for real-schema IDs, undirected input,
missing source_file nodes, and non-import relations
Three compounding bugs caused ~30-50% of semantic chunks to come back
as 'hollow responses' on the claude-cli backend, triggering adaptive
bisection that doubled or tripled the number of subprocess calls.
Root causes
-----------
1. _parse_llm_json only stripped markdown fences when raw.startswith('```').
Claude frequently prepends a short preamble before the fence
('Here are the extracted entities:\n\n```json\n{...}```'), making
the check fail. json.loads then drops the chunk. Each bisected half
may exhibit the same failure, so cost compounds.
2. _call_claude_cli used --append-system-prompt, which layers graphify's
extraction prompt on top of Claude Code's default interactive-agent
prompt ('use markdown formatting', 'output text to communicate with
the user'). The conflicting instructions explain ~50% of the
preambles and fences from (1). Switching to --system-prompt (replace)
eliminates the conflict at the source.
3. claude-cli defaults to Opus, which is overkill for the structured
JSON extraction graphify performs. New GRAPHIFY_CLAUDE_CLI_MODEL env
var lets users opt into haiku / sonnet for big builds. Default
behaviour unchanged when the env var is unset.
Fix
---
- Robust _parse_llm_json: strips fences regardless of position, with a
balanced-brace fallback that scans for the first complete JSON object
in the response. Handles preambles, trailing prose, prose-wrapped
JSON without fences. Diagnostic log on terminal failure includes the
first 200 chars of the response.
- _call_claude_cli switches to --system-prompt.
- _call_claude_cli respects GRAPHIFY_CLAUDE_CLI_MODEL when set.
Tests (tests/test_llm_parser.py)
--------------------------------
- The four PR-body failure modes: preamble+fence, prose+JSON, raw JSON,
total refusal.
- Bonus: uppercase fence tag, unclosed fence, empty response.
- argv shape: --system-prompt present, --append-system-prompt absent.
- argv shape: --model added iff GRAPHIFY_CLAUDE_CLI_MODEL is set.
19/19 tests pass (9 pre-existing in test_claude_cli_backend.py +
10 new). Verified end-to-end on a 800-file repo: 0 hollow responses
after, vs ~30-50% before; output tokens -93%; wall time 44 min -> 4 min.
* chore: declare pytest as a uv dev dependency
The contributing guide currently tells contributors `pip install pytest`
as a separate step, and CI does the same. Move pytest into PEP 735
`[dependency-groups]` so it's declared in pyproject.toml and `uv sync`
installs it by default (no `--with` workaround, no separate install
line). Update CI to use astral-sh/setup-uv + `uv sync` + `uv run pytest`,
and refresh the Contributing section of the README to match.
`[dependency-groups]` is the right home (vs `[project.optional-dependencies]`)
because pytest is dev-only and shouldn't appear in the published wheel's
optional features list alongside things like `pdf` or `mcp`.
* remove uv.lock from gitignore