- Universal -h/--help/-? guard after cmd dispatch: any help flag anywhere in
argv stops execution and prints "Run 'graphify --help'" instead of triggering
the subcommand — cursor/kiro/gemini install --help no longer silently installs;
benchmark --help no longer crashes with FileNotFoundError (#821)
- --version / -v / version subcommand: print graphify {__version__} and exit (#818)
- GRAPHIFY_OLLAMA_NUM_CTX=<invalid> now falls through to auto-derived num_ctx
instead of hardcoding 131072 (the cap that causes OOM on constrained VRAM);
pinned num_ctx < estimated input now triggers an explicit truncation warning
with a suggested --token-budget correction (#820)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Gemini is often the cheaper available quota for low-stakes semantic graph extraction, while OpenAI is a useful fallback. Extend the direct extraction backend registry, CLI validation, docs, and tests so headless extraction can use GEMINI_API_KEY, GOOGLE_API_KEY, or OPENAI_API_KEY without changing the existing Claude and Kimi paths.
Constraint: Gemini supports OpenAI-compatible chat completions at the Google generative-language endpoint
Rejected: Native google-genai integration | higher dependency and response-shape churn for the same chat-completions path
Confidence: medium
Scope-risk: moderate
Directive: Keep backend detection explicit and test every accepted API-key environment variable before adding new providers
Tested: uv run --directory vendor/graphify pytest tests/test_llm_backends.py tests/test_chunking.py -q
Not-tested: Live Gemini/OpenAI API calls; no GEMINI_API_KEY or OPENAI_API_KEY present in this environment
- Register .groovy and .gradle in CODE_EXTENSIONS, _DISPATCH, and collect_files
- Add _GROOVY_CONFIG (reuses Java import handler)
- Add regex-based _extract_spock_fallback for Spock spec files where
tree-sitter-groovy wraps the body in ERROR nodes due to def-string methods
- _is_spock_file detects via regex scan (def "...") instead of node-label
heuristic, avoiding false negatives on classes whose name differs from stem
- Fallback retains only file node + import edges from tree-sitter pass to
prevent orphaned constructor/method nodes
- Add tree-sitter-groovy>=0.1.2 dependency
- Add 11 tests covering plain Groovy and Spock paths, including apostrophe
in feature method names
- Add Fortran support (26th language): .f/.F/.f90/.F90/.f95/.F95/.f03/.F03/.f08/.F08
via tree-sitter-fortran; capital-F files preprocessed with cpp -w -P
- Add graphify export {html,obsidian,wiki,svg,graphml,neo4j} CLI subcommands
- Add graphify query/path/explain CLI subcommands
- Reduce skill.md from 63KB to 47KB by replacing Python heredocs with CLI calls
- Extend to_html() with node_limit param for auto-aggregation on large graphs
- Add integration tests for all export/query/path/explain subcommands
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
This commit adds full VB.NET language support to graphify, raising the
supported language count from 25 to 26. The implementation follows the
established LanguageConfig pattern used by all other tree-sitter-backed
extractors.
New dependency:
- Adds optional extra [vbnet] backed by tree-sitter-vbnet (published
to PyPI at https://pypi.org/project/tree-sitter-vbnet/0.1.0/).
Install with: pip install graphifyy[vbnet]
graphify/detect.py:
- Added .vb to CODE_EXTENSIONS so VB.NET files are discovered during
corpus ingestion and file-system watching.
graphify/extract.py:
- _import_vbnet(): import handler for imports_statement nodes; emits
imports edges using the namespace_name child text.
- _vbnet_extra_walk(): extra-walk hook that intercepts namespace_block
nodes, emits a namespace node, and recurses.
- _VBNET_CONFIG: full LanguageConfig covering class_block / module_block /
structure_block / interface_block as class types; method_declaration /
constructor_declaration / property_declaration as function types;
invocation call nodes with target/member_access fields.
- VB.NET-specific branches in _extract_generic:
* Class body: VB.NET has no wrapper body node; inherits and implements
are named fields directly on the class_block. Emits separate inherits
and implements edges for each base type, stripping generic arguments.
* Constructor name: constructor_declaration carries no name field in
the grammar; always resolves to New.
* Function body: uses the declaration node itself as body sentinel so
the call-graph pass can find invocations inside methods.
- extract_vbnet(path): public wrapper that delegates to _extract_generic.
- _DISPATCH['.vb']: routes .vb files to extract_vbnet.
pyproject.toml:
- Added vbnet = ['tree-sitter-vbnet'] optional dependency group.
- Added 'tree-sitter-vbnet' to the all extra.
tests/fixtures/sample.vb:
- New fixture file exercising: Imports statements, Namespace block,
Interface, Class with Inherits + Implements, Module, Structure,
Sub/Function/Property methods, and method calls.
tests/test_languages.py:
- Added 13 tests covering: no-error, class/interface/module/structure
detection, method detection, imports relation, inherits edge,
implements edge, and no-dangling-edges invariant.
README.md:
- Updated language count 25 to 26.
- Added VB.NET to language list and file-extension table.
- extract_sql(): deterministic tree-sitter extraction of tables, views,
functions, foreign key references, and FROM/JOIN reads_from edges
- .sql added to CODE_EXTENSIONS and dispatch table
- tree-sitter-sql added as optional dep under [sql] extra
- xlsx_extract_structure(): extracts sheet/table/column nodes from .xlsx
(utility — pipeline wiring in follow-up)
- 6 new SQL tests, 447 total passing
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Three independent improvements to extract_corpus_parallel:
1. Token-aware chunking. Replaces `chunk_size=20` static packing with
a greedy packer keyed on `token_budget` (default 60_000), grouped
by parent directory so related artefacts share a chunk. Pass
`token_budget=None` to fall back to fixed-count packing.
2. Optional tiktoken (added to the [kimi] extra). When available,
`_estimate_file_tokens` uses cl100k_base for accurate counts;
without it, the existing chars/4 heuristic kicks in. Kimi-K2 ships
a tiktoken-based tokenizer so estimates against Moonshot are very
close to truth.
3. True parallelism. The function name said "parallel" but the body
was a sequential for-loop. Now uses ThreadPoolExecutor capped at
`max_concurrency` (default 4 — conservative against provider rate
limits). `on_chunk_done(idx, total, result)` still fires once per
chunk with the original submission idx so progress UIs work
unchanged. `max_concurrency=1` skips the pool to preserve
sequential semantics.
Plus failure tolerance: a chunk raising is now caught, logged to
stderr, and the run continues. Other chunks' results merge as normal.
On a 162-file repo (~125k words), the same work that took ~36 min
sequential under the old code finishes in ~7 min.