102 Commits
Author SHA1 Message Date
abhay-codes07 7c5203dde4 fix(llm): estimate PDF tokens from the extracted text, not the container (#2903)
Token estimation for a PDF read the raw container bytes, which are mostly binary and
bear no relation to the extractable text, so a small-text PDF could be judged oversized
(or vice versa). Estimate from the extracted text instead, memoized on path+size+mtime.
2026-08-21 16:50:44 +01:00
rajarshidattapy 69e2c0deae fix(llm): retry hollow responses instead of bisecting them (#2880)
A response that parses but carries no symbols (a "hollow" reply) was routed into the
truncation-bisection path, which just re-split a chunk the model had already answered
emptily, wasting calls. Give hollow its own finish_reason and route it through a bounded
same-chunk retry with backoff instead; on persistent hollow, give up loudly and mark the
files partial. GRAPHIFY_MAX_RETRY_DEPTH=0 now disables the hollow retry too, so a chunk
costs exactly one call. The #2866 timeout and the truncation paths still bisect.
2026-08-20 15:45:18 +01:00
safishamsi e8bef863f0 fix(llm): gate the reasoning-JSON winner on sanitized content, not raw value (#2882)
A reasoning sketch that lists ids as bare strings (`{"nodes": ["A", "B"]}`) has truthy
arrays but no node/edge objects. The winner gate tested the raw parsed value, so the
sketch won and then sanitized to empty, shadowing the real fragment that followed and
re-triggering the #2880 hollow-response bisection. Sanitize before the emptiness gate so
such a sketch is demoted to the empty-fragment tier. Also document the keyed-candidate
cap crowd-out as a known gap.
2026-08-20 15:44:56 +01:00
rajarshidattapy bce2c9b3d3 fix(llm): recover JSON from reasoning-first model replies (#2882)
Reasoning-first models think out loud, fence the answer, or both, so the reply is not a
bare JSON object and the strict parse dropped the whole chunk. Try the whole reply first
(unchanged fast path), then each balanced-brace candidate, preferring braces that carry
the extraction keys so narration braces do not shadow the answer; a shape restatement
that carries the keys but no content is kept only as a last resort.
2026-08-20 15:44:56 +01:00
hopstreax 26092ce5b1 fix: bisect extraction chunks on timeout (#2866)
A timeout on a multi-file extraction chunk failed the whole chunk. Route recognized
timeouts (subprocess.TimeoutExpired, SDK APITimeoutError, botocore read/connect
timeouts) through the same bounded bisection/merge path already used for
context-window-exceeded errors, so a single slow file no longer takes its chunk-mates
down. A single unsplittable file that times out is left unstamped and retried next run,
as before.
2026-08-19 14:05:58 +01:00
mdshzb04andClaude Opus 4.8 55822b0399 fix(claude): read text blocks after a leading ThinkingBlock (#2697)
Extended thinking is on by default on current Claude models, so a response's
content[0] is a ThinkingBlock and the old content[0].text raised AttributeError
on the SDK claude backend. A shape-aware helper (mirroring the existing
_bedrock_response_text) skips leading non-text blocks and returns the first
text block; a thinking-only response falls back to the default. Only the two
SDK claude call sites are touched; the claude-cli JSON-envelope path is unaffected.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-08-18 22:19:29 +01:00
rajashidattapy 1aab181d15 test: Windows path portability in tests; llm emits POSIX source_file to the model (#2620) 2026-08-12 20:56:28 +01:00
himanshupatro-334 91e43c7678 fix(llm): add rationale guidance to the API extraction prompt (#2482) 2026-08-11 16:54:35 +01:00
safishamsiandClaude Opus 4.8 09a34ad87a fix(update,llm): retry failed extractions; surface claude-cli envelope errors; bump to 0.9.37 (#2543, #2554)
#2543 (adopts PR #2546, thanks @michaelxer): a failed extraction is no
longer stamped in the incremental manifest as up-to-date, so graphify
update retries it instead of skipping it forever; a manifest already
poisoned by the old behavior is healed on the next run; genuinely
unchanged files are not re-processed. Extended to the watch save_manifest
paths too.

#2554 (adopts PR #2555, thanks @annieyii): the claude-cli backend now
inspects the stdout envelope for an is_error result (e.g. a rate limit
returned with exit code 0) and raises it on both the zero and non-zero
exit paths, instead of parsing it as an empty success and bisecting
against a live rate limit.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-08-08 23:37:49 +01:00
safishamsiandClaude Opus 4.8 6ba0868228 fix(cli): surface four silent success-exit failures (#2534)
cluster-only warns when --backend/--model/--batch-size are ignored on the
label-reuse path; the community-label prompt key no longer collides with
the discard sentinel (an echoed key was silently dropped); tree --root
exits non-zero when it matches no source file instead of flattening the
tree; and cluster-only stamps built_at_commit from the analysed graph, not
the shell cwd. Also folds in the cluster-only refused-write guard from
PR #2522 (thanks @aniJani).

Thanks @elecnix for the report.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-08-07 22:49:25 +01:00
safishamsiandClaude Opus 4.8 07b9143d4b fix(hyperedge,skill): merge/load hyperedge integrity + community labels; bump to 0.9.34
#2486 (thanks @adminwat): normalize dict-shaped hyperedge members to ids
(or drop with a warning) so a malformed hyperedge can't abort a completed
merge with a TypeError.
#2484 (thanks @sortakool; approach from @oleksii-tumanov's #1691):
merge-graphs relabels hyperedge member ids and ids with the repo prefix,
unions both inputs' hyperedges instead of clobbering, and writes both
persistence slots.
#2485 (thanks @sortakool): build_from_json reads hyperedges from the
top-level and nested slots; a full validation wipeout is reported loudly.
#2490 (thanks @PapiScholz): the skill Step-5 flow passes curated
community_labels to to_json, so graph.json ships community_name.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-08-05 22:08:29 +01:00
safishamsiandClaude Opus 4.8 ffc6dc0207 fix(llm): correct bedrock max_attempts semantics + stub botocore.config in the reasoning test (follow-up to #2283/#2288)
botocore max_attempts counts the initial call, so GRAPHIFY_MAX_RETRIES must
map to _resolve_max_retries() + 1 (a value of 6 -> 7 total attempts; 0 ->
1, i.e. no retry). Also stub botocore.config in the #2288 reasoning-model
test, which broke once #2283 added the botocore.config import.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-29 17:25:08 +01:00
Zhi Yan Liu a428378ad3 fix(llm): honor GRAPHIFY_API_TIMEOUT in the bedrock backend
The two bedrock-runtime clients (primary extraction in _call_bedrock and
the secondary dispatch path in _call_llm) were built with no botocore
config, so Converse used botocore's 60s default read timeout and ignored
GRAPHIFY_API_TIMEOUT / --api-timeout entirely. A long opus-class
generation then died with "Read timeout on endpoint URL" no matter how
high the timeout was set.

Both client constructions now pass a botocore.config.Config wiring
read_timeout to _resolve_api_timeout() (default 600s), a 10s
connect_timeout, and retries from _resolve_max_retries() in adaptive
mode. This mirrors the fixes that closed the same gap for the claude-cli
subprocess (#1112/#1111) and the secondary LLM dispatch path (#1442) --
bedrock was the last cloud backend still ignoring the knob.

Also updates the README env-var row, which listed the timeout as
applying to HTTP/claude-cli/Anthropic only, and the _fake_boto3 test
fixture to register botocore.config and capture the client config so the
new coverage can assert the timeout is wired.
2026-07-29 17:01:49 +01:00
Zhi Yan Liu d47f3ea152 fix(llm): read the first text block of a bedrock Converse response
Converse returns output.message.content as a list of blocks and does not
promise a text block is first. Reasoning-capable models emit a
reasoningContent block ahead of the answer, and toolUse or future block
types can precede it too, but both bedrock call sites indexed position 0:

    content", [{}])[0].get("text", "{}")

For those models the default was returned on every call, so _parse_llm_json
saw an empty object, _response_is_hollow reported a hollow result,
finish_reason was rewritten to "length", and the adaptive retry bisected the
chunk. Splitting could not converge because the position assumption fails
identically at every chunk size, and raising GRAPHIFY_MAX_OUTPUT_TOKENS did
nothing because output length was never the constraint. stopReason on those
responses was end_turn, i.e. the model had answered correctly.

Selection now keys on the block's shape rather than its position, at both
_call_bedrock and the bedrock branch of _call_llm. A response whose first
block is already text -- every non-reasoning model today -- is unaffected.

On a 48-document corpus the hollow warnings and the bisection to the
recursion cap disappear, the 17 files previously reported as producing no
nodes are extracted, and output tokens drop from 217,538 to 53,274 as the
wasted retries stop.

Fixes #2287
2026-07-29 17:00:15 +01:00
safishamsi 16315f1956 fix(llm): parse claude-cli structured_output, not the prose result field (follow-up to #2095) 2026-07-22 15:57:54 +01:00
Yyunozor 7116a6f5f5 fix(llm): pin claude-cli output to a JSON schema when supported (#2076)
The claude-cli backend delivers the extraction schema in the user turn and
trusts the model to emit raw JSON. Newer Claude Code releases treat that
prompt as an agentic task and report the result in prose instead ("Knowledge
graph extracted — 21 nodes, 20 edges…"), so the graph parses empty, reads as
truncation, and adaptive-retry bisects without ever converging.

Pass --json-schema (structured output) when the CLI advertises it — probed
once via `claude --help` and cached — so the object shape is constrained
regardless of prompt framing. Older CLIs that predate the flag keep the
user-turn prompt as a fallback. The `result` envelope still carries the JSON
string, so the parse path is unchanged.
2026-07-22 15:51:03 +01:00
safishamsiandClaude Opus 4.8 868f75de38 fix(llm): OLLAMA_HOST fallback in detect_backend + default port (follow-up to #2019)
The PR resolved OLLAMA_HOST for the client base_url but not in detect_backend(),
so the headline #1940 case (OLLAMA_HOST set, no --backend) still errored with
"no LLM API key found". detect_backend() now uses _resolve_ollama_base_url (empty
default stays falsy, so ollama remains opt-in and never shadows a paid key). Also
default the port to 11434 when OLLAMA_HOST omits it (a bare host would otherwise
resolve to port 80), and handle bare-port / :port forms.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-20 11:41:43 +01:00
김재현 220a675265 fix(llm): fall back to OLLAMA_HOST when OLLAMA_BASE_URL is unset
Ollama's own server has no OLLAMA_BASE_URL concept — it reads
OLLAMA_HOST (https://docs.ollama.com/faq#how-do-i-configure-ollama-server).
graphify only ever read OLLAMA_BASE_URL, so anyone who configured Ollama
the way its own docs describe (OLLAMA_HOST) had graphify silently ignore
it and fall back to the localhost:11434 default.

Add _resolve_ollama_base_url(default): OLLAMA_BASE_URL still wins when
set (unchanged behavior, graphify's existing convention for OpenAI-
compatible base_url overrides across backends). Otherwise falls back to
OLLAMA_HOST, normalized into an OpenAI-compatible URL (adds a scheme if
missing, ensures a trailing /v1). Wired into the BACKENDS dict's ollama
entry and the two ollama_url call sites that read the env var directly.

Fixes #1940.
2026-07-20 11:29:22 +01:00
shazeb 08166306ba fix(cache): anchor semantic cache writes to cache_root so --out round-trips (#1990, #1991)
With `graphify extract --out <dir>`, the semantic cache write and read
sides disagreed on both location and key anchoring, breaking the cache
round-trip in two ways:

- Checkpoints (#1990): `_checkpoint_chunk` called `save_semantic_cache`
  with only `root=target`, so per-chunk recovery checkpoints were written
  under `<corpus>/graphify-out/` while the reader consulted
  `<out>/graphify-out/` — creating an unwanted graphify-out/ inside the
  analyzed source tree and making every interrupted run re-extract (and
  re-bill) completed chunks.

- Final save (#1991): cli.py passed `root=out_root`, so corpus-relative
  `source_file` paths resolved against the --out directory, failed
  `p.is_file()`, and every result group was silently skipped — the cache
  the reader would consult was never populated at all, with no warning.

Fix, following the split the AST cache already uses (#1774):

- `save_semantic_cache` and `check_semantic_cache` gain a `cache_root`
  parameter mirroring `load_cached`/`save_cached`: `root` stays the
  source-key anchor (content-hash keys, source_file resolution and
  relativization), `cache_root` selects where cache files live. Omitting
  it keeps `root` for both, so existing callers are unchanged.
- `extract_corpus_parallel` plumbs `cache_root` into `_checkpoint_chunk`.
- cli.py extract passes `root=target, cache_root=out_root` at the cache
  read, the checkpoint path, and the final save, and re-anchors the
  prune sweep's live hashes to `target` (keys anchored to out_root would
  mismatch every entry and sweep the fresh cache as orphaned).
- `save_semantic_cache` now warns loudly when every result group is
  dropped because its source_file does not resolve to a real file — the
  silent-0-writes failure mode #1991 asked to surface.

Regression tests cover: checkpoint written under cache_root (not the
corpus, no corpus graphify-out/ created), recovery read finds the
checkpoint via the same root/cache_root split, the final-save call shape
writes entries where the reader looks, the all-groups-dropped warning,
and backward compatibility when cache_root is omitted.

Fixes #1990
Fixes #1991
2026-07-18 19:09:53 +01:00
safishamsiandClaude Opus 4.8 8bc477f2e2 fix(llm): move the unverifiable-node flag to a real, consumed field (follow-up to #1949)
The PR flagged unverifiable code nodes as confidence="UNVERIFIED", but node
confidence is read by nothing and UNVERIFIED isn't in the validated (edge-only)
confidence vocabulary {EXTRACTED,INFERRED,AMBIGUOUS} — a dead, colliding field.

- Move the flag to a dedicated verification="unverified" node field, off the
  confidence key. A node the model already hedged (INFERRED/AMBIGUOUS) is left
  untouched, as before.
- Also check the node id's identifiers (not just the label): ids carry the
  verbatim symbol, cutting false flags on prettified labels.
- Skip the (potentially PDF-re-extracting) source read entirely when a result
  has no code-typed node with a source_file.
- Wire a consumer: diagnose_extraction now counts and reports unverified nodes,
  so the persisted flag is actually surfaced.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-17 00:09:02 +01:00
tpateeq 741876b7f2 fix(llm): downgrade unverifiable code-typed semantic nodes to UNVERIFIED
The semantic (LLM) extraction runs on documents/papers/images; code files are
handled by the deterministic AST engine and never reach the model. A node the
model tags file_type="code" is therefore a symbol it surfaced from within a
document (a name in a fenced code block, an API referenced in a paper), and it
enters graph.json today with no check that the symbol actually appears in the
source the model read. `_out_of_scope` (#1895) only rejects nodes attributed to a
file that was NOT dispatched, so a fabricated symbol on a dispatched file passes.

`extract_files_direct` now verifies every file_type=="code" node whose
source_file was dispatched in the call: if no identifier from its label occurs
(case-insensitive substring) in that file's source bytes, the node's confidence
is downgraded to "UNVERIFIED" (never dropped) and the count is reported to
stderr. Concept/document nodes, nodes without a source_file, nodes attributed to
undispatched files, and labels with no checkable identifier are left untouched.
Best-effort; never aborts extraction; one unreadable file is skipped, not fatal.
2026-07-17 00:02:15 +01:00
safishamsiandClaude Opus 4.8 479f1af455 fix(cache): mark files partial on empty-parse truncation (follow-up to #1950)
The _partial item-marker approach couldn't fire on the most common truncation
shape: a mid-JSON cut parses to zero items, so a sliced document whose second
slice truncated empty was still stamped complete (the PR's own headline case).

- The adaptive-retry give-up sites now record the chunk's own source files in a
  result-level _partial_files list, independent of parsed items; it propagates
  through _merge_two / the recursion merges / _merge_into so it reaches both the
  per-chunk checkpoint and the run-level manifest stamp.
- _partial_source_files unions _partial_files with the item markers.
- save_semantic_cache seeds an empty group for a named partial file with no
  items so its entry is stamped partial, and carries a partial prev entry's flag
  forward so a later clean slice merging over it can't re-promote it to complete.
- the CLI final save now passes partial_source_files (computed before the save)
  so an empty-parse file isn't written back as a complete entry.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-17 00:02:15 +01:00
tpateeq 90dcb6a2ef fix(cache): don't promote truncated LLM chunks to the semantic cache as complete
A chunk whose LLM response is truncated (`finish_reason="length"`) and can't be
recovered by splitting, or that hits the adaptive-retry depth cap, returns a
partial node set. Today that set is checkpointed and written to the content-hash
semantic cache + manifest-stamped as complete, so the incomplete nodes are
served forever until the file content changes or `--force`.

Truncated give-up results are now tagged with an internal `_partial` marker.
`save_semantic_cache` stamps the affected file's entry `partial: True` (detected
from the marker or an explicit `partial_source_files` arg), and `load_cached`
treats a partial entry as a cache MISS, so the file re-dispatches and retries.
The file is also left unstamped in the manifest (like a failed chunk, #933) so
detect_incremental re-queues it on the next incremental run — not only on a full
/ `--force` / content-change run. A file sliced across chunks accumulates via a
partial-aware `merge_existing` peek (`load_cached(allow_partial=True)`) so a
truncated slice is never dropped or silently promoted to complete. Self-heals: a
later complete extraction overwrites the same key with a non-partial entry. The
marker is stripped after the final save so it never leaks into graph.json.
2026-07-16 23:48:07 +01:00
SinghAman21 0d018b4c6d fix(cache): key semantic cache on the extraction prompt (#1939)
The semantic cache keyed entries on sha256(file content + path) alone, with
no component for the extraction prompt that produced them. After an upgrade
that changed the prompt, every unchanged file was a cache hit and replayed
the older prompt's extraction: the run exited 0, cost.json looked cheap, and
the graph silently carried two prompt generations side by side. The reporter
saw 506 of 512 docs replay an older vintage on a rebuild they expected to be
cold; deleting the whole cache was the only workaround.

output, and invalidating them on every release would re-bill extraction for
unchanged files. Fingerprinting the prompt itself keeps both properties:
entries survive releases that don't touch the prompt, and invalidate only
when it actually changed.

Semantic entries now live under cache/semantic/p{fingerprint}/, mirroring the
AST cache's v{version}/ layout. Both extraction paths pass their prompt: the
Python/CLI path from llm.py's _EXTRACTION_SYSTEM (shared by every backend),
and the skill path via a new prompt_file argument in Step B0/B3 naming the
references/extraction-spec.md the subagents were handed. The fingerprint
normalizes line endings so a CRLF checkout isn't mistaken for a new prompt.

Pre-existing entries predate fingerprinting and have unknowable vintage, so
they are still served rather than re-billing a whole corpus on upgrade — but
check_semantic_cache now warns with the count, turning "no signal at all"
into a visible one. merge_existing refuses to fuse such an entry into a
current-vintage write, which would mix two prompts inside one entry and then
attest the result to a prompt that produced half of it.

Old-fingerprint entries are pruned by liveness only, never swept wholesale
the way stale AST versions are: two hosts with different prompts (verbose vs
compact extraction-spec) can share one graphify-out/, and a wholesale sweep
would have each run delete the other's entries and re-bill on every
alternation. prune/clear/cached_files glob recursively so fingerprinted
entries can't become unprunable orphans (the #1527 failure mode).

The two monolith skills (aider, devin) inline their prompt instead of
shipping a spec sidecar and stay on the unfingerprinted path for now.
2026-07-16 16:38:12 +01:00
safishamsiandClaude Opus 4.8 b7ddee3c28 fix(extract): make --mode deep effective over a warm cache; add --force (#1894)
`graphify extract --mode deep` over a warm tree was a silent no-op, for
three stacked reasons:

1. The semantic cache ignored mode: deep runs were served standard-mode
   entries (and vice versa). check_semantic_cache/save_semantic_cache now
   take `mode` (default None, byte-identical when omitted so older
   installed callers keep working) and map it to a namespaced kind —
   cache/semantic/ for None, cache/semantic-{mode}/ otherwise. The
   per-chunk checkpoint in llm.extract_corpus_parallel and the extract /
   cache-check consumers thread the run's mode through (cache-check grows
   --mode/--deep). cached_files, clear_cache, and prune_semantic_cache
   sweep BOTH namespaces; prune uses the same live-hash set for both
   (liveness is content-based, mode-independent) so semantic-deep/ can't
   regrow the #1527 unbounded-orphan problem and inherits the
   files_by_type-derived exclusion gating for free.

2. extract had no --force — the flag was silently swallowed by the
   parser's unknown-arg fallthrough. It is now real (plus GRAPHIFY_FORCE
   env parity with `update`): force disables the incremental gate so
   detection is a full scan and skips the semantic cache READ so every
   semantic file re-dispatches, while the post-run save and manifest
   stamping still happen.

3. The incremental gate dispatched zero files on a warm unchanged tree
   before the cache was ever consulted, so namespacing alone couldn't fix
   the repro. In deep+incremental runs the semantic pass now widens to the
   full live doc/paper/image set from detect_incremental's files_by_type
   (already exclusion-filtered, #1908/#1909) and lets the mode-namespaced
   cache decide hits/misses, with a loud count line so the first deep
   run's full re-dispatch is visible.

Skill-side threading of mode is deliberately deferred to PR-2; mode
defaults keep generated skills byte-compatible.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-15 14:14:13 +01:00
safishamsiandClaude Opus 4.8 0cacd708e9 fix(llm): drop out-of-scope nodes from the merged extraction result (#1895)
The #1757 cache guard refuses to write a cache entry for a node whose
source_file resolves to a real corpus file that was not dispatched, but
the node itself still flowed into merged["nodes"] and landed in
graph.json (plus any edges/hyperedges built on it). extract_corpus_
parallel now filters the merged result right before the #1890
dispatched-vs-returned reconciliation: a node is dropped when its
source_file resolves (against root, same normalization as #1890) to an
existing file (.is_file(), mirroring the #1757 condition) outside the
dispatched set. Non-file source_files (concepts, model-invented anchors)
pass through untouched. Edges whose endpoint and hyperedges whose member
is a dropped node id go with it, plus any edge/hyperedge itself
attributed to an undispatched real file. One summary warning names the
offending files and merged["out_of_scope_dropped"] records the count.
Running before the reconciliation keeps covered/uncovered reflecting the
post-filter graph.

Regression tests: a chunk over A.md+C.md returning a stray B.py node
loses the stray (and its edge/hyperedge) while sibling and concept
attributions survive; a clean run records a zero count and no warning.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-15 00:41:36 +01:00
safishamsiandClaude Opus 4.8 a4ab6ed3f6 fix(llm): reconcile dispatched vs returned files in semantic extract (#1890)
A semantic chunk can return a clean, non-empty response that omits some
of the documents it was given. Those docs vanished from the graph with
no node, no warning, and no cache/manifest stamp, so they were silently
re-dispatched (and typically re-omitted) on every run. extract_corpus_
parallel now diffs the dispatched file set against the source_files that
returned, records the gap in merged["uncovered_files"], and prints a loud
warning naming the omitted files. Smallest visibility guard; routing docs
through the deterministic extractor for a guaranteed file node is a
separate, larger change.

Adds a regression test: a chunk that omits odd-numbered docs is caught
and warned, not silently dropped.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-14 23:45:46 +01:00
safishamsiandClaude Opus 4.8 cfc7cf2c93 fix(cache): resolve FileSlice via unit_path in checkpoint allowlist (#1870)
The #1757 batch-scoping followup built the per-chunk allowlist by reading
FileSlice.rel, which does not exist (a FileSlice carries its parent file
in .path). So every chunk containing a sliced oversized document leaked
the FileSlice object into the allowlist, save_semantic_cache raised
TypeError on Path(FileSlice), and the best-effort except swallowed it:
extraction finished but those chunks were never checkpointed, so a
re-run or a crash/rate-limit resume re-billed them.

Resolve each unit through the canonical unit_path() helper so a slice
maps to its parent file. Adds a regression test that slices a real
oversized .md and asserts the checkpoint writes without swallowing a
TypeError.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-14 00:25:34 +01:00
safishamsiandClaude Opus 4.8 da9616d99e fix(cache): scope the incremental checkpoint write too (#1757 followup)
The #1835 fix scoped save_semantic_cache's final CLI write to an
allowed_source_files allowlist, but the per-chunk incremental checkpoint
in llm.py `_checkpoint_chunk` — the write that actually runs on every
`graphify extract`/`update` via extract_corpus_parallel — still called
save_semantic_cache with no allowlist. A chunk whose model result
mis-attributes a node's source_file to another corpus file would merge
that stray fragment into the victim's cache entry (merge_existing=True).

Scope the checkpoint write to the chunk's own dispatched files (FileSlice
-> .rel, bare Path -> the relative source_file). Also hoist the
`import warnings` in cache.py to module level.

Adds an extract_corpus_parallel integration test: a chunk dispatching
only A.py that returns a node attributed to already-cached B.py must
leave B.py's cache entry untouched.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-13 00:47:10 +01:00
A.Levin 2ea37734c0 Checkpoint semantic cache per chunk so interrupted runs resume
Semantic extraction only wrote to the cache once, at the very end of the
run (save_semantic_cache in __main__ after extract_corpus_parallel returns).
A run interrupted partway — a crash, a kill, or a claude-cli/API run that
exits when it hits a rate limit — therefore lost every completed chunk and
restarted from scratch. On a large corpus with a slow local backend this can
throw away many hours of work.

Persist each chunk's results to the semantic cache as soon as it completes,
in both the serial and threaded paths of extract_corpus_parallel. Add a
merge_existing option to save_semantic_cache so a file split into slices
across several chunks accumulates its slices instead of the later chunk
overwriting the earlier one. The checkpoint is best-effort (a cache write
error never aborts extraction) and can be disabled with
GRAPHIFY_NO_INCREMENTAL_CACHE. Default behaviour of save_semantic_cache is
unchanged (merge_existing defaults to False).
2026-07-08 01:10:44 +01:00
safishamsiandClaude Opus 4.8 5d0137388e Feat: opt-in GRAPHIFY_DISABLE_THINKING; correct deepseek thinking default (#1621)
@sub4biz verified against the live DeepSeek API that deepseek-v4-flash (and
v4-pro) have thinking ENABLED by default, contradicting the built-in config's
stale "non-thinking" comment (now corrected).

The naive fix (mirror the kimi branch and force thinking off) is the wrong
call: @sub4biz's production testing on real corpora found that disabling
thinking removes a rare reasoning-leak failure — which the adaptive
extraction/labeling retry already recovers from — but trades it for far more
frequent benign truncation AND measurably lower extraction quality and file
coverage, confirmed by a blind second reviewer.

So thinking stays ON by default (quality/coverage), with a documented opt-in
`GRAPHIFY_DISABLE_THINKING=1` for users who prefer run-to-run stability. Applies
to reasoning-capable OpenAI-compatible backends at both extra_body sites
(extraction + labeling). An explicit providers.json extra_body still wins, and
the moonshot/kimi branch is unchanged (it must disable thinking or content is empty).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-06 16:02:45 +01:00
safishamsiandClaude Opus 4.8 21b851b3d2 Fix: salvage truncated labeling replies and account for labeling token cost (#1690, #1694)
Two related fixes in the community-labeling path:

#1690 (thanks @vdgbcrypto): a truncated or slightly malformed reply no
longer discards the whole batch with "Expecting value: line 1 column 6".
`_parse_label_response` now salvages the complete `"id": "name"` pairs from
a reply that failed a strict `json.loads` (e.g. one truncated mid-object),
raising only when no pairs can be recovered. The per-batch token budget was
also raised (256 + 48*n, was 64 + 24*n) so models that prepend a short
preamble have headroom to finish the JSON. The exact provider truncation
could not be reproduced without a live key; the parser and budget address
the mechanism.

#1694 (thanks @sub4biz): cluster-only mode reported a hardcoded
`0 input * 0 output` token cost because the labeling LLM calls were never
accounted for. `_call_llm` now accumulates per-response usage into an
optional accumulator threaded through the labeling path and surfaced in
GRAPH_REPORT.md. Backends that do not return usage (the Claude Code CLI)
contribute nothing, which is honest rather than estimated.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-06 12:46:12 +01:00
safishamsiandClaude Opus 4.8 b78248f22a Fix: cap Ollama client-side retries so a hung request cannot multiply the stall (#1686)
A stalled local model wedged for `timeout * (max_retries + 1)`, which with
the default 6 retries turned one long stall into a ~21-minute block with no
progress. `_call_openai_compat` now defaults the Ollama backend to zero
client-side retries (a local model that stalls will not un-stall on retry);
set `GRAPHIFY_MAX_RETRIES` to opt back in. Other backends are unchanged.

This bounds the wait; the underlying stall is driven by the model server
and is non-deterministic, so it is not eliminated. Thanks @Kyzcreig for the report.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-06 12:44:47 +01:00
safishamsiandClaude Opus 4.8 0ff584f070 Fix: tolerate tiktoken special-token text in token estimation (#1685)
`_TOKENIZER.encode(content)` raises ValueError by default when the text
contains a special token such as `<|endoftext|>`, so a doc or corpus that
merely mentions these strings crashed the entire semantic pass. Both
`encode` sites in `_estimate_file_tokens` now pass `disallowed_special=()`
so such text is tokenized as ordinary bytes. Thanks @Kyzcreig for the report.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-06 12:42:39 +01:00
safishamsiandClaude Opus 4.8 e2ef4ef3d1 fix: harden semantic extraction and kill phantom import edges (#1631, #1638, #1632)
#1631: a malformed LLM chunk (a stray non-dict entry in edges/nodes/hyperedges)
crashed the AST+semantic merge and the semantic-cache write with
`AttributeError: 'list' object has no attribute 'get'`, discarding every
successful chunk and writing no graph.json. `_parse_llm_json` now sanitizes each
fragment at the single parse chokepoint (dict entries only; non-list values
coerced to []), protecting the cache writer, the adaptive-retry merge, and the
CLI merge in one place.

#1638: an unresolved bare npm import (`import colors from "tailwindcss/colors"`)
emitted an imports_from edge to the bare id `colors`, which build.py's
pre-migration alias index then remapped onto an unrelated local file of that
stem (backend/utils/colors.py) - a confident EXTRACTED cross-language phantom
edge, one per importing file. The external-import fallback now namespaces its
target with the `ref` prefix (the J-4 convention), so it can never collapse to a
local node id; the ref target has no node, so build drops it as an external
reference.

#1632: with a parallel LLM backend, extract_corpus_parallel merged chunk results
in completion order, so which network call returned first reordered nodes/edges
run-to-run even when the model returned identical content - churning graph.json.
Chunks are now merged in deterministic submission order after the pool drains
(matching the serial path); the progress callback still fires in completion
order. The model's own content variance is unchanged (irreducible).

Full suite: 2882 passed, 3 skipped. Validated end-to-end via a local wheel build
on a mixed TS+Python corpus: `explain colors.py` shows only the real importer,
and graph.json is byte-identical across repeated runs.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-04 03:12:23 +01:00
Tok6Flow0 009a98b6dd Contain symlinked extraction inputs 2026-07-02 22:29:32 +01:00
Jeisson 32ff6d6fb3 fix(claude-cli): deliver extraction instructions in the user turn
The claude-cli backend passed the extraction schema via --system-prompt with
only the raw file dump in the user turn, assuming a replacement system prompt
is the model's sole authority. Claude Code >= ~2.1 (verified on 2.1.197) does
not honour that: it still layers in the local coding-agent context
(CLAUDE.md/AGENTS.md in cwd, skills, MCP) and, given a user turn that is just a
file with no request, replies conversationally ("I see the file, but there's no
actual request attached"). That prose parses to zero nodes/edges, so
_response_is_hollow flags it as truncation and the adaptive-retry path bisects
the chunk indefinitely (94 -> 47 -> 23 -> ...), never converging and never
writing graph.json.

Move the full extraction schema plus an explicit imperative into the user turn
and drop --system-prompt, so the CLI emits the JSON object directly. The
<untrusted_source> prompt-injection guardrails are carried verbatim; model
override, --add-dir image handling, timeout, and token accounting are untouched.
2026-07-02 22:29:30 +01:00
DhruvTilvaandClaude Opus 4.8 4e4935a64a fix(llm): enforce API timeout in the secondary LLM dispatch path (#1442)
_call_llm (used by the dedup LLM tiebreaker) built its Anthropic and
OpenAI-compatible clients with max_retries but no timeout, so requests
on this path silently ignored GRAPHIFY_API_TIMEOUT — unlike the primary
extraction paths (_call_openai_compat / _call_claude) which already pass
both. Add timeout=_resolve_api_timeout() to both constructors.

The PR branch self-neutralized: a v8 merge resolved the conflict in
favor of the max_retries-bearing line and dropped the original one-line
fix, so it is re-applied here on top of current v8 with max_retries
preserved. Adds regression coverage for both _call_llm branches, which
were previously untested.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-29 09:45:03 +01:00
safishamsiandClaude Opus 4.8 64c1f21070 fix(llm): retry rate-limited (429) requests instead of dropping the chunk (#1523)
On strict per-org concurrency/RPM caps (notably Moonshot/kimi), a parallel
`graphify extract --force` 429'd, and because the provider SDKs default to only
max_retries=2, the chunk gave up after two attempts, logged `chunk N failed`, and
was silently dropped — an incomplete graph plus the screen full of
rate_limit_reached_error spam in the report. The SDKs already back off
exponentially and honor Retry-After; they just need more attempts to outlast the
rate window.

The OpenAI-compatible, Azure, and Anthropic clients are now built with
max_retries=_resolve_max_retries() (default 6, override via GRAPHIFY_MAX_RETRIES;
0 disables). For very tight accounts, --max-concurrency 1 further cuts the
concurrency that trips org-level limits.

Reported by @bercedev (#1523). Tests: _resolve_max_retries default/env/invalid,
and the OpenAI-compatible client is constructed with retries.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-28 20:11:18 +01:00
safishamsiandClaude Opus 4.8 b46634ef7a fix(ids): node IDs include the full repo-relative path (#1504, #1509)
BREAKING node-ID format change. The stem that prefixes every node id was the
immediate parent dir + filename, so same-named files in different directories
collided into one last-writer-wins node and silently dropped graph content
(docs/v1/api/README.md and docs/v2/api/README.md both -> api_readme). The stem is
now the full repo-relative path (docs_v1_api_readme vs docs_v2_api_readme);
top-level files are unchanged (setup.py -> setup).

- extractors/base.py::_file_stem -> full path (as_posix; make_id collapses
  separators). The two hand-copied stems (symbol_resolution, mcp_ingest) now
  import the canonical one, so they can't drift again.
- llm.py system prompt + extraction-spec fragments aligned to the same rule,
  fixing the #1509 AST<->LLM divergence (prompt used zero parent dirs -> ghosts).
  Regenerated + blessed the per-host specs.
- build_from_json deterministically re-keys every non-AST node id from its
  source_file via the new stem (and registers old-stem aliases), so a cached or
  pre-migration semantic fragment carrying an old short id reconciles with the
  AST node instead of spawning a ghost. The semantic cache is unversioned, so
  this code-side re-key (not LLM prose) is what makes it survive the format change
  with no re-bill. AST nodes are already canonical and skipped.
- Migration: existing graphs migrate automatically on the next build/update (the
  re-key runs in build_from_json, including the whole graph fed back through
  build_merge); `graphify extract --force` for a clean rebuild. Neo4j persisted
  stores need a re-import; GraphML layouts go stale.

Full suite green (2505 passed); #1504 collision fixed and old-id fragments
re-keyed verified by smoke. Lands on v9 for review before a 0.9.0 release.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-28 16:57:00 +01:00
nuthalapativarunandClaude Opus 4.8 0e8d92cf5f fix(llm): tolerate non-UTF8 claude-cli output on Windows GBK systems (#1505)
On Windows where claude.cmd emits GBK/cp936 bytes, the claude-cli subprocess
decoding raised UnicodeDecodeError and crashed extraction. Both claude-cli
subprocess.run sites (_call_claude_cli and the claude-cli branch of _call_llm) now
pass errors="replace", so incidental non-UTF8 console chatter decodes to
replacement chars instead of crashing — the structured JSON payload (ASCII/UTF-8
on stdout) is unaffected. capture_output=True means the single errors= covers both
stdout and stderr.

Ported from PR #1507 by @nuthalapativarun (dropped an unrelated .gitignore change).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-28 10:32:57 +01:00
jiangyq9andClaude Opus 4.8 ff47316b8a fix(llm): force non-streaming on OpenAI-compatible calls (#1223)
Some OpenAI-compatible gateways default to SSE streaming when `stream` is omitted,
but graphify always reads the result as a single resp.choices[0]. The call would
then fail against those gateways. Pass `stream: False` explicitly.

Ported from PR #1482 by @jiangyq9 (covers the extraction dispatch path,
_call_openai_compat). Maintainer fix on top: applied the same `stream: False` to
the second OpenAI-compatible call site, _call_llm, which feeds the --dedup-llm
tiebreaker and had the identical bug.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-27 10:19:15 +01:00
jc2shileandClaude Opus 4.8 68dba89a99 feat(llm): honor *_BASE_URL for kimi/gemini/deepseek backends (#1458)
The kimi, gemini, and deepseek backends hardcoded their base_url, so users
behind an OpenAI-compatible proxy/gateway or running a self-hosted relay had no
way to redirect them (unlike ollama/openai, which already read *_BASE_URL). Each
backend now reads KIMI_BASE_URL / GEMINI_BASE_URL / DEEPSEEK_BASE_URL and falls
back to its official default when unset, so behavior is unchanged for anyone who
doesn't set the variable.

Ported from PR #1458 by @jc2shile onto current v8. The PR branch carried 624
unrelated files from a stale base; this lands just the clean 16-line llm.py
change. Added subprocess-based tests covering both the override and the default
for all three backends (BACKENDS reads the env at import time, so each case runs
in a fresh interpreter).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-25 12:31:19 +01:00
safishamsiandClaude Opus 4.8 22a58ffc20 feat: parallel community labeling via --max-concurrency / --batch-size (#1390)
label_communities ran batches one LLM call at a time, so a large graph needed
hundreds of sequential calls even on backends that allow heavy concurrency. It
now fans batches out across a thread pool, mirroring extract_corpus_parallel:
results are returned per batch and merged on the main thread (labels dict is
never mutated concurrently, no lock), and workers==1 keeps the original
sequential path verbatim. ollama and claude-cli are forced serial unless the
matching GRAPHIFY_*_PARALLEL env opt-in is set (same guard as extract).

generate_community_labels threads max_concurrency + batch_size through, and the
cluster-only/label CLI parses --max-concurrency and --batch-size (both `--flag N`
and `--flag=N` forms; the space form is parsed explicitly so the value is not
mistaken for the positional scan path by the arg-walk's catch-all).

Output is deterministic regardless of concurrency (keyed by community id). Tests:
parallel == sequential result, batch-size controls batch count, batches actually
run concurrently, ollama forced serial, and the CLI parses both new flags. Full
suite 2393 passed; ruff clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-24 13:53:33 +01:00
safishamsiandClaude Opus 4.8 aad3b47098 Request hyperedges in the native-backend extraction prompt (#1418 follow-up)
`graphify extract --backend <gemini|claude|claude-cli|openai|kimi|...>` produced
zero hyperedges for any corpus: llm._EXTRACTION_SYSTEM only showed
"hyperedges":[] in its output schema and never described what a hyperedge is, so
every model returned the empty array. Meanwhile the agent/skill path, whose
references/extraction-spec.md fully documents hyperedges ("3 or more nodes
participate together..."), produced them — the two prompts had drifted.

Bring the native prompt in line with the skill spec: add the hyperedge
instruction and a populated schema example. The parse/merge side already handled
hyperedges, so this is prompt-only. Verified with a real claude-cli run — a doc
that previously yielded 0 hyperedges now yields one, correctly relativized (#1418).

Adds two guard tests: the native prompt must request hyperedges with a populated
example, and it must share the skill spec's hyperedge wording so they can't drift
apart again.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-22 22:20:40 +01:00
SafiandClaude Opus 4.8 7d07a24708 Don't Path()-coerce FileSlice units in the extract entry points (#1397, #1399)
The 0.8.43 str-path coercion (#1386) ran Path(f) over every item in
extract_files_direct, but extract_corpus_parallel feeds it FileSlice units from
the oversized-doc slicing (#1369), and Path(FileSlice) raises TypeError -- so
semantic extraction of any Markdown file larger than _FILE_CHAR_CAP crashed.
Coerce only non-Path/non-FileSlice entries. The #1386 tests used small files, so
slicing never ran; added a regression test with a file that actually slices.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-19 15:03:38 +01:00
SafiandClaude Opus 4.8 5038129c7e Accept str paths in the semantic extract entry points (#1386)
extract_corpus_parallel and extract_files_direct are typed list[Path] but
crashed with AttributeError on str paths (f.suffix in slicing/partition, f.parent
in packing). Coerce files = [Path(f) for f in files] at both public entry points;
the AST extract() already coerced.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-19 10:01:31 +01:00
ab1e0ec588 Adaptive split-and-retry for community labeling on parse failure (#1280, #1278)
label_communities logged-and-skipped any batch whose LLM response was malformed
JSON, silently losing ~100 names per failed batch on large graphs. Split the
batch at the midpoint and retry each half (smaller prompt -> smaller output),
mirroring _extract_with_adaptive_retry; the base case re-raises so the caller
skips just that batch. Removed leftover scaffolding and corrected the docstring.

Co-Authored-By: CJdev232 <CJdev232@users.noreply.github.com>
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-18 22:50:52 +01:00
SafiandClaude Opus 4.8 4f539e729f Slice oversized text documents so the whole file is extracted (#1369)
_read_files capped every file at 20,000 chars, so a Markdown/text/rST document
longer than that lost everything past the cap with no recovery. Oversized
splittable-text files are now split at heading/paragraph boundaries into units
that each fit the cap and together cover the whole file; every slice reports its
parent file as source_file so the graph isn't fragmented per-slice, and a slice
that still overflows output is bisected and retried. Code and PDFs are never
sliced.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-18 01:55:56 +01:00
SafiandClaude Opus 4.8 5b0c154828 Honour configured output-token cap for OpenAI-compatible backends (#1365)
ollama/openai/deepseek/kimi set max_tokens in their backend config, but the
openai-compat dispatch read only max_completion_tokens (which only gemini
defines), so their output silently capped at the 8192 fallback and truncated
deep-mode JSON. Read either key and give the openai config an explicit cap;
GRAPHIFY_MAX_OUTPUT_TOKENS still overrides.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-17 14:35:58 +01:00