The modal submitted the active tag as `tags=<uuid>`, but the watchlist filters on
`tag` - nothing reads `tags`. So searching from inside a tag view silently
searched every watch, despite the label promising "URL or Title in '<tag>'".
The test pulls the hidden field straight out of the rendered modal and feeds it
back to the watchlist, so the field name can't drift from the arg again.
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* UI - Search - Fix search modal navigating to the host root on sub-path deployments
base_path was referenced by search-modal.js but never defined, so searching
always jumped to the host root instead of the X-Forwarded-Prefix sub-path.
* UI - Search modal - submit natively instead of rebuilding the URL in JS
Alternative to defining a `base_path` JS global: give the search form a
server-rendered `action`, so url_for() supplies the reverse-proxy sub-path the
same way every other link on the page already does, and let the browser submit.
Drops the submit handler and the Enter handler from search-modal.js - Enter in
the input reaches the footer's submit button via implicit submission, which also
runs the `required` validation the synthetic `new Event('submit')` skipped.
The hidden tag field is only rendered when a tag is active, so a plain search no
longer carries an empty value.
Also drops the nginx-job grep for the rendered markup - test_search.py already
covers the sub-path case, and asserting on an exact HTML attribute string from a
shell grep breaks on any unrelated edit to that tag.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: ponstream24 <87808547+ponstream24@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Browser fetchers - Judge a fetch on the document we end up extracting, not the first navigation
The goal is to compare the text of the page the browser lands on, even when the site navigates
again after the first response. Both fetchers were bound to the first navigation, which shows up
as two different bugs:
1. pyppeteer hangs until the hard processing timeout. Its navigation watcher is bound to the
loaderId of the navigation it started, so when the site replaces that document the 'load' it
waits for never arrives for that loaderId. With timeout=0 and setDefaultNavigationTimeout(0)
there is nothing to break the wait, so goto() blocks until
PUPPETEER_MAX_PROCESSING_TIMEOUT_SECONDS (180s) kills the fetch and the watch records an empty
xpath_data - while the browser is sitting on a fully loaded page. Traced on slated.com:
0.24s goto start
0.74s main frame networkIdle loaderId=A0E0E0B2 <- never gets 'load'
2.72s main frame init loaderId=56D2B5EF <- re-navigated to get.slated.com
3.49s main frame load loaderId=56D2B5EF <- fires for the new document
25.2s goto still hanging, frame._loaderId is now 56D2B5EF
Now the navigation races goto() against the main frame firing 'load', bounded by
BROWSER_NAVIGATION_TIMEOUT_SECONDS (default 30), and falls back to the document we can see.
slated.com / getastra.com / addupsolutions.com went from a 180s timeout with no content to
200 with full content in 5-35s.
2. Both fetchers reported the status of the interstitial. A site that gates unseen visitors with
an error status plus a client-side redirect (reported against fotokoch.de: 503 + meta refresh,
then a 200 with the real page) failed the watch even though the content was present, and the
only workaround was ignore_status_codes, which also hides genuine 404s and 500s forever.
The fetchers now keep the latest main-frame document response and judge on that - the refresh
lands during the existing extra_wait, so the 200 wins.
Playwright also waits for a settled load state before extracting, which is what produced
"Execution context was destroyed, most likely because of a navigation" when the refresh collided
with extraction.
The navigation-response tracker is installed once per page and shared between the fetcher and
action_goto_url() rather than each navigation adding its own listener - 'response' fires once per
HTTP response, hundreds of times on a heavy page, so the callbacks are worth not duplicating.
Verified one listener remains after an install plus four navigations.
Selenium is unaffected either way - it hardcodes status_code = 200 because WebDriver cannot see
the HTTP status.
Tested: new test_renavigation.py covers the interstitial case end to end and was checked to fail
without the fix and pass with it, on both fetchers. The test endpoint gates on last-seen time
rather than a hit count, because a counter lets the second check see a clean 200 and the test
then passes without the fix. Full browser suite 11 passed on playwright and on pyppeteer, 494
unit tests pass.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Puppeteer fetcher - One content-ready deadline instead of a stopLoading watchdog per frame event
Page.stopLoading is what stops a page that would otherwise load forever waiting on a subresource
that never answers, so that we can still screenshot and scrape what rendered. That intent was
right, but it was implemented as a fire-and-forget task armed by every frame event, which measured
on a single fetch of an iframe-heavy page came to:
14 watchdog tasks spawned
11 page-wide Page.stopLoading calls
3 tasks outliving the fetch and firing against a closed page
Page.stopLoading takes no frame or loader argument - it is the Stop button, and it stops the whole
page. Verified directly: one call stopped a pending main frame and a pending iframe in the same
instant. So the other 10 calls were redundant, and because they landed at arbitrary later times
they could stop a *subsequent* navigation we actually wanted - which is the likeliest reason the
same URL fetched in 8s on one run and 35s on the next.
Replaced with a single deadline, awaited inline so nothing can outlive the fetch (there is no
create_task left in this file at all):
navigate (bounded) -> wait the configured delay -> Page.stopLoading -> extract
The delay is measured from when navigation finished, not from when it started. Anchoring it to the
start would quietly rob a slow-loading page of its settle time, and letting JS-rendered content
appear after load is the whole point of the setting. Verified with a server that takes 5s to answer
and renders via JS 2s after load: total 9.6s for a 4s delay, and the late content is captured.
Because a page is never reliably "finished" - many sites navigate as part of their normal design -
the delay restarts when the MAIN frame replaces its document, so a redirect or interstitial gets
the same settle time the first document got. Iframes do not restart it, and it is capped by
BROWSER_CONTENT_READY_MAX_RESETS (default 2).
Only the existing "wait n seconds before extracting text" stays user-facing;
BROWSER_NAVIGATION_TIMEOUT_SECONDS is a safety net with a sane default rather than a second knob
for users to reason about. This matches what other scrapers do: bound the navigation, do not fail
when it times out, settle, then extract.
Timings are also more predictable now - the four reported URLs went from 5-35s of variance to
4.6-7.2s at a 3s delay, all with full content and a 200.
Tested: 11 passed pyppeteer browser suite, 7 passed playwright, 494 unit + llm.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Compare against a resolved data_dir in the history path-traversal test
Watch.history resolves entries with os.path.realpath, so
test_normal_snapshot_entry_is_accepted compared a resolved path against an
unresolved data_dir. On macOS the datastore lives under /tmp, which is a
symlink to /private/tmp, so the assertion fails for a path that is in fact
inside the directory. The guard is correct; the test was not.
Resolve both sides, matching what the production code does.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* Add a test that actually exercises the history containment check
Disabling the containment check in Watch.history left every test in
TestHistoryPathTraversal passing. os.path.basename() reduces both traversal
fixtures ('/etc/passwd', '../../etc/passwd') to 'passwd', so neither reaches
the check — they stop at the os.path.exists() test below it.
A bare '..' survives basename() and resolves to the parent of data_dir, which
exists, so the containment check is what rejects it. With that check disabled
this new test is the only one in the class that fails.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: GG5533 <285285461+GG5533@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* fix(conditions): support zero values in condition filtering and json logic conversion
- Fix filter_complete_rules dropping rules where value is 0/0.0 due to 0 == False in Python
- Fix convert_to_jsonlogic raising EmptyConditionRuleRowNotUsable on value=0 due to truthiness check
- Fix str != 'None' type comparison typo in convert_to_jsonlogic
- Add comprehensive unit tests covering zero value condition filtering, conversion, and execution
* chore: re-trigger CI checks
---------
Co-authored-by: Andrew Peabody <apeabody@users.noreply.github.com>
`tag` on POST /watch was documented as taking a tag UUID, but the value went to
add_tag(title): a UUID silently created a junk tag *titled* with that UUID and never
applied the tag the caller asked for. `tags` (UUIDs) was the only thing that worked.
- `tag=` now resolves an existing tag UUID to that tag, still falling back to title
matching/creation for names. A UUID-shaped value matching nothing is skipped with a
warning rather than becoming a group named after a UUID.
- Blank tokens ("One,,Two,") no longer store False in watch['tags'] - add_tag() returns
False for an empty title and it was appended unguarded. Consumers tolerate it
(get_all_tags_for_watch() dictfilt()s over known tags) but it is not valid data.
- add_tag()'s title search is extracted to tag_uuid_for_title(), so existence can be
tested without creating as a side effect. add_tag()'s contract is unchanged.
- api-spec: `tag` is marked `deprecated: true` (so Redoc renders the badge) and states
plainly that it takes names, not UUIDs. `tags` now says what it really does - applied
verbatim, never creates, unknown UUIDs stored as dangling refs. Rendered docs rebuilt.
Every claim in the new field docs is asserted in test_api_tags.py against the real API.
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Env var - PAGE_WATCH_LIMIT enhancements
* Rebuild API docs
* Bump APi doc version
* test: cover PAGE_WATCH_LIMIT across every add path
- API create returns 429; API import returns 429 and refuses the batch whole
- quick-add and the UI importer flash the limit (importer once per file, not per row)
and hand unimported URLs back
- clone at the limit no longer KeyErrors
- an instance already over the limit still loads from disk and stays editable, only
new watches are refused
- the Info tab shows the limit only when one is set
- add_watch() with no request context returns None instead of raising from flash()
- an absent, empty or unparseable env var all mean unlimited
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
- Define _NO_TEMPERATURE_MODEL_KEYWORDS constant and omit temperature=0 on initial request
- Omit thinkingConfig in _thinking_extra_body for flash-lite models
- Unconditionally strip rejected sampling params and extra_body thinkingConfig on BadRequestError (handling Gemini 400 INVALID_ARGUMENT without named details)
- Add comprehensive unit tests in test_llm_client_gemini.py
* UI - LLM section tidyup
* Rebuild translatiosn
* UI - Fixing language for prompt adjustement to be more clear
* UI - clarify field action
* UI - clarify layout for sections
* WIP
* test tweaks
* Scheduler+API Bug - if an invalid timezone was set (through edit of watch or API) it could have crashed the scheduler, Added `timezone` to the official API docs
* adding missing files
Follow-up to #4340. The <think> stripper added there only matches a closed pair, but a
reasoning scratchpad routinely contains JSON of its own ("initially I thought
{"important": false}, but..."), so any leftover scratchpad lets _extract_json lock onto
a discarded intermediate answer. Three shapes slipped through, all of which inverted the
verdict to important=false and therefore silently suppressed the notification:
- opener stripped by the provider/chat template, only </think> comes back over the wire
- <thinking> spelled out in full
- unterminated block, i.e. the response was cut off mid-thought by max_tokens
The first two are now stripped. The third raises ValueError, because a truncated response
holds no answer at all - only the abandoned guess. Raising routes it to the existing
handler in evaluator.py, which passes the change through as important rather than dropping
it; parse_eval_response deliberately does not catch ValueError, since its own fallback
(important=False) would suppress the notification instead.
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
- Simplify _to_bool using changedetectionio.strtobool.strtobool
- Strip <think>...</think> reasoning blocks in _extract_json for reasoning models
- Fix _annotate_moved_lines short-circuit so standalone relative timestamps are always annotated
- Add comprehensive unit tests in test_response_parser and test_prompt_builder
* Restock detection - fix inverted condition that skipped the OpenGraph availability fallback
get_itemprop_availability() only dug through OpenGraph properties when
price was missing OR when availability was ALREADY found - so on pages
that expose price via JSON-LD but stock state only via the OpenGraph
commerce tags (<meta property="product:availability">), the availability
(and currency) was silently dropped.
The inner loop already guards each field with 'if not value.get(...)',
so the outer gate now matches that intent: dig OpenGraph when any of
price/availability/currency is still missing.
Co-Authored-By: Claude <noreply@anthropic.com>
* Normalise availability after the OpenGraph fallback, not before
An availability found via OpenGraph was left raw, so the Facebook
commerce spelling 'in stock' never matched the 'instock' test and the
product read as out of stock.
---------
Co-authored-by: Claude <noreply@anthropic.com>