On any upgraded install update_35 has migrated the old per-engine request timeout and
User-Agent into browser configs keyed by the engine name ('html_requests',
'html_webdriver'), so those ids exist in BOTH the built-in engine list and browsers.json.
list_watch_browser_choices() concatenated the two sources, so the watch edit page rendered
the same browser as two radios with the same label - and the group override select did the
same, from its own copy of that concatenation.
Deduplicated by value at the source, keeping each value's first position (built-ins stay in
engine order) with its last label - the saved one, since it is the same browser under the
name the user can actually change. The group override select and the Add-Watch list now
derive from that single list instead of rebuilding or re-deduplicating it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
An "extra browser" was a name + a ws(s):// endpoint in settings.requests.extra_browsers,
selected by a watch as the magic string 'extra_browser_<name>'. That string resolved to
html_webdriver plus a custom connection URL, which meant the protocol the endpoint was
spoken to came from env vars rather than from the entry: CDP over a WebSocket with
PLAYWRIGHT_DRIVER_URL set, CDP via pyppeteer with FAST_PUPPETEER_CHROME_FETCHER, and the
W3C WebDriver protocol over HTTP on a Selenium-only install - where a wss:// URL cannot
work at all. The form only ever accepted ws:// / wss://, so the feature was silently
broken on exactly the installs that could not honour it.
So it becomes an engine, html_external_cdp, which pins the protocol: a subclass of the
Playwright fetcher that takes its endpoint from the watch's browser config
(FetcherConfig.connection_url) instead of the environment. It is base-only
(ready_to_use=False) because an endpoint is required, so each endpoint is one browser
config ("variation") on the Browsers page - which is what the old settings list was.
update_36 migrates each extra_browsers row to such a variation, keyed by the SAME
'extra_browser_<name>' string watches already hold, so no watch, group override, API value
or global default needs rewriting; the legacy selector simply becomes a real browser-config
id. A row whose endpoint the model rejects is logged and skipped rather than taking the
update chain, and with it startup, down.
Knock-on cleanups, all of which delete a special case rather than add one:
- The proxy opt-out for custom endpoints is now Fetcher.ignores_proxy_setting, asked of
the engine, instead of a string-prefix test in call_browser().
- A live browser-steps / visual-selector session asks the engine where to connect
(Fetcher.browser_steps_connection_url, overridden by html_external_cdp) and refuses an
engine whose supports_browser_steps is False, instead of reading the env var itself and
silently stepping a browser the watch does not check with. That refusal is real: on a
Selenium install html_webdriver cannot drive a live session.
- is_valid_browser_selector() answers "may a watch store this in fetch_backend?" in one
place; the API (create/update/import), the quick-add form validator and the bulk "set
browser" operation each had their own copy, which is how they came to disagree about
whether a browser-config id was acceptable.
- api-spec.yaml's fetch_backend pattern enumerated extra_browser_* while rejecting
browser-config ids and every engine newer than html_webdriver. Valid values are
per-install, so the schema now bounds the string and the handlers do the real check.
- html_external_cdp registers unconditionally (unlike html_playwright_builtin): migrated
configs name it, so it must resolve even without the playwright library, or those
watches would quietly fetch with the plain HTTP client. The library is imported lazily
inside run(), and an unavailable engine now warns instead of falling back silently.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A browser config could carry fields its engine ignores, and in two cases a *wrong*
value: saving a plain-HTTP-client variation wrote browser_type='chromium' (the
unrendered SelectField's own default) and delete_created_files=False (an unchecked,
unrendered BooleanField), because the form mapped every field regardless of what the
engine can honour.
The capability a field needs now rides on the field itself - _needs() wraps
Field(json_schema_extra={'capability': ...}) - so there is one declaration site
instead of a parallel name->flag table to keep in step. FetcherConfig.applicable_fields()
reads it back and drives BOTH halves: which fields _browser_config_fields.html renders,
and which ones FetcherConfig.from_submitted() accepts from a POST. So a field can no
longer be rendered-but-unsaveable or hidden-but-written-from-its-widget-default, and a
crafted POST cannot put a setting on a browser that ignores it.
Also rejects control characters in the per-profile User-Agent at the model, because that
value is written straight into an outbound request header. urllib3 would refuse it at
request time anyway, but now it can never be persisted in browsers.json at all.
Reads stay deliberately tolerant (pydantic extra='ignore' plus the coercion pass in
BrowserConfigStore.all()), so a browsers.json written by another version still loads.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Several places call gc.collect() after a check to keep C-level memory
(pyppeteer buffers, libxml2 documents, PIL, brotli) from accumulating.
Individually each is reasonable; run concurrently by many fetch workers they
become a storm. Every gc.collect() is a full stop-the-world pass that walks the
whole heap holding the GIL, so at FETCH_WORKERS=50 the process spends most of
its time stopped in the collector - which also starves each worker's asyncio
loop, leaving its CDP websocket unread.
Measured on 153 real puppeteer checks of a live site at FETCH_WORKERS=50, with
an `//div` include filter so the lxml document tree is realistic:
collects objects freed gc time checks/sec CPU/check RSS
one per call site 790 2,594,098 91.3s 0.766 1.373s 279.6MB
debounced to 1s 48 2,219,010 7.3s 1.433 0.731s 275.3MB
none at all 0 0 0.0s 1.503 0.685s 287.7MB
Debouncing keeps 86% of the reclamation for 6% of the collections: 1.9x the
throughput, half the CPU per check, gc down from 30.8% to 6.4% of wall time,
and a lower resident plateau than collecting every time. It works because the
collector is process-wide - any worker's collection breaks every other worker's
cycles too, so with many workers the calls are overwhelmingly redundant
duplicates rather than independently necessary.
Removing them entirely is slightly faster still, but it was the only
configuration whose RSS had not plateaued by the end of the run, so it is not
the default. EXPLICIT_GC_MIN_INTERVAL=0 restores the previous behaviour.
Collecting a younger generation was measured and rejected: gen 0 freed 1,136
objects against the full pass's 2,594,098, because objects surviving a 10-30s
fetch have already been promoted out of gen 0.
Two call sites are additionally fixed because they could never reclaim anything:
- Watch.py brotli: brotli.Compressor is not gc-tracked, so the collector cannot
see it - `del` frees it by refcount. Over 60 x 2.2MB compressions, RSS growth
was +0.9MB with neither mechanism, +0.2MB with gc.collect() alone, and +0.0MB
with malloc_trim() alone or with both. malloc_trim is the load-bearing line
and is kept; the collect cost ~31ms of stop-the-world per snapshot save for no
reclamation.
- puppeteer quit(): runs twice per check (run()'s finally, then the worker's
safety net) and nulls self.page/self.browser in its own finally blocks, so the
second call closes nothing and breaks no cycles yet still paid for a full
collection - 88 calls across 51 checks. Now only collects when it actually
closed something.
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* UI - LLM section tidyup
* Rebuild translatiosn
* UI - Fixing language for prompt adjustement to be more clear
* UI - clarify field action
* UI - clarify layout for sections
* WIP
* test tweaks