Files
changedetection.io/changedetectionio
dgtlmoonandClaude Opus 5 d56ec682fe Puppeteer fetcher - Re-navigation cap must not mean "extract right now" (#4439)
BROWSER_CONTENT_READY_MAX_RESETS exists so a page that re-navigates in a loop cannot extend a
fetch forever. On hitting that cap the content-ready wait broke straight out into
Page.stopLoading and extraction, giving the document we end up on zero settle time - the exact
opposite of what the wait is for.

Measured against a page that hops every 500ms and then renders via JS 2s after the final load:

  before:   2.7s, 130 bytes of an intermediate hop, no final document, no JS-rendered content
  after:   14.1s, final document, JS-rendered content present

It also logged "Content-ready wait of 12s elapsed" immediately before extracting, having waited
0s, which is why this reads as a fetcher that ignores the setting.

Note 0.60.4 could not do this: its wait was an unconditional `await asyncio.sleep(1 + extra_wait)`
after goto(), so every fetch got its settle time no matter how the page behaved.

Now the cap stops the wait from being *restarted*, and the delay is spent one final time before
extracting. Total stays bounded at (max_resets + 2) * extra_wait, and whatever we extract has had
the same settle time every other fetch gets. The stopLoading log line no longer claims a wait
that may not have happened.

Unrelated to #4437 - found while reading #4426 for that investigation, which turned out to be a
locale/collation bug in the filter layer, not a fetcher problem.

Tested: new reset-cap probe checked to fail before and pass after; normal (non-re-navigating)
path unchanged at 12.5s with content intact; 3 passed browser fetcher suite on pyppeteer, 323
unit tests pass.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 20:31:48 +02:00
..
…
2026-09-14 14:49:55 +02:00
2026-09-14 14:49:55 +02:00