Files
changedetection.io/changedetectionio/html_tools.py
T
dgtlmoonandClaude Opus 5 7102544310 xPath filters - Stop LC_COLLATE deciding what contains() means (#4437) (#4438)
Every watch whose xPath filter used contains() started reporting "Warning, no filters were
found" on pages whose HTML plainly contained the target. 18 unrelated watches broke in the same
hour, browser and plain-requests fetchers alike, and the saved snapshot was perfect every time.

elementpath implements the XPath string functions on top of locale.strxfrm:

    def contains(self, a, b):  return self.strxfrm(b) in self.strxfrm(a)

Under LC_COLLATE=C, strxfrm() is the identity function and that is an ordinary substring test.
Under a real locale it returns a binary collation key, and a substring of a collation key is not
the collation key of the substring - so contains(), starts-with(), ends-with() and
substring-before/after() return false for EVERY input. Reduced to one line, no document needed:

    LC_COLLATE=C            contains("xx month xx", "month") -> True
    LC_COLLATE=en_US.UTF-8  contains("xx month xx", "month") -> False

Name tests, axes and '=' are untouched, which is exactly why it read as "did the page change
layout?" - //div kept working while //div[contains(.,"month")] returned nothing.

No code change caused this. Generating the image's locales (#4429) made ENV LC_ALL=en_US.UTF-8
satisfiable for the first time; flask_app's setlocale(LC_ALL, ...) had been raising locale.Error
and leaving us in C, and once it succeeded it took LC_COLLATE with it. That is why the bug
survived a bisect to 0.60.2 and why reinstalling the exact pip set from a working install did not
shift it - it could only be found by diffing the two containers. Same image base, same Python
3.11.16, lxml 6.1.3, libxml2 2.14.6, elementpath 5.1.1, same HTML:

    good-old 0.60.4   setlocale FAILED: unsupported locale setting   67 matches
    bad-new  0.60.5   setlocale en_US.UTF-8                           0 matches

Fixed in both places, because either alone leaves a hole:

 - flask_app sets LC_CTYPE/LC_NUMERIC/LC_MONETARY/LC_TIME individually instead of LC_ALL. This
   block exists to make prices render correctly and it still does - 1234567 is still "1,234,567"
   - it just no longer touches collation.

 - html_tools.xpath_filter() pins the Unicode codepoint collation per evaluation, so a filter
   means the same thing whatever an operator puts in LANG/LC_ALL, and does not depend on a
   distant module's locale bookkeeping. forms.py's XPath validation pins it too, so validation
   cannot accept an expression that then behaves differently at check time.

Per XPath 3.1 the default collation is codepoint and must not consult LC_COLLATE, so the
underlying behaviour is arguably an elementpath bug; the pin above holds regardless.

Verified end to end inside the failing container: LC_COLLATE=C, LC_NUMERIC=en_US.UTF-8,
thousands separators intact, and the reported filter back from 0 to 1190797 chars of output.

Tested: new unit test covers contains/starts-with/ends-with under a UTF-8 collation and was
checked to fail without the fix; 325 unit tests pass.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 20:33:35 +02:00

973 lines
43 KiB
Python

from functools import lru_cache, wraps
from loguru import logger
from typing import List
import html
import json
import os
import re
import threading
# ---------------------------------------------------------------------------------------
# libxml2 (lxml) concurrency guard
#
# Watch checks run run_changedetection() on a shared ThreadPoolExecutor
# (worker_pool.queue_executor), so extraction for different watches parses HTML in parallel
# threads. A single check uses TWO different lxml parsers: xpath_filter() builds its own
# etree.HTMLParser(), while html_to_text() -> inscriptis parses with lxml's process-global
# html_parser. lxml locks per parser OBJECT, so many threads on the SAME parser serialise
# safely - but that lock does not span two DIFFERENT parsers, and they still share libxml2's
# interned-string dictionary. Concurrent parses corrupt it, and the rendered text of one
# watch's page ends up inside another watch's snapshot.
#
# lxml's FAQ sanctions exactly two patterns: "the default parser (which is replicated for
# each thread) or create a parser for each thread yourself". A parser per CALL is neither -
# but note that a parser per THREAD does not fix this either (measured: still leaks), because
# inscriptis picks its own parser internally and we cannot redirect it. Serialising is the
# only configuration measured clean.
#
# Serialising is measured to be free: extraction is ~80ms mean / 204ms p95 over real data,
# a ceiling of ~12 checks/sec against the ~0.5 checks/sec that browser fetches can feed.
# It also HALVES peak RSS on large pages (481MB -> 232MB on a 3.85MB page across 7 threads),
# because only one libxml2 document is ever live.
#
# EVERY lxml entry point in this process must go through here - a single unguarded parse in
# any thread (Flask request, worker) is enough to reintroduce the corruption.
#
# This is a threading lock, not an asyncio one, and that is deliberate: nothing here is called
# from a coroutine. The async workers hand run_changedetection() to a ThreadPoolExecutor
# (worker.py: "Run change detection in executor to avoid blocking event loop"), the requests
# fetcher does the same, and Flask is synchronous. If you ever call a guarded function directly
# from a coroutine you will stall that worker's event loop for the duration of someone else's
# parse - bounded and deadlock-free (the holder is pure CPU and never awaits), but avoid it.
#
# Re-entrant: these helpers call into each other (element_removal -> subtractive_xpath_selector).
#
# LXML_LOCK_DISABLED=true removes the guard. It exists so the corruption can be reproduced on
# demand (see tests/test_lxml_concurrency.py) - it is NOT a tuning option, it reinstates a
# data-integrity bug that silently writes one watch's page text into another watch's snapshot.
def _build_lxml_lock():
from changedetectionio.strtobool import strtobool
if not strtobool(os.getenv('LXML_LOCK_DISABLED', 'false')):
return threading.RLock()
logger.warning("LXML_LOCK_DISABLED=true - lxml parsing is unguarded. Under concurrency this "
"is known to leak one watch's rendered page text into another watch's "
"snapshot. For testing only, never run this in production.")
class _NullLock:
def __enter__(self): return self
def __exit__(self, *a): return False
return _NullLock()
_LXML_LOCK = _build_lxml_lock()
def lxml_guarded(fn):
"""Serialise a function's libxml2 work - parse, xpath traversal, tostring() and clear()
all touch the document, so the whole call is held, not just the parse."""
@wraps(fn)
def wrapper(*args, **kwargs):
with _LXML_LOCK:
return fn(*args, **kwargs)
return wrapper
def lxml_guard():
"""Context manager for code OUTSIDE this module that drives lxml directly, so it shares the
same lock (e.g. XPath validation in forms.py, which runs on Flask request threads)."""
return _LXML_LOCK
def lxml_html_parser():
"""A fresh parser for each call, freed immediately afterwards. Never lxml's process-global
default parser, which is shared across every thread."""
from lxml import etree
return etree.HTMLParser()
# ---------------------------------------------------------------------------------------
# HTML added to be sure each result matching a filter (.example) gets converted to a new line by Inscriptis
TEXT_FILTER_LIST_LINE_SUFFIX = "<br>"
TRANSLATE_WHITESPACE_TABLE = str.maketrans('', '', '\r\n\t ')
PERL_STYLE_REGEX = r'^/(.*?)/([a-z]*)?$'
TITLE_RE = re.compile(r"<title[^>]*>(.*?)</title>", re.I | re.S)
META_CS = re.compile(r'<meta[^>]+charset=["\']?\s*([a-z0-9_\-:+.]+)', re.I)
# jq builtins that can leak sensitive data or cause harm when user-supplied expressions are executed.
# env/$ENV reads all process environment variables (passwords, API keys, etc.)
# include/import can read arbitrary files from disk
# input/inputs reads beyond the supplied JSON data
# debug/stderr leaks data to stderr
# halt/halt_error terminates the process (DoS)
_JQ_BLOCKED_PATTERNS = [
(re.compile(r'\benv\b'), 'env (reads environment variables)'),
(re.compile(r'\$ENV\b'), '$ENV (reads environment variables)'),
(re.compile(r'\binclude\b'), 'include (reads files from disk)'),
(re.compile(r'\bimport\b'), 'import (reads files from disk)'),
(re.compile(r'\binputs?\b'), 'input/inputs (reads beyond provided data)'),
(re.compile(r'\bdebug\b'), 'debug (leaks data to stderr)'),
(re.compile(r'\bstderr\b'), 'stderr (leaks data to stderr)'),
(re.compile(r'\bhalt(?:_error)?\b'), 'halt/halt_error (terminates the process)'),
(re.compile(r'\$__loc__\b'), '$__loc__ (leaks file path information)'),
(re.compile(r'\bbuiltins\b'), 'builtins (enumerates available functions)'),
(re.compile(r'\bmodulemeta\b'), 'modulemeta (leaks module information)'),
(re.compile(r'\$JQ_BUILD_CONFIGURATION\b'), '$JQ_BUILD_CONFIGURATION (leaks build information)'),
]
def validate_jq_expression(expression: str) -> None:
"""Raise ValueError if the jq expression uses any dangerous builtin.
User-supplied jq expressions are executed server-side. Without this check,
builtins like `env` expose every process environment variable (SALTED_PASS,
proxy credentials, API keys, etc.) as watch output.
"""
from changedetectionio.strtobool import strtobool
if strtobool(os.getenv('JQ_ALLOW_RISKY_EXPRESSIONS', 'false')):
return
for pattern, description in _JQ_BLOCKED_PATTERNS:
if pattern.search(expression):
msg = f"jq expression uses disallowed builtin: {description}"
logger.critical(f"Security: blocked jq expression containing '{description}' - expression: {expression!r}")
raise ValueError(msg)
META_CT = re.compile(r'<meta[^>]+http-equiv=["\']?content-type["\']?[^>]*content=["\'][^>]*charset=([a-z0-9_\-:+.]+)', re.I)
# 'price' , 'lowPrice', 'highPrice' are usually under here
# All of those may or may not appear on different websites - I didnt find a way todo case-insensitive searching here
LD_JSON_PRODUCT_OFFER_SELECTORS = ["json:$..offers", "json:$..Offers"]
class JSONNotFound(ValueError):
def __init__(self, msg):
ValueError.__init__(self, msg)
_DEFAULT_UNSAFE_XPATH3_FUNCTIONS = [
'unparsed-text',
'unparsed-text-lines',
'unparsed-text-available',
'doc',
'doc-available',
'json-doc',
'json-doc-available',
'collection', # XPath 2.0+: loads XML node collections from arbitrary URIs
'uri-collection', # XPath 3.0+: enumerates URIs from resource collections
'transform', # XPath 3.1: XSLT transformation (currently raises, block proactively)
'load-xquery-module', # XPath 3.1: loads XQuery modules (currently raises, block proactively)
'environment-variable',
'available-environment-variables',
]
# XPath 3.1 says the default collation is Unicode codepoint collation. elementpath instead leaves
# its collation functions pointing at locale.strxfrm / locale.strcoll, so the process-wide
# LC_COLLATE decides what contains() means:
#
# def contains(self, a, b): return self.strxfrm(b) in self.strxfrm(a)
#
# With LC_COLLATE=C, strxfrm() is the identity and that is an ordinary substring test. With
# LC_COLLATE=en_US.UTF-8 it returns a binary collation key, and a substring of a collation key is
# not the collation key of the substring - so contains(), starts-with(), ends-with() and
# substring-before/after() return false for every input, and a filter that matches 67 elements
# matches 0 (#4437). Name tests, axes and '=' are unaffected, which is what made it look like the
# page had changed layout.
#
# flask_app.py keeps LC_COLLATE in "C" for this reason, but a filter must not depend on a distant
# module's locale bookkeeping, nor on what an operator puts in LANG/LC_ALL. Pinning the collation
# per evaluation makes the filter mean the same thing in every deployment.
XPATH_CODEPOINT_COLLATION = 'http://www.w3.org/2005/xpath-functions/collation/codepoint'
def get_safe_xpath3_parser():
"""Return an XPath3Parser subclass with filesystem/environment access functions removed.
XPath 3.0 includes functions that can read arbitrary files or environment variables:
- unparsed-text / unparsed-text-lines / unparsed-text-available (file read)
- doc / doc-available (XML fetch from URI)
- environment-variable / available-environment-variables (env var leakage)
Subclassing gives us an independent symbol_table copy (not shared with the parent class),
so removing entries here does not affect XPath3Parser itself.
Override the blocked list via the XPATH_BLOCKED_FUNCTIONS environment variable
(comma-separated, e.g. "unparsed-text,doc,environment-variable").
"""
import os
from elementpath.xpath3 import XPath3Parser
class SafeXPath3Parser(XPath3Parser):
pass
env_override = os.getenv('XPATH_BLOCKED_FUNCTIONS')
if env_override is not None:
blocked = [f.strip() for f in env_override.split(',') if f.strip()]
else:
blocked = _DEFAULT_UNSAFE_XPATH3_FUNCTIONS
for _fn in blocked:
SafeXPath3Parser.symbol_table.pop(_fn, None)
return SafeXPath3Parser
# Doesn't look like python supports forward slash auto enclosure in re.findall
# So convert it to inline flag "(?i)foobar" type configuration
@lru_cache(maxsize=100)
def perl_style_slash_enclosed_regex_to_options(regex):
res = re.search(PERL_STYLE_REGEX, regex, re.IGNORECASE)
if res:
flags = res.group(2) if res.group(2) else 'i'
regex = f"(?{flags}){res.group(1)}"
else:
# Fall back to just ignorecase as an option
regex = f"(?i){regex}"
return regex
# Given a CSS Rule, and a blob of HTML, return the blob of HTML that matches
def include_filters(include_filters, html_content, append_pretty_line_formatting=False):
from bs4 import BeautifulSoup
soup = BeautifulSoup(html_content, "html.parser")
html_block = ""
r = soup.select(include_filters, separator="")
for element in r:
# When there's more than 1 match, then add the suffix to separate each line
# And where the matched result doesn't include something that will cause Inscriptis to add a newline
# (This way each 'match' reliably has a new-line in the diff)
# Divs are converted to 4 whitespaces by inscriptis
if append_pretty_line_formatting and len(html_block) and not element.name in (['br', 'hr', 'div', 'p']):
html_block += TEXT_FILTER_LIST_LINE_SUFFIX
html_block += str(element)
return html_block
def subtractive_css_selector(css_selector, content):
from bs4 import BeautifulSoup
soup = BeautifulSoup(content, "html.parser")
# So that the elements dont shift their index, build a list of elements here which will be pointers to their place in the DOM
elements_to_remove = soup.select(css_selector)
if not elements_to_remove:
# Better to return the original that rebuild with BeautifulSoup
return content
# Then, remove them in a separate loop
for item in elements_to_remove:
item.decompose()
return str(soup)
@lxml_guarded
def subtractive_xpath_selector(selectors: List[str], html_content: str) -> str:
from lxml import etree
# Parse the HTML content using lxml. Own parser, not etree.HTML()'s process-global default.
html_tree = etree.fromstring(html_content, parser=lxml_html_parser())
# First, collect all elements to remove
elements_to_remove = []
# Iterate over the list of XPath selectors
for selector in selectors:
# Collect elements for each selector
elements_to_remove.extend(html_tree.xpath(selector))
# If no elements were found, return the original HTML content
if not elements_to_remove:
return html_content
# Then, remove them in a separate loop
for element in elements_to_remove:
if element.getparent() is not None: # Ensure the element has a parent before removing
element.getparent().remove(element)
# Convert the modified HTML tree back to a string
modified_html = etree.tostring(html_tree, method="html").decode("utf-8")
return modified_html
def element_removal(selectors: List[str], html_content):
"""Removes elements that match a list of CSS or XPath selectors."""
modified_html = html_content
css_selectors = []
xpath_selectors = []
for selector in selectors:
if selector.strip().startswith(('xpath:', 'xpath1:', '//')):
# Handle XPath selectors separately
xpath_selector = selector.removeprefix('xpath:').removeprefix('xpath1:')
xpath_selectors.append(xpath_selector)
else:
# Collect CSS selectors as one "hit", see comment in subtractive_css_selector
css_selectors.append(selector.strip().strip(","))
if xpath_selectors:
modified_html = subtractive_xpath_selector(xpath_selectors, modified_html)
if css_selectors:
# Remove duplicates, then combine all CSS selectors into one string, separated by commas
# This stops the elements index shifting
unique_selectors = list(set(css_selectors)) # Ensure uniqueness
combined_css_selector = " , ".join(unique_selectors)
modified_html = subtractive_css_selector(combined_css_selector, modified_html)
return modified_html
def elementpath_tostring(obj):
"""
change elementpath.select results to string type
# The MIT License (MIT), Copyright (c), 2018-2021, SISSA (Scuola Internazionale Superiore di Studi Avanzati)
# https://github.com/sissaschool/elementpath/blob/dfcc2fd3d6011b16e02bf30459a7924f547b47d0/elementpath/xpath_tokens.py#L1038
"""
import elementpath
from decimal import Decimal
import math
if obj is None:
return ''
# https://elementpath.readthedocs.io/en/latest/xpath_api.html#elementpath.select
elif isinstance(obj, elementpath.XPathNode):
return obj.string_value
elif isinstance(obj, bool):
return 'true' if obj else 'false'
elif isinstance(obj, Decimal):
value = format(obj, 'f')
if '.' in value:
return value.rstrip('0').rstrip('.')
return value
elif isinstance(obj, float):
if math.isnan(obj):
return 'NaN'
elif math.isinf(obj):
return str(obj).upper()
value = str(obj)
if '.' in value:
value = value.rstrip('0').rstrip('.')
if '+' in value:
value = value.replace('+', '')
if 'e' in value:
return value.upper()
return value
return str(obj)
# Return str Utf-8 of matched rules
@lxml_guarded
def xpath_filter(xpath_filter, html_content, append_pretty_line_formatting=False, is_xml=False):
"""
:param xpath_filter:
:param html_content:
:param append_pretty_line_formatting:
:param is_xml: set to true if is XML or is RSS (RSS is XML)
:return:
"""
from lxml import etree, html
import elementpath
parser = lxml_html_parser()
tree = None
try:
if is_xml:
# So that we can keep CDATA for cdata_in_document_to_text() to process
parser = etree.XMLParser(strip_cdata=False, resolve_entities=False, no_network=True)
# For XML/RSS content, use etree.fromstring to properly handle XML declarations
tree = etree.fromstring(html_content.encode('utf-8') if isinstance(html_content, str) else html_content, parser=parser)
else:
tree = html.fromstring(html_content, parser=parser)
html_block = ""
# Build namespace map for XPath queries
namespaces = {'re': 'http://exslt.org/regular-expressions'}
# Handle default namespace in documents (common in RSS/Atom feeds, but can occur in any XML)
# XPath spec: unprefixed element names have no namespace, not the default namespace
# Solution: Register the default namespace with empty string prefix in elementpath
# This is primarily for RSS/Atom feeds but works for any XML with default namespace
if hasattr(tree, 'nsmap') and tree.nsmap and None in tree.nsmap:
# Register the default namespace with empty string prefix for elementpath
# This allows //title to match elements in the default namespace
namespaces[''] = tree.nsmap[None]
r = elementpath.select(tree, xpath_filter.strip(), namespaces=namespaces,
parser=get_safe_xpath3_parser(),
default_collation=XPATH_CODEPOINT_COLLATION)
#@note: //title/text() now works with default namespaces (fixed by registering '' prefix)
#@note: //title/text() wont work where <title>CDATA.. (use cdata_in_document_to_text first)
if type(r) != list:
r = [r]
for element in r:
# When there's more than 1 match, then add the suffix to separate each line
# And where the matched result doesn't include something that will cause Inscriptis to add a newline
# (This way each 'match' reliably has a new-line in the diff)
# Divs are converted to 4 whitespaces by inscriptis
if append_pretty_line_formatting and len(html_block) and (not hasattr( element, 'tag' ) or not element.tag in (['br', 'hr', 'div', 'p'])):
html_block += TEXT_FILTER_LIST_LINE_SUFFIX
if type(element) == str:
html_block += element
elif issubclass(type(element), etree._Element) or issubclass(type(element), etree._ElementTree):
# Use 'xml' method for RSS/XML content, 'html' for HTML content
# parser will be XMLParser if we detected XML content
method = 'xml' if (is_xml or isinstance(parser, etree.XMLParser)) else 'html'
html_block += etree.tostring(element, pretty_print=True, method=method, encoding='unicode')
else:
html_block += elementpath_tostring(element)
# Drop element references before the finally block so tree.clear() can release
# the libxml2 document immediately (elements pin the C-level doc via refcount).
del r
return html_block
finally:
# Explicitly clear the tree to free memory
# lxml trees can hold significant memory, especially with large documents
if tree is not None:
tree.clear()
# Return str Utf-8 of matched rules
# 'xpath1:'
@lxml_guarded
def xpath1_filter(xpath_filter, html_content, append_pretty_line_formatting=False, is_xml=False):
from lxml import etree, html
# Own parser, not lxml's process-global default (which parser=None would select).
parser = lxml_html_parser()
tree = None
try:
if is_xml:
# So that we can keep CDATA for cdata_in_document_to_text() to process
parser = etree.XMLParser(strip_cdata=False, resolve_entities=False, no_network=True)
# For XML/RSS content, use etree.fromstring to properly handle XML declarations
tree = etree.fromstring(html_content.encode('utf-8') if isinstance(html_content, str) else html_content, parser=parser)
else:
tree = html.fromstring(html_content, parser=parser)
html_block = ""
# Build namespace map for XPath queries
namespaces = {'re': 'http://exslt.org/regular-expressions'}
# NOTE: lxml's native xpath() does NOT support empty string prefix for default namespace
# For documents with default namespace (RSS/Atom feeds), users must use:
# - local-name(): //*[local-name()='title']/text()
# - Or use xpath_filter (not xpath1_filter) which supports default namespaces
# XPath spec: unprefixed element names have no namespace, not the default namespace
r = tree.xpath(xpath_filter.strip(), namespaces=namespaces)
#@note: xpath1 (lxml) does NOT automatically handle default namespaces
#@note: Use //*[local-name()='element'] or switch to xpath_filter for default namespace support
#@note: //title/text() wont work where <title>CDATA.. (use cdata_in_document_to_text first)
for element in r:
# When there's more than 1 match, then add the suffix to separate each line
# And where the matched result doesn't include something that will cause Inscriptis to add a newline
# (This way each 'match' reliably has a new-line in the diff)
# Divs are converted to 4 whitespaces by inscriptis
if append_pretty_line_formatting and len(html_block) and (not hasattr(element, 'tag') or not element.tag in (['br', 'hr', 'div', 'p'])):
html_block += TEXT_FILTER_LIST_LINE_SUFFIX
# Some kind of text, UTF-8 or other
if isinstance(element, (str, bytes)):
html_block += element
else:
# Return the HTML/XML which will get parsed as text
# Use 'xml' method for RSS/XML content, 'html' for HTML content
# parser will be XMLParser if we detected XML content
method = 'xml' if (is_xml or isinstance(parser, etree.XMLParser)) else 'html'
html_block += etree.tostring(element, pretty_print=True, method=method, encoding='unicode')
return html_block
finally:
# Explicitly clear the tree to free memory
# lxml trees can hold significant memory, especially with large documents
if tree is not None:
tree.clear()
# Extract/find element
def extract_element(find='title', html_content=''):
from bs4 import BeautifulSoup
#Re #106, be sure to handle when its not found
element_text = None
soup = BeautifulSoup(html_content, 'html.parser')
result = soup.find(find)
if result and result.string:
element_text = result.string.strip()
return element_text
# Any surrogate still present after json.loads() is a lone one - the decoder already folds
# well-formed pairs (😀 -> a single emoji character) before we see them. Lone ones
# come from escapes like \uD800 in the source document, which travel as plain ASCII and so
# pass straight through the fetched-content sanitizer in processors/base.py.
_LONE_SURROGATE_RE = re.compile(r'[\ud800-\udfff]')
def _has_lone_surrogate(value):
if isinstance(value, str):
return bool(_LONE_SURROGATE_RE.search(value))
if isinstance(value, dict):
return any(_has_lone_surrogate(k) or _has_lone_surrogate(v) for k, v in value.items())
if isinstance(value, list):
return any(_has_lone_surrogate(v) for v in value)
return False
def _sanitize_lone_surrogate_str(value: str) -> str:
return _LONE_SURROGATE_RE.sub('\ufffd', value)
def _sanitize_lone_surrogates(value):
if isinstance(value, str):
return _sanitize_lone_surrogate_str(value)
if isinstance(value, dict):
# JSON object keys are always strings, so they only need the substitution - recursing on a
# key would (in theory) hand back an unhashable dict/list
return {_sanitize_lone_surrogate_str(k) if isinstance(k, str) else k: _sanitize_lone_surrogates(v)
for k, v in value.items()}
if isinstance(value, list):
return [_sanitize_lone_surrogates(v) for v in value]
return value
#
def _parse_json(json_data, json_filter):
from jsonpath_ng.ext import parse
# Replace lone surrogates with U+FFFD, matching what processors/base.py already does for
# fetched content. Two separate failures otherwise, neither of which is confined to the
# offending value: jq re-serializes the whole document for its own C parser, which rejects
# a lone high surrogate, so every jq:/jqraw: filter on the page errors even when it points
# at an unrelated field; and on the json: path the surrogate reaches the caller intact and
# raises UnicodeEncodeError later in checksums and history writes.
# See: https://github.com/dgtlmoon/changedetection.io/issues/4273
if _has_lone_surrogate(json_data):
logger.warning("Lone unicode surrogate(s) in JSON document, replacing with U+FFFD before filtering")
json_data = _sanitize_lone_surrogates(json_data)
if json_filter.startswith("json:"):
jsonpath_expression = parse(json_filter.replace('json:', ''))
match = jsonpath_expression.find(json_data)
return _get_stripped_text_from_json_match(match)
if json_filter.startswith("jq:") or json_filter.startswith("jqraw:"):
try:
import jq
except ModuleNotFoundError:
# `jq` requires full compilation in windows and so isn't generally available
raise Exception("jq not support not found")
if json_filter.startswith("jq:"):
expr = json_filter.removeprefix("jq:")
validate_jq_expression(expr)
jq_expression = jq.compile(expr)
match = jq_expression.input(json_data).all()
return _get_stripped_text_from_json_match(match)
if json_filter.startswith("jqraw:"):
expr = json_filter.removeprefix("jqraw:")
validate_jq_expression(expr)
jq_expression = jq.compile(expr)
match = jq_expression.input(json_data).all()
return '\n'.join(str(item) for item in match)
def _get_stripped_text_from_json_match(match):
s = []
# More than one result, we will return it as a JSON list.
if len(match) > 1:
for i in match:
s.append(i.value if hasattr(i, 'value') else i)
# Single value, use just the value, as it could be later used in a token in notifications.
if len(match) == 1:
s = match[0].value if hasattr(match[0], 'value') else match[0]
# Re #257 - Better handling where it does not exist, in the case the original 's' value was False..
if not match:
# Re 265 - Just return an empty string when filter not found
return ''
# Ticket #462 - allow the original encoding through, usually it's UTF-8 or similar
stripped_text_from_html = json.dumps(s, indent=4, ensure_ascii=False)
return stripped_text_from_html
def extract_json_blob_from_html(content, ensure_is_ldjson_info_type, json_filter):
from bs4 import BeautifulSoup
stripped_text_from_html = ''
# Foreach <script json></script> blob.. just return the first that matches json_filter
# As a last resort, try to parse the whole <body>
soup = BeautifulSoup(content, 'html.parser')
if ensure_is_ldjson_info_type:
bs_result = soup.find_all('script', {"type": "application/ld+json"})
else:
bs_result = soup.find_all('script')
bs_result += soup.find_all('body')
bs_jsons = []
for result in bs_result:
# result.text is how bs4 magically strips JSON from the body
content_start = result.text.lstrip("\ufeff").strip()[:100] if result.text else ''
# Skip empty tags, and things that dont even look like JSON
if not result.text or not (content_start[0] == '{' or content_start[0] == '['):
continue
try:
json_data = json.loads(result.text)
bs_jsons.append(json_data)
except json.JSONDecodeError:
# Skip objects which cannot be parsed
continue
if not bs_jsons:
raise JSONNotFound("No parsable JSON found in this document")
for json_data in bs_jsons:
stripped_text_from_html = _parse_json(json_data, json_filter)
if ensure_is_ldjson_info_type:
# Could sometimes be list, string or something else random
if isinstance(json_data, dict):
# If it has LD JSON 'key' @type, and @type is 'product', and something was found for the search
# (Some sites have multiple of the same ld+json @type='product', but some have the review part, some have the 'price' part)
# @type could also be a list although non-standard ("@type": ["Product", "SubType"],)
# LD_JSON auto-extract also requires some content PLUS the ldjson to be present
# 1833 - could be either str or dict, should not be anything else
t = json_data.get('@type')
if t and stripped_text_from_html:
if isinstance(t, str) and t.lower() == ensure_is_ldjson_info_type.lower():
break
# The non-standard part, some have a list
elif isinstance(t, list):
if ensure_is_ldjson_info_type.lower() in [x.lower().strip() for x in t]:
break
elif stripped_text_from_html:
break
return stripped_text_from_html
# content - json
# json_filter - ie json:$..price
# ensure_is_ldjson_info_type - str "product", optional, "@type == product" (I dont know how to do that as a json selector)
def extract_json_as_string(content, json_filter, ensure_is_ldjson_info_type=None):
stripped_text_from_html = False
# https://github.com/dgtlmoon/changedetection.io/pull/2041#issuecomment-1848397161w
# Try to parse/filter out the JSON, if we get some parser error, then maybe it's embedded within HTML tags
# Looks like clean JSON, dont bother extracting from HTML
content_start = content.lstrip("\ufeff").strip()[:100]
if content_start[0] == '{' or content_start[0] == '[':
try:
# .lstrip("\ufeff") strings ByteOrderMark from UTF8 and still lets the UTF work
stripped_text_from_html = _parse_json(json.loads(content.lstrip("\ufeff")), json_filter)
except json.JSONDecodeError as e:
logger.warning(f"Error processing JSON {content[:20]}...{str(e)})")
else:
# Check for JSONP wrapper: someCallback({...}) or some.namespace({...})
# Server may claim application/json but actually return JSONP
jsonp_match = re.match(r'^\w[\w.]*\s*\((.+)\)\s*;?\s*$', content.lstrip("\ufeff").strip(), re.DOTALL)
if jsonp_match:
try:
inner = jsonp_match.group(1).strip()
logger.warning(f"Content looks like JSONP, attempting to extract inner JSON for filter '{json_filter}'")
stripped_text_from_html = _parse_json(json.loads(inner), json_filter)
except json.JSONDecodeError as e:
logger.warning(f"Error processing JSONP inner content {content[:20]}...{str(e)})")
if not stripped_text_from_html:
# Probably something else, go fish inside for it
try:
stripped_text_from_html = extract_json_blob_from_html(content=content,
ensure_is_ldjson_info_type=ensure_is_ldjson_info_type,
json_filter=json_filter)
except json.JSONDecodeError as e:
logger.warning(f"Error processing JSON while extracting JSON from HTML blob {content[:20]}...{str(e)})")
if not stripped_text_from_html:
# Re 265 - Just return an empty string when filter not found
return ''
return stripped_text_from_html
# Mode - "content" return the content without the matches (default)
# - "line numbers" return a list of line numbers that match (int list)
#
# wordlist - list of regex's (str) or words (str)
# Preserves all linefeeds and other whitespacing, its not the job of this to remove that
def strip_ignore_text(content, wordlist, mode="content"):
ignore_text = []
ignore_regex = []
ignore_regex_multiline = []
ignored_lines = []
if not content:
return ''
for k in wordlist:
# Skip empty strings to avoid matching everything
if not k or not k.strip():
continue
# Is it a regex?
res = re.search(PERL_STYLE_REGEX, k, re.IGNORECASE)
if res:
res = re.compile(perl_style_slash_enclosed_regex_to_options(k))
if res.flags & re.DOTALL or res.flags & re.MULTILINE:
ignore_regex_multiline.append(res)
else:
ignore_regex.append(res)
else:
ignore_text.append(k.strip())
for r in ignore_regex_multiline:
for match in r.finditer(content):
content_lines = content[:match.end()].splitlines(keepends=True)
match_lines = content[match.start():match.end()].splitlines(keepends=True)
end_line = len(content_lines)
start_line = end_line - len(match_lines)
if end_line - start_line <= 1:
# Match is empty or in the middle of the line
ignored_lines.append(start_line)
else:
for i in range(start_line, end_line):
ignored_lines.append(i)
line_index = 0
lines = content.splitlines(keepends=True)
for line in lines:
# Always ignore blank lines in this mode. (when this function gets called)
got_match = False
for l in ignore_text:
if l.lower() in line.lower():
got_match = True
if not got_match:
for r in ignore_regex:
if r.search(line):
got_match = True
if got_match:
ignored_lines.append(line_index)
line_index += 1
ignored_lines = set([i for i in ignored_lines if i >= 0 and i < len(lines)])
# Used for finding out what to highlight
if mode == "line numbers":
return [i + 1 for i in ignored_lines]
output_lines = set(range(len(lines))) - ignored_lines
return ''.join([lines[i] for i in output_lines])
def cdata_in_document_to_text(html_content: str, render_anchor_tag_content=False) -> str:
from xml.sax.saxutils import escape as xml_escape
pattern = '<!\[CDATA\[(\s*(?:.(?<!\]\]>)\s*)*)\]\]>'
def repl(m):
text = m.group(1)
return xml_escape(html_to_text(html_content=text)).strip()
return re.sub(pattern, repl, html_content)
# NOTE!! ANYTHING LIBXML, HTML5LIB ETC WILL CAUSE SOME SMALL MEMORY LEAK IN THE LOCAL "LIB" IMPLEMENTATION OUTSIDE PYTHON
def html_to_text(html_content: str, render_anchor_tag_content=False, is_rss=False, timeout=10) -> str:
"""
Convert HTML content to plain text using inscriptis.
Thread-Safety: inscriptis.get_text() parses with lxml's process-global default parser.
That is fine on its own, but NOT alongside the xpath helpers' own parser - see the
_LXML_LOCK comment at the top of this module. So the get_text() call is guarded.
Do NOT "fix" this by giving get_text() its own parser instead. That was tried
(bee1130c6, reverted by 272e68ad2) and measured at ~10x WORSE cross-document leakage,
because then both sides use unlocked per-call parsers and lxml's per-parser lock stops
helping at all.
"""
from inscriptis import get_text
from inscriptis.model.config import ParserConfig
if render_anchor_tag_content:
parser_config = ParserConfig(
annotation_rules={"a": ["hyperlink"]},
display_links=True
)
else:
parser_config = None
if is_rss:
html_content = re.sub(r'<title([\s>])', r'<h1\1', html_content)
html_content = re.sub(r'</title>', r'</h1>', html_content)
else:
# Use BS4 html.parser to strip bloat — SPA's often dump 10MB+ of CSS/JS into <head>,
# causing inscriptis to silently give up. Regex-based stripping is unsafe because tags
# can appear inside JSON data attributes with JS-escaped closing tags (e.g. <\/script>),
# causing the regex to scan past the intended close and eat real page content.
from bs4 import BeautifulSoup
soup = BeautifulSoup(html_content, 'html.parser')
# Strip tags that inscriptis cannot render as meaningful text and which can be very large.
# svg/math: produce path-data/MathML garbage; canvas/iframe/template: no inscriptis handlers.
# video/audio/picture are kept — they may contain meaningful fallback text or captions.
for tag in soup.find_all(['head', 'script', 'style', 'noscript', 'svg',
'math', 'canvas', 'iframe', 'template']):
tag.decompose()
# SPAs often use <body style="display:none"> to hide content until JS loads.
# inscriptis respects CSS display rules, so strip hiding styles from the body tag.
body_tag = soup.find('body')
if body_tag and body_tag.get('style'):
style = body_tag['style']
if re.search(r'\b(?:display\s*:\s*none|visibility\s*:\s*hidden)\b', style, re.IGNORECASE):
logger.debug(f"html_to_text: Removing hiding styles from body tag (found: '{style}')")
del body_tag['style']
html_content = str(soup)
# Only the lxml part is guarded - the BeautifulSoup stripping above is pure Python.
with _LXML_LOCK:
text_content = get_text(html_content, config=parser_config)
return text_content
# Does LD+JSON exist with a @type=='product' and a .price set anywhere?
def has_ldjson_product_info(content):
try:
# Better than .lower() which can use a lot of ram
if (re.search(r'application/ld\+json', content, re.IGNORECASE) and
re.search(r'"price"', content, re.IGNORECASE) and
re.search(r'"pricecurrency"', content, re.IGNORECASE)):
return True
# On some pages this is really terribly expensive when they dont really need it
# (For example you never want price monitoring, but this runs on every watch to suggest it)
# for filter in LD_JSON_PRODUCT_OFFER_SELECTORS:
# pricing_data += extract_json_as_string(content=content,
# json_filter=filter,
# ensure_is_ldjson_info_type="product")
except Exception as e:
# OK too
return False
return False
def workarounds_for_obfuscations(content):
"""
Some sites are using sneaky tactics to make prices and other information un-renderable by Inscriptis
This could go into its own Pip package in the future, for faster updates
"""
# HomeDepot.com style <span>$<!-- -->90<!-- -->.<!-- -->74</span>
# https://github.com/weblyzard/inscriptis/issues/45
if not content:
return content
content = re.sub('<!--\s+-->', '', content)
return content
def get_triggered_text(content, trigger_text):
triggered_text = []
result = strip_ignore_text(content=content,
wordlist=trigger_text,
mode="line numbers")
i = 1
for p in content.splitlines():
if i in result:
triggered_text.append(p)
i += 1
return triggered_text
def extract_title(data: bytes | str, sniff_bytes: int = 2048, scan_chars: int = 8192) -> str | None:
"""Extract the <title> from an HTML document.
Rather than decoding/scanning a fixed prefix of the whole document, we first
locate the raw ``<title`` marker and then decode only a small window around
it. This handles pages (e.g. Amazon) where large ``<head>`` sections push
the title tag well past the old 8 192-character scan limit.
"""
# Maximum bytes/chars to extract after (and including) the opening <title tag.
# The regex needs to see </title>, so the window must cover the full content.
# The return value is always capped at 2 000 chars; titles beyond that are
# rare but possible. We read up to 128 KiB from the tag onwards to handle
# even pathological cases without scanning the whole document.
_TITLE_WINDOW = 131072
try:
match data:
case bytes() if data.startswith((b"\xff\xfe\x00\x00", b"\x00\x00\xfe\xff")):
# UTF-32: locate the tag in the raw bytes, then decode the window.
tag_pos = data.lower().find(b"<\x00\x00\x00t\x00\x00\x00")
if tag_pos == -1:
return None
chunk = data[tag_pos: tag_pos + _TITLE_WINDOW * 4].decode("utf-32", errors="replace")
prefix = chunk
case bytes() if data.startswith((b"\xff\xfe", b"\xfe\xff")):
# UTF-16: simple byte-pair search is tricky; fall back to decoding
# a reasonable head chunk and let the regex do the rest.
prefix = data[: max(scan_chars * 2, _TITLE_WINDOW)].decode("utf-16", errors="replace")
case bytes():
# UTF-8 / legacy 8-bit: find the tag cheaply in raw bytes.
tag_pos = data.lower().find(b"<title")
if tag_pos == -1:
return None
raw_chunk = data[tag_pos: tag_pos + _TITLE_WINDOW]
try:
chunk = raw_chunk.decode("utf-8")
except UnicodeDecodeError:
try:
head = data[:sniff_bytes].decode("ascii", errors="ignore")
if m := (META_CS.search(head) or META_CT.search(head)):
enc = m.group(1).lower()
else:
enc = "cp1252"
chunk = raw_chunk.decode(enc, errors="replace")
except Exception as e:
logger.error(f"Title extraction encoding detection failed: {e}")
return None
prefix = chunk
case str():
tag_pos = data.lower().find("<title")
if tag_pos == -1:
return None
prefix = data[tag_pos: tag_pos + _TITLE_WINDOW]
case _:
logger.error(f"Title extraction received unsupported data type: {type(data)}")
return None
# Search only in the (now tag-anchored) prefix
if m := TITLE_RE.search(prefix):
title = html.unescape(" ".join(m.group(1).split())).strip()
# Some safe limit
return title[:2000]
return None
except Exception as e:
logger.error(f"Title extraction failed: {e}")
return None