mirror of
https://github.com/dgtlmoon/changedetection.io.git
synced 2026-09-28 08:16:03 +00:00
Every watch whose xPath filter used contains() started reporting "Warning, no filters were
found" on pages whose HTML plainly contained the target. 18 unrelated watches broke in the same
hour, browser and plain-requests fetchers alike, and the saved snapshot was perfect every time.
elementpath implements the XPath string functions on top of locale.strxfrm:
def contains(self, a, b): return self.strxfrm(b) in self.strxfrm(a)
Under LC_COLLATE=C, strxfrm() is the identity function and that is an ordinary substring test.
Under a real locale it returns a binary collation key, and a substring of a collation key is not
the collation key of the substring - so contains(), starts-with(), ends-with() and
substring-before/after() return false for EVERY input. Reduced to one line, no document needed:
LC_COLLATE=C contains("xx month xx", "month") -> True
LC_COLLATE=en_US.UTF-8 contains("xx month xx", "month") -> False
Name tests, axes and '=' are untouched, which is exactly why it read as "did the page change
layout?" - //div kept working while //div[contains(.,"month")] returned nothing.
No code change caused this. Generating the image's locales (#4429) made ENV LC_ALL=en_US.UTF-8
satisfiable for the first time; flask_app's setlocale(LC_ALL, ...) had been raising locale.Error
and leaving us in C, and once it succeeded it took LC_COLLATE with it. That is why the bug
survived a bisect to 0.60.2 and why reinstalling the exact pip set from a working install did not
shift it - it could only be found by diffing the two containers. Same image base, same Python
3.11.16, lxml 6.1.3, libxml2 2.14.6, elementpath 5.1.1, same HTML:
good-old 0.60.4 setlocale FAILED: unsupported locale setting 67 matches
bad-new 0.60.5 setlocale en_US.UTF-8 0 matches
Fixed in both places, because either alone leaves a hole:
- flask_app sets LC_CTYPE/LC_NUMERIC/LC_MONETARY/LC_TIME individually instead of LC_ALL. This
block exists to make prices render correctly and it still does - 1234567 is still "1,234,567"
- it just no longer touches collation.
- html_tools.xpath_filter() pins the Unicode codepoint collation per evaluation, so a filter
means the same thing whatever an operator puts in LANG/LC_ALL, and does not depend on a
distant module's locale bookkeeping. forms.py's XPath validation pins it too, so validation
cannot accept an expression that then behaves differently at check time.
Per XPath 3.1 the default collation is codepoint and must not consult LC_COLLATE, so the
underlying behaviour is arguably an elementpath bug; the pin above holds regardless.
Verified end to end inside the failing container: LC_COLLATE=C, LC_NUMERIC=en_US.UTF-8,
thousands separators intact, and the reported filter back from 0 to 1190797 chars of output.
Tested: new unit test covers contains/starts-with/ends-with under a UTF-8 collation and was
checked to fail without the fix; 325 unit tests pass.
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
973 lines
43 KiB
Python
973 lines
43 KiB
Python
from functools import lru_cache, wraps
|
|
|
|
from loguru import logger
|
|
from typing import List
|
|
import html
|
|
import json
|
|
import os
|
|
import re
|
|
import threading
|
|
|
|
# ---------------------------------------------------------------------------------------
|
|
# libxml2 (lxml) concurrency guard
|
|
#
|
|
# Watch checks run run_changedetection() on a shared ThreadPoolExecutor
|
|
# (worker_pool.queue_executor), so extraction for different watches parses HTML in parallel
|
|
# threads. A single check uses TWO different lxml parsers: xpath_filter() builds its own
|
|
# etree.HTMLParser(), while html_to_text() -> inscriptis parses with lxml's process-global
|
|
# html_parser. lxml locks per parser OBJECT, so many threads on the SAME parser serialise
|
|
# safely - but that lock does not span two DIFFERENT parsers, and they still share libxml2's
|
|
# interned-string dictionary. Concurrent parses corrupt it, and the rendered text of one
|
|
# watch's page ends up inside another watch's snapshot.
|
|
#
|
|
# lxml's FAQ sanctions exactly two patterns: "the default parser (which is replicated for
|
|
# each thread) or create a parser for each thread yourself". A parser per CALL is neither -
|
|
# but note that a parser per THREAD does not fix this either (measured: still leaks), because
|
|
# inscriptis picks its own parser internally and we cannot redirect it. Serialising is the
|
|
# only configuration measured clean.
|
|
#
|
|
# Serialising is measured to be free: extraction is ~80ms mean / 204ms p95 over real data,
|
|
# a ceiling of ~12 checks/sec against the ~0.5 checks/sec that browser fetches can feed.
|
|
# It also HALVES peak RSS on large pages (481MB -> 232MB on a 3.85MB page across 7 threads),
|
|
# because only one libxml2 document is ever live.
|
|
#
|
|
# EVERY lxml entry point in this process must go through here - a single unguarded parse in
|
|
# any thread (Flask request, worker) is enough to reintroduce the corruption.
|
|
#
|
|
# This is a threading lock, not an asyncio one, and that is deliberate: nothing here is called
|
|
# from a coroutine. The async workers hand run_changedetection() to a ThreadPoolExecutor
|
|
# (worker.py: "Run change detection in executor to avoid blocking event loop"), the requests
|
|
# fetcher does the same, and Flask is synchronous. If you ever call a guarded function directly
|
|
# from a coroutine you will stall that worker's event loop for the duration of someone else's
|
|
# parse - bounded and deadlock-free (the holder is pure CPU and never awaits), but avoid it.
|
|
#
|
|
# Re-entrant: these helpers call into each other (element_removal -> subtractive_xpath_selector).
|
|
#
|
|
# LXML_LOCK_DISABLED=true removes the guard. It exists so the corruption can be reproduced on
|
|
# demand (see tests/test_lxml_concurrency.py) - it is NOT a tuning option, it reinstates a
|
|
# data-integrity bug that silently writes one watch's page text into another watch's snapshot.
|
|
def _build_lxml_lock():
|
|
from changedetectionio.strtobool import strtobool
|
|
|
|
if not strtobool(os.getenv('LXML_LOCK_DISABLED', 'false')):
|
|
return threading.RLock()
|
|
|
|
logger.warning("LXML_LOCK_DISABLED=true - lxml parsing is unguarded. Under concurrency this "
|
|
"is known to leak one watch's rendered page text into another watch's "
|
|
"snapshot. For testing only, never run this in production.")
|
|
|
|
class _NullLock:
|
|
def __enter__(self): return self
|
|
def __exit__(self, *a): return False
|
|
|
|
return _NullLock()
|
|
|
|
|
|
_LXML_LOCK = _build_lxml_lock()
|
|
|
|
|
|
def lxml_guarded(fn):
|
|
"""Serialise a function's libxml2 work - parse, xpath traversal, tostring() and clear()
|
|
all touch the document, so the whole call is held, not just the parse."""
|
|
@wraps(fn)
|
|
def wrapper(*args, **kwargs):
|
|
with _LXML_LOCK:
|
|
return fn(*args, **kwargs)
|
|
return wrapper
|
|
|
|
|
|
def lxml_guard():
|
|
"""Context manager for code OUTSIDE this module that drives lxml directly, so it shares the
|
|
same lock (e.g. XPath validation in forms.py, which runs on Flask request threads)."""
|
|
return _LXML_LOCK
|
|
|
|
|
|
def lxml_html_parser():
|
|
"""A fresh parser for each call, freed immediately afterwards. Never lxml's process-global
|
|
default parser, which is shared across every thread."""
|
|
from lxml import etree
|
|
return etree.HTMLParser()
|
|
|
|
# ---------------------------------------------------------------------------------------
|
|
|
|
# HTML added to be sure each result matching a filter (.example) gets converted to a new line by Inscriptis
|
|
TEXT_FILTER_LIST_LINE_SUFFIX = "<br>"
|
|
TRANSLATE_WHITESPACE_TABLE = str.maketrans('', '', '\r\n\t ')
|
|
PERL_STYLE_REGEX = r'^/(.*?)/([a-z]*)?$'
|
|
|
|
TITLE_RE = re.compile(r"<title[^>]*>(.*?)</title>", re.I | re.S)
|
|
META_CS = re.compile(r'<meta[^>]+charset=["\']?\s*([a-z0-9_\-:+.]+)', re.I)
|
|
|
|
# jq builtins that can leak sensitive data or cause harm when user-supplied expressions are executed.
|
|
# env/$ENV reads all process environment variables (passwords, API keys, etc.)
|
|
# include/import can read arbitrary files from disk
|
|
# input/inputs reads beyond the supplied JSON data
|
|
# debug/stderr leaks data to stderr
|
|
# halt/halt_error terminates the process (DoS)
|
|
_JQ_BLOCKED_PATTERNS = [
|
|
(re.compile(r'\benv\b'), 'env (reads environment variables)'),
|
|
(re.compile(r'\$ENV\b'), '$ENV (reads environment variables)'),
|
|
(re.compile(r'\binclude\b'), 'include (reads files from disk)'),
|
|
(re.compile(r'\bimport\b'), 'import (reads files from disk)'),
|
|
(re.compile(r'\binputs?\b'), 'input/inputs (reads beyond provided data)'),
|
|
(re.compile(r'\bdebug\b'), 'debug (leaks data to stderr)'),
|
|
(re.compile(r'\bstderr\b'), 'stderr (leaks data to stderr)'),
|
|
(re.compile(r'\bhalt(?:_error)?\b'), 'halt/halt_error (terminates the process)'),
|
|
(re.compile(r'\$__loc__\b'), '$__loc__ (leaks file path information)'),
|
|
(re.compile(r'\bbuiltins\b'), 'builtins (enumerates available functions)'),
|
|
(re.compile(r'\bmodulemeta\b'), 'modulemeta (leaks module information)'),
|
|
(re.compile(r'\$JQ_BUILD_CONFIGURATION\b'), '$JQ_BUILD_CONFIGURATION (leaks build information)'),
|
|
]
|
|
|
|
def validate_jq_expression(expression: str) -> None:
|
|
"""Raise ValueError if the jq expression uses any dangerous builtin.
|
|
|
|
User-supplied jq expressions are executed server-side. Without this check,
|
|
builtins like `env` expose every process environment variable (SALTED_PASS,
|
|
proxy credentials, API keys, etc.) as watch output.
|
|
"""
|
|
from changedetectionio.strtobool import strtobool
|
|
if strtobool(os.getenv('JQ_ALLOW_RISKY_EXPRESSIONS', 'false')):
|
|
return
|
|
|
|
for pattern, description in _JQ_BLOCKED_PATTERNS:
|
|
if pattern.search(expression):
|
|
msg = f"jq expression uses disallowed builtin: {description}"
|
|
logger.critical(f"Security: blocked jq expression containing '{description}' - expression: {expression!r}")
|
|
raise ValueError(msg)
|
|
|
|
META_CT = re.compile(r'<meta[^>]+http-equiv=["\']?content-type["\']?[^>]*content=["\'][^>]*charset=([a-z0-9_\-:+.]+)', re.I)
|
|
|
|
# 'price' , 'lowPrice', 'highPrice' are usually under here
|
|
# All of those may or may not appear on different websites - I didnt find a way todo case-insensitive searching here
|
|
LD_JSON_PRODUCT_OFFER_SELECTORS = ["json:$..offers", "json:$..Offers"]
|
|
|
|
class JSONNotFound(ValueError):
|
|
def __init__(self, msg):
|
|
ValueError.__init__(self, msg)
|
|
|
|
|
|
_DEFAULT_UNSAFE_XPATH3_FUNCTIONS = [
|
|
'unparsed-text',
|
|
'unparsed-text-lines',
|
|
'unparsed-text-available',
|
|
'doc',
|
|
'doc-available',
|
|
'json-doc',
|
|
'json-doc-available',
|
|
'collection', # XPath 2.0+: loads XML node collections from arbitrary URIs
|
|
'uri-collection', # XPath 3.0+: enumerates URIs from resource collections
|
|
'transform', # XPath 3.1: XSLT transformation (currently raises, block proactively)
|
|
'load-xquery-module', # XPath 3.1: loads XQuery modules (currently raises, block proactively)
|
|
'environment-variable',
|
|
'available-environment-variables',
|
|
]
|
|
|
|
|
|
# XPath 3.1 says the default collation is Unicode codepoint collation. elementpath instead leaves
|
|
# its collation functions pointing at locale.strxfrm / locale.strcoll, so the process-wide
|
|
# LC_COLLATE decides what contains() means:
|
|
#
|
|
# def contains(self, a, b): return self.strxfrm(b) in self.strxfrm(a)
|
|
#
|
|
# With LC_COLLATE=C, strxfrm() is the identity and that is an ordinary substring test. With
|
|
# LC_COLLATE=en_US.UTF-8 it returns a binary collation key, and a substring of a collation key is
|
|
# not the collation key of the substring - so contains(), starts-with(), ends-with() and
|
|
# substring-before/after() return false for every input, and a filter that matches 67 elements
|
|
# matches 0 (#4437). Name tests, axes and '=' are unaffected, which is what made it look like the
|
|
# page had changed layout.
|
|
#
|
|
# flask_app.py keeps LC_COLLATE in "C" for this reason, but a filter must not depend on a distant
|
|
# module's locale bookkeeping, nor on what an operator puts in LANG/LC_ALL. Pinning the collation
|
|
# per evaluation makes the filter mean the same thing in every deployment.
|
|
XPATH_CODEPOINT_COLLATION = 'http://www.w3.org/2005/xpath-functions/collation/codepoint'
|
|
|
|
|
|
def get_safe_xpath3_parser():
|
|
"""Return an XPath3Parser subclass with filesystem/environment access functions removed.
|
|
|
|
XPath 3.0 includes functions that can read arbitrary files or environment variables:
|
|
- unparsed-text / unparsed-text-lines / unparsed-text-available (file read)
|
|
- doc / doc-available (XML fetch from URI)
|
|
- environment-variable / available-environment-variables (env var leakage)
|
|
|
|
Subclassing gives us an independent symbol_table copy (not shared with the parent class),
|
|
so removing entries here does not affect XPath3Parser itself.
|
|
|
|
Override the blocked list via the XPATH_BLOCKED_FUNCTIONS environment variable
|
|
(comma-separated, e.g. "unparsed-text,doc,environment-variable").
|
|
"""
|
|
import os
|
|
from elementpath.xpath3 import XPath3Parser
|
|
|
|
class SafeXPath3Parser(XPath3Parser):
|
|
pass
|
|
|
|
env_override = os.getenv('XPATH_BLOCKED_FUNCTIONS')
|
|
if env_override is not None:
|
|
blocked = [f.strip() for f in env_override.split(',') if f.strip()]
|
|
else:
|
|
blocked = _DEFAULT_UNSAFE_XPATH3_FUNCTIONS
|
|
|
|
for _fn in blocked:
|
|
SafeXPath3Parser.symbol_table.pop(_fn, None)
|
|
|
|
return SafeXPath3Parser
|
|
|
|
|
|
# Doesn't look like python supports forward slash auto enclosure in re.findall
|
|
# So convert it to inline flag "(?i)foobar" type configuration
|
|
@lru_cache(maxsize=100)
|
|
def perl_style_slash_enclosed_regex_to_options(regex):
|
|
|
|
res = re.search(PERL_STYLE_REGEX, regex, re.IGNORECASE)
|
|
|
|
if res:
|
|
flags = res.group(2) if res.group(2) else 'i'
|
|
regex = f"(?{flags}){res.group(1)}"
|
|
else:
|
|
# Fall back to just ignorecase as an option
|
|
regex = f"(?i){regex}"
|
|
|
|
return regex
|
|
|
|
# Given a CSS Rule, and a blob of HTML, return the blob of HTML that matches
|
|
def include_filters(include_filters, html_content, append_pretty_line_formatting=False):
|
|
from bs4 import BeautifulSoup
|
|
soup = BeautifulSoup(html_content, "html.parser")
|
|
html_block = ""
|
|
r = soup.select(include_filters, separator="")
|
|
|
|
for element in r:
|
|
# When there's more than 1 match, then add the suffix to separate each line
|
|
# And where the matched result doesn't include something that will cause Inscriptis to add a newline
|
|
# (This way each 'match' reliably has a new-line in the diff)
|
|
# Divs are converted to 4 whitespaces by inscriptis
|
|
if append_pretty_line_formatting and len(html_block) and not element.name in (['br', 'hr', 'div', 'p']):
|
|
html_block += TEXT_FILTER_LIST_LINE_SUFFIX
|
|
|
|
html_block += str(element)
|
|
|
|
return html_block
|
|
|
|
def subtractive_css_selector(css_selector, content):
|
|
from bs4 import BeautifulSoup
|
|
soup = BeautifulSoup(content, "html.parser")
|
|
|
|
# So that the elements dont shift their index, build a list of elements here which will be pointers to their place in the DOM
|
|
elements_to_remove = soup.select(css_selector)
|
|
|
|
if not elements_to_remove:
|
|
# Better to return the original that rebuild with BeautifulSoup
|
|
return content
|
|
|
|
# Then, remove them in a separate loop
|
|
for item in elements_to_remove:
|
|
item.decompose()
|
|
|
|
return str(soup)
|
|
|
|
@lxml_guarded
|
|
def subtractive_xpath_selector(selectors: List[str], html_content: str) -> str:
|
|
from lxml import etree
|
|
# Parse the HTML content using lxml. Own parser, not etree.HTML()'s process-global default.
|
|
html_tree = etree.fromstring(html_content, parser=lxml_html_parser())
|
|
|
|
# First, collect all elements to remove
|
|
elements_to_remove = []
|
|
|
|
# Iterate over the list of XPath selectors
|
|
for selector in selectors:
|
|
# Collect elements for each selector
|
|
elements_to_remove.extend(html_tree.xpath(selector))
|
|
|
|
# If no elements were found, return the original HTML content
|
|
if not elements_to_remove:
|
|
return html_content
|
|
|
|
# Then, remove them in a separate loop
|
|
for element in elements_to_remove:
|
|
if element.getparent() is not None: # Ensure the element has a parent before removing
|
|
element.getparent().remove(element)
|
|
|
|
# Convert the modified HTML tree back to a string
|
|
modified_html = etree.tostring(html_tree, method="html").decode("utf-8")
|
|
return modified_html
|
|
|
|
|
|
def element_removal(selectors: List[str], html_content):
|
|
"""Removes elements that match a list of CSS or XPath selectors."""
|
|
modified_html = html_content
|
|
css_selectors = []
|
|
xpath_selectors = []
|
|
|
|
for selector in selectors:
|
|
if selector.strip().startswith(('xpath:', 'xpath1:', '//')):
|
|
# Handle XPath selectors separately
|
|
xpath_selector = selector.removeprefix('xpath:').removeprefix('xpath1:')
|
|
xpath_selectors.append(xpath_selector)
|
|
else:
|
|
# Collect CSS selectors as one "hit", see comment in subtractive_css_selector
|
|
css_selectors.append(selector.strip().strip(","))
|
|
|
|
if xpath_selectors:
|
|
modified_html = subtractive_xpath_selector(xpath_selectors, modified_html)
|
|
|
|
if css_selectors:
|
|
# Remove duplicates, then combine all CSS selectors into one string, separated by commas
|
|
# This stops the elements index shifting
|
|
unique_selectors = list(set(css_selectors)) # Ensure uniqueness
|
|
combined_css_selector = " , ".join(unique_selectors)
|
|
modified_html = subtractive_css_selector(combined_css_selector, modified_html)
|
|
|
|
|
|
return modified_html
|
|
|
|
def elementpath_tostring(obj):
|
|
"""
|
|
change elementpath.select results to string type
|
|
# The MIT License (MIT), Copyright (c), 2018-2021, SISSA (Scuola Internazionale Superiore di Studi Avanzati)
|
|
# https://github.com/sissaschool/elementpath/blob/dfcc2fd3d6011b16e02bf30459a7924f547b47d0/elementpath/xpath_tokens.py#L1038
|
|
"""
|
|
|
|
import elementpath
|
|
from decimal import Decimal
|
|
import math
|
|
|
|
if obj is None:
|
|
return ''
|
|
# https://elementpath.readthedocs.io/en/latest/xpath_api.html#elementpath.select
|
|
elif isinstance(obj, elementpath.XPathNode):
|
|
return obj.string_value
|
|
elif isinstance(obj, bool):
|
|
return 'true' if obj else 'false'
|
|
elif isinstance(obj, Decimal):
|
|
value = format(obj, 'f')
|
|
if '.' in value:
|
|
return value.rstrip('0').rstrip('.')
|
|
return value
|
|
|
|
elif isinstance(obj, float):
|
|
if math.isnan(obj):
|
|
return 'NaN'
|
|
elif math.isinf(obj):
|
|
return str(obj).upper()
|
|
|
|
value = str(obj)
|
|
if '.' in value:
|
|
value = value.rstrip('0').rstrip('.')
|
|
if '+' in value:
|
|
value = value.replace('+', '')
|
|
if 'e' in value:
|
|
return value.upper()
|
|
return value
|
|
|
|
return str(obj)
|
|
|
|
# Return str Utf-8 of matched rules
|
|
@lxml_guarded
|
|
def xpath_filter(xpath_filter, html_content, append_pretty_line_formatting=False, is_xml=False):
|
|
"""
|
|
|
|
:param xpath_filter:
|
|
:param html_content:
|
|
:param append_pretty_line_formatting:
|
|
:param is_xml: set to true if is XML or is RSS (RSS is XML)
|
|
:return:
|
|
"""
|
|
from lxml import etree, html
|
|
import elementpath
|
|
|
|
parser = lxml_html_parser()
|
|
tree = None
|
|
try:
|
|
if is_xml:
|
|
# So that we can keep CDATA for cdata_in_document_to_text() to process
|
|
parser = etree.XMLParser(strip_cdata=False, resolve_entities=False, no_network=True)
|
|
# For XML/RSS content, use etree.fromstring to properly handle XML declarations
|
|
tree = etree.fromstring(html_content.encode('utf-8') if isinstance(html_content, str) else html_content, parser=parser)
|
|
else:
|
|
tree = html.fromstring(html_content, parser=parser)
|
|
html_block = ""
|
|
|
|
# Build namespace map for XPath queries
|
|
namespaces = {'re': 'http://exslt.org/regular-expressions'}
|
|
|
|
# Handle default namespace in documents (common in RSS/Atom feeds, but can occur in any XML)
|
|
# XPath spec: unprefixed element names have no namespace, not the default namespace
|
|
# Solution: Register the default namespace with empty string prefix in elementpath
|
|
# This is primarily for RSS/Atom feeds but works for any XML with default namespace
|
|
if hasattr(tree, 'nsmap') and tree.nsmap and None in tree.nsmap:
|
|
# Register the default namespace with empty string prefix for elementpath
|
|
# This allows //title to match elements in the default namespace
|
|
namespaces[''] = tree.nsmap[None]
|
|
|
|
r = elementpath.select(tree, xpath_filter.strip(), namespaces=namespaces,
|
|
parser=get_safe_xpath3_parser(),
|
|
default_collation=XPATH_CODEPOINT_COLLATION)
|
|
#@note: //title/text() now works with default namespaces (fixed by registering '' prefix)
|
|
#@note: //title/text() wont work where <title>CDATA.. (use cdata_in_document_to_text first)
|
|
|
|
if type(r) != list:
|
|
r = [r]
|
|
|
|
for element in r:
|
|
# When there's more than 1 match, then add the suffix to separate each line
|
|
# And where the matched result doesn't include something that will cause Inscriptis to add a newline
|
|
# (This way each 'match' reliably has a new-line in the diff)
|
|
# Divs are converted to 4 whitespaces by inscriptis
|
|
if append_pretty_line_formatting and len(html_block) and (not hasattr( element, 'tag' ) or not element.tag in (['br', 'hr', 'div', 'p'])):
|
|
html_block += TEXT_FILTER_LIST_LINE_SUFFIX
|
|
|
|
if type(element) == str:
|
|
html_block += element
|
|
elif issubclass(type(element), etree._Element) or issubclass(type(element), etree._ElementTree):
|
|
# Use 'xml' method for RSS/XML content, 'html' for HTML content
|
|
# parser will be XMLParser if we detected XML content
|
|
method = 'xml' if (is_xml or isinstance(parser, etree.XMLParser)) else 'html'
|
|
html_block += etree.tostring(element, pretty_print=True, method=method, encoding='unicode')
|
|
else:
|
|
html_block += elementpath_tostring(element)
|
|
|
|
# Drop element references before the finally block so tree.clear() can release
|
|
# the libxml2 document immediately (elements pin the C-level doc via refcount).
|
|
del r
|
|
return html_block
|
|
finally:
|
|
# Explicitly clear the tree to free memory
|
|
# lxml trees can hold significant memory, especially with large documents
|
|
if tree is not None:
|
|
tree.clear()
|
|
|
|
# Return str Utf-8 of matched rules
|
|
# 'xpath1:'
|
|
@lxml_guarded
|
|
def xpath1_filter(xpath_filter, html_content, append_pretty_line_formatting=False, is_xml=False):
|
|
from lxml import etree, html
|
|
|
|
# Own parser, not lxml's process-global default (which parser=None would select).
|
|
parser = lxml_html_parser()
|
|
tree = None
|
|
try:
|
|
if is_xml:
|
|
# So that we can keep CDATA for cdata_in_document_to_text() to process
|
|
parser = etree.XMLParser(strip_cdata=False, resolve_entities=False, no_network=True)
|
|
# For XML/RSS content, use etree.fromstring to properly handle XML declarations
|
|
tree = etree.fromstring(html_content.encode('utf-8') if isinstance(html_content, str) else html_content, parser=parser)
|
|
else:
|
|
tree = html.fromstring(html_content, parser=parser)
|
|
html_block = ""
|
|
|
|
# Build namespace map for XPath queries
|
|
namespaces = {'re': 'http://exslt.org/regular-expressions'}
|
|
|
|
# NOTE: lxml's native xpath() does NOT support empty string prefix for default namespace
|
|
# For documents with default namespace (RSS/Atom feeds), users must use:
|
|
# - local-name(): //*[local-name()='title']/text()
|
|
# - Or use xpath_filter (not xpath1_filter) which supports default namespaces
|
|
# XPath spec: unprefixed element names have no namespace, not the default namespace
|
|
|
|
r = tree.xpath(xpath_filter.strip(), namespaces=namespaces)
|
|
#@note: xpath1 (lxml) does NOT automatically handle default namespaces
|
|
#@note: Use //*[local-name()='element'] or switch to xpath_filter for default namespace support
|
|
#@note: //title/text() wont work where <title>CDATA.. (use cdata_in_document_to_text first)
|
|
|
|
for element in r:
|
|
# When there's more than 1 match, then add the suffix to separate each line
|
|
# And where the matched result doesn't include something that will cause Inscriptis to add a newline
|
|
# (This way each 'match' reliably has a new-line in the diff)
|
|
# Divs are converted to 4 whitespaces by inscriptis
|
|
if append_pretty_line_formatting and len(html_block) and (not hasattr(element, 'tag') or not element.tag in (['br', 'hr', 'div', 'p'])):
|
|
html_block += TEXT_FILTER_LIST_LINE_SUFFIX
|
|
|
|
# Some kind of text, UTF-8 or other
|
|
if isinstance(element, (str, bytes)):
|
|
html_block += element
|
|
else:
|
|
# Return the HTML/XML which will get parsed as text
|
|
# Use 'xml' method for RSS/XML content, 'html' for HTML content
|
|
# parser will be XMLParser if we detected XML content
|
|
method = 'xml' if (is_xml or isinstance(parser, etree.XMLParser)) else 'html'
|
|
html_block += etree.tostring(element, pretty_print=True, method=method, encoding='unicode')
|
|
|
|
return html_block
|
|
finally:
|
|
# Explicitly clear the tree to free memory
|
|
# lxml trees can hold significant memory, especially with large documents
|
|
if tree is not None:
|
|
tree.clear()
|
|
|
|
# Extract/find element
|
|
def extract_element(find='title', html_content=''):
|
|
from bs4 import BeautifulSoup
|
|
|
|
#Re #106, be sure to handle when its not found
|
|
element_text = None
|
|
|
|
soup = BeautifulSoup(html_content, 'html.parser')
|
|
result = soup.find(find)
|
|
if result and result.string:
|
|
element_text = result.string.strip()
|
|
|
|
return element_text
|
|
|
|
# Any surrogate still present after json.loads() is a lone one - the decoder already folds
|
|
# well-formed pairs (😀 -> a single emoji character) before we see them. Lone ones
|
|
# come from escapes like \uD800 in the source document, which travel as plain ASCII and so
|
|
# pass straight through the fetched-content sanitizer in processors/base.py.
|
|
_LONE_SURROGATE_RE = re.compile(r'[\ud800-\udfff]')
|
|
|
|
def _has_lone_surrogate(value):
|
|
if isinstance(value, str):
|
|
return bool(_LONE_SURROGATE_RE.search(value))
|
|
if isinstance(value, dict):
|
|
return any(_has_lone_surrogate(k) or _has_lone_surrogate(v) for k, v in value.items())
|
|
if isinstance(value, list):
|
|
return any(_has_lone_surrogate(v) for v in value)
|
|
return False
|
|
|
|
def _sanitize_lone_surrogate_str(value: str) -> str:
|
|
return _LONE_SURROGATE_RE.sub('\ufffd', value)
|
|
|
|
def _sanitize_lone_surrogates(value):
|
|
if isinstance(value, str):
|
|
return _sanitize_lone_surrogate_str(value)
|
|
if isinstance(value, dict):
|
|
# JSON object keys are always strings, so they only need the substitution - recursing on a
|
|
# key would (in theory) hand back an unhashable dict/list
|
|
return {_sanitize_lone_surrogate_str(k) if isinstance(k, str) else k: _sanitize_lone_surrogates(v)
|
|
for k, v in value.items()}
|
|
if isinstance(value, list):
|
|
return [_sanitize_lone_surrogates(v) for v in value]
|
|
return value
|
|
|
|
#
|
|
def _parse_json(json_data, json_filter):
|
|
from jsonpath_ng.ext import parse
|
|
|
|
# Replace lone surrogates with U+FFFD, matching what processors/base.py already does for
|
|
# fetched content. Two separate failures otherwise, neither of which is confined to the
|
|
# offending value: jq re-serializes the whole document for its own C parser, which rejects
|
|
# a lone high surrogate, so every jq:/jqraw: filter on the page errors even when it points
|
|
# at an unrelated field; and on the json: path the surrogate reaches the caller intact and
|
|
# raises UnicodeEncodeError later in checksums and history writes.
|
|
# See: https://github.com/dgtlmoon/changedetection.io/issues/4273
|
|
if _has_lone_surrogate(json_data):
|
|
logger.warning("Lone unicode surrogate(s) in JSON document, replacing with U+FFFD before filtering")
|
|
json_data = _sanitize_lone_surrogates(json_data)
|
|
|
|
if json_filter.startswith("json:"):
|
|
jsonpath_expression = parse(json_filter.replace('json:', ''))
|
|
match = jsonpath_expression.find(json_data)
|
|
return _get_stripped_text_from_json_match(match)
|
|
|
|
if json_filter.startswith("jq:") or json_filter.startswith("jqraw:"):
|
|
|
|
try:
|
|
import jq
|
|
except ModuleNotFoundError:
|
|
# `jq` requires full compilation in windows and so isn't generally available
|
|
raise Exception("jq not support not found")
|
|
|
|
if json_filter.startswith("jq:"):
|
|
expr = json_filter.removeprefix("jq:")
|
|
validate_jq_expression(expr)
|
|
jq_expression = jq.compile(expr)
|
|
match = jq_expression.input(json_data).all()
|
|
return _get_stripped_text_from_json_match(match)
|
|
|
|
if json_filter.startswith("jqraw:"):
|
|
expr = json_filter.removeprefix("jqraw:")
|
|
validate_jq_expression(expr)
|
|
jq_expression = jq.compile(expr)
|
|
match = jq_expression.input(json_data).all()
|
|
return '\n'.join(str(item) for item in match)
|
|
|
|
def _get_stripped_text_from_json_match(match):
|
|
s = []
|
|
# More than one result, we will return it as a JSON list.
|
|
if len(match) > 1:
|
|
for i in match:
|
|
s.append(i.value if hasattr(i, 'value') else i)
|
|
|
|
# Single value, use just the value, as it could be later used in a token in notifications.
|
|
if len(match) == 1:
|
|
s = match[0].value if hasattr(match[0], 'value') else match[0]
|
|
|
|
# Re #257 - Better handling where it does not exist, in the case the original 's' value was False..
|
|
if not match:
|
|
# Re 265 - Just return an empty string when filter not found
|
|
return ''
|
|
|
|
# Ticket #462 - allow the original encoding through, usually it's UTF-8 or similar
|
|
stripped_text_from_html = json.dumps(s, indent=4, ensure_ascii=False)
|
|
|
|
return stripped_text_from_html
|
|
|
|
def extract_json_blob_from_html(content, ensure_is_ldjson_info_type, json_filter):
|
|
from bs4 import BeautifulSoup
|
|
stripped_text_from_html = ''
|
|
|
|
# Foreach <script json></script> blob.. just return the first that matches json_filter
|
|
# As a last resort, try to parse the whole <body>
|
|
soup = BeautifulSoup(content, 'html.parser')
|
|
|
|
if ensure_is_ldjson_info_type:
|
|
bs_result = soup.find_all('script', {"type": "application/ld+json"})
|
|
else:
|
|
bs_result = soup.find_all('script')
|
|
bs_result += soup.find_all('body')
|
|
|
|
bs_jsons = []
|
|
|
|
for result in bs_result:
|
|
# result.text is how bs4 magically strips JSON from the body
|
|
content_start = result.text.lstrip("\ufeff").strip()[:100] if result.text else ''
|
|
# Skip empty tags, and things that dont even look like JSON
|
|
if not result.text or not (content_start[0] == '{' or content_start[0] == '['):
|
|
continue
|
|
try:
|
|
json_data = json.loads(result.text)
|
|
bs_jsons.append(json_data)
|
|
except json.JSONDecodeError:
|
|
# Skip objects which cannot be parsed
|
|
continue
|
|
|
|
if not bs_jsons:
|
|
raise JSONNotFound("No parsable JSON found in this document")
|
|
|
|
for json_data in bs_jsons:
|
|
stripped_text_from_html = _parse_json(json_data, json_filter)
|
|
|
|
if ensure_is_ldjson_info_type:
|
|
# Could sometimes be list, string or something else random
|
|
if isinstance(json_data, dict):
|
|
# If it has LD JSON 'key' @type, and @type is 'product', and something was found for the search
|
|
# (Some sites have multiple of the same ld+json @type='product', but some have the review part, some have the 'price' part)
|
|
# @type could also be a list although non-standard ("@type": ["Product", "SubType"],)
|
|
# LD_JSON auto-extract also requires some content PLUS the ldjson to be present
|
|
# 1833 - could be either str or dict, should not be anything else
|
|
|
|
t = json_data.get('@type')
|
|
if t and stripped_text_from_html:
|
|
|
|
if isinstance(t, str) and t.lower() == ensure_is_ldjson_info_type.lower():
|
|
break
|
|
# The non-standard part, some have a list
|
|
elif isinstance(t, list):
|
|
if ensure_is_ldjson_info_type.lower() in [x.lower().strip() for x in t]:
|
|
break
|
|
|
|
elif stripped_text_from_html:
|
|
break
|
|
|
|
return stripped_text_from_html
|
|
|
|
# content - json
|
|
# json_filter - ie json:$..price
|
|
# ensure_is_ldjson_info_type - str "product", optional, "@type == product" (I dont know how to do that as a json selector)
|
|
def extract_json_as_string(content, json_filter, ensure_is_ldjson_info_type=None):
|
|
|
|
stripped_text_from_html = False
|
|
# https://github.com/dgtlmoon/changedetection.io/pull/2041#issuecomment-1848397161w
|
|
# Try to parse/filter out the JSON, if we get some parser error, then maybe it's embedded within HTML tags
|
|
|
|
# Looks like clean JSON, dont bother extracting from HTML
|
|
|
|
content_start = content.lstrip("\ufeff").strip()[:100]
|
|
|
|
if content_start[0] == '{' or content_start[0] == '[':
|
|
try:
|
|
# .lstrip("\ufeff") strings ByteOrderMark from UTF8 and still lets the UTF work
|
|
stripped_text_from_html = _parse_json(json.loads(content.lstrip("\ufeff")), json_filter)
|
|
except json.JSONDecodeError as e:
|
|
logger.warning(f"Error processing JSON {content[:20]}...{str(e)})")
|
|
else:
|
|
# Check for JSONP wrapper: someCallback({...}) or some.namespace({...})
|
|
# Server may claim application/json but actually return JSONP
|
|
jsonp_match = re.match(r'^\w[\w.]*\s*\((.+)\)\s*;?\s*$', content.lstrip("\ufeff").strip(), re.DOTALL)
|
|
if jsonp_match:
|
|
try:
|
|
inner = jsonp_match.group(1).strip()
|
|
logger.warning(f"Content looks like JSONP, attempting to extract inner JSON for filter '{json_filter}'")
|
|
stripped_text_from_html = _parse_json(json.loads(inner), json_filter)
|
|
except json.JSONDecodeError as e:
|
|
logger.warning(f"Error processing JSONP inner content {content[:20]}...{str(e)})")
|
|
|
|
if not stripped_text_from_html:
|
|
# Probably something else, go fish inside for it
|
|
try:
|
|
stripped_text_from_html = extract_json_blob_from_html(content=content,
|
|
ensure_is_ldjson_info_type=ensure_is_ldjson_info_type,
|
|
json_filter=json_filter)
|
|
except json.JSONDecodeError as e:
|
|
logger.warning(f"Error processing JSON while extracting JSON from HTML blob {content[:20]}...{str(e)})")
|
|
|
|
if not stripped_text_from_html:
|
|
# Re 265 - Just return an empty string when filter not found
|
|
return ''
|
|
|
|
return stripped_text_from_html
|
|
|
|
# Mode - "content" return the content without the matches (default)
|
|
# - "line numbers" return a list of line numbers that match (int list)
|
|
#
|
|
# wordlist - list of regex's (str) or words (str)
|
|
# Preserves all linefeeds and other whitespacing, its not the job of this to remove that
|
|
def strip_ignore_text(content, wordlist, mode="content"):
|
|
ignore_text = []
|
|
ignore_regex = []
|
|
ignore_regex_multiline = []
|
|
ignored_lines = []
|
|
|
|
if not content:
|
|
return ''
|
|
|
|
for k in wordlist:
|
|
# Skip empty strings to avoid matching everything
|
|
if not k or not k.strip():
|
|
continue
|
|
# Is it a regex?
|
|
res = re.search(PERL_STYLE_REGEX, k, re.IGNORECASE)
|
|
if res:
|
|
res = re.compile(perl_style_slash_enclosed_regex_to_options(k))
|
|
if res.flags & re.DOTALL or res.flags & re.MULTILINE:
|
|
ignore_regex_multiline.append(res)
|
|
else:
|
|
ignore_regex.append(res)
|
|
else:
|
|
ignore_text.append(k.strip())
|
|
|
|
for r in ignore_regex_multiline:
|
|
for match in r.finditer(content):
|
|
content_lines = content[:match.end()].splitlines(keepends=True)
|
|
match_lines = content[match.start():match.end()].splitlines(keepends=True)
|
|
|
|
end_line = len(content_lines)
|
|
start_line = end_line - len(match_lines)
|
|
|
|
if end_line - start_line <= 1:
|
|
# Match is empty or in the middle of the line
|
|
ignored_lines.append(start_line)
|
|
else:
|
|
for i in range(start_line, end_line):
|
|
ignored_lines.append(i)
|
|
|
|
line_index = 0
|
|
lines = content.splitlines(keepends=True)
|
|
for line in lines:
|
|
# Always ignore blank lines in this mode. (when this function gets called)
|
|
got_match = False
|
|
for l in ignore_text:
|
|
if l.lower() in line.lower():
|
|
got_match = True
|
|
|
|
if not got_match:
|
|
for r in ignore_regex:
|
|
if r.search(line):
|
|
got_match = True
|
|
|
|
if got_match:
|
|
ignored_lines.append(line_index)
|
|
|
|
line_index += 1
|
|
|
|
ignored_lines = set([i for i in ignored_lines if i >= 0 and i < len(lines)])
|
|
|
|
# Used for finding out what to highlight
|
|
if mode == "line numbers":
|
|
return [i + 1 for i in ignored_lines]
|
|
|
|
output_lines = set(range(len(lines))) - ignored_lines
|
|
return ''.join([lines[i] for i in output_lines])
|
|
|
|
def cdata_in_document_to_text(html_content: str, render_anchor_tag_content=False) -> str:
|
|
from xml.sax.saxutils import escape as xml_escape
|
|
pattern = '<!\[CDATA\[(\s*(?:.(?<!\]\]>)\s*)*)\]\]>'
|
|
def repl(m):
|
|
text = m.group(1)
|
|
return xml_escape(html_to_text(html_content=text)).strip()
|
|
|
|
return re.sub(pattern, repl, html_content)
|
|
|
|
|
|
# NOTE!! ANYTHING LIBXML, HTML5LIB ETC WILL CAUSE SOME SMALL MEMORY LEAK IN THE LOCAL "LIB" IMPLEMENTATION OUTSIDE PYTHON
|
|
|
|
|
|
def html_to_text(html_content: str, render_anchor_tag_content=False, is_rss=False, timeout=10) -> str:
|
|
"""
|
|
Convert HTML content to plain text using inscriptis.
|
|
|
|
Thread-Safety: inscriptis.get_text() parses with lxml's process-global default parser.
|
|
That is fine on its own, but NOT alongside the xpath helpers' own parser - see the
|
|
_LXML_LOCK comment at the top of this module. So the get_text() call is guarded.
|
|
|
|
Do NOT "fix" this by giving get_text() its own parser instead. That was tried
|
|
(bee1130c6, reverted by 272e68ad2) and measured at ~10x WORSE cross-document leakage,
|
|
because then both sides use unlocked per-call parsers and lxml's per-parser lock stops
|
|
helping at all.
|
|
"""
|
|
from inscriptis import get_text
|
|
from inscriptis.model.config import ParserConfig
|
|
|
|
if render_anchor_tag_content:
|
|
parser_config = ParserConfig(
|
|
annotation_rules={"a": ["hyperlink"]},
|
|
display_links=True
|
|
)
|
|
else:
|
|
parser_config = None
|
|
if is_rss:
|
|
html_content = re.sub(r'<title([\s>])', r'<h1\1', html_content)
|
|
html_content = re.sub(r'</title>', r'</h1>', html_content)
|
|
else:
|
|
# Use BS4 html.parser to strip bloat — SPA's often dump 10MB+ of CSS/JS into <head>,
|
|
# causing inscriptis to silently give up. Regex-based stripping is unsafe because tags
|
|
# can appear inside JSON data attributes with JS-escaped closing tags (e.g. <\/script>),
|
|
# causing the regex to scan past the intended close and eat real page content.
|
|
from bs4 import BeautifulSoup
|
|
soup = BeautifulSoup(html_content, 'html.parser')
|
|
# Strip tags that inscriptis cannot render as meaningful text and which can be very large.
|
|
# svg/math: produce path-data/MathML garbage; canvas/iframe/template: no inscriptis handlers.
|
|
# video/audio/picture are kept — they may contain meaningful fallback text or captions.
|
|
for tag in soup.find_all(['head', 'script', 'style', 'noscript', 'svg',
|
|
'math', 'canvas', 'iframe', 'template']):
|
|
tag.decompose()
|
|
|
|
# SPAs often use <body style="display:none"> to hide content until JS loads.
|
|
# inscriptis respects CSS display rules, so strip hiding styles from the body tag.
|
|
body_tag = soup.find('body')
|
|
if body_tag and body_tag.get('style'):
|
|
style = body_tag['style']
|
|
if re.search(r'\b(?:display\s*:\s*none|visibility\s*:\s*hidden)\b', style, re.IGNORECASE):
|
|
logger.debug(f"html_to_text: Removing hiding styles from body tag (found: '{style}')")
|
|
del body_tag['style']
|
|
|
|
html_content = str(soup)
|
|
|
|
# Only the lxml part is guarded - the BeautifulSoup stripping above is pure Python.
|
|
with _LXML_LOCK:
|
|
text_content = get_text(html_content, config=parser_config)
|
|
return text_content
|
|
|
|
# Does LD+JSON exist with a @type=='product' and a .price set anywhere?
|
|
def has_ldjson_product_info(content):
|
|
try:
|
|
# Better than .lower() which can use a lot of ram
|
|
if (re.search(r'application/ld\+json', content, re.IGNORECASE) and
|
|
re.search(r'"price"', content, re.IGNORECASE) and
|
|
re.search(r'"pricecurrency"', content, re.IGNORECASE)):
|
|
return True
|
|
|
|
# On some pages this is really terribly expensive when they dont really need it
|
|
# (For example you never want price monitoring, but this runs on every watch to suggest it)
|
|
# for filter in LD_JSON_PRODUCT_OFFER_SELECTORS:
|
|
# pricing_data += extract_json_as_string(content=content,
|
|
# json_filter=filter,
|
|
# ensure_is_ldjson_info_type="product")
|
|
except Exception as e:
|
|
# OK too
|
|
return False
|
|
|
|
return False
|
|
|
|
|
|
|
|
def workarounds_for_obfuscations(content):
|
|
"""
|
|
Some sites are using sneaky tactics to make prices and other information un-renderable by Inscriptis
|
|
This could go into its own Pip package in the future, for faster updates
|
|
"""
|
|
|
|
# HomeDepot.com style <span>$<!-- -->90<!-- -->.<!-- -->74</span>
|
|
# https://github.com/weblyzard/inscriptis/issues/45
|
|
if not content:
|
|
return content
|
|
|
|
content = re.sub('<!--\s+-->', '', content)
|
|
|
|
return content
|
|
|
|
|
|
def get_triggered_text(content, trigger_text):
|
|
triggered_text = []
|
|
result = strip_ignore_text(content=content,
|
|
wordlist=trigger_text,
|
|
mode="line numbers")
|
|
|
|
i = 1
|
|
for p in content.splitlines():
|
|
if i in result:
|
|
triggered_text.append(p)
|
|
i += 1
|
|
|
|
return triggered_text
|
|
|
|
|
|
def extract_title(data: bytes | str, sniff_bytes: int = 2048, scan_chars: int = 8192) -> str | None:
|
|
"""Extract the <title> from an HTML document.
|
|
|
|
Rather than decoding/scanning a fixed prefix of the whole document, we first
|
|
locate the raw ``<title`` marker and then decode only a small window around
|
|
it. This handles pages (e.g. Amazon) where large ``<head>`` sections push
|
|
the title tag well past the old 8 192-character scan limit.
|
|
"""
|
|
# Maximum bytes/chars to extract after (and including) the opening <title tag.
|
|
# The regex needs to see </title>, so the window must cover the full content.
|
|
# The return value is always capped at 2 000 chars; titles beyond that are
|
|
# rare but possible. We read up to 128 KiB from the tag onwards to handle
|
|
# even pathological cases without scanning the whole document.
|
|
_TITLE_WINDOW = 131072
|
|
|
|
try:
|
|
match data:
|
|
case bytes() if data.startswith((b"\xff\xfe\x00\x00", b"\x00\x00\xfe\xff")):
|
|
# UTF-32: locate the tag in the raw bytes, then decode the window.
|
|
tag_pos = data.lower().find(b"<\x00\x00\x00t\x00\x00\x00")
|
|
if tag_pos == -1:
|
|
return None
|
|
chunk = data[tag_pos: tag_pos + _TITLE_WINDOW * 4].decode("utf-32", errors="replace")
|
|
prefix = chunk
|
|
case bytes() if data.startswith((b"\xff\xfe", b"\xfe\xff")):
|
|
# UTF-16: simple byte-pair search is tricky; fall back to decoding
|
|
# a reasonable head chunk and let the regex do the rest.
|
|
prefix = data[: max(scan_chars * 2, _TITLE_WINDOW)].decode("utf-16", errors="replace")
|
|
case bytes():
|
|
# UTF-8 / legacy 8-bit: find the tag cheaply in raw bytes.
|
|
tag_pos = data.lower().find(b"<title")
|
|
if tag_pos == -1:
|
|
return None
|
|
raw_chunk = data[tag_pos: tag_pos + _TITLE_WINDOW]
|
|
try:
|
|
chunk = raw_chunk.decode("utf-8")
|
|
except UnicodeDecodeError:
|
|
try:
|
|
head = data[:sniff_bytes].decode("ascii", errors="ignore")
|
|
if m := (META_CS.search(head) or META_CT.search(head)):
|
|
enc = m.group(1).lower()
|
|
else:
|
|
enc = "cp1252"
|
|
chunk = raw_chunk.decode(enc, errors="replace")
|
|
except Exception as e:
|
|
logger.error(f"Title extraction encoding detection failed: {e}")
|
|
return None
|
|
prefix = chunk
|
|
case str():
|
|
tag_pos = data.lower().find("<title")
|
|
if tag_pos == -1:
|
|
return None
|
|
prefix = data[tag_pos: tag_pos + _TITLE_WINDOW]
|
|
case _:
|
|
logger.error(f"Title extraction received unsupported data type: {type(data)}")
|
|
return None
|
|
|
|
# Search only in the (now tag-anchored) prefix
|
|
if m := TITLE_RE.search(prefix):
|
|
title = html.unescape(" ".join(m.group(1).split())).strip()
|
|
# Some safe limit
|
|
return title[:2000]
|
|
return None
|
|
|
|
except Exception as e:
|
|
logger.error(f"Title extraction failed: {e}")
|
|
return None |