fix dupe import

cleanup of backup creation
Use a rglob that will work cross-platform more reliably
2025-12-09 09:35:33 +00:00 · 2022-10-27 11:23:46 +02:00 · 2022-10-27 11:21:55 +02:00 · 2022-10-27 11:05:39 +02:00 · 2022-10-27 10:56:10 +02:00 · 2022-10-27 10:29:28 +02:00
27 changed files with 396 additions and 141 deletions
--- a/.github/workflows/test-container-build.yml
+++ b/.github/workflows/test-container-build.yml
@@ -0,0 +1,55 @@
+name: ChangeDetection.io Container Build Test
+
+# Triggers the workflow on push or pull request events
+
+# This line doesnt work, even tho it is the documented one
+#on: [push, pull_request]
+
+on:
+  push:
+    paths:
+      - requirements.txt
+      - Dockerfile
+
+  pull_request:
+    paths:
+      - requirements.txt
+      - Dockerfile
+
+  # Changes to requirements.txt packages and Dockerfile may or may not always be compatible with arm etc, so worth testing
+  # @todo: some kind of path filter for requirements.txt and Dockerfile
+jobs:
+  test-container-build:
+    runs-on: ubuntu-latest
+    steps:
+        - uses: actions/checkout@v2
+        - name: Set up Python 3.9
+          uses: actions/setup-python@v2
+          with:
+            python-version: 3.9
+
+        # Just test that the build works, some libraries won't compile on ARM/rPi etc
+        - name: Set up QEMU
+          uses: docker/setup-qemu-action@v1
+          with:
+            image: tonistiigi/binfmt:latest
+            platforms: all
+
+        - name: Set up Docker Buildx
+          id: buildx
+          uses: docker/setup-buildx-action@v1
+          with:
+            install: true
+            version: latest
+            driver-opts: image=moby/buildkit:master
+
+        - name: Test that the docker containers can build
+          id: docker_build
+          uses: docker/build-push-action@v2
+          # https://github.com/docker/build-push-action#customizing
+          with:
+            context: ./
+            file: ./Dockerfile
+            platforms: linux/arm/v7,linux/arm/v6,linux/amd64,linux/arm64,
+            cache-from: type=local,src=/tmp/.buildx-cache
+            cache-to: type=local,dest=/tmp/.buildx-cache
--- a/.github/workflows/test-only.yml
+++ b/.github/workflows/test-only.yml
@@ -1,28 +1,25 @@
-name: ChangeDetection.io Test
+name: ChangeDetection.io App Test

 # Triggers the workflow on push or pull request events
 on: [push, pull_request]

 jobs:
-  test-build:
+  test-application:
    runs-on: ubuntu-latest
    steps:
-
      - uses: actions/checkout@v2
      - name: Set up Python 3.9
        uses: actions/setup-python@v2
        with:
          python-version: 3.9

-      - name: Show env vars
-        run: set
-
      - name: Install dependencies
        run: |
          python -m pip install --upgrade pip
          pip install flake8 pytest
          if [ -f requirements.txt ]; then pip install -r requirements.txt; fi
          if [ -f requirements-dev.txt ]; then pip install -r requirements-dev.txt; fi
+
      - name: Lint with flake8
        run: |
          # stop the build if there are Python syntax errors or undefined names
@@ -39,7 +36,4 @@ jobs:
          # Each test is totally isolated and performs its own cleanup/reset
          cd changedetectionio; ./run_all_tests.sh

-      # https://github.com/docker/build-push-action/blob/master/docs/advanced/test-before-push.md ?
-      # https://github.com/docker/buildx/issues/59 ? Needs to be one platform?

-      # https://github.com/docker/buildx/issues/495#issuecomment-918925854
--- a/CONTRIBUTING.md
+++ b/CONTRIBUTING.md
@@ -6,7 +6,7 @@ Otherwise, it's always best to PR into the `dev` branch.

 Please be sure that all new functionality has a matching test!

-Use `pytest` to validate/test, you can run the existing tests as `pytest tests/test_notifications.py` for example
+Use `pytest` to validate/test, you can run the existing tests as `pytest tests/test_notification.py` for example

 ```
 pip3 install -r requirements-dev
--- a/15
+++ b/15
@@ -5,13 +5,14 @@ FROM python:3.8-slim as builder
 ARG CRYPTOGRAPHY_DONT_BUILD_RUST=1

 RUN apt-get update && apt-get install -y --no-install-recommends \
-    libssl-dev \
-    libffi-dev \
+    g++ \
    gcc \
    libc-dev \
+    libffi-dev \
+    libssl-dev \
    libxslt-dev \
-    zlib1g-dev \
-    g++
+    make \
+    zlib1g-dev

 RUN mkdir /install
 WORKDIR /install
@@ -25,6 +26,11 @@ RUN pip install --target=/dependencies -r /requirements.txt
 RUN pip install --target=/dependencies playwright~=1.26 \
    || echo "WARN: Failed to install Playwright. The application can still run, but the Playwright option will be disabled."

+
+RUN pip install --target=/dependencies jq~=1.3 \
+    || echo "WARN: Failed to install JQ. The application can still run, but the Jq: filter option will be disabled."
+
+
 # Final image stage
 FROM python:3.8-slim

@@ -58,6 +64,7 @@ EXPOSE 5000

 # The actual flask app
 COPY changedetectionio /app/changedetectionio
+
 # The eventlet server wrapper
 COPY changedetection.py /app/changedetection.py

--- a/MANIFEST.in
+++ b/MANIFEST.in
@@ -2,6 +2,7 @@ recursive-include changedetectionio/api *
 recursive-include changedetectionio/templates *
 recursive-include changedetectionio/static *
 recursive-include changedetectionio/model *
+recursive-include changedetectionio/tests *
 include changedetection.py
 global-exclude *.pyc
 global-exclude node_modules
--- a/README.md
+++ b/README.md
@@ -121,8 +121,8 @@ See the wiki for more information https://github.com/dgtlmoon/changedetection.io


 ## Filters
-XPath, JSONPath, jq, and CSS support comes baked in! You can be as specific as you need, use XPath exported from various XPath element query creation tools.

+XPath, JSONPath, jq, and CSS support comes baked in! You can be as specific as you need, use XPath exported from various XPath element query creation tools. 
 (We support LXML `re:test`, `re:math` and `re:replace`.)

 ## Notifications
@@ -161,46 +161,14 @@ This will re-parse the JSON and apply formatting to the text, making it super ea

 ### JSONPath or jq?

-For more complex parsing, filtering, and modifying of JSON data, jq is recommended due to the built-in operators and functions. Refer to the [documentation](https://stedolan.github.io/jq/manual/) for more information on jq.
+For more complex parsing, filtering, and modifying of JSON data, jq is recommended due to the built-in operators and functions. Refer to the [documentation](https://stedolan.github.io/jq/manual/) for more specifc information on jq.

-The example below adds the price in dollars to each item in the JSON data, and then filters to only show items that are greater than 10.
+One big advantage of `jq` is that you can use logic in your JSON filter, such as filters to only show items that have a value greater than/less than etc.

-#### Sample input data from API
-```
-{
-    "items": [
-        {
-           "name": "Product A",
-           "priceInCents": 2500
-        },
-        {
-           "name": "Product B",
-           "priceInCents": 500
-        },
-        {
-           "name": "Product C",
-           "priceInCents": 2000
-        }
-    ]
-}
-```
+See the wiki https://github.com/dgtlmoon/changedetection.io/wiki/JSON-Selector-Filter-help for more information and examples

-#### Sample jq
-`jq:.items[] | . + { "priceInDollars": (.priceInCents / 100) } | select(.priceInDollars > 10)`
+Note: `jq` library must be added separately (`pip3 install jq`)

-#### Sample output data
-```
-{
-  "name": "Product A",
-  "priceInCents": 2500,
-  "priceInDollars": 25
-}
-{
-  "name": "Product C",
-  "priceInCents": 2000,
-  "priceInDollars": 20
-}
-```

 ### Parse JSON embedded in HTML!

--- a/changedetectionio/init.py
+++ b/changedetectionio/init.py
@@ -33,7 +33,7 @@ from flask_wtf import CSRFProtect
 from changedetectionio import html_tools
 from changedetectionio.api import api_v1

-__version__ = '0.39.20'
+__version__ = '0.39.20.4'

 datastore = None

@@ -194,6 +194,9 @@ def changedetection_app(config=None, datastore_o=None):
    watch_api.add_resource(api_v1.Watch, '/api/v1/watch/<string:uuid>',
                           resource_class_kwargs={'datastore': datastore, 'update_q': update_q})

+    watch_api.add_resource(api_v1.SystemInfo, '/api/v1/systeminfo',
+                           resource_class_kwargs={'datastore': datastore, 'update_q': update_q})
+



@@ -636,20 +639,27 @@ def changedetection_app(config=None, datastore_o=None):
            # Only works reliably with Playwright
            visualselector_enabled = os.getenv('PLAYWRIGHT_DRIVER_URL', False) and default['fetch_backend'] == 'html_webdriver'

+            # JQ is difficult to install on windows and must be manually added (outside requirements.txt)
+            jq_support = True
+            try:
+                import jq
+            except ModuleNotFoundError:
+                jq_support = False

            output = render_template("edit.html",
-                                     uuid=uuid,
-                                     watch=datastore.data['watching'][uuid],
-                                     form=form,
-                                     has_empty_checktime=using_default_check_time,
-                                     has_default_notification_urls=True if len(datastore.data['settings']['application']['notification_urls']) else False,
-                                     using_global_webdriver_wait=default['webdriver_delay'] is None,
                                     current_base_url=datastore.data['settings']['application']['base_url'],
                                     emailprefix=os.getenv('NOTIFICATION_MAIL_BUTTON_PREFIX', False),
+                                     form=form,
+                                     has_default_notification_urls=True if len(datastore.data['settings']['application']['notification_urls']) else False,
+                                     has_empty_checktime=using_default_check_time,
+                                     jq_support=jq_support,
+                                     playwright_enabled=os.getenv('PLAYWRIGHT_DRIVER_URL', False),
                                     settings_application=datastore.data['settings']['application'],
+                                     using_global_webdriver_wait=default['webdriver_delay'] is None,
+                                     uuid=uuid,
                                     visualselector_data_is_ready=visualselector_data_is_ready,
                                     visualselector_enabled=visualselector_enabled,
-                                     playwright_enabled=os.getenv('PLAYWRIGHT_DRIVER_URL', False)
+                                     watch=datastore.data['watching'][uuid],
                                     )

        return output
@@ -809,8 +819,10 @@ def changedetection_app(config=None, datastore_o=None):

        newest_file = history[dates[-1]]

+        # Read as binary and force decode as UTF-8
+        # Windows may fail decode in python if we just use 'r' mode (chardet decode exception)
        try:
-            with open(newest_file, 'r') as f:
+            with open(newest_file, 'r', encoding='utf-8', errors='ignore') as f:
                newest_version_file_contents = f.read()
        except Exception as e:
            newest_version_file_contents = "Unable to read {}.\n".format(newest_file)
@@ -823,7 +835,7 @@ def changedetection_app(config=None, datastore_o=None):
            previous_file = history[dates[-2]]

        try:
-            with open(previous_file, 'r') as f:
+            with open(previous_file, 'r', encoding='utf-8', errors='ignore') as f:
                previous_version_file_contents = f.read()
        except Exception as e:
            previous_version_file_contents = "Unable to read {}.\n".format(previous_file)
@@ -900,7 +912,7 @@ def changedetection_app(config=None, datastore_o=None):
        timestamp = list(watch.history.keys())[-1]
        filename = watch.history[timestamp]
        try:
-            with open(filename, 'r') as f:
+            with open(filename, 'r', encoding='utf-8', errors='ignore') as f:
                tmp = f.readlines()

                # Get what needs to be highlighted
@@ -975,9 +987,6 @@ def changedetection_app(config=None, datastore_o=None):

        # create a ZipFile object
        backupname = "changedetection-backup-{}.zip".format(int(time.time()))
-
-        # We only care about UUIDS from the current index file
-        uuids = list(datastore.data['watching'].keys())
        backup_filepath = os.path.join(datastore_o.datastore_path, backupname)

        with zipfile.ZipFile(backup_filepath, "w",
@@ -993,12 +1002,12 @@ def changedetection_app(config=None, datastore_o=None):
            # Add the flask app secret
            zipObj.write(os.path.join(datastore_o.datastore_path, "secret.txt"), arcname="secret.txt")

-            # Add any snapshot data we find, use the full path to access the file, but make the file 'relative' in the Zip.
-            for txt_file_path in Path(datastore_o.datastore_path).rglob('*.txt'):
-                parent_p = txt_file_path.parent
-                if parent_p.name in uuids:
-                    zipObj.write(txt_file_path,
-                                 arcname=str(txt_file_path).replace(datastore_o.datastore_path, ''),
+            # Add any data in the watch data directory.
+            for uuid, w in datastore.data['watching'].items():
+                for f in Path(w.watch_data_dir).glob('*'):
+                    zipObj.write(f,
+                                 # Use the full path to access the file, but make the file 'relative' in the Zip.
+                                 arcname=os.path.join(f.parts[-2], f.parts[-1]),
                                 compress_type=zipfile.ZIP_DEFLATED,
                                 compresslevel=8)

--- a/changedetectionio/api/api_v1.py
+++ b/changedetectionio/api/api_v1.py
@@ -122,3 +122,37 @@ class CreateWatch(Resource):
            return {'status': "OK"}, 200

        return list, 200
+
+class SystemInfo(Resource):
+    def __init__(self, **kwargs):
+        # datastore is a black box dependency
+        self.datastore = kwargs['datastore']
+        self.update_q = kwargs['update_q']
+
+    @auth.check_token
+    def get(self):
+        import time
+        overdue_watches = []
+
+        # Check all watches and report which have not been checked but should have been
+
+        for uuid, watch in self.datastore.data.get('watching', {}).items():
+            # see if now - last_checked is greater than the time that should have been
+            # this is not super accurate (maybe they just edited it) but better than nothing
+            t = watch.threshold_seconds()
+            if not t:
+                # Use the system wide default
+                t = self.datastore.threshold_seconds
+
+            time_since_check = time.time() - watch.get('last_checked')
+
+            # Allow 5 minutes of grace time before we decide it's overdue
+            if time_since_check - (5 * 60) > t:
+                overdue_watches.append(uuid)
+
+        return {
+                   'queue_size': self.update_q.qsize(),
+                   'overdue_watches': overdue_watches,
+                   'uptime': round(time.time() - self.datastore.start_time, 2),
+                   'watch_count': len(self.datastore.data.get('watching', {}))
+               }, 200
--- a/changedetectionio/changedetection.py
+++ b/changedetectionio/changedetection.py
@@ -102,6 +102,14 @@ def main():
                    has_password=datastore.data['settings']['application']['password'] != False
                    )

+    # Monitored websites will not receive a Referer header
+    # when a user clicks on an outgoing link.
+    @app.after_request
+    def hide_referrer(response):
+        if os.getenv("HIDE_REFERER", False):
+            response.headers["Referrer-Policy"] = "no-referrer"
+        return response
+
    # Proxy sub-directory support
    # Set environment var USE_X_SETTINGS=1 on this script
    # And then in your proxy_pass settings
--- a/changedetectionio/content_fetcher.py
+++ b/changedetectionio/content_fetcher.py
@@ -575,6 +575,11 @@ class html_requests(Fetcher):
            ignore_status_codes=False,
            current_css_filter=None):

+        # Make requests use a more modern looking user-agent
+        if not 'User-Agent' in request_headers:
+            request_headers['User-Agent'] = os.getenv("DEFAULT_SETTINGS_HEADERS_USERAGENT",
+                                                      'Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/87.0.4280.66 Safari/537.36')
+
        proxies = {}

        # Allows override the proxy on a per-request basis
--- a/changedetectionio/download.zip
+++ b/changedetectionio/download.zip
--- a/changedetectionio/fetch_site_status.py
+++ b/changedetectionio/fetch_site_status.py
@@ -35,6 +35,8 @@ class perform_site_check():


    def run(self, uuid):
+        from jinja2 import Environment
+
        changed_detected = False
        screenshot = False  # as bytes
        stripped_text_from_html = ""
@@ -65,7 +67,11 @@ class perform_site_check():
            request_headers['Accept-Encoding'] = request_headers['Accept-Encoding'].replace(', br', '')

        timeout = self.datastore.data['settings']['requests'].get('timeout')
-        url = watch.get('url')
+
+        # Jinja2 available in URLs along with https://pypi.org/project/jinja2-time/
+        jinja2_env = Environment(extensions=['jinja2_time.TimeExtension'])
+        url = str(jinja2_env.from_string(watch.get('url')).render())
+
        request_body = self.datastore.data['watching'][uuid].get('body')
        request_method = self.datastore.data['watching'][uuid].get('method')
        ignore_status_codes = self.datastore.data['watching'][uuid].get('ignore_status_codes', False)
--- a/changedetectionio/forms.py
+++ b/changedetectionio/forms.py
@@ -303,12 +303,16 @@ class ValidateCSSJSONXPATHInput(object):

                # Re #265 - maybe in the future fetch the page and offer a
                # warning/notice that its possible the rule doesnt yet match anything?
-
-            if 'jq:' in line:
                if not self.allow_json:
                    raise ValidationError("jq not permitted in this field!")

-                import jq
+            if 'jq:' in line:
+                try:
+                    import jq
+                except ModuleNotFoundError:
+                    # `jq` requires full compilation in windows and so isn't generally available
+                    raise ValidationError("jq not support not found")
+
                input = line.replace('jq:', '')

                try:
--- a/changedetectionio/html_tools.py
+++ b/changedetectionio/html_tools.py
@@ -1,12 +1,11 @@
-import json
-from typing import List

 from bs4 import BeautifulSoup
-from jsonpath_ng.ext import parse
-import jq
-import re
 from inscriptis import get_text
 from inscriptis.model.config import ParserConfig
+from jsonpath_ng.ext import parse
+from typing import List
+import json
+import re

 class FilterNotFoundInResponse(ValueError):
    def __init__(self, msg):
@@ -85,9 +84,18 @@ def _parse_json(json_data, json_filter):
        jsonpath_expression = parse(json_filter.replace('json:', ''))
        match = jsonpath_expression.find(json_data)
        return _get_stripped_text_from_json_match(match)
+
    if 'jq:' in json_filter:
+
+        try:
+            import jq
+        except ModuleNotFoundError:
+            # `jq` requires full compilation in windows and so isn't generally available
+            raise Exception("jq not support not found")
+
        jq_expression = jq.compile(json_filter.replace('jq:', ''))
        match = jq_expression.input(json_data).all()
+
        return _get_stripped_text_from_json_match(match)

 def _get_stripped_text_from_json_match(match):
--- a/changedetectionio/model/App.py
+++ b/changedetectionio/model/App.py
@@ -13,10 +13,6 @@ class model(dict):
            'watching': {},
            'settings': {
                'headers': {
-                    'User-Agent': getenv("DEFAULT_SETTINGS_HEADERS_USERAGENT", 'Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/87.0.4280.66 Safari/537.36'),
-                    'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,image/apng,*/*;q=0.8,application/signed-exchange;v=b3;q=0.9',
-                    'Accept-Encoding': 'gzip, deflate',  # No support for brolti in python requests yet.
-                    'Accept-Language': 'en-GB,en-US;q=0.9,en;'
                },
                'requests': {
                    'timeout': int(getenv("DEFAULT_SETTINGS_REQUESTS_TIMEOUT", "45")),  # Default 45 seconds
--- a/changedetectionio/model/Watch.py
+++ b/changedetectionio/model/Watch.py
@@ -1,6 +1,8 @@
-import os
-import uuid as uuid_builder
 from distutils.util import strtobool
+import logging
+import os
+import time
+import uuid

 minimum_seconds_recheck_time = int(os.getenv('MINIMUM_SECONDS_RECHECK_TIME', 60))
 mtable = {'seconds': 1, 'minutes': 60, 'hours': 3600, 'days': 86400, 'weeks': 86400 * 7}
@@ -22,7 +24,7 @@ class model(dict):
            #'newest_history_key': 0,
            'title': None,
            'previous_md5': False,
-            'uuid': str(uuid_builder.uuid4()),
+            'uuid': str(uuid.uuid4()),
            'headers': {},  # Extra headers to send
            'body': None,
            'method': 'GET',
@@ -60,7 +62,7 @@ class model(dict):
        self.update(self.__base_config)
        self.__datastore_path = kw['datastore_path']

-        self['uuid'] = str(uuid_builder.uuid4())
+        self['uuid'] = str(uuid.uuid4())

        del kw['datastore_path']

@@ -82,10 +84,9 @@ class model(dict):
        return False

    def ensure_data_dir_exists(self):
-        target_path = os.path.join(self.__datastore_path, self['uuid'])
-        if not os.path.isdir(target_path):
-            print ("> Creating data dir {}".format(target_path))
-            os.mkdir(target_path)
+        if not os.path.isdir(self.watch_data_dir):
+            print ("> Creating data dir {}".format(self.watch_data_dir))
+            os.mkdir(self.watch_data_dir)

    @property
    def label(self):
@@ -109,16 +110,40 @@ class model(dict):

    @property
    def history(self):
+        """History index is just a text file as a list
+            {watch-uuid}/history.txt
+
+            contains a list like
+
+            {epoch-time},{filename}\n
+
+            We read in this list as the history information
+
+        """
        tmp_history = {}
-        import logging
-        import time

        # Read the history file as a dict
-        fname = os.path.join(self.__datastore_path, self.get('uuid'), "history.txt")
+        fname = os.path.join(self.watch_data_dir, "history.txt")
        if os.path.isfile(fname):
            logging.debug("Reading history index " + str(time.time()))
            with open(fname, "r") as f:
-                tmp_history = dict(i.strip().split(',', 2) for i in f.readlines())
+                for i in f.readlines():
+                    if ',' in i:
+                        k, v = i.strip().split(',', 2)
+
+                        # The index history could contain a relative path, so we need to make the fullpath
+                        # so that python can read it
+                        if not '/' in v and not '\'' in v:
+                            v = os.path.join(self.watch_data_dir, v)
+                        else:
+                            # It's possible that they moved the datadir on older versions
+                            # So the snapshot exists but is in a different path
+                            snapshot_fname = v.split('/')[-1]
+                            proposed_new_path = os.path.join(self.watch_data_dir, snapshot_fname)
+                            if not os.path.exists(v) and os.path.exists(proposed_new_path):
+                                v = proposed_new_path
+
+                        tmp_history[k] = v

        if len(tmp_history):
            self.__newest_history_key = list(tmp_history.keys())[-1]
@@ -129,7 +154,7 @@ class model(dict):

    @property
    def has_history(self):
-        fname = os.path.join(self.__datastore_path, self.get('uuid'), "history.txt")
+        fname = os.path.join(self.watch_data_dir, "history.txt")
        return os.path.isfile(fname)

    # Returns the newest key, but if theres only 1 record, then it's counted as not being new, so return 0.
@@ -148,31 +173,27 @@ class model(dict):
    # Save some text file to the appropriate path and bump the history
    # result_obj from fetch_site_status.run()
    def save_history_text(self, contents, timestamp):
-        import uuid
-        import logging
-
-        output_path = "{}/{}".format(self.__datastore_path, self['uuid'])

        self.ensure_data_dir_exists()
+        snapshot_fname = "{}.txt".format(str(uuid.uuid4()))

-        snapshot_fname = "{}/{}.stripped.txt".format(output_path, uuid.uuid4())
-        logging.debug("Saving history text {}".format(snapshot_fname))
-
-        with open(snapshot_fname, 'wb') as f:
+        # in /diff/ and /preview/ we are going to assume for now that it's UTF-8 when reading
+        # most sites are utf-8 and some are even broken utf-8
+        with open(os.path.join(self.watch_data_dir, snapshot_fname), 'wb') as f:
            f.write(contents)
            f.close()

        # Append to index
        # @todo check last char was \n
-        index_fname = "{}/history.txt".format(output_path)
+        index_fname = os.path.join(self.watch_data_dir, "history.txt")
        with open(index_fname, 'a') as f:
            f.write("{},{}\n".format(timestamp, snapshot_fname))
            f.close()

        self.__newest_history_key = timestamp
-        self.__history_n+=1
+        self.__history_n += 1

-        #@todo bump static cache of the last timestamp so we dont need to examine the file to set a proper ''viewed'' status
+        # @todo bump static cache of the last timestamp so we dont need to examine the file to set a proper ''viewed'' status
        return snapshot_fname

    @property
@@ -205,14 +226,14 @@ class model(dict):
        return not local_lines.issubset(existing_history)

    def get_screenshot(self):
-        fname = os.path.join(self.__datastore_path, self['uuid'], "last-screenshot.png")
+        fname = os.path.join(self.watch_data_dir, "last-screenshot.png")
        if os.path.isfile(fname):
            return fname

        return False

    def __get_file_ctime(self, filename):
-        fname = os.path.join(self.__datastore_path, self['uuid'], filename)
+        fname = os.path.join(self.watch_data_dir, filename)
        if os.path.isfile(fname):
            return int(os.path.getmtime(fname))
        return False
@@ -237,9 +258,14 @@ class model(dict):
    def snapshot_error_screenshot_ctime(self):
        return self.__get_file_ctime('last-error-screenshot.png')

+    @property
+    def watch_data_dir(self):
+        # The base dir of the watch data
+        return os.path.join(self.__datastore_path, self['uuid'])
+    
    def get_error_text(self):
        """Return the text saved from a previous request that resulted in a non-200 error"""
-        fname = os.path.join(self.__datastore_path, self['uuid'], "last-error.txt")
+        fname = os.path.join(self.watch_data_dir, "last-error.txt")
        if os.path.isfile(fname):
            with open(fname, 'r') as f:
                return f.read()
@@ -247,7 +273,7 @@ class model(dict):

    def get_error_snapshot(self):
        """Return path to the screenshot that resulted in a non-200 error"""
-        fname = os.path.join(self.__datastore_path, self['uuid'], "last-error-screenshot.png")
+        fname = os.path.join(self.watch_data_dir, "last-error-screenshot.png")
        if os.path.isfile(fname):
            return fname
        return False
--- a/changedetectionio/run_all_tests.sh
+++ b/changedetectionio/run_all_tests.sh
@@ -9,6 +9,8 @@
 # exit when any command fails
 set -e

+SCRIPT_DIR=$( cd -- "$( dirname -- "${BASH_SOURCE[0]}" )" &> /dev/null && pwd )
+
 find tests/test_*py -type f|while read test_name
 do
  echo "TEST RUNNING $test_name"
@@ -23,6 +25,13 @@ export BASE_URL="https://really-unique-domain.io"
 pytest tests/test_notification.py


+## JQ + JSON: filter test
+# jq is not available on windows and we should just test it when the package is installed
+# this will re-test with jq support
+pip3 install jq~=1.3
+pytest tests/test_jsonpath_jq_selector.py
+
+
 # Now for the selenium and playwright/browserless fetchers
 # Note - this is not UI functional tests - just checking that each one can fetch the content

@@ -38,7 +47,9 @@ docker kill $$-test_selenium

 echo "TESTING WEBDRIVER FETCH > PLAYWRIGHT/BROWSERLESS..."
 # Not all platforms support playwright (not ARM/rPI), so it's not packaged in requirements.txt
-pip3 install playwright~=1.24
+PLAYWRIGHT_VERSION=$(grep -i -E "RUN pip install.+" "$SCRIPT_DIR/../Dockerfile" | grep --only-matching -i -E "playwright[=><~+]+[0-9\.]+")
+echo "using $PLAYWRIGHT_VERSION"
+pip3 install "$PLAYWRIGHT_VERSION"
 docker run -d --name $$-test_browserless -e "DEFAULT_LAUNCH_ARGS=[\"--window-size=1920,1080\"]" --rm  -p 3000:3000  --shm-size="2g"  browserless/chrome:1.53-chrome-stable
 # takes a while to spin up
 sleep 5
--- a/changedetectionio/store.py
+++ b/changedetectionio/store.py
@@ -30,14 +30,14 @@ class ChangeDetectionStore:
    def __init__(self, datastore_path="/datastore", include_default_watches=True, version_tag="0.0.0"):
        # Should only be active for docker
        # logging.basicConfig(filename='/dev/stdout', level=logging.INFO)
-        self.needs_write = False
+        self.__data = App.model()
        self.datastore_path = datastore_path
        self.json_store_path = "{}/url-watches.json".format(self.datastore_path)
+        self.needs_write = False
        self.proxy_list = None
+        self.start_time = time.time()
        self.stop_thread = False

-        self.__data = App.model()
-
        # Base definition for all watchers
        # deepcopy part of #569 - not sure why its needed exactly
        self.generic_definition = deepcopy(Watch.model(datastore_path = datastore_path, default={}))
@@ -575,3 +575,11 @@ class ChangeDetectionStore:
                continue
        return

+
+    # We incorrectly used common header overrides that should only apply to Requests
+    # These are now handled in content_fetcher::html_requests and shouldnt be passed to Playwright/Selenium
+    def update_7(self):
+        # These were hard-coded in early versions
+        for v in ['User-Agent', 'Accept', 'Accept-Encoding', 'Accept-Language']:
+            if self.data['settings']['headers'].get(v):
+                del self.data['settings']['headers'][v]
--- a/changedetectionio/templates/edit.html
+++ b/changedetectionio/templates/edit.html
@@ -40,7 +40,8 @@
                <fieldset>
                    <div class="pure-control-group">
                        {{ render_field(form.url, placeholder="https://...", required=true, class="m-d") }}
-                        <span class="pure-form-message-inline">Some sites use JavaScript to create the content, for this you should <a href="https://github.com/dgtlmoon/changedetection.io/wiki/Fetching-pages-with-WebDriver">use the Chrome/WebDriver Fetcher</a></span>
+                        <span class="pure-form-message-inline">Some sites use JavaScript to create the content, for this you should <a href="https://github.com/dgtlmoon/changedetection.io/wiki/Fetching-pages-with-WebDriver">use the Chrome/WebDriver Fetcher</a></span><br/>
+                        <span class="pure-form-message-inline">You can use variables in the URL, perfect for inserting the current date and other logic, <a href="https://github.com/dgtlmoon/changedetection.io/wiki/Handling-variables-in-the-watched-URL">help and examples here</a></span><br/>
                    </div>
                    <div class="pure-control-group">
                        {{ render_field(form.title, class="m-d") }}
@@ -184,10 +185,14 @@ User-Agent: wonderbra 1.0") }}
                        <span class="pure-form-message-inline">
                    <ul>
                        <li>CSS - Limit text to this CSS rule, only text matching this CSS rule is included.</li>
-                        <li>JSON - Limit text to this JSON rule, using either <a href="https://pypi.org/project/jsonpath-ng/" target="new">JSONPath</a> or <a href="https://stedolan.github.io/jq/" target="new">jq</a>.
+                        <li>JSON - Limit text to this JSON rule, using either <a href="https://pypi.org/project/jsonpath-ng/" target="new">JSONPath</a> or <a href="https://stedolan.github.io/jq/" target="new">jq</a> (if installed).
                            <ul>
                                <li>JSONPath: Prefix with <code>json:</code>, use <code>json:$</code> to force re-formatting if required,  <a href="https://jsonpath.com/" target="new">test your JSONPath here</a>.</li>
+                                {% if jq_support %}
                                <li>jq: Prefix with <code>jq:</code> and <a href="https://jqplay.org/" target="new">test your jq here</a>. Using <a href="https://stedolan.github.io/jq/" target="new">jq</a> allows for complex filtering and processing of JSON data with built-in functions, regex, filtering, and more. See examples and documentation <a href="https://stedolan.github.io/jq/manual/" target="new">here</a>.</li>
+                                {% else %}
+                                <li>jq support not installed</li>
+                                {% endif %}
                            </ul>
                        </li>
                        <li>XPath - Limit text to this XPath rule, simply start with a forward-slash,
@@ -198,7 +203,7 @@ User-Agent: wonderbra 1.0") }}
                            </ul>
                            </li>
                    </ul>
-                    Please be sure that you thoroughly understand how to write CSS, JSONPath, XPath, or jq selector rules before filing an issue on GitHub! <a
+                    Please be sure that you thoroughly understand how to write CSS, JSONPath, XPath{% if jq_support %}, or jq selector{%endif%} rules before filing an issue on GitHub! <a
                                href="https://github.com/dgtlmoon/changedetection.io/wiki/CSS-Selector-help">here for more CSS selector help</a>.<br/>
                </span>
                    </div>
--- a/changedetectionio/tests/test_api.py
+++ b/changedetectionio/tests/test_api.py
@@ -147,6 +147,16 @@ def test_api_simple(client, live_server):
    # @todo how to handle None/default global values?
    assert watch['history_n'] == 2, "Found replacement history section, which is in its own API"

+    # basic systeminfo check
+    res = client.get(
+        url_for("systeminfo"),
+        headers={'x-api-key': api_key},
+    )
+    info = json.loads(res.data)
+    assert info.get('watch_count') == 1
+    assert info.get('uptime') > 0.5
+
+
    # Finally delete the watch
    res = client.delete(
        url_for("watch", uuid=watch_uuid),
--- a/changedetectionio/tests/test_backup.py
+++ b/changedetectionio/tests/test_backup.py
@@ -1,18 +1,31 @@
 #!/usr/bin/python3

-import time
+from .util import set_original_response, set_modified_response, live_server_setup
 from flask import url_for
 from urllib.request import urlopen
-from . util import set_original_response, set_modified_response, live_server_setup
+from zipfile import ZipFile
+import re
+import time


 def test_backup(client, live_server):
-
    live_server_setup(live_server)

+    set_original_response()
+
    # Give the endpoint time to spin up
    time.sleep(1)

+    # Add our URL to the import page
+    res = client.post(
+        url_for("import_page"),
+        data={"urls": url_for('test_endpoint', _external=True)},
+        follow_redirects=True
+    )
+
+    assert b"1 Imported" in res.data
+    time.sleep(3)
+
    res = client.get(
        url_for("get_backup"),
        follow_redirects=True
@@ -20,6 +33,19 @@ def test_backup(client, live_server):

    # Should get the right zip content type
    assert res.content_type == "application/zip"
+
    # Should be PK/ZIP stream
    assert res.data.count(b'PK') >= 2

+    # ZipFile from buffer seems non-obvious, just save it instead
+    with open("download.zip", 'wb') as f:
+        f.write(res.data)
+
+    zip = ZipFile('download.zip')
+    l = zip.namelist()
+    uuid4hex = re.compile('^[a-f0-9]{8}-?[a-f0-9]{4}-?4[a-f0-9]{3}-?[89ab][a-f0-9]{3}-?[a-f0-9]{12}.*txt', re.I)
+    newlist = list(filter(uuid4hex.match, l))  # Read Note below
+
+    # Should be two txt files in the archive (history and the snapshot)
+    assert len(newlist) == 2
+
--- a/changedetectionio/tests/test_jinja2.py
+++ b/changedetectionio/tests/test_jinja2.py
@@ -0,0 +1,33 @@
+#!/usr/bin/python3
+
+import time
+from flask import url_for
+from .util import live_server_setup
+
+
+# If there was only a change in the whitespacing, then we shouldnt have a change detected
+def test_jinja2_in_url_query(client, live_server):
+    live_server_setup(live_server)
+
+    # Give the endpoint time to spin up
+    time.sleep(1)
+
+    # Add our URL to the import page
+    test_url = url_for('test_return_query', _external=True)
+
+    # because url_for() will URL-encode the var, but we dont here
+    full_url = "{}?{}".format(test_url,
+                              "date={% now 'Europe/Berlin', '%Y' %}.{% now 'Europe/Berlin', '%m' %}.{% now 'Europe/Berlin', '%d' %}", )
+    res = client.post(
+        url_for("form_quick_watch_add"),
+        data={"url": full_url, "tag": "test"},
+        follow_redirects=True
+    )
+    assert b"Watch added" in res.data
+    time.sleep(3)
+    # It should report nothing found (no new 'unviewed' class)
+    res = client.get(
+        url_for("preview_page", uuid="first"),
+        follow_redirects=True
+    )
+    assert b'date=2' in res.data
--- a/changedetectionio/tests/test_jsonpath_jq_selector.py
+++ b/changedetectionio/tests/test_jsonpath_jq_selector.py
@@ -5,7 +5,12 @@ import time
 from flask import url_for, escape
 from . util import live_server_setup
 import pytest
+jq_support = True

+try:
+    import jq
+except ModuleNotFoundError:
+    jq_support = False

 def test_setup(live_server):
    live_server_setup(live_server)
@@ -40,22 +45,24 @@ and it can also be repeated
    assert text == "23.5"

    # also check for jq
-    text = html_tools.extract_json_as_string(content, "jq:.offers.price")
-    assert text == "23.5"
+    if jq_support:
+        text = html_tools.extract_json_as_string(content, "jq:.offers.price")
+        assert text == "23.5"
+
+        text = html_tools.extract_json_as_string('{"id":5}', "jq:.id")
+        assert text == "5"

    text = html_tools.extract_json_as_string('{"id":5}', "json:$.id")
    assert text == "5"

-    text = html_tools.extract_json_as_string('{"id":5}', "jq:.id")
-    assert text == "5"
-
    # When nothing at all is found, it should throw JSONNOTFound
    # Which is caught and shown to the user in the watch-overview table
    with pytest.raises(html_tools.JSONNotFound) as e_info:
        html_tools.extract_json_as_string('COMPLETE GIBBERISH, NO JSON!', "json:$.id")

-    with pytest.raises(html_tools.JSONNotFound) as e_info:
-        html_tools.extract_json_as_string('COMPLETE GIBBERISH, NO JSON!', "jq:.id")
+    if jq_support:
+        with pytest.raises(html_tools.JSONNotFound) as e_info:
+            html_tools.extract_json_as_string('COMPLETE GIBBERISH, NO JSON!', "jq:.id")

 def set_original_ext_response():
    data = """
@@ -271,7 +278,8 @@ def test_check_jsonpath_filter(client, live_server):
    check_json_filter('json:boss.name', client, live_server)

 def test_check_jq_filter(client, live_server):
-    check_json_filter('jq:.boss.name', client, live_server)
+    if jq_support:
+        check_json_filter('jq:.boss.name', client, live_server)

 def check_json_filter_bool_val(json_filter, client, live_server):
    set_original_response()
@@ -329,7 +337,8 @@ def test_check_jsonpath_filter_bool_val(client, live_server):
    check_json_filter_bool_val("json:$['available']", client, live_server)

 def test_check_jq_filter_bool_val(client, live_server):
-    check_json_filter_bool_val("jq:.available", client, live_server)
+    if jq_support:
+        check_json_filter_bool_val("jq:.available", client, live_server)

 # Re #265 - Extended JSON selector test
 # Stuff to consider here
@@ -408,4 +417,5 @@ def test_check_jsonpath_ext_filter(client, live_server):
    check_json_ext_filter('json:$[?(@.status==Sold)]', client, live_server)

 def test_check_jq_ext_filter(client, live_server):
-    check_json_ext_filter('jq:.[] | select(.status | contains("Sold"))', client, live_server)
+    if jq_support:
+        check_json_ext_filter('jq:.[] | select(.status | contains("Sold"))', client, live_server)
--- a/changedetectionio/tests/util.py
+++ b/changedetectionio/tests/util.py
@@ -159,5 +159,10 @@ def live_server_setup(live_server):
        ret = " ".join([auth.username, auth.password, auth.type])
        return ret

+    # Just return some GET var
+    @live_server.app.route('/test-return-query', methods=['GET'])
+    def test_return_query():
+        return request.query_string
+
    live_server.start()

--- a/changedetectionio/tests/visualselector/test_fetch_data.py
+++ b/changedetectionio/tests/visualselector/test_fetch_data.py
@@ -13,9 +13,9 @@ def test_visual_selector_content_ready(client, live_server):
    live_server_setup(live_server)
    time.sleep(1)

-    # Add our URL to the import page, maybe better to use something we control?
-    # We use an external URL because the docker container is too difficult to setup to connect back to the pytest socket
-    test_url = 'https://news.ycombinator.com'
+    # Add our URL to the import page, because the docker container (playwright/selenium) wont be able to connect to our usual test url
+    test_url = "https://changedetection.io/ci-test/test-runjs.html"
+
    res = client.post(
        url_for("form_quick_watch_add"),
        data={"url": test_url, "tag": '', 'edit_and_watch_submit_button': 'Edit > Watch'},
@@ -25,13 +25,27 @@ def test_visual_selector_content_ready(client, live_server):

    res = client.post(
        url_for("edit_page", uuid="first", unpause_on_save=1),
-        data={"css_filter": ".does-not-exist", "url": test_url, "tag": "", "headers": "", 'fetch_backend': "html_webdriver"},
+        data={
+              "url": test_url,
+              "tag": "",
+              "headers": "",
+              'fetch_backend': "html_webdriver",
+              'webdriver_js_execute_code': 'document.querySelector("button[name=test-button]").click();'
+        },
        follow_redirects=True
    )
    assert b"unpaused" in res.data
    time.sleep(1)
    wait_for_all_checks(client)
    uuid = extract_UUID_from_client(client)
+
+    # Check the JS execute code before extract worked
+    res = client.get(
+        url_for("preview_page", uuid="first"),
+        follow_redirects=True
+    )
+    assert b'I smell JavaScript' in res.data
+
    assert os.path.isfile(os.path.join('test-datastore', uuid, 'last-screenshot.png')), "last-screenshot.png should exist"
    assert os.path.isfile(os.path.join('test-datastore', uuid, 'elements.json')), "xpath elements.json data should exist"

--- a/docker-compose.yml
+++ b/docker-compose.yml
@@ -45,6 +45,9 @@ services:
  #        Respect proxy_pass type settings, `proxy_set_header Host "localhost";` and `proxy_set_header X-Forwarded-Prefix /app;`
  #        More here https://github.com/dgtlmoon/changedetection.io/wiki/Running-changedetection.io-behind-a-reverse-proxy-sub-directory
  #      - USE_X_SETTINGS=1
+  #
+  #        Hides the `Referer` header so that monitored websites can't see the changedetection.io hostname.
+  #      - HIDE_REFERER=true

      # Comment out ports: when using behind a reverse proxy , enable networks: etc.
      ports:
--- a/requirements.txt
+++ b/requirements.txt
@@ -1,8 +1,8 @@
-flask~= 2.0
+flask ~= 2.0
 flask_wtf
-eventlet>=0.31.0
+eventlet >= 0.31.0
 validators
-timeago ~=1.0
+timeago ~= 1.0
 inscriptis ~= 2.2
 feedgen ~= 0.9
 flask-login ~= 0.5
@@ -10,13 +10,17 @@ flask_restful
 pytz

 # Set these versions together to avoid a RequestsDependencyWarning
-requests[socks] ~= 2.26
+# >= 2.26 also adds Brotli support if brotli is installed
+brotli ~= 1.0
+requests[socks] ~= 2.28
+
 urllib3 > 1.26
 chardet > 2.3.0

 wtforms ~= 3.0
 jsonpath-ng ~= 1.5.3
-jq ~= 1.3.0
+
+# jq not available on Windows so must be installed manually

 # Notification library
 apprise ~= 1.1.0
@@ -42,4 +46,9 @@ selenium ~= 4.1.0
 # need to revisit flask login versions
 werkzeug ~= 2.0.0

+# Templating, so far just in the URLs but in the future can be for the notifications also
+jinja2
+jinja2-time
+
 # playwright is installed at Dockerfile build time because it's not available on all platforms
+
Author	SHA1	Message	Date
dgtlmoon	cb7b12ab9c	fix dupe import	2022-10-27 11:23:46 +02:00
dgtlmoon	197a3e6366	cleanup of backup creation	2022-10-27 11:21:55 +02:00
dgtlmoon	9c1b18ec22	Use a rglob that will work cross-platform more reliably	2022-10-27 11:05:39 +02:00
dgtlmoon	02b460768c	Support moving the data set	2022-10-27 10:56:10 +02:00
dgtlmoon	50edc36cc4	Improve test	2022-10-27 10:29:28 +02:00
dgtlmoon	3cf2a798cf	Remove dupe code	2022-10-26 11:09:49 +02:00
dgtlmoon	13daa92c92	Make snapshots path relative so the datadir is portable	2022-10-26 10:49:37 +02:00
dgtlmoon	53f4472fb8	Merge branch 'master' of github.com:dgtlmoon/changedetection.io into history-txt-snapshot-fix	2022-10-25 11:45:18 +02:00
dgtlmoon	724cb17224	Re #1052 - Dynamic URLs, use variables in the URL (such as the current date, the date in a month, and other logic see https://github.com/dgtlmoon/changedetection.io/wiki/Handling-variables-in-the-watched-URL ) (#1057 )	2022-10-24 23:20:39 +02:00
dgtlmoon	a1bdf7e43f	Backup zip - always use correct relative path of the snapshot text instead of the datadir path	2022-10-24 15:53:25 +02:00
dgtlmoon	4eb4b401a1	API - system info - allow 5 minutes grace before watch is considered 'overdue'	2022-10-23 23:12:28 +02:00
dgtlmoon	5d40e16c73	API - Adding basic system info/system state API (#1051 )	2022-10-23 19:15:11 +02:00
dgtlmoon	492bbce6b6	Build - Fix syntax in container build test (#1050 )	2022-10-23 16:02:13 +02:00
dgtlmoon	0394a56be5	Building - Test container build on PR	2022-10-23 15:54:19 +02:00
Entepotenz	7839551d6b	Testing - Use same version of playwright while running tests as in production builds (#1047 )	2022-10-23 11:26:32 +02:00
Entepotenz	9c5588c791	update path for validation in the CONTRIBUTING.md (#1046 )	2022-10-23 11:25:29 +02:00
dgtlmoon	5a43a350de	History index safety check - Be sure that only valid history index lines are read (#1042 )	2022-10-19 22:41:13 +02:00
Michael McMillan	3c31f023ce	Option to Hide the Referer header from monitored websites. (#996 )	2022-10-18 09:16:22 +02:00
dgtlmoon	4cbcc59461	0.39.20.4	2022-10-17 18:36:47 +02:00
dgtlmoon	4be0260381	Better cross platform file handling in diff and preview (#1034 )	2022-10-17 18:36:22 +02:00
dgtlmoon	957a3c1c16	0.39.20.3	2022-10-17 17:43:35 +02:00
dgtlmoon	85897e0bf9	Windows - diff file handling improvements (#1031 )	2022-10-17 17:40:28 +02:00
dgtlmoon	63095f70ea	Also include tests in pip build	2022-10-17 17:13:15 +02:00
dgtlmoon	8d5b0b5576	Update README.md	2022-10-12 10:51:39 +02:00
dgtlmoon	1b077abd93	0.39.20.2	2022-10-12 09:53:59 +02:00
dgtlmoon	32ea1a8721	Windows - JQ - Make library optional so it doesnt break Windows pip installs (#1009 )	2022-10-12 09:53:16 +02:00
dgtlmoon	fff32cef0d	Adding test - Test the 'execute JS before changedetection' (#1006 )	2022-10-11 14:40:36 +02:00
dgtlmoon	8fb146f3e4	0.39.20.1	2022-10-09 23:05:35 +02:00
dgtlmoon	770b0faa45	Code - check containers build when Dockerfile or requirements.txt changes (#1005 )	2022-10-09 22:58:01 +02:00
dgtlmoon	f6faa90340	Adding `make` to Dockerfile build as required by jq for ARM devices	2022-10-09 22:29:18 +02:00
dgtlmoon	669fd3ae0b	Dont use default Requests `user-agent` and `accept` headers in playwright+selenium requests, breaks sites such as united.com. (#1004 )	2022-10-09 18:25:36 +02:00