Files
changedetection.io/changedetectionio/translations
dgtlmoonandClaude Opus 5 ba2280ba81 External CDP Browser engine - replaces the extra_browser_* hackery
An "extra browser" was a name + a ws(s):// endpoint in settings.requests.extra_browsers,
selected by a watch as the magic string 'extra_browser_<name>'. That string resolved to
html_webdriver plus a custom connection URL, which meant the protocol the endpoint was
spoken to came from env vars rather than from the entry: CDP over a WebSocket with
PLAYWRIGHT_DRIVER_URL set, CDP via pyppeteer with FAST_PUPPETEER_CHROME_FETCHER, and the
W3C WebDriver protocol over HTTP on a Selenium-only install - where a wss:// URL cannot
work at all. The form only ever accepted ws:// / wss://, so the feature was silently
broken on exactly the installs that could not honour it.

So it becomes an engine, html_external_cdp, which pins the protocol: a subclass of the
Playwright fetcher that takes its endpoint from the watch's browser config
(FetcherConfig.connection_url) instead of the environment. It is base-only
(ready_to_use=False) because an endpoint is required, so each endpoint is one browser
config ("variation") on the Browsers page - which is what the old settings list was.

update_36 migrates each extra_browsers row to such a variation, keyed by the SAME
'extra_browser_<name>' string watches already hold, so no watch, group override, API value
or global default needs rewriting; the legacy selector simply becomes a real browser-config
id. A row whose endpoint the model rejects is logged and skipped rather than taking the
update chain, and with it startup, down.

Knock-on cleanups, all of which delete a special case rather than add one:

 - The proxy opt-out for custom endpoints is now Fetcher.ignores_proxy_setting, asked of
   the engine, instead of a string-prefix test in call_browser().
 - A live browser-steps / visual-selector session asks the engine where to connect
   (Fetcher.browser_steps_connection_url, overridden by html_external_cdp) and refuses an
   engine whose supports_browser_steps is False, instead of reading the env var itself and
   silently stepping a browser the watch does not check with. That refusal is real: on a
   Selenium install html_webdriver cannot drive a live session.
 - is_valid_browser_selector() answers "may a watch store this in fetch_backend?" in one
   place; the API (create/update/import), the quick-add form validator and the bulk "set
   browser" operation each had their own copy, which is how they came to disagree about
   whether a browser-config id was acceptable.
 - api-spec.yaml's fetch_backend pattern enumerated extra_browser_* while rejecting
   browser-config ids and every engine newer than html_webdriver. Valid values are
   per-install, so the schema now bounds the string and the handlers do the real check.
 - html_external_cdp registers unconditionally (unlike html_playwright_builtin): migrated
   configs name it, so it must resolve even without the playwright library, or those
   watches would quietly fetch with the plain HTTP client. The library is imported lazily
   inside run(), and an unavailable engine now warns instead of falling back silently.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-17 19:39:16 +02:00
..

Translators Guide

This document is for contributors who write templates (HTML) and for translators who maintain .po files. It exists because fragmented msgids — splitting a single sentence across multiple _() calls — cause systematic translation breakage across many languages. Follow the patterns here to prevent that.


Terminology

  • Always use "monitor" or "watcher" for the concept of watching a URL — never the bare word "watch", which translates to "clock" (e.g. hodinky in Czech, 시계 in Korean, 時計 in Japanese).
  • Use the shortest suitable wording for each language. If a language naturally uses the English derivative, prefer that.

Template rules: do not fragment msgids

Why fragments break translation

The GNU gettext manual is explicit on this:

Entire sentences: Translatable strings should be entire sentences. Because gender/number declension depends on other parts of the sentence, half-sentence "dumb string concatenation" breaks in many languages other than English.

No string concatenation: Placing adjacent _() calls is semantically equivalent to runtime strcat concatenation, so the same guideline applies. The manual also notes that "in some languages the translator might want to swap the order" of components.

No embedded URLs: URLs should not be written directly inside msgids; they should be injected via %(name)s placeholders and values passed as kwargs.

No unusual markup: "HTML markup, however, is common enough that it's probably ok to use in translatable strings."

Fragments break differently depending on language family:

Language family How fragmentation breaks it
SOV (Japanese, Korean, Turkish) Verb-final word order can't be achieved when verb and subject are in separate fragments
Germanic (German) Gender/case agreement between article and noun is lost across fragment boundaries
Romance (French, Spanish, Italian, Portuguese) Adjective placement, subjunctive mood, verb agreement can't be maintained
Slavic (Czech, Ukrainian) Case (driven by preposition/verb relationships) is easy to get wrong
CJK (Chinese, Japanese, Korean) Modifier position and SVO-vs-topic-prominent differences can't be applied at fragment level

A past workaround was redistributing translations across adjacent fragments and using msgstr " " (a single space) to suppress unused fragments. This is fragile: as soon as the same short msgid is reused in a different template, the redistributed translation is applied verbatim and breaks the new context.


The four correct patterns

Pattern 1 — Inline HTML embedding

Keep markup inside the msgid. Render with | safe. This also lets CJK translators decide how to handle <i> (see CJK section below).

{# BAD: three fragments; CJK translators can't see the <i> at all #}
{{ _('Helps reduce changes detected caused by sites shuffling lines around, combine with') }}
<i>{{ _('check unique lines') }}</i>
{{ _('below.') }}

{# GOOD: one msgid, rendered with |safe #}
{{ _('Helps reduce changes detected caused by sites shuffling lines around, combine with <i>check unique lines</i> below.') | safe }}

Pattern 2 — URL as kwarg

Pass URLs via %(name)s so translators can freely reorder them.

{# BAD: URL hardcoded between three fragments #}
{{ _('Use') }}
<a target="newwindow" href="https://github.com/caronc/apprise">{{ _('AppRise Notification URLs') }}</a>
{{ _('for notification to just about any service!') }}

{# GOOD: URL passed as kwarg, <a> embedded in the msgid #}
{{ _('Use <a target="newwindow" href="%(url)s">AppRise Notification URLs</a> for notification to just about any service!',
     url='https://github.com/caronc/apprise') | safe }}

Pattern 3 — Literal {{}} escape as kwarg

Jinja2 would double-interpolate {{token}} inside a _() call. Pass it as a kwarg instead.

{# BAD: literal {{token}} in the middle forces splitting #}
{{ _('Accepts the') }} <code>{{ '{{token}}' }}</code> {{ _('placeholders listed below') }}

{# GOOD: literal passed as kwarg; msgid stays as an entire sentence #}
{{ _('Accepts the <code>%(token)s</code> placeholders listed below', token='{{token}}') | safe }}

Pattern 4 — {% if %} outside the msgid

Move conditional branches outside _() so each branch is a complete sentence, not a fragment.

{# BAD: three fragments; SOV languages can't reorder %(title)s relative to "URL or Title" #}
{{ _('URL or Title') }}{% if active_tag_uuid %} {{ _('in') }} '{{ active_tag.title }}'{% endif %}

{# GOOD: branch between two complete msgids; each language can freely reorder %(title)s #}
{% if active_tag_uuid %}
  {{ _("URL or Title in '%(title)s'", title=active_tag.title) }}
{% else %}
  {{ _('URL or Title') }}
{% endif %}

CJK italic policy

CJK fonts typically have no true italic cut — <i> falls back to a mechanical slant that reduces legibility. Now that <i> is inside msgids, CJK translators can handle it per-locale. Apply this policy for ja / zh / zh_Hant_TW:

Context Action
<i> used for general emphasis Replace with <strong>, or drop if the emphasis is self-evident
<strong><i>...</i></strong> Collapse to <strong>...</strong>
<i> wrapping a UI term (e.g. "check unique lines") Wrap in locale-conventional quotation marks: 「」 for ja/zh_Hant_TW, "" for zh

Translation workflow

Always use these commands — they read consistent settings from setup.cfg and produce minimal diffs:

python setup.py extract_messages   # Extract translatable strings from source
python setup.py update_catalog     # Propagate new msgids to all .po files
python setup.py compile_catalog    # Compile .po files to binary .mo files

Running pybabel directly without the project options causes reordering, rewrapping, and line-number churn that makes diffs hard to review.

Configuration

All translation settings are in setup.cfg (single source of truth):

[extract_messages]
mapping_file = babel.cfg
output_file = changedetectionio/translations/messages.pot
input_paths = changedetectionio
keywords = _ _l gettext
sort_by_file = true       # Keeps entries ordered by file path
width = 120               # Consistent line width (prevents rewrapping)
add_location = file       # Show file path only (not line numbers)

[update_catalog]
input_file = changedetectionio/translations/messages.pot
output_dir = changedetectionio/translations
domain = messages
width = 120
no_fuzzy_matching = true  # Avoids incorrect automatic matches

[compile_catalog]
directory = changedetectionio/translations
domain = messages

Multi-language fix process

When you find a translation error in any language, you must check all others for the same msgid:

for lang in cs de en_GB en_US es fr it ja ko pt_BR ru tr uk zh zh_Hant_TW; do
  echo "=== $lang ===" && grep -A1 'msgid "YourString"' changedetectionio/translations/$lang/LC_MESSAGES/messages.po
done
  1. Identify every language with the same problem
  2. Fix all affected .po files in the same session
  3. Recompile: python setup.py compile_catalog

Never fix one language and move on.


Supported languages

Code Language
cs Czech (Čeština)
de German (Deutsch)
en_GB English (UK)
en_US English (US)
es Spanish (Español)
fr French (Français)
id Indonesian (Bahasa Indonesia)
it Italian (Italiano)
ja Japanese (日本語)
ko Korean (한국어)
pt_BR Portuguese (Brasil)
ru Russian (Русский)
tr Turkish (Türkçe)
uk Ukrainian (Українська)
zh Chinese Simplified (中文简体)
zh_Hant_TW Chinese Traditional (繁體中文)

Adding a new language

pybabel init -i changedetectionio/translations/messages.pot \
             -d changedetectionio/translations \
             -l NEW_LANG_CODE
# Reset POT-Creation-Date to the sentinel so it matches the other catalogs
sed -i 's|^"POT-Creation-Date: .*\\n"$|"POT-Creation-Date: 1970-01-01 00:00+0000\\n"|' \
  changedetectionio/translations/NEW_LANG_CODE/LC_MESSAGES/messages.po
python setup.py compile_catalog

Babel auto-discovers the new language on subsequent runs.


Dennis linter

We use mozilla/dennis to enforce technical correctness in .po and .pot files. See the Table of Warnings and Errors for the full list of rules.

Running the linter locally

To match the CI checks, run the following commands:

# Check for errors only (always enforced)
dennis-cmd lint --errorsonly changedetectionio/translations/

# Check for warnings (excluding W302 unchanged translations)
dennis-cmd lint --excluderules=W302 changedetectionio/translations/

Common problems and resolutions

HTML tag mismatch (W303)

The W303 rule ensures that HTML tags in the msgstr match the msgid. This is crucial for catching broken markup (e.g., missing closing tags).

Handling intentional deviations

Some W303 warnings are intentional. Use the dennis-ignore: W303 comment in the source files (templates or Python code) within a TRANSLATORS comment to suppress these warnings. This ensures the ignore instruction is extracted into the .po files.

  • CJK italic policy: When replacing <i> with locale-conventional quotation marks, tags will no longer match.

Examples in Jinja2 templates:

{# TRANSLATORS: CJK fonts lack native italics; allow substitution with conventional local styling. dennis-ignore: W303 #}
<p>{{ _('These settings are <strong><i>added</i></strong> to any existing watch configurations.')|safe }}</p>

Example in Python source:

# dennis-ignore: W303 - CJK fonts lack native italics; allow substitution with conventional local styling.
message = StringField(_l('This is <i>experimental</i> and may change'))

CI linter

A GitHub Actions job (lint-template-i18n) checks for adjacent {{ _(...) }} calls on the same line separated only by HTML — the primary symptom of fragmented msgids. It enforces a declining baseline: the count of existing violations may only go down, never up. When you fix a template, lower the BASELINE_LIMIT in .github/workflows/test-only.yml by the number of lines you fixed.

See issue #4074 for full background and PR #4076 for worked consolidation examples.