* feat: allow browser extension origins in auth requests
Add chrome-extension://, moz-extension://, safari-extension://, safari-web-extension://, and extension:// to the protocol allow-list so browser extensions can obtain app tokens via /auth/get-user-app-token.
Extract WEB_AND_EXTENSION_PROTOCOLS constant in validation.js so the allow-list is defined once and shared by AuthService, AppStore, and AppDriver.
* fix: harden the extension-origin allow-list
Review follow-ups on the extension-origin change:
- Require a host in `validateUrl`. Only "special" schemes need an
authority, so `chrome-extension:` parsed with an empty hostname and
slipped past the reserved-system-host guard in AppDriver.
- Lowercase the host in `AuthService#normalizedOrigin`. `new URL()`
lowercases http(s) hosts but leaves opaque ones alone, so one
extension in two spellings resolved to two app uids — two AppData
trees, two permission sets — and missed the origin blocklist.
- Drop `extension:`. No browser emits it, and it accepted
`extension://evil.com` as an app origin.
- Freeze the allow-list and derive `validateUrl`'s http(s) default from
`WEB_PROTOCOLS` so the two spellings can't drift apart.
- Pin the tests to the real uid derivation, use unique extension ids so
a row left by another test can't mask the bootstrap path, and cover
the host-less and near-miss schemes.
- Fix the prettier/eslint failure in AppStore.js.
---------
Co-authored-by: Daniel Salazar <daniel.salazar@puter.com>
* fix: new cache costs for claude
* fix: refuse app-issued scoped tokens on events handler and worker routes
An access token an app mints (e.g. a getReadURL() token) resolves
effectiveApp to the issuing app, so #handlerApp and listEventsWorkers
treated it as the app: it could publish, remove and list handlers, and
list or destroy the app's worker, regardless of its permission manifest.
Refuse any scoped token there, as the kv-handle routes already do.
* test: freeze the clock before the first backoff read in workerInvoker
The first hold was read on real time, so a slow run measured the 2s
backoff as 1s.
A public origin-bootstrap app row planted on an alternate hosting domain
could resolve ahead of the owner's private app and skip the private
access gate (PUT-1883).
- get-user-app-token canonicalizes hosted-subdomain origins before uid
derivation and bootstrap, so every hosting variant shares one app row
- the app create/update conflict check crosses hosting-domain variants,
so an existing alt-host stub is absorbed by the owner's create
- resolveOwnedAppForHostedSite and the canonical-uid lookup order
matches private-first with a deterministic id tiebreak
* feat(ai): sync GPT-6 Sol, GPT-6 Luna, and Claude Opus 5.5
* chore: remove documentation changes from model sync
* fix(ai): bill OpenAI cache writes and long-context pricing
GPT-5.6 and later bill prompt-cache writes at 1.25x input and report them
in `cache_write_tokens`, inside the input total. Both OpenAI calculators now
split them out of the prompt count and meter them under their own key, and
the six GPT-5.6+ models carry the rate.
GPT-6, GPT-5.6, GPT-5.5 and GPT-5.4 (incl. Pro) bill a request with more
than 272K input tokens at 2x input (cached reads and cache writes included)
and 1.5x output for the whole request. Models declare this as
`long_context_pricing`, and the ledger overrides, the reported `usd_cents`
and the credit gate all apply the multipliers.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* fix(ai): bill GPT-5.6 Sol at OpenAI's promotional pricing
The catalog carried GPT-5.5's $5/$0.50/$30, but OpenAI bills GPT-5.6 Sol
at $4/$0.40/$20 per million input/cached/output tokens, promotional
through at least 2026-11-21. PUT-1943 tracks re-checking before then.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* fix(ai): send stable, non-sequential user identifiers to AI providers
A precedence bug in the AI providers' identifier expression made every
request send `user: ":undefined"` (the ternary bound the app-uid suffix to
the whole `actor.user.id + actor.app?.uid` sum instead of just the suffix),
or read `actor.user.id` on a missing user. The same expression also shipped
the sequential internal user id, letting AI vendors correlate a single
account across apps and sessions.
All eight OpenAI-, Azure-, xAI-, Meta- and ZAI-style providers now build
the identifier through one shared helper, `aiUserIdentifier()`:
- `puter-<user-uuid>[-<app-token>]`: the random user UUID is always
preserved in full; `maxLength` constrains only the app-bearing form
- app attribution reads `effectiveApp`, so access-token requests name the
issuing app instead of looking like direct user traffic
- the app token is truncated to fit the budget, and omitted entirely when
the remaining budget is below 8 chars, where a truncation could collide
with another app's uid
- nothing is sent for the system actor
- Meta and ZAI keep a caller-supplied `safety_identifier` / `user_id`
override, applied before the helper result
`user` is deprecated by OpenAI; the SDK types direct callers to
`safety_identifier` (abuse detection) and `prompt_cache_key` (cache-hit
bucketing). The four OpenAI/Azure chat providers and MetaProvider now send
`prompt_cache_key` as well, defaulting it to the same per-user identifier
unless the caller supplies one; Azure's Grok branch drops both fields,
matching its rejection of unknown args. The cap comment cites only verified
limits: OpenAI's 64 for `safety_identifier` (from the SDK types) and Z.AI's
6-128 for `user_id` (from Z.AI's docs); Meta and xAI document none, so none
is claimed.
The xAI image `#edit` path now carries the identifier like generation, and
takes a named-options param so `user` cannot be transposed with the
adjacent same-typed `aspectRatio`.
Tests share a four-actor matrix (`user` / `user+app` / `access token` /
`system`) with `assertActorMatrixIdentifiers()` across the six
OpenAI-style suites; the helper has exact-string and boundary coverage
(size caps, zero-budget and sub-base cases, no dangling separator, UUID
never truncated, collision guard); the Azure Grok assertions run under a
real user actor so they cannot pass vacuously. 212 provider-suite tests
pass; typecheck and ESLint are clean.
* fix(ai): lock the vendor identifier down and keep vitest out of the test util
Meta and Z.AI no longer let `custom` override the abuse identifier; it is
Puter's attribution, not the caller's. The shared test util exposes pure
field pickers instead of importing vitest into a file the production
tsconfig compiles. The helper's length-cap comment now matches vendor docs
(Meta does cap `safety_identifier` at 64), the redundant budget branch and
the unused export are gone, and the per-user `prompt_cache_key` trade-off is
stated once in the helper instead of five times in providers.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
---------
Co-authored-by: 404oops <me@404oops.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
* fix(ai-image): sync image catalogs, unify request shape, document models
Model sync against every vendor listing and a live generation sweep:
Gemini stable ids replace the retired preview spellings (2.5 Flash Image
delisted ahead of its 2026-10-02 shutdown), OpenAI gpt-image-1/-1-mini/-1.5
are delisted but routable until their shutdown dates with gpt-image-2 as
the default, xAI gains grok-imagine-image-2.0 with its quality tiers,
BytePlus gains the 3K/4K tiers and per-model pixel bounds, Cloudflare
gains SDXL Lightning/Base and SD 1.5 Inpainting, Replicate gains a
schema-driven catalog of 88 additional models with version-pinned
community predictions, and every Together image route is excluded for
the third-party data-sharing requirement.
Request normalization: the driver collapses ratio/width/height/aspect_ratio
into one imageSize with aspect-versus-pixel intent, validates prompt,
quality and resolution once, resolves provider hints (short names or
full driver ids) and aliases with exact ids winning over resellers, and
hands each provider an immutable copy of the caller's args. Providers
share prompt validation, aspect snapping, and content-sniffed data URIs
so responses carry the right MIME type. Replicate predictions are created
once, cancelled on abort or deadline, and bounded in input fan-out.
SDK: txt2img copies caller options, rejects blank prompts with
prompt_required in every call form, and documents that normalize has no
effect because the image return shape is already uniform.
Docs: txt2img options per provider, provider defaults and discovery,
availability notes, a new image model catalog and pricing page, and the
image-generation limits in the quotas page.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* fix(ai-image): sniff SVG outputs correctly, reject hex/exponent dimensions
imageDataUri only decoded the first 24 base64 chars (18 bytes) before
sniffImageMime, which cannot see an <svg> root behind an XML prolog -
SVG outputs from recraft-v*-svg models were being labeled image/png.
Decode enough of the payload to cover the 8 KB SVG sniff window.
dimension() accepted string number literals ('0x10', '1e3') as if they
were decimal dimensions; restrict to plain decimal notation.
* docs(ai): link to the model directory instead of a static catalog
The image model catalog and pricing page duplicated the always up to date
model directory at developer.puter.com/ai/models. Link to the directory
from txt2img-related docs and keep the Together data-sharing exclusion
note self-contained.
---------
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Chat and video drivers keyed their provider and model maps on plain objects,
so `model` or `provider` values such as `__proto__` reached Object.prototype
and surfaced as 500s. Both maps are now null-prototype objects and reject
those names as ordinary unknown models.
txt2vid now rejects blank or non-string prompts with `prompt_required` before
any request, tolerates `null` in either argument slot, and copies the caller's
options before resolving the `duration` alias and output path so frozen
option objects work and caller objects are never mutated.
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Emits ai.cost.multiplier.<driver>.<provider>:<model> before recording AI
usage, so what a model costs to charge is policy an extension owns rather
than a number in core. Nothing listening records the provider cost.
MeteringService.withAiCostMultiplier(driver) returns a view of the service
whose recording paths scale costOverride by the hook's answer; every AI
driver hands that view to its providers, so all of them are covered without
touching provider code.
* fix: list a link share only while the owner's plan covers it
* fix: fs limits for signed urls
* feat: a paid plan counts as a verified card for the sharing gate
feat: the team seat experience — no email required, forced password change, team label, and plan-based limits (PUT-1792)
Note: Bypassing the code owners rule, since there are couple approvals in place for this.
Review catch by @Salazareo: `org_seat_free` reached `FREE_SUBSCRIPTION_IDS`
and so the `requireSubscription` gate, but two other surfaces decide on
plan and neither consults that set.
`bySubscription` maps name `user_free` and `temp_free`. A plan that
matches no key fell through to the top-level `limit` -- the paid cap --
so a seat outranked an ordinary free account: 240 event listings a minute
against their 120, and the same shape across the kv, notification,
subdomain and worker drivers. Both resolvers now fall back to the
`user_free` entry for anything in the free set, which covers every driver
at once and any free plan added later.
`subscriberOnly` compared against the two named ids, so a seat could
reach a paid-only model. It asks the set now.
A paid plan that names no cap of its own still takes the base, and an
unresolved plan still takes the base; there are tests for both so the
fallback cannot widen into "free by default".
- PUT-1800: gate `createWorkerSessionToken` on actor type, so an app or an
access token can no longer mint an app-less, root-shaped worker session;
`WorkerDriver` binds on `effectiveApp` instead of `app`.
- PUT-1799: add `isAccountContext` and read it where "no app" was being read
as "the account" — handler publish, events-worker listing, kv handle
mint/revoke/list. A scoped API token is no longer an account session.
- PUT-1802: re-authorize a durable row before its backlog drains, settling it
permanently when the grant is gone. Covers an ancestor-level unshare, which
the revoke settle deliberately leaves to the delivery re-check.
- PUT-1803: mask the owner's absolute path out of deliveries and subscription
anchors on a foreign node, the way every FS surface already does.
- PUT-1804: let a revoke reach rows already suspended for a resumable reason,
re-stamping them so a resume cannot hand over the held backlog.
- PUT-1805: apply the subscribe path's audience gate to `/events/fetch` before
the query, so a cursor can no longer count and name invisible notifications.
- PUT-1807: refuse `mode: 'manage'` from any actor holding an app — inside its
own AppData the ACL short-circuit would otherwise supply the reach.
- PUT-1808: take the sending peer from the verified signature header rather
than the request body.
- PUT-1810: re-base a kv share-handle row's stored match filter on the handle,
so the owner's absolute key prefix stays hidden.
- PUT-1814: escape LIKE wildcards and anchor the issuer-prefix queries on a
segment; anchor `manage:` stripping; reject a backslash in a share prefix;
assert a resolved actor in `subscribeDurable`.
- PUT-1815: bound the char/varchar columns behind `event_subscriptions` and
`kv_share_handles` at the store layer.
Infron sells the same model at several service tiers. `min_prompt_price`
and `min_completion_price` are the floor across all of them, so every
model with a flex tier was advertised at a batch-job price nobody gets
by default: 25 of 286 chat models, among them gpt-6-astra at $5/$25
against the $7.5/$37.5 a default request actually bills.
Price each tier from its own row in `providers[]` and pin that tier on
the request, so the price quoted is the price charged. Tiers beyond the
default are listed under their own `<model>:<tier>` ids, letting callers
opt into flex or priority by model name:
puter.ai.chat(prompt, { model: 'infron:openai/gpt-6-astra:flex' })
Suffix parsing matches the catalog exactly before reading a trailing
segment as a tier, since catalog ids can carry a colon of their own
(`deepseek/deepseek-v4-flash:free`). Tier variants take the context
window of their own offering, which differs from the model-level figure
for 11 of them.
Billing was never wrong — it bills Infron's reported `cost` — but the
understated prices fed the credit gate's output cap, which let a request
run roughly 50% past the balance it was gated against, and the fallback
path that prices per token when a response carries no cost.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
OpenAI shuts the Sora Videos API down on 2026-09-24 and sora-2 was the
default txt2vid model, so the default moves to Veo 3.1 Lite on Gemini
and the OpenAI video provider goes. While there, the video catalogs are
brought in line with what each vendor serves today, the request options
are unified across providers, and the txt2vid docs are rewritten.
- driver: default provider gemini-video-generation with
veo-3.1-lite-generate-preview; a request under the generic `ai-video`
driver name lands on the default instead of the first-registered
provider; `WIDTHxHEIGHT` sizes map onto tier catalogs by the shorter
side and fill width/height
- openai video provider, the `openai-video-generation` alias, its
config template and migration entries, and Together's openai/sora-2*
rows removed
- gemini: Veo 3.1 Fast rates 10/12/30 cents per second for
720p/1080p/4K, Veo 3.1 Lite accepts reference images, URL image
inputs are fetched server-side through the SSRF-guarded fetch
- together: drop nine models retired upstream, add eighteen from the
live listing; per-second models are estimated from the catalog rate,
clamped to remaining credit and billed at the cost Together reports
on the job; tier-sized models take resolution/ratio;
input_reference/last_frame map onto keyframes; generate_audio is
forwarded
- byteplus: Seedance 2.5 (dreamina-seedance-2-5-260628) with per-model
reference-image caps
- util/imageInput: string-level image helpers shared by the image and
video drivers; ai-image/inputImage re-exports them unchanged
- puter.js types: provider and generate_audio options; docs: txt2vid
page rewritten with per-provider model tables, unified options and
four new playground examples
Known follow-up: Veo returns a key-protected Google file URL, so the
default clip cannot be played directly by a browser until the provider
fetches it server-side.
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
When the chat fallback chain is exhausted, the driver already records
each attempt (model, provider, status, code, message, timeout) in the
error's `fields.attempts`, but the alarm keyed on the classified message
alone and the alarm client printed the error at inspect depth 2, so the
log and Slack line read `internal_error:All providers failed` with
nothing about which providers failed or why. Deduped repeats printed
only a count.
- The HTTP alarm gate now attaches an HttpError's `fields` to the alarm
under a single `details` key. One key can't shadow the gate's own
request fields, and a repeat from another thrower on a shared id
replaces it instead of merging into it.
- The chat driver logs one warn line per exhausted chain with the
completion id, the resolved route, the classified code and the
attempts as JSON, so every occurrence is greppable by trace even when
the alarm dedupes it. Chains marked `noAlarm` don't log.
- Docs: `fields` reaches both the client and the alarm, so it has to be
safe to show the caller.
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
The SDK shorthand sends `{ image_url: { url } }` with no type. Claude
forwarded it untouched (Anthropic: "content.1.type: Field required"),
the OpenAI/Azure Responses providers sent a Chat Completions part the
Responses API rejects, and Mistral emitted snake_case `image_url` where
its SDK validates `imageUrl`. All three only appeared to work because
the fallback loop re-served them through OpenRouter.
- utils/mediaParts: canonicalise every inbound media part (untyped
shorthand, bare string URLs, Responses input_image, Anthropic image
blocks, Gemini inline_data) to the Chat Completions form in the
driver, so providers translate from one shape
- Claude → Anthropic image blocks (url / base64 sources), copy-on-write
- Responses processor → input_image with a string URL and detail, and
copies messages instead of mutating the caller's objects
- Mistral → camelCase imageUrl; the test that asserted the old shape
is corrected
- driver: 400 up front when the catalog says the model has no image or
video input; resolve puter_path parts per attempt for providers
without their own upload path (Claude keeps the Files API)
- move the Moonshot http→data-URL inliner to utils/inlineImages and
apply it to Grok-via-Azure and Gemini, whose upstreams cannot fetch
http image URLs
Fixes#3409
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
- puter.email.sendTransactional is the new name; send stays as a
deprecated alias with the same arguments and result
- fileInput: openFileInputStream / resolveFileInputEntry expose the
ACL-checked FS read as a stream; loadFileInput wraps them
- EmailAttachment accepts a `path` the transport streams on its own
Google shut down imagen-4.0-fast/standard/ultra in the Gemini API on
2026-08-17; live calls now return 404 "not found ... or is not supported
for predict". Drop the three entries, the generateImages code path that
only they used, and their tests; the integration test moves to
gemini-2.5-flash-image.
Moonshot's live /models listing and the Kimi pricing docs now carry only
kimi-k3, kimi-k2.7-code(-highspeed) and kimi-k2.6. Drop kimi-k2.5 and the
whole moonshot-v1-* family and repoint the unit and integration tests at
current models. kimi-k2.5 still resolves through OpenRouter.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* fix: harden events dispatch, single delivery and KV share handles
Dispatch: a filtered subscription used the anchor path stored at subscribe
time, so renaming or moving the anchor folder silently ended its deliveries;
dispatch now resolves the anchor's live path from the event's own ancestor
chain. A move out of a watched folder now reaches that folder's subscribers,
with `from` only for rows that watched the source side. Gap markers are
authorized like deliveries and coalesced per subscription and subject instead
of fanning per lost event. Session subscriptions: the per-socket cap decides
on the write, not before it; an orphaned watched-set token heals on refresh;
durable rows keep their watch window when a session subscribe touches the
same keys. `self` is false when the acting user is unknown.
Single delivery: a subscription in backoff or suspended with a backlog pinned
the sweeper's head and starved everyone behind it — the sweep now defers it.
Only a settled handler run bills a delivery. A socket-only account row no
longer wedges after two attempts nobody received. The lease is twice the
handler timeout; remote candidates have their own attempt counter; the region
depth reconcile runs once a minute region-wide with a bounded scan.
KV share handles: a grantee no longer sees the owner's namespace and absolute
prefix on the subscribe answer or listing, nor in the delivery token; revoking
a wider handle retires the handles it covers; minting the same handle twice
returns the existing one, after the delegation check; a row whose event
cannot be re-based onto its handle is dropped rather than delivered raw.
* fix: presence survives replication, long sessions and region churn
One presence item per (user, app) with per-region map fields lost a region
whenever two regions joined inside the replication window, and nothing ever
put it back. Presence is now one item per (user, app, region): each region
writes only its own, a leave or repair retires it conditionally on its own
write stamp, and a read is a prefix query. Items carry a 48 h ttl refreshed by
a claim-gated write off the existing socket renew path, at most once per
12 h, so a tab that stays connected keeps its region in the row. A region
that answered "no socket" or completed a leave releases a shared pin, so a
reconnect on another node rejoins and a flapping client cannot force a
replicated write per cycle. Cached rows expire after a minute; unaddressable
region names are filtered and pruned; relayed acks settle under a bounded
concurrency; the forward queue is bounded in bytes as well as items.
* feat: indexes for the event_subscriptions hot queries
Handler publish, remove and listing, and the hourly expiry and suspension
sweeps, all scanned `event_subscriptions`. Adds (app_uid, handler_name),
(expires_at) and (suspended_at, id), guarded on every engine. Existing
migrations: the postgres widens are now guarded so a boot does not take an
exclusive lock for a no-op, the kv_share_handles grantee FK gets an index,
the sqlite notification rebuild is transactional and idempotent.
* fix: notification writes go through the registry
The driver's `create` bypassed the type registry, producing uncatalogued
rows with no size bound; it now requires a registered type, caps the payload,
and answers 400 rather than 500 for a bad one. `mark_acknowledged` emits the
ack other tabs listen for, and only when a row was actually changed.
* fix: the handler scanner, unsubscribe, and the in-tab handler environment
The free-variable scanner skipped arrows inside a declaration's initializer,
so `const ids = event.items.map(x => x.id)` was refused, and treated a name
after a comma in a nested initializer as bound, so a real free variable slipped
through to fail on first delivery. `unsubscribe()` now drops the durable
routing entry so the events socket can close. A broadcast handler running in
the tab gets `user` and `fetch` like the worker gives it. `single` without
a handler name is refused before the round trip.
* docs: events limits, error codes and the background-workers section
Retention is deployment-configured rather than a fixed 14 days, and the
template no longer ships it armed. Documents `events_terminal`, the two
per-event gap reasons, the subject length and listing caps, the `from` field
on moves, and the handle-relative anchor. The sessions manager hides the
background-workers section when the server has none to show.
* feat: a background handler acts as the app does for its user
A handler's `user` was a five-minute access token scoped to the subscription's
`list` grant, which could stat the changed file but not read it, and could
not reach the app's KV or AppData — so an app told that a file was written
could do nothing with it. It now runs with the same authority the app has for
that user in a tab: an app-under-user worker session, one row per (user, app)
named `events:handlers`, visible and revocable in the sessions list. The
`events:background` consent is what authorizes running it unattended, and is
re-checked before every mint.
The wider token exposed two things: puter.js opens a filesystem socket the
moment it has a token, which would have parked the isolate in the app's own
delivery room and steered deliveries at it; the events client now opts out of
sockets (and the per-open bookkeeping) before construction, and is memoized
per token in the isolate. And four filesystem operations assumed a socket
exists; they no longer do.
* feat: bake published handlers into a generated events worker
* feat: deploy and address the per-app events worker behind a flag
* test: single delivery end to end through a real local worker
* feat: events workers run their own runtime, in their own namespace
An events worker was being deployed as an ordinary worker: default dispatch
namespace, a `subdomains` row, the router preamble, and an app-scoped worker
token baked in. The public dispatcher resolves any script in that namespace
straight off the hostname, so the worker answered at `<name>.puter.work`, and
the only thing in front of it was an unguessable name plus a check that a
`puter-auth` header was present — which the router never validates. Anyone who
learned the hostname could run an app's handlers with a body of their choosing,
in an isolate holding the owner's token as `me`.
Instead:
- Handlers run on their own runtime (`src/worker/src/events-runtime.js`), which
provides no `router` and no `me`, owns the single invoke route, and hands a
handler only `{ event, ctx, user, fetch, ack }`. `user` is built from the
invocation's delivery token, so a handler acts as the subscriber whose
delivery it is and nothing wider. The preamble build emits one bundle per
runtime; the shared half of the template is now included by both.
- The deploy target carries the runtime to prepend, the source to deploy, and
whether to mint a worker token at all, so an events worker deploys into the
`events` dispatch namespace from generated source with no token binding, no
`subdomains` row, and no claim on the owner's worker quota or worker list.
- An invocation carries a key derived from the deployment secret and the script
name, bound as a secret and checked in constant time inside the isolate,
which reads it once and drops it before handler code runs.
- Scripts are named after the handler set they contain, so publishing writes
rows and deploys nothing: a set is deployed the first time a delivery needs
it, and a changed set is a new script rather than an overwrite of a running
one. Publish responses keep the shape they had before the runtime existed.
- Invocations reach a worker only through the events dispatcher, which has no
zone route and requires the internal secret; the backend's own deploy path is
the rehydrate route the dispatcher calls on a namespace miss. Locally there is
no dispatcher, so the controller hands the service an in-process transport
that deploys on miss itself.
The SDK stops allowlisting `puter` as a handler global — a handler that reaches
for an ambient SDK is now refused at publish time, naming `user` instead, rather
than passing the scan and failing on its first delivery.
Requires `events.workerNamespace`, `events.dispatcherUrl` and
`events.internalSecret`; without them nothing is addressable and background
deliveries stay retriable, as they did with the runtime off.
* fix: a handler's delivery token gets through the read routes
An events handler acts as the subscriber through the access token its
invocation carried, but every FS read route refused scoped access tokens
outright, so `user.fs.stat(event.path)` — the design's own example — answered
403 inside the worker. The read-side routes now admit them; the ACL each
handler already runs intersects the token's grant with its issuer's, which is
the check that keeps a token to what it was minted for. The end-to-end suite
asserts the stat from inside the isolate.
* fix: shorthand-method handlers publish as functions
`{ ingest({ event }) { … } }` stringifies without the `function` keyword, so
its source is not an expression and the events worker baked it as a broken
stub — every delivery a retriable 500 until the subscription suspended, with
nothing at publish time to say why. The SDK now gives a shorthand method the
keyword before hashing and sending; getters, setters and computed names are
left for the server-side check to refuse.
* feat: an app's events worker is listable and destroyable
An app with published handlers has an events worker, and hosted deployments
bill it monthly per app, so its owner needs to see it and be able to take it
down. The core announces the lifecycle on the bus — `events.worker.create`
when an app's first handler is published, `events.worker.destroy` when its last
one goes — with the owner as the actor, so pricing can plug in from outside.
`GET /events/workers` lists the caller's workers (paginated, with the script
each set deploys as) and `POST /events/workers/destroy` removes every handler
of an app under the same owner scoping as the handler routes, suspending the
subscriptions bound to them. `puter.events.workers.list/destroy` in the SDK,
a docs page, and a 5 MB cap on an app's combined handler source
(`events_worker_too_large`) so a set that publishes can always deploy.
* fix: harden the events worker runtime for production
- A 4xx is terminal only when it carries the handled marker the runtime (and
the dispatcher) stamp on every answer that came from a script; an unmarked
4xx — an edge 404 for a wrong dispatcher hostname, a WAF page — stays
retriable and is logged, once per script per minute, with the runtime's
reason header.
- Script names are scoped to this backend's exposed API origin, so two
backends sharing a namespace never resolve one script with the wrong
endpoint binding or key. Shape unchanged.
- Each handler is validated in the exact context it is emitted into and the
whole generated file is compiled once; a source that would break the script
marks every handler broken instead of deploying a SyntaxError.
- Locally, events scripts live under their own registry key: the public local
worker host cannot reach them and an ordinary worker cannot take their name.
- A suspended or deleted app owner stops invocations; deploys are throttled
per app per hour; in-flight deploys are keyed by app and script; the
upstream deploy call times out; the generated source is size-capped with a
margin over the publish cap; boot fails when the runtime is on but its
preamble is not built. Byte-length secret compare, appUid shape check,
dispatcher URL prefix preserved, wider connection pool.
* feat: background workers are listed in the sessions manager
A user paying for an app's events worker needs somewhere to see it and take it
down. The sessions manager gets a section listing the apps that run event
handlers in the background, with a Destroy action that removes their published
handlers.
* fix(ai): make chat fallback reach streamed Claude calls and rank Azure explicitly
- ClaudeProvider opens the upstream stream and awaits its connection
before returning the populator, so an overloaded or rate-limited route
throws from complete() and reaches the driver's fallback loop instead
of surfacing as an error frame on a 200
- the OpenAI-compatible chat and completions routes only pin a provider
when the caller sent one, so they get the same preferred healthy route
puter.js callers do and unhealthy-route skipping applies to their
first attempt
- Azure is ranked ahead of the vendors it fronts by an explicit tier in
modelRouting rather than a price tie plus registration order
- drop Together's synthetic always-failing model-fallback-test-1 entry
- test that a 4xx leaves a route in rotation while a 503 marks it
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* fix(ai): surface swallowed Claude stream errors, keep compat-route defaults
Review follow-ups on the fallback work.
- The pre-created event iterator only receives an error if a reader is
already waiting on it, so a failure landing between the connect and the
populator's first pull ended the stream cleanly — truncated content
billed and reported as a success. Rethrow when the stream is errored.
- A refused stream deleted its Anthropic uploads but left the caller's
message parts pointing at those file ids, so the fallback route was
handed handles it cannot resolve. processPuterPathUploads now returns a
restore() that both failure paths call.
- The OpenAI-compat routes keep pinning OpenAI when the caller sends no
model at all, so the default model stays put instead of moving to
Azure's.
- Say why /openai/v1/responses and /anthropic/v1/messages stay pinned:
each translates one provider's native shape by hand.
- PREFERRED_PROVIDERS is unexported and its doc now states the rank is
unconditional; the duplicated hidden-model list is one constant.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(ai): undo the puter_path rewrite by field instead of snapshotting the part
Copying the content part kept whatever the caller sent on it — a large
inline `source` or `text` alongside `puter_path` — reachable until the
request ended, where overwriting the field used to make it garbage right
away. The only fields this function writes are `type`/`source` on success
and `type`/`text` on failure, and the fallback uploader keys off
`puter_path` alone, so restore undoes those three by name.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Widen the notif: match filter and fetch scope for a session's own
generic developer/app-user subscribe: today it pins ref to the
session's own uuid, so a row naming an app (handler-suspension
notices, app-bound worker deploys) never matches live and never
replays on reconnect, even though the audience predicate already
grants the holder every such row it owns. The predicate is the
authority and already reruns per row/page after the match, so
widening the filter to it (account is unaffected — it never names an
app) adds no exposure.
An actor holding an app reads the `app-user` rows naming that app, plus its
`developer` rows when the holder owns it. `account` rows reach no app, and a
slice an actor may not see comes back empty rather than refused. The audience
predicate becomes the enforced read path in the same change that lifts the
blanket app-actor 403, layered behind an audience/app_uid SQL scope; two-segment
`notif:` subjects expand server-side from the actor's own app, so an app can
never name another app's uid.
No feature flag: `audience` defaults to 'account', so every pre-registry row is
default-denied to app actors and the backfill can only narrow.
A video job that outlived its poll window, an SDK request that timed out,
or a Veo operation that finished with an error all reached the HTTP error
handler as plain Errors. Each became an unhandled 500 with critical
severity and paged on-call for what is the provider's pace or the
provider's fault.
Video providers now share one poll loop that gives up with a 504
`upstream_timeout`, treats a transient poll failure (timeout, dropped
connection, 408/429/5xx) as a missed poll rather than a failed job, and
stops polling with a 400 `client_aborted` when the caller disconnects, so
nothing is metered for a clip nobody will receive. The driver controller
exposes the disconnect as an `abortSignal` on the request context. The
window is ten minutes for every provider; Together and BytePlus move up
from five.
Failed jobs are classified: content-filter refusals become a 400
`bad_request` with `errorCode: moderation_flagged`, rejected parameters a
400 `upstream_bad_request`, and anything else a 502 `upstream_failed`,
each carrying the provider's own code. Veo's filtered output keeps
`disallowed_value` and gains the same `errorCode`. The sanitizer and
content-filter pattern move from the Replicate provider into a shared
util so image and video agree.
Status-less SDK connection timeouts are translated to a 504
`upstream_timeout` at the driver boundary, and the chat driver records
them per attempt so an all-timeout chain is a 504 and a mixed chain is
`upstream_failed` instead of an `internal_error` 500. The Together chat
client gets the same ten-minute request timeout as the other providers.
The OpenAI video provider is left alone beyond an import path: its API is
scheduled to shut down on 2026-09-24.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
OpenRouter and Together reject a request whose prompt plus max_tokens
overflows the model's context window. Both providers retried by deleting
max_tokens, which threw away the output cap the credit gate had sized to
the caller's remaining balance and let the retry run to the model's full
output limit with no second gate and no new hold.
The retry now goes through a shared helper that sizes a new cap from the
window and input count the rejection reports, falling back to the model's
declared context and a doubled prompt estimate, and never exceeds the cap
the gate set. When no window can be determined or no output fits, the
original rejection is rethrown instead of retrying uncapped. The rejected
params are copied rather than mutated, so the first attempt's record is
not rewritten after the fact.
Also corrects the estimator's own comment, which described the mean of
two approximations as a deliberate halving.
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
A Replicate prediction that ran and ended `failed` reaches the provider as a
plain Error with no HTTP status, so the driver-boundary translator could not
classify it and it surfaced as an unhandled 500, a critical alarm, and an
on-call page. Most of these are the model's content filter refusing the
user's prompt.
Wrap the run call and classify the failure: content-filter refusals become
a 400 with `errorCode: moderation_flagged` (the code chat refusals already
use); anything else becomes a 502 `upstream_failed`, which the alarm gate
skips. Status-bearing SDK errors pass through untouched so the boundary
translator keeps handling them. Upstream messages are stripped of markup and
bounded so an HTML error page can no longer ride into a response body or an
alarm signature.
Documents the codes callers can now act on in the txt2img reference.
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Register claude-fable-5-1 in the Claude catalog and gate it in the
provider the same way as Fable 5: no sampling params, effort via
output_config, adaptive thinking with summarized display. The bare
claude-fable / claude-fable-latest aliases move to 5.1, matching how
the Opus aliases moved when Opus 5 landed.
Fable 5.1 keeps Fable 5's $10/$50 per MTok, 1M context and 128K output,
but bills cache reads at 0.025x input instead of the 0.1x every other
Claude model uses, so the catalog row carries its own rate and a test
pins it.
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
A live audit of every keyed provider's /models endpoint against the hardcoded
catalogs found no stale entries but large gaps. This backfills them under four
rules: nothing vendor-deprecated, nothing without a price confirmed on the
vendor's official pricing page (each entry's source was recorded during
review), nothing absent from the live /models listing, and nothing that fails
a live routing probe.
Added: 46 Alibaba entries (qwen3/3.5/3.7/3.8 families, VL/omni/MT lines, and
Model Studio's hosted GLM/DeepSeek/Kimi third-party models) plus 9 dated
aliases; OpenAI chat-latest and gpt-4o-2024-11-20 plus 16 snapshot aliases;
Gemini gemma-4-31b-it and gemma-4-26b-a4b-it (vendor-documented free tier)
plus rolling -latest aliases; Mistral-hosted zai-glm-5-2 and a
mistral-medium-3.5 alias; deepseek-v4-flash-vision-exp; glm-5.3-flash.
Culled by the rules: 15 vendor-deprecated OpenAI entries (the 3.5/4/4-turbo
legacy line, gpt-4o-2024-05-13, o1-pro, four chat-latest predecessors, the
5.x codex line — deprecations page, most shut down 2026-10-23) and dated
aliases onto the deprecated o1/o3-mini/o4-mini; qwen3-vl-flash-2025-10-15
(live routing probe returned upstream 400 twice). Tiered Alibaba prices are
encoded at the base tier and busy-hour rates where time-of-day priced, noted
in comments.
Every surviving addition was verified end-to-end through a local deployment:
58/58 answered a live prompt, including all 46 Alibaba entries and every
spot-checked alias.
Not changed, flagged for maintainers: pre-existing o1, o3-mini and o4-mini
entries are now vendor-deprecated (shutdown 2026-10-23); the pre-existing
deepseek-v4-flash/-pro prices no longer match DeepSeek's current pricing
page; gemma-4 emits its own <thought> markup inline in content.
Co-Authored-By: Claude Fable 5 (1M context) <noreply@anthropic.com>
Live-probing magistral-small-latest showed the model inlines its reasoning as
answer prose in a flat string — no ThinkChunk content, no markers, nothing a
client can separate. Mistral's chunked thinking shape is requested via
`prompt_mode: 'reasoning'`, which the provider previously dropped on the
floor: there was no way to even ask for it.
`custom.prompt_mode` now forwards to the SDK's `promptMode`, following the
BytePlus custom-params precedent. Opt-in rather than a default because the
API rejects the mode where the account/model lacks it ('Reasoning prompt
mode is not enabled for this model', code 3051) — verified end-to-end: the
3051 travels back through the stack, which also proves the parameter is
delivered. The moment Mistral enables the mode, the ThinkChunk content flows
into the existing splitter and comes out as `message.reasoning` and
`reasoning` stream chunks with no further changes.
Co-Authored-By: Claude Fable 5 (1M context) <noreply@anthropic.com>
Creating a hosted subdomain gated `root_dir` on `write`, and hosting serves
everything under that directory with the ACL deliberately bypassed. So a
recipient of a `write` share could point a `*.puter.site` subdomain at the
owner's folder and make the subtree world-readable — continuously, covering
files the owner added later, with the row under the recipient's account where
nothing the owner can list would show it. `update` had the same gate for a
changed `root_dir`.
`#checkPublishAccess` now decides both: the actor's own tree still takes
`write`, anyone else's takes `manage` — "Can edit & share", the level that
delegates the decision.
Keyed on who owns the entry rather than asking for `manage` outright, which is
what the ticket proposed. `manage`'s is-owner implicator declines to answer for
app actors, so a flat `manage` would refuse every app publishing a directory
its user handed it, with no way for the app to obtain the grant. The write
check still runs first — it is what masks a directory the caller cannot see as
a 404 — and `manage` satisfies every lower mode, so the order costs a
manage-holder nothing.
The GUI's Publish As Website item reuses the own-it-or-`manage` answer it
already computes for sharing, so it is not offered where this would refuse.
Docs state the rule on `hosting.create()` and in `share()`'s level list.
Regression tests fail without the driver change: a write-share recipient is
refused on create and on repointing an existing subdomain, while `manage` and
the actor's own directory are accepted.
Review of the previous change found three of its claims unmet.
The image-generation crash it reported fixed is still reachable. The assert was
scattered across three helpers, and Gemini and OpenAI call `isHttpUrl` directly
on `input_images` without going through any of them — so two of seven providers
still 500 on a non-string. `isHttpUrl` now refuses a non-string itself, and the
shape is settled once in `ImageGenerationDriver.generate`, where the driver call
arrives, rather than per helper. That also covers `input_images` that isn't an
array, which produced a different crash per provider.
The sixth case in the ticket, previously unlocated, is
`Messages.js` reading `tool_call.function.name` with no guard — reachable with
`{"messages":[{"role":"assistant","tool_calls":[{"id":"x"}]}]}`. Guarded, along
with the same shape in `make_claude_tools`: a TypeError there carries no status,
so the retry loop reads it as a provider failure and marks the route unhealthy
for every caller.
`#hardExpiryFromExpiresIn` returning null for a bad type moved the failure past
the session INSERT, leaving an orphaned non-expiring row and still answering
500. Reverted; the controller guard is the fix, now covering fractions,
negatives and unparseable durations rather than only wrong types.
Also: the batch write handlers check that the body is an array but not what is
in it, so a null element 500s the same way; `#requireObjectBody` accepted an
array despite its name; `handleCreateAccessToken` destructured a body that may
be absent; and two AGPL notices had been rewrapped with a Markdown link.
Each of these read a field off caller input that wasn't the shape the code
assumed, threw a TypeError, and was served as a 500 with a critical page.
- POST /fs/write and /startWrite: a request whose body never parsed left
`req.body` undefined, and the first read of `fileMetadata` threw.
- POST /drivers/call, image generation: `input_image` / `input_images` entries
are documented as strings but nothing checked, so a number or an object
reached `.startsWith`. Type confusion on caller input, so reachable on
demand rather than by accident.
- POST /auth/create-access-token: `expiresIn` went to the expiry parser
unvalidated, where anything but a string or a number has no `.trim`.
- POST /login/wait: destructuring `session` out of an absent body threw.
Same class as the POST /login body ticket; that one covers /login itself.
The comment claimed the fallback loop cannot be driven from a test — false
since ChatCompletionDriver.test.ts gained a two-attempt fallback test that
drives the actual loop. Say where that test lives instead.
Co-Authored-By: Claude Fable 5 (1M context) <noreply@anthropic.com>