* fix: charge stat's return_shares against the share-listing budget
`return_shares` on /fs/stat and the legacy /stat runs the same work as
GET /share/shares, but was only metered under fs:stat's far more
generous limit — and the two scopes stacked instead of sharing one
counter.
Adds consumeRouteRateLimit(req, spec), an imperative charge that
resolves the key and per-subscription limit exactly as rateLimitGate
does, so a handler can conditionally spend a second scope when a
request flag makes the route expensive. Both stat handlers now charge
share:list before doing the listing work; SHARE_LIST_LIMIT moves to a
shared share/limits.ts so all callers pin the same spec.
Closes PUT-1597.
* feat: consumeRouteRateLimit takes the array spec form too
Review follow-up on #3870: a multi-window spec passed whole would have
read an undefined window and silently never pruned. Charge each window
in order instead, refusing on the first refusal, matching the gate.
* fix: list a link share only while the owner's plan covers it
* fix: fs limits for signed urls
* feat: a paid plan counts as a verified card for the sharing gate
* feat(email): inline cid attachments and Puter mailbox delivery for sendTransactional
EmailAttachment gains cid/contentDisposition so the transactional driver can send inline images. The SDK's EmailAttachment typedef now comes from types.js, which already carried cid. Docs describe delivery to <username>@puter.email recipients and the not_found code.
* feat: share with anyone with the link (PUT-1580); gate sharing on a verified phone or card
feat: the team seat experience — no email required, forced password change, team label, and plan-based limits (PUT-1792)
Note: Bypassing the code owners rule, since there are couple approvals in place for this.
`teams_allowed_email_domains` limits who may enter the teams surface;
unset keeps today's behavior. Gated on the two routes that constitute
entry — creating a team, and the listing that shows the tab — with the
same 404 a teams-off deployment answers, so the GUI needs no change and
a staged rollout is indistinguishable from the feature being off.
Members of an existing team always pass, whatever their domain: an
allowed owner brought them in, and the surface follows the team.
Deleting a team suspends every provisioned seat, but every member route
resolved live teams only — so the suspended seats could never be
deleted afterwards, stranding the accounts and, on the billing side,
their paused subscriptions. deleteMember now resolves the team
soft-deleted or not, matching the audit reader.
It also gains the guard enableMember always had: only the team's own
suspension qualifies. A platform-suspended seat could previously be
cascade-deleted by the owner, destroying an account Puter had frozen.
From the adversarial review of this PR.
The forced password-change gate could deadlock: it POSTed to the
cookie-only route with a bare fetch, and both initgui call sites open it
before update_auth_data mints the session cookie — a fresh browser with a
token URL 401'd every submit inside a non-dismissible loop. It uses the
session-cookie retry wrapper now, taking the caller's token because
window.auth_token does not exist yet on that path. It also gains the
logout footer its sibling gates have; a lost temporary password was a
hard lock with devtools as the only exit.
Password recovery refused for seats: the address is admin-supplied and
never verified, so whoever holds that inbox could take the seat over at
any later time. A seat's recovery channel is its admin's reset.
change-email gets the same seat guard as change-username and deletion —
the address is where admin-issued credentials go.
The login response now carries `team` alongside requires_password_change:
no-reload logins store that payload as window.user verbatim, and every
seat restriction keys on it.
Smaller: the team-badge tooltip no longer double-encodes; the create-token
hint for an emailless account stops pointing at a verification it can
never perform; the quotas doc records the halved org_seat_free allowance;
the config template tells upgrading operators how to keep the old flat
cap; the SDK suite covers emailless provisioning and the owner-only uuid.
Review catch by @Salazareo: `org_seat_free` reached `FREE_SUBSCRIPTION_IDS`
and so the `requireSubscription` gate, but two other surfaces decide on
plan and neither consults that set.
`bySubscription` maps name `user_free` and `temp_free`. A plan that
matches no key fell through to the top-level `limit` -- the paid cap --
so a seat outranked an ordinary free account: 240 event listings a minute
against their 120, and the same shape across the kv, notification,
subdomain and worker drivers. Both resolvers now fall back to the
`user_free` entry for anything in the free set, which covers every driver
at once and any free plan added later.
`subscriberOnly` compared against the two named ids, so a seat could
reach a paid-only model. It asks the set now.
A paid plan that names no cap of its own still takes the base, and an
unresolved plan still takes the base; there are tests for both so the
fallback cannot widen into "free by default".
- PUT-1800: gate `createWorkerSessionToken` on actor type, so an app or an
access token can no longer mint an app-less, root-shaped worker session;
`WorkerDriver` binds on `effectiveApp` instead of `app`.
- PUT-1799: add `isAccountContext` and read it where "no app" was being read
as "the account" — handler publish, events-worker listing, kv handle
mint/revoke/list. A scoped API token is no longer an account session.
- PUT-1802: re-authorize a durable row before its backlog drains, settling it
permanently when the grant is gone. Covers an ancestor-level unshare, which
the revoke settle deliberately leaves to the delivery re-check.
- PUT-1803: mask the owner's absolute path out of deliveries and subscription
anchors on a foreign node, the way every FS surface already does.
- PUT-1804: let a revoke reach rows already suspended for a resumable reason,
re-stamping them so a resume cannot hand over the held backlog.
- PUT-1805: apply the subscribe path's audience gate to `/events/fetch` before
the query, so a cursor can no longer count and name invisible notifications.
- PUT-1807: refuse `mode: 'manage'` from any actor holding an app — inside its
own AppData the ACL short-circuit would otherwise supply the reach.
- PUT-1808: take the sending peer from the verified signature header rather
than the request body.
- PUT-1810: re-base a kv share-handle row's stored match filter on the handle,
so the owner's absolute key prefix stays hidden.
- PUT-1814: escape LIKE wildcards and anchor the issuer-prefix queries on a
segment; anchor `manage:` stripping; reject a backslash in a share prefix;
assert a resolved actor in `subscribeDurable`.
- PUT-1815: bound the char/varchar columns behind `event_subscriptions` and
`kv_share_handles` at the store layer.
A provisioned account could rename itself, delete itself, and see a
Billing tab for a subscription it does not hold — all of it the team's,
not the account's. Each is now refused server-side and dropped from the
UI, keyed on one predicate: whoami reports a team only for an org-owned
seat, and the owner joined their own team.
The Teams tab showed a seat nothing but its own audit rows, so a member
could not see who else was on the team they were told they shared it
with. It now lists them; listMembers was already membership-gated and
already withholds from a member what is not theirs.
The dashboard share modal had no way to reach a team, though the desktop
dialog has had one since team sharing shipped and the SDK has always
taken `{ team }`. Same control, same copy, same helper. The access list
needed a team bucket to go with it: a team share names no holder, so the
aggregate dropped it and a team you had just shared with vanished.
The per-account subscription button did nothing. `listMembers` never
returned a member's uuid, the SDK's `toMember` dropped it, and the Plan
column read `seatTiers[undefined]`, so every seat rendered as Free and
the action dispatched with no seat.
The uuid is owner-only: it is what billing keys a seat's plan on, and one
member has no business identifying another.
The backend already refused every route with `password_change_required` until a
provisioned account replaced the password its administrator chose, and
`/user-protected/change-password` was already exempt so the account could act.
The client half was missing entirely: nothing in the GUI referenced that code,
so a seat signed in and then failed at everything with no prompt and no way out.
Two gaps, both closed here.
The flag never reached the client. Neither the login response nor `/whoami`
carried `requires_password_change`, so the GUI could not have known even if it
wanted to. It ships from both now, alongside the other three verification flags
the whoami extension already describes as "the flags the GUI acts on".
There was no window to show. `UIWindowChangePassword` is a settings dialog --
closable, and it never resolves on success -- so it cannot act as a gate.
`UIWindowPasswordChangeRequired` mirrors the existing `*Required` windows: it
resolves true only once the change lands, and `initgui` loops on it. It runs
last in the boot chain, matching the server order in assertVerifiedAccount.
It refuses a new password equal to the current one. Without that the account
stays on the credential its administrator still holds, which is the entire thing
the gate exists to end.
Falsified twice: dropping the same-password guard fails "refuses to reuse the
password the admin handed over", and resolving on a rejected response fails
"stays open on a rejected change, so the gate cannot be escaped" -- each
breaking only its own test.
175 backend tests, 326 GUI/SDK tests, typecheck clean.
PUT-1792. Two separate problems, both from treating a provisioned account like
a self-registered one.
`email` was required, so an admin creating ten seats had to invent ten addresses
and then keep track of ten uniqueness constraints -- for accounts that sign in
by username and never use the address. It is now optional at every layer, and
the add-account form does not ask for it at all: username is the only thing a
seat needs.
`requires_email_confirmation` was set to true, with the reasoning that an
admin-supplied address is unverified. True, but `requireVerifiedAccount` turns
away on exactly `requires_email_confirmation && !email_confirmed`, so a
freshly created seat was asked to confirm an address it may not hold and could
not use the product until it did. The team creating the account is the trust
anchor, not the mailbox, so this is now false either way.
An address is still accepted and still stored when given, because the notices
are worth delivering. `#notifyUser` already returned early on a missing
address, so `team_account_created`, `team_account_disabled`,
`team_password_reset` and `team_closed` degrade quietly with no new branching --
the temporary password is in the API response, which is the documented delivery.
`idx_user_owned_email` is partial and skips password-null rows, so omitting the
address sidesteps it rather than creating a collision surface. Two seats with no
address do not conflict, and there is a test for it.
Docs now say an emailless seat is recoverable only through its team's owner.
That falls out of the design rather than being a limitation of this change, but
it should be written down rather than discovered.
Falsified: putting `requires_email_confirmation: true` back fails
"never demands confirmation, with or without an address" with
`expected true to be false`, and nothing else.
164 team tests, 40 SDK tests, typecheck clean.
Members can already enumerate each other: `/teams/:uid/members` needs a user
actor and nothing more. The only thing this adds is admitting an app actor to
the same names, so an app can offer colleagues without the member driving it.
That is the whole risk, so it is off until the team owner turns it on.
`group.directory_enabled` defaults to 0, and a team that has not opted in
answers 404 rather than 403 -- whether a team has this on is not something an
app should be able to probe for either.
Three things bound what an app sees. The membership tested is always the
person's, never the app's, so an app installed by a member of one team can
never read another's. The page carries username and uuid and nothing else --
no email, activation state, usage or role. And suspended accounts and ones
that never took up their credential are left out, since offering someone who
cannot sign in is noise and their existence is not this list's to disclose.
Activation is the forced-change flag clearing, not the password existing: a
provisioned seat holds its temporary password from birth, so testing
`password IS NOT NULL` would have leaked exactly the accounts meant to be
excluded. A test covers that distinction.
Turning the directory on or off writes an audit row, because it changes who
can read the member list and that is not something a team should be able to
alter silently. Setting it to the value it already has records nothing.
The toggle lives in TabTeams, and turning it on asks for confirmation while
turning it off does not -- one grants access, the other only takes it away.
Closes PUT-1736.
A disabled account persists indefinitely. Removing it is an explicit request,
never a timer, and it is refused on a live account with
`account_must_be_disabled_first` — which puts a reversible step in front of the
only irreversible operation in the feature.
The audit row is written before `cascadeDelete` runs. The FKs are ON DELETE SET
NULL and the `_keep` columns carry the identifiers, so the record of what was
done survives the account it names.
No second billing emit here: `cascadeDelete` already captures the seat and
fires `team.account.deleted` through UserAccountService, and emitting again
would close the storage charge twice. Disabling closed the per-account charge;
this closes the storage one, and it is the only thing that does.
There is no restore window, and none was wanted: the reversible step already
exists earlier at disable, a disabled account costs only the bytes it holds so
nothing pressures a hasty delete, and a restore promise means retaining data
the team explicitly asked to be rid of.
Published in rate-limits-and-quotas.md alongside the team-deletion note,
since the two are easy to confuse and only one of them frees a seat.
Closes PUT-1732.
Two halves of the same question -- what did the team do to me, and how do
I find out. One is pull, the other push.
The activity view (PUT-1746) was already user-scoped and already 404'd a
non-member; what was missing is the load-bearing half. A member is told a reset
happened, but a reset only matters alongside who signed in afterwards, so
`SessionStore.listSignIns` merges their own sign-ins into the same stream,
newest first, with a per-stream cursor. Without that row the audit says a
credential was issued and never says whether it was used.
The notices (PUT-1733) cover the two things a member cannot discover for
themselves: their account being disabled, and their team being closed.
Deliberately one each -- closing a team disables every account in it, so
sending both would tell one person twice about one event. Both say plainly that
nothing was deleted; the accounts persist, suspended, holding their files.
Squashed because they are one change to a reviewer: same audience, same
purpose, seven files, overlapping only in TeamService. Neither touches shared
platform code.
`user.requires_password_change` shipped with the team columns but nothing
enforced it and nothing ever cleared it, so a provisioned seat kept its
administrator-issued password indefinitely and `reissueCredential`'s
"already activated" 409 was unreachable.
Adds the fourth clause to `assertVerifiedAccount`, the only place a
verification gate may live -- WebDAV builds its own actor and calls that
function directly, so a second implementation would bypass it the way the
phone and card gates once were bypassed.
A gate that refuses everything also refuses the endpoint that clears it,
so `/user-protected/change-password` opts out with `allowUnconfirmed`.
That widens the route: an account pending email, phone or card
verification can now change its password, which it could not before. The
caller is authenticated and proves the current password, so this is
benign, but it is a behaviour change to a shared route.
Also here, because the gate is worthless without them:
- change-password and the recovery-token path clear the flag, and record
an `activate` entry when the account is a seat.
- Reset takes a live account back with a fresh credential, capped at 20
per day and audited as `reset_member_password` with no credential in
the row. Re-issue is audited the same way; it stays closed once a seat
has chosen its own password.
- An issued credential expires after 24h (new `temp_password_expires_at`
column, three dialects) and login refuses it after that, so an unused
reset dies instead of becoming a standing credential.
- 2FA is untouched by a reset, so a reset alone is not takeover.
Covers PUT-1726, PUT-1727 and PUT-1729. Together because a recipient without
the listing work produces a share that is created correctly, resolves
correctly, and never appears in "Shared with me" — a state nobody would ship.
The recipient. `ShareRecipient` gains `team` (uid) and `teamHandle`, resolved
before the email and username branches and never falling through to them.
Passing both is an error rather than a precedence rule, so a call site always
shows which was chosen: handles are released on soft delete and can be
reclaimed by an unrelated team, and a scripted share to a handle would
silently retarget. There is no bare-string spelling, since that would change
how existing strings are interpreted.
ACLService.setUserGroup mirrors setUserUser: same read-modify-write, same
one-mode-per-node rule, under a node lock keyed on the group.
One grant against the team, not one per member, so membership changes
apply without touching the grant and a share spends one unit of the daily
quota however many members there are. A member who joins afterwards gets
access, which is asserted.
The index row carries `holder_group_id` and leaves `holder_user_id` NULL, so
`0077`'s group index constrains it rather than the user-holder one.
Listing. `listByHolder` and `countByHolder` union the caller's teams into
the same keyset page — `ORDER BY id` still holds — and `#grantEvidence` gains
group grants as a third source. Without that third source the share is
filtered out of every listing as dead: nothing errors, the share simply is not
there, which is the one place a user would look for it.
Unsharing revokes the grant as well as deleting the index row. Deleting the
row alone would hide the share while leaving every member holding real access.
Closes PUT-1726, PUT-1727 and PUT-1729.
A seat is a real Puter account: it takes a name from the global username pool
and gets a home directory. Nothing charges for one — that is prod's job — so
until it does, the only bound on creation is the request rate limit, which
bounds the rate and not the total.
max_teams_per_user default 1
max_seats_per_team default 50
Ordering is the substance of both checks. The team cap is tested before
the handle, so a capped user is told they are capped rather than that the name
they picked was unusable. The seat cap is tested before any account state
exists, so a refused provision does not burn a global username.
Soft-deleted teams do not count toward the owner's cap, so deleting frees
the slot — which does mean create, provision, delete, repeat still consumes
usernames over time, bounded by the daily rate limit. The caps raise the cost
and make the cycle audited; they do not close it.
Lowering the seat limit blocks new provisioning and disables nobody.
Both limits are published in rate-limits-and-quotas.md, and both keys are
documented in config.template.jsonc and config.default.json.
Closes PUT-1758.
`addMember` refuses to turn an account that already has a password into a
team seat. No service path did this — `provisionAccount` always creates —
but the store permitted it, and the design rules out existing accounts joining
a team.
Provisioning passes the guard because it admits the account before setting its
temporary password.
There is no bypass parameter. Both HTTP suites now provision a real seat and
authenticate as it, using the same token-minting the harness uses for
`POST /login`. That surfaced something worth knowing: an unactivated seat
cannot call the API at all. Provisioning leaves `requires_email_confirmation`
set and `requireVerified` rejects it, so the suites activate the seat first —
which is the state a member is actually in when making requests.
`listMembers` and `getMembership` also return `u.uuid`, which the billing
events need in order to name the account without a second lookup.
The dispatcher's rehydrate callback reaches whichever backend answers the
API's public hostname. A backend with the runtime flag on that is not behind
that hostname, or the only one in the fleet with it on, could never get its
scripts deployed that way — the callback answered "disabled". The invoking
backend already knows the app and script, so on a dispatcher miss it deploys
the set itself and retries once, telling the dispatcher to skip its callback
and negative cache. The callback stays the path for evicted scripts.
* feat: bake published handlers into a generated events worker
* feat: deploy and address the per-app events worker behind a flag
* test: single delivery end to end through a real local worker
* feat: events workers run their own runtime, in their own namespace
An events worker was being deployed as an ordinary worker: default dispatch
namespace, a `subdomains` row, the router preamble, and an app-scoped worker
token baked in. The public dispatcher resolves any script in that namespace
straight off the hostname, so the worker answered at `<name>.puter.work`, and
the only thing in front of it was an unguessable name plus a check that a
`puter-auth` header was present — which the router never validates. Anyone who
learned the hostname could run an app's handlers with a body of their choosing,
in an isolate holding the owner's token as `me`.
Instead:
- Handlers run on their own runtime (`src/worker/src/events-runtime.js`), which
provides no `router` and no `me`, owns the single invoke route, and hands a
handler only `{ event, ctx, user, fetch, ack }`. `user` is built from the
invocation's delivery token, so a handler acts as the subscriber whose
delivery it is and nothing wider. The preamble build emits one bundle per
runtime; the shared half of the template is now included by both.
- The deploy target carries the runtime to prepend, the source to deploy, and
whether to mint a worker token at all, so an events worker deploys into the
`events` dispatch namespace from generated source with no token binding, no
`subdomains` row, and no claim on the owner's worker quota or worker list.
- An invocation carries a key derived from the deployment secret and the script
name, bound as a secret and checked in constant time inside the isolate,
which reads it once and drops it before handler code runs.
- Scripts are named after the handler set they contain, so publishing writes
rows and deploys nothing: a set is deployed the first time a delivery needs
it, and a changed set is a new script rather than an overwrite of a running
one. Publish responses keep the shape they had before the runtime existed.
- Invocations reach a worker only through the events dispatcher, which has no
zone route and requires the internal secret; the backend's own deploy path is
the rehydrate route the dispatcher calls on a namespace miss. Locally there is
no dispatcher, so the controller hands the service an in-process transport
that deploys on miss itself.
The SDK stops allowlisting `puter` as a handler global — a handler that reaches
for an ambient SDK is now refused at publish time, naming `user` instead, rather
than passing the scan and failing on its first delivery.
Requires `events.workerNamespace`, `events.dispatcherUrl` and
`events.internalSecret`; without them nothing is addressable and background
deliveries stay retriable, as they did with the runtime off.
* fix: a handler's delivery token gets through the read routes
An events handler acts as the subscriber through the access token its
invocation carried, but every FS read route refused scoped access tokens
outright, so `user.fs.stat(event.path)` — the design's own example — answered
403 inside the worker. The read-side routes now admit them; the ACL each
handler already runs intersects the token's grant with its issuer's, which is
the check that keeps a token to what it was minted for. The end-to-end suite
asserts the stat from inside the isolate.
* fix: shorthand-method handlers publish as functions
`{ ingest({ event }) { … } }` stringifies without the `function` keyword, so
its source is not an expression and the events worker baked it as a broken
stub — every delivery a retriable 500 until the subscription suspended, with
nothing at publish time to say why. The SDK now gives a shorthand method the
keyword before hashing and sending; getters, setters and computed names are
left for the server-side check to refuse.
* feat: an app's events worker is listable and destroyable
An app with published handlers has an events worker, and hosted deployments
bill it monthly per app, so its owner needs to see it and be able to take it
down. The core announces the lifecycle on the bus — `events.worker.create`
when an app's first handler is published, `events.worker.destroy` when its last
one goes — with the owner as the actor, so pricing can plug in from outside.
`GET /events/workers` lists the caller's workers (paginated, with the script
each set deploys as) and `POST /events/workers/destroy` removes every handler
of an app under the same owner scoping as the handler routes, suspending the
subscriptions bound to them. `puter.events.workers.list/destroy` in the SDK,
a docs page, and a 5 MB cap on an app's combined handler source
(`events_worker_too_large`) so a set that publishes can always deploy.
* fix: harden the events worker runtime for production
- A 4xx is terminal only when it carries the handled marker the runtime (and
the dispatcher) stamp on every answer that came from a script; an unmarked
4xx — an edge 404 for a wrong dispatcher hostname, a WAF page — stays
retriable and is logged, once per script per minute, with the runtime's
reason header.
- Script names are scoped to this backend's exposed API origin, so two
backends sharing a namespace never resolve one script with the wrong
endpoint binding or key. Shape unchanged.
- Each handler is validated in the exact context it is emitted into and the
whole generated file is compiled once; a source that would break the script
marks every handler broken instead of deploying a SyntaxError.
- Locally, events scripts live under their own registry key: the public local
worker host cannot reach them and an ordinary worker cannot take their name.
- A suspended or deleted app owner stops invocations; deploys are throttled
per app per hour; in-flight deploys are keyed by app and script; the
upstream deploy call times out; the generated source is size-capped with a
margin over the publish cap; boot fails when the runtime is on but its
preamble is not built. Byte-length secret compare, appUid shape check,
dispatcher URL prefix preserved, wider connection pool.
* feat: background workers are listed in the sessions manager
A user paying for an app's events worker needs somewhere to see it and take it
down. The sessions manager gets a section listing the apps that run event
handlers in the background, with a Destroy action that removes their published
handlers.
* fix(ai): make chat fallback reach streamed Claude calls and rank Azure explicitly
- ClaudeProvider opens the upstream stream and awaits its connection
before returning the populator, so an overloaded or rate-limited route
throws from complete() and reaches the driver's fallback loop instead
of surfacing as an error frame on a 200
- the OpenAI-compatible chat and completions routes only pin a provider
when the caller sent one, so they get the same preferred healthy route
puter.js callers do and unhealthy-route skipping applies to their
first attempt
- Azure is ranked ahead of the vendors it fronts by an explicit tier in
modelRouting rather than a price tie plus registration order
- drop Together's synthetic always-failing model-fallback-test-1 entry
- test that a 4xx leaves a route in rotation while a 503 marks it
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* fix(ai): surface swallowed Claude stream errors, keep compat-route defaults
Review follow-ups on the fallback work.
- The pre-created event iterator only receives an error if a reader is
already waiting on it, so a failure landing between the connect and the
populator's first pull ended the stream cleanly — truncated content
billed and reported as a success. Rethrow when the stream is errored.
- A refused stream deleted its Anthropic uploads but left the caller's
message parts pointing at those file ids, so the fallback route was
handed handles it cannot resolve. processPuterPathUploads now returns a
restore() that both failure paths call.
- The OpenAI-compat routes keep pinning OpenAI when the caller sends no
model at all, so the default model stays put instead of moving to
Azure's.
- Say why /openai/v1/responses and /anthropic/v1/messages stay pinned:
each translates one provider's native shape by hand.
- PREFERRED_PROVIDERS is unexported and its doc now states the rank is
unconditional; the duplicated hidden-model list is one constant.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(ai): undo the puter_path rewrite by field instead of snapshotting the part
Copying the content part kept whatever the caller sent on it — a large
inline `source` or `text` alongside `puter_path` — reachable until the
request ended, where overwriting the field used to make it garbage right
away. The only fields this function writes are `type`/`source` on success
and `type`/`text` on failure, and the fallback uploader keys off
`puter_path` alone, so restore undoes those three by name.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
* feat: presence and cross-region event forwarding (PUT-1679)
* fix: fan cache bumps to sibling nodes and stop the forward shed cascading (PUT-1679)
`outer.events.generationBumped` and `outer.events.presenceBumped` rode
`outer.*`, which the broadcast service only webhooks to peer regions;
only `outer.pubsub.*` also fans over Redis to a region's other nodes.
Both caches are per-process maps with no expiry, so a bump landing on
one node left its siblings stale until that user's next transition.
Renamed onto `outer.pubsub.events.*`; the listeners already accept the
`from_outside` copy the Redis re-emit carries.
`PeerForwardQueue.push` called `onOverflow` synchronously and the
handler pushed markers straight back, each of which re-tripped the
bound and shed the next item: one item over a 5000 bound recursed ~2200
deep, threw a RangeError, and turned ~2200 queued deliveries into gap
markers. It also re-summed `bytes` over the whole queue per drop. The
handler now returns its markers and the queue appends them past the
bound check, sheds deliveries before markers, keeps one pending marker
per (peer, subscription), and subtracts bytes per dropped item.
App payloads carried only the /app-icon endpoint URL, which 302s to the
icons hosting subdomain. Networks that mangle that redirect render no icon
at all, and every icon load pays a round trip for the hop.
Ship the direct subdomain URL as `iconCdnUrl` alongside it (taskbar items,
installedApps, recent/recommended launch apps, suggested apps), and have the
GUI load that first with the endpoint URL as a one-shot retry - desktop
taskbar, start menu, dashboard app grid and recents. Only rows whose `icon`
column is already an http(s) URL get one: a data: column means the resize
pipeline has not written anything to the subdomain yet.
Also folds the four copies of the generated-size list into one exported
APP_ICON_SIZES.
Covers PUT-1708, PUT-1709 and PUT-1743.
Twelve routes, every one setting requireUserActor -- that option is what
installs requireAuthGate, requireVerifiedAccount and requireNonAccessTokenGate,
because server.ts derives `needsAuth` from the route options. Reads need it as
much as writes: without an auth option a route gets no suspension check and
admits access tokens, so a just-disabled member could still read the roster and
a scoped third-party token could read the audit log.
Authority is checked before anything observable. Validating the body first made
POST /members answer 400 before 403, and resolving :username first turned the
member routes into a global username-existence oracle.
Provisioning applies the same username and email rules as signup rather than
its own -- USERNAME_REGEX, USERNAME_MAX_LENGTH, RESERVED_USERNAMES and
validator.isEmail, now exported from AuthController. Without them a workspace
could mint accounts signup would refuse, claim unregistered reserved names, and
mail arbitrary unvalidated addresses.
Handle problems are 400 or 409 rather than a bare Error, which the server turns
into a 500 and a deduped critical alarm -- an uppercase handle should not page
on-call.
Disable drops sessions through SessionStore.removeByUuid rather than a raw
DELETE. The store invalidates every composite cache key; without that a
disabled member kept authenticating from cache for the session TTL, which is
exactly the "takes effect on the next request, not after a cache TTL" property
disable is supposed to have. Revoking also preserves last_ip/last_user_agent,
which the member-facing audit view reads.
Audit writes live in TeamService at the point of each action rather than in the
route, so a caller reaching the service directly cannot skip them, and the SQL
lives in TeamStore. Audit reads map internal user ids to usernames, and remain
readable by the owner after the workspace is soft-deleted -- otherwise the
delete_team entry was written and immediately unreachable.
teams_enabled gates route registration through an optional isEnabled() the
server honours, so with it off the paths do not exist rather than existing and
refusing. It does not gate DDL.
TeamIsolation.http.test.ts asserts the negative the feature rests on: the
workspace manages accounts and cannot read them, including through a
full-access token and after the member is disabled. It asserts outcomes rather
than the absence of an implicator.
Covers PUT-1705. The master account supplies { username, email }; the account
is created with no password, gets the default filesystem tree, joins with
org_owned = 1, and receives a one-shot activation link.
Activation reuses password recovery rather than new token machinery: the same
pass_recovery_token, the same one-hour purpose-scoped JWT, the same
/action/set-new-password link. No team_activation table, no new token type,
and no unauthenticated endpoint on the team surface. Activation state needs no
column either -- an unactivated account is one with no password.
Applies the same username and email rules as signup rather than its own:
USERNAME_REGEX, USERNAME_MAX_LENGTH, RESERVED_USERNAMES and validator.isEmail,
now exported from AuthController. Without them a workspace could mint accounts
signup would refuse -- the username becomes the /username home-directory
segment -- claim unregistered reserved names, and send activation mail to
arbitrary unvalidated addresses at the route's daily limit.
Usernames come from Puter's global pool, so a taken one is refused with free
alternatives rather than silently modified: a suffixed name would appear in
every share dialog that person ever sees, and they never agreed to it. The
check runs before any write, so a rejected provision leaves no orphaned user
row -- asserted by a test on the workspace's member count.
The new account carries requires_email_confirmation, since the address came
from the administrator rather than its holder.
Adds a team_account_activation email template stating what the workspace can
and cannot do -- including that it can reset the password, which the design
requires be said rather than only claiming files are private.
free_storage stamping and the billing event are phase 3.
A video job that outlived its poll window, an SDK request that timed out,
or a Veo operation that finished with an error all reached the HTTP error
handler as plain Errors. Each became an unhandled 500 with critical
severity and paged on-call for what is the provider's pace or the
provider's fault.
Video providers now share one poll loop that gives up with a 504
`upstream_timeout`, treats a transient poll failure (timeout, dropped
connection, 408/429/5xx) as a missed poll rather than a failed job, and
stops polling with a 400 `client_aborted` when the caller disconnects, so
nothing is metered for a clip nobody will receive. The driver controller
exposes the disconnect as an `abortSignal` on the request context. The
window is ten minutes for every provider; Together and BytePlus move up
from five.
Failed jobs are classified: content-filter refusals become a 400
`bad_request` with `errorCode: moderation_flagged`, rejected parameters a
400 `upstream_bad_request`, and anything else a 502 `upstream_failed`,
each carrying the provider's own code. Veo's filtered output keeps
`disallowed_value` and gains the same `errorCode`. The sanitizer and
content-filter pattern move from the Replicate provider into a shared
util so image and video agree.
Status-less SDK connection timeouts are translated to a 504
`upstream_timeout` at the driver boundary, and the chat driver records
them per attempt so an all-timeout chain is a 504 and a mixed chain is
`upstream_failed` instead of an `internal_error` 500. The Together chat
client gets the same ten-minute request timeout as the other providers.
The OpenAI video provider is left alone beyond an import path: its API is
scheduled to shut down on 2026-09-24.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* feat: events metering and quotas (PUT-1683)
* fix: drop the standing subscription charge and price single deliveries at 100 µ¢
An idle durable row costs nothing worth billing; the plan quotas bound how
many an account holds. Removing the daily line also removes the global
day-claim, the whole-table scan and the timer it rode on.
* test: stop asserting on documentation pages
The limits test read rate-limits-and-quotas.md and grepped it for numbers,
so every rewording of the page failed the backend suite. Docs are kept in
step by the PR and checked in review; AGENTS.md now says so.
* feat: delivery re-check cache, revocation and anchor settle (PUT-1677)
* fix: authorize re-anchors, settle each row once, purge revoked backlog (PUT-1677)
- A path-form row whose anchor is deleted only climbs to an ancestor its
holder may still watch under the mode it subscribed with; otherwise it ends
with `anchor_deleted`. It used to land on any surviving ancestor (a guest's
row on the owner's home), where the re-check denied every delivery but the
row still held an anchor slot and a filter evaluation there.
- After a climb the new anchor is re-verified and the climb repeated if a
recursive delete took that level too, instead of leaving the row on a dead uid.
- suspend() is one conditional write per row and reports which rows it was the
one to suspend; concurrent settles of the same grant (an unshare revokes
several strings) no longer each purge, forget and notify the same rows.
- One "subscriptions ended" notification per holder and app, carrying the count
and subjects, instead of one per row.
- A revoke that removed nothing no longer announces; the sweeper purges (not
defers) the backlog of a permission_revoked row; the reap purges pending
entries with the row.
- The delivery auth cache indexes entries by subscription so forget() is not a
scan of the whole cache.
* feat: pending event deliveries and delivery-class invariants (PUT-1676)
* fix: make pending delivery claims and drains atomic, keep the region under its ceiling (PUT-1676)
- claim() and the drain-time reindex run as Lua over the subscription's own
{subId}-tagged keys. Two claimers can no longer both lease the head, and a
drain that finds the queue empty deletes it in the same step it checks, so a
concurrent enqueue is never wiped between the two.
- An append writes the entry and its queue position in one MULTI (same slot),
with the index seeded before it and corrected after, so an entry is never
visible without its position and never left out of the sweeper's index.
- Pipelines no longer mix slots (index/counter vs. per-subscription keys), so
the store works on a multi-shard cluster, not only a single-shard one.
- Region shedding counts the marker it leaves behind; it used to stop one over
the ceiling and convert a real event into a marker on every enqueue after.
- A claimed or suspended subscription moves to the back of the sweeper's index,
so a delivery nobody settles cannot hold the head against every other backlog.
- `single` rows must carry a `worker` target: with sockets exhausted and no
handler, an unacknowledged delivery would sit at the head forever.
- A gap marker for a row with no socket target is dropped rather than counted
as a delivery of nothing.
- Backlog keys carry a 7-day TTL, refreshed by every claim, as a backstop for
keys a purge/enqueue race left unindexed.
* feat: durable event subscriptions store, cache, and routes (PUT-1673)
* fix: durable subscription hardening (PUT-1673)
- Expired rows stop delivering at dispatch time and no longer count toward
the per-account cap, instead of waiting for the sweep.
- The expiry sweep runs hourly with a jittered first pass shortly after boot;
a 24 h interval never fired on a fleet that redeploys more often than that.
- Only a durable generation bump marks peer regions cold. A session
subscribe/unsubscribe in one region used to force a primary read in every
other region on its next dispatch.
- `subject`/`anchor_path` widen to varchar(4096) to match `fsentries.path`,
and subjects longer than that are refused with `invalid_subject` rather than
failing the insert on MySQL/Postgres.
- The dispatch and durable integration suites wait for the specific delivery
they expect and assert only within their own folder; the old any-delivery
`settle()` let a late event from a previous test satisfy or pollute the
next one under CI load.