The suite was CPU-bound on one runner, and about half of that CPU went to
re-importing the backend for every test file.
- Run the suite as 4 shards and merge their blob reports for coverage.
`test (pr)` stays as the required check and fails if any shard fails.
- Drop the per-PR base leg; pushes to main publish the coverage PRs
compare against.
- Files that boot a server without module mocks or extensions run in a
`shared-worker` vitest project (`isolate: false`), so the backend is
imported once per worker. testSharedWorkerSetup.ts resets the redis
mock, configContainer and extensionStore between files.
* auth: per-account 2FA code limit, single-use otp-login tokens, recovery mail limits
- /login/otp and /login/recovery-code share a per-account bucket charged on
wrong codes; an otp-login token completes one sign-in.
- /send-pass-recovery-email adds an IP bucket, a per-account cap and captcha,
and reuses the outstanding recovery token instead of rotating it.
- Reauth tokens carry the rejection reason; /signup re-attaches a temp
account only from an expired session.
* auth: app-issued access tokens older than their app stop authenticating
An access token whose row or parent session predates the app now holding its
uid is refused, the same check app and worker sessions already get.
* teams: signup can take an address a seat never confirmed
Password signup, save-account and a provider-verified OIDC sign-up take an
address a team seat holds unconfirmed; the seat's copy is cleared at the
write. A seat address that was confirmed still blocks.
* dev-center: encode dropped directory child names
* fix(events): bill reached deliveries only, bound dispatch reads, seal cursors
- A broadcast socket copy is billed only when a connection for its holder
exists on this node, anywhere in this region (connection count) or, for a
durable row, in a region presence names. Without peers it bills as before.
- Compiled match filters are held in a bounded LRU with a TTL, and dropped
when another process reports a subscription ended.
- GET /events/fetch reads the continuation from `cursor`, falling back to
`after`.
- Cursors from /events/fetch, /events/subscriptions and /events/kv-handles
are encrypted with sealCursor/openCursor; plain cursors still read.
- A durable subscribe counts again after its insert and removes its own row
when over the cap; the generation bump for that removal is published.
- Subscribe checks access before reading the anchor's owner, and an anchor
gone in between answers with the subject the caller sent.
- Dispatch reads at most the rows per token routing can use, scanning large
token hashes; a removal still settles every row on the removed anchor.
- puter.events.unsubscribe() stops routing after the server ends the
subscription or reports it gone; fetch() accepts `cursor`.
* fix(events): forward kv values only to regions with rows that ask
- A region counts its session rows asking for kv values per token and
announces the token's value variant (`v#<token>`) over the existing
`watch` item on the first such row, withdrawing it with the last. Add,
remove, reap, reanchor and refresh all keep it, reading `includeValue` off
the row as it is dropped.
- A kv write reads the value variants alongside the watch index and attaches
the value only to forwards bound for a region that announced one.
* fix(pagination): seal id-keyed cursors on every list endpoint
Apps, subdomains, workers, team (members, directory, teams, audit, own
audit), shares (inbound, outbound), the user-to-user permission audit and
readdir/descendant listings now hand out cursors sealed with sealCursor and
read them with openCursor, so a cursor no longer carries a row id. Plain
cursors issued before still read.
* test(events): pass socket ids to noteConnect in the billing tests
* fix(puter-js): free one-shot callbacks and listeners, guard message handlers
- setAuthToken starts the GUI cache refresh loop once instead of per call
- IPC reply callbacks are dropped after their reply; dehydrated callbacks stay
- AppConnection removes its window listener when the target app closes
- reply branches for app data and instance count drop their callback
- UI, Debug, tool and GUI-token message listeners no longer reject; an
unknown or failing tool answers toolResponse with an error
- parseResponse handles a blob response with no content-type
- retry waits remove their abort listener when the wait ends
- EmailConfirmationDialog escapes its message
- EventListener uses a real Map and snapshots listeners per emit
- Context get/set/runWithContext route every known key the same way
* fix(puter-js): time out requests that make no progress
- buildXhr starts an idle clock on send, reset by state changes and body
progress (upload progress only on requests already carrying a bearer)
- read-safe requests get 60 s while progress is observable, everything
else 15 min; timeout: 0 on fetchUrl/driverCall/sendWithRetry turns it off
- a timed-out read replays once; writes and AI calls never replay
- timeouts reject with code request_timeout on fetchUrl (TypeError),
driverCall, streams and the legacy XHR path
- document the limits in rate-limits-and-quotas and the chat stream note
* fix: only sign-in and picker popups hand the opener a token
Popup actions are an allowlist now: sign-in popups (no action, sign-in,
login, signup) deliver at boot, pickers deliver with the pick, and
request-permission keeps its exchange. Any other action runs no exchange
and posts no token. The consent check in postAuthActions applies to every
action that delivers at boot instead of only no-action and sign-in.
* fix: a dismissed account picker hands the opener no token
* fix(events): keep each region's presence item written only by that region
- A region that receives a forward for a pair it holds no socket for
retires its own presence item through the pin-ordered leave path,
gated by a region-shared claim. The sender no longer retires another
region's item, and the forward reply no longer carries noSocket.
- Presence reads skip items for regions outside the peer list instead
of retiring them.
- removeConnection decrements and deletes the count in one script.
- A failed join releases the join pin, and any connect that finds the
pin free writes the join.
- A single delivery treats a socket on any node of the region as local,
using the region-wide connection count, and picks the next remote
candidate by the regions already tried instead of by index.
- Correct the connection-count TTL comment.
* fix(events): count presence by live socket and honor the leave window region-wide
- A region's presence for a pair is a sorted set of socket ids scored by
last renewal under a new key, replacing the integer count. A socket
counts while renewed within the concurrency-slot window; add, renew and
remove each prune lapsed members and report the live count in one
script, so a node that dies stops holding the pair.
- The last socket going marks the pair as leaving region-wide for the
grace window, in the same script; a connect clears it. A forward that
lands on another node inside the window does not retire the item.
- Add the events.presence.write counter, labelled by op (join, leave,
refresh, retire).
* fix(events): resolve relative fs: paths like fs operations do
* docs(events): use bare relative paths in examples, test both relative forms
* docs(events): clarify that relative fs: subjects come back resolved
---------
Co-authored-by: Daniel Salazar <daniel.salazar@puter.com>
* fix: a picker popup hands its opener a token once the user has picked
A picker opens because the site asked, not because the user agreed to
anything, so the hand-off now waits for the answer the popup was opened
to get. The exchange still runs at boot — the app row and host_app_uid
are what the dialog is opened against — but the token travels with
fileOpenPicked, directoryPicked or fileSaved. Cancelling sends nothing.
deliversTokenAtBoot is the one gate every boot-time hand-off reads, so
the plain exchange, the cross-origin-isolated /login/set path, temp-user
creation and manual signup all follow the same rule. A failed exchange
still reports and closes, having nothing to defer.
* fix: the review's findings on the picker hand-off
Double-clicking a file answers the opener from openItem.js, not from the
UIWindow button, so that pick returned the item and no token. The
deliverer is a global now and every path that answers an opener calls
it: both UIWindow handlers, openItem, and the save flow.
A failed exchange posted `{success: false, token: null}`, and the SDK's
handler sets that as readily as a real token, so a picker that failed
wiped whatever the site already held. The failure now goes only where a
boot-time token would have.
* fix: settle the url file-access gate from a grant held on a folder above
The gate asks `/auth/check-permissions` with `app_uid`, which resolved
through the permission scan. `fs:` access inherits down a tree and only
ACL walks that, so a grant on a folder never answered for a file inside
it and the gate asked again, writing another per-file row on each Allow.
The `app_uid` branch now routes `fs:<uid>:<mode>` through
`ACLService.check` as the app-under-user actor, which is the same
authority the operation itself will face. Everything else keeps the
scan, including `puter.perms.holds`, which sends no `app_uid`.
* fix: the review's findings on the consent gate
A permission that is not a string reached `startsWith` and faulted the
whole request, where it used to answer false inside the catch; both
parsers now refuse a non-string, which also covers the other callers
that read a request's list.
The path spelling of a permission took the scan while the uid spelling
took ACL, so one file could answer two ways. Both spellings resolve to
their entry and take the same route.
The checks run together rather than one after another: the caller gives
up after five seconds and prompts, and a deep path is several reads.
* fix: pin the token sender, scope the grant flag, and stop two stragglers
Five of the SDK auth-plumbing items.
**The global `puter.token` handler pinned origin but not source.** Its three
siblings all pin both, and each says why: the GUI origin is shared by every
frame on that domain, so origin alone identifies a domain, not a sender.
`signIn` pins its popup and binds `msg_id`, the consent dialog pins its popup,
and the reauth wait pins the embedder. This one pinned the embedder only when
`env === 'app'`. Since `authenticateWithPuter` now hands off to `signIn`, which
adopts the token itself, the embedder is the only sender this handler is left
serving, so it pins in every environment.
**`is_grant_user_app_permission` was a request-wide flag held across an await.**
It tells the `app-root-dir` rewriter that a row is being written, and while it
is set that rewriter resolves to a real `fs:` permission instead of
`PERMISSION_FOR_NOTHING_IN_PARTICULAR`. Anything evaluating an `app-root-dir:`
permission beside the grant -- the grant route runs one under `Promise.all`
with fs work -- read the same flag. It now lives in a scope derived for that
call, so what the write needs to say about itself is not also said to its
neighbours. `runInDerivedContext` is new: `runWithContext` starts an empty
store, which would drop the `actor` the rewriter reads.
**`fs.stat` and `fs.readdir` deduped on parameters alone**, so two identities
in one realm could share a result. They now key on origin and token as
`os/user.js` does, which documents the reason.
**`signIn`'s relay poll never stopped.** Closing the popup settles the promise;
the loop kept asking `/login/wait` every second for the life of the page.
**`triggerReauth`'s five-minute timeout was never cleared**, so a successful
reauth left a timer to fire into a settled promise.
The concurrency test parks the rewrite and reads the flag while it is held --
my first attempt read before the flag was even set and passed either way.
* fix: scope the fs cache per identity, and retarget two paths that never ran
The remaining SDK auth-plumbing items.
**`puter._cache` outlived the identity that filled it.** `_clearAuthToken`
dropped the token and the stored session but left the cache and `whoami` as
they were, and its keys were raw paths -- `item:/alice/Documents` and a
`~`-relative key alike, identical whoever asked. A GUI user switch without a
reload could be served the previous user's stat or readdir. Keys now carry the
origin and token they were read as, built in one place so the writers, the
readers, the invalidations and the warm-up probes cannot drift apart, and
signing out drops the cache and `whoami` with the token.
**The driver permission prompt could not fire.** Both paths waited for a 200
carrying `{success:false, error:{code:'permission_denied'}}`. The driver route
answers `success: true` or throws, so a refusal arrives as a 403 -- the shape
the retry now reads. Its two tests synthesised the old envelope, so they passed
against a response the API does not send; they use the real one now, and the
prompt goes from never called to called once.
This is a behaviour change: a driver call refused for permissions will now
prompt, which is what the path was written to do. Happy to delete it instead if
the prompt is not wanted.
The half-written copy of the same idea in `utils.js` is gone -- it prompted for
one hard-coded interface and its retry was a `// todo`.
**The verification-gate single-flight was keyed on nothing**, so a phone-gated
request in flight beside an email-gated one took the email answer, replayed,
failed again and surfaced raw without ever asking. Keyed per gate: same gate
still shares one dialog, a different gate gets its own.
* test: ask the cache the same way the code does
The API suites built `item:<path>` keys by hand, so they stopped finding
anything once the real keys carried the identity they were read as. They go
through `fsCacheKey` now, which is the point: a test that spells the key itself
passes until the moment the spelling changes, and then fails for a reason that
has nothing to do with the behaviour it is guarding.
These run against a live server, so a plain `vitest run` does not reach them --
they were green locally and red in CI.
* fix: the review's findings on this branch
Eight, and three of them were regressions this branch introduced.
**Pinning the token sender to `globalThis.parent` in every environment broke
the case it was meant to protect.** On a top-level page `parent` is the page
itself, and the token arrives from a popup, so the handler stopped accepting
anything. The GUI posts `puter.token` to `window.opener` for the file pickers
too, not only for sign-in, and this handler is their only consumer -- a site
calling `showOpenFilePicker()` would have signed the user in, dropped the token
and run the following read with none. It now pins the embedder when framed and
a window this SDK opened when not; `UI.js` already tracked its picker popups
for the same reason, so that set is shared rather than duplicated.
**The GUI still spelled a cache key by hand.** `updateSubdomainsForItems` built
`item:<path>`, which no longer matches what the SDK writes, so its `get` could
not hit and its `set` wrote an orphan nothing reads and nothing evicts -- a
published website's badge quietly stopped appearing on a cached render. It goes
through `fsCacheKey` now, and that is the last hand-spelled key in the tree.
**Retargeting the driver permission prompt to 403 was wrong and is reverted to
the deletion the ticket asked for.** The route refuses with `forbidden`, not
`permission_denied`, so the retarget still missed it; meanwhile a driver that
raises `permission_denied` for its own reasons would have opened a prompt that
had nothing to do with the refusal. And the prompt asks for
`driver:<iface>:<method>` while the route checks `service:<driver>:ii:<iface>`,
so granting it changes nothing. The path cannot work as written.
The rest were pre-existing or latent:
`cacheWhoami_` wrote its answer without checking the token it was issued under
was still current, so a reauth landing mid-flight could repopulate `whoami`
with the previous identity. `_dropIdentityCaches` also left `whoamiCache_`,
which other call sites await.
`setAuthToken` now drops the identity caches when the token actually changes:
keys carry the token and the store is persisted with no TTL, so a rotation that
did not go through `_clearAuthToken` orphaned a whole key space with nothing to
reclaim it.
The verification-gate entry was registered after the body could already have
deleted it -- a synchronous throw stranded a resolved "not verified" in the map
for the life of the page. It registers first and clears on a microtask.
`signIn`'s poll is bounded as well as settle-checked: on a cross-origin
isolated page opened from a gesture, nothing watches the popup, so closing it
never settled the promise and the loop ran on exactly as before.
`runInDerivedContext` says in its docblock that callee writes do not survive;
only the one flag needs the scope, and a future caller should not expect to
hand a value back through it.
* fix: the review's findings on the SDK auth plumbing
**The relay bound stopped asking without saying so.** The deadline left the
`signIn` promise pending and its listener attached. On an isolated page opened
from a gesture the relay is the only thing that can settle it -- nothing watches
the popup there -- so a user who took longer than the deadline was left with a
promise that never resolved, and `authenticateWithPuter` queues every later
implicit-auth call behind it. It rejects now.
**Pinning the token source left `signIn`'s own popup untracked.** Only the
picker popups registered, so on a top-level page the global handler bailed on a
sign-in it had opened itself: `puter.onAuth` never fired, and
`puterAuthState.authGranted` was never set. `openAuthPopup` is the one place
auth popups are opened, so they register there.
**A closed popup is ours too.** Rejecting one that had already closed loses a
message from a popup that posts and then closes itself, which is why `UI.js`
tolerates the same case. Kept, with the set bounded instead.
**The gate key carried the factors**, so one route naming them and another not
opened two dialogs for the same gate. It keys on the gate, as `ctx.done` does.
**`fsCacheKey` left `~` unresolved**, so a writer and an invalidation spelling
the same place differently missed each other -- the drift the one key builder
was meant to end.
The dead `opts.permission` went with the path that used to read it.
Four workerd-runner failures are left: `ai > chat stream adapter supports
cancellation`, `kv > a GUI boot key written by another client reads fresh` and
two cache-warmer ones. All four fail the same way on `origin/main`, and the node
and browser runners pass.
* chore: one-line the comments added here
* chore: reflow the comments the removed retry cause left ragged
The 2FA spinner showed "verifying", the files selection bar showed a
"done" button, and two login and signup errors showed their key names,
because i18n() returns the key itself when en.js has no entry for it.
The login error now uses the existing login_email_username_required key,
which already has translations. A test checks that every literal i18n()
key in the GUI exists in en.js.
The in-flight claim's lease (2x the invoke timeout) is also a future score, so
polling heldForMs > 0 could pass before the failed attempt's hold was written:
the backoff test read 600 instead of 2000, and the 404/hang/429/suspended-owner
tests ended while their attempt was still in flight, letting it land in the
next test. Wait on the subscription's failure counter instead, which only moves
after the hold or gap marker is written.
Durable subscribe, list, fetch and unsubscribe now apply the same scoped
token check as the handler, worker and kv-handle routes:
`isAccessTokenActor(actor) && !isAccountContext(actor)`. An access token
an app issued (effectiveApp = the app) is treated the same as one the
account issued: subscribe answers 403 `events_durable_requires_account`,
list and fetch return an empty page, unsubscribe answers 404.
`rowInActorScope` returns false for these tokens, which covers durable
unsubscribe and ack. The session verbs that also use it already refuse
access tokens at the socket handshake.
Full-access tokens, app tokens and sessions are unchanged.
The duplicate-username check in the signup placeholder claim, change-username
and save-account runs before the user-row write, so a concurrent write can
take the name in between and the unique index rejects ours. Each path now
re-reads the holder from the primary on a unique violation and returns the
400 its pre-check returns for a taken name, instead of a 500. In all three,
the home-directory rename, rename budget charge and username-changed event
run after the user-row write, so nothing has happened yet when it fails.
The boot-key batch now starts only when puter.env is 'gui'. In every other
environment a get() of one of those names reads that key on its own, like
any other key, instead of fetching all 14 and serving them for 4 s.
- signup: a username taken by a concurrent signup between the duplicate
check and the insert now returns the pre-check's 400 instead of
surfacing the unique-key error as a 500. The holder is read from the
primary.
- fs move: when the UPDATE trips the (parent_id, name) unique key after
the collision probe missed the occupant, re-read it from the primary and
resolve it the same way the probe would have (overwrite, dedupe, or the
409 item_with_same_name_exists), then retry the move once.
- server health: a check that hits the per-check timeout reports how long
the event loop was blocked during it, like the database-liveness latency
failure already does.
- flush follows LastEvaluatedKey until the namespace is empty and rejects
when a batch delete fails; cached reads for every key a page returned are
still invalidated and the flushed event still fires.
- decr validates the amount map before negating it, so it accepts exactly
what incr accepts.
- del checks the key like every other store method (oversized key is a 400).
- list offset emulation treats a missing continuation key as end of data
instead of restarting from the first key.
- Array get maps results through a Map instead of a linear find per key.
- The driver gates includeTotal on credit like value listings.
- New driver caps in drivers/kv/limits.ts: 1,000 keys per array get, 1,000
items per batchPut (both rejected past the cap), list limit lowered to
1,000. Published in rate-limits-and-quotas.md and the set/list pages.
- puter.kv.list() no longer throws synchronously for invalid options; it
returns a rejected promise, or for stream: true an iterator whose first
next() rejects. Error codes unchanged.
- Writes through puter.kv (set, del, incr, decr, add, remove, update,
expire, expireAt, flush) stop the boot-key batch from serving the keys
they touch, so the next get() reads the store. Untouched boot keys are
still served from the batch.
- Client-side key/value size checks measure UTF-8 bytes, and values are
measured as their JSON encoding, matching the store. Same codes and
messages.
- stat: only path-addressed results are cached; uid-addressed stats no
longer share one cache key.
- readdir: overloads declare the page result for recursive + limit/offset
and for the positional page/stream forms.
- move: only a not-found destination lookup means "new path"; other
lookup failures reject. A uid destination skips the lookup.
- getAbsolutePathForApp always resolves paths. read/copy/move/delete use
getAbsolutePathOrUidForApp, which keeps sending UID-shaped strings as uids.
- Docs for read, copy, move, delete and readdir.
Two findings from review, both of which contradict what the last commit
claimed.
**The backfill could merge keys, not only split them.** I asserted it "can only
split rows apart, so it cannot collide". Not true, and reproducible: a row
predating the `clean_email` column has a NULL there, so its index key is
`lower(email)` -- filling it in folds `+tag`, and `john+news@me.com` lands on a
key `john@me.com` already owns. A legacy row holding an un-lowercased
`clean_email` collides the same way once lowercased. Either throws, and the two
engines that re-run every file on every boot would then refuse to start until
someone operated on the database by hand.
Both updates now fire only where the stored value is exactly what the old rule
produced and differs from the correct one. That makes "splits only" true: the
sole change is dots coming back. NULLs and legacy values are left alone.
The same predicate makes them write-idempotent. Postgres has no no-op-update
suppression, so without it every boot rewrote a tuple per Apple row forever.
**The pending-invite residue misdelivers rather than just going missing.**
`claimPendingShares` looks the row up by the canonical address, so whoever
confirms `jsmith@icloud.com` claims the invites addressed to
`j.smith@icloud.com` -- someone else's file, handed over on confirmation. I had
written this off in the PR as unfindable invites that "expire on their own";
`share` has no expiry column and the lookup has no time predicate, so they are
claimable indefinitely.
Fixed here rather than later: the typed address is kept as `invitedAddress`
whenever it differed from the canonical, which is exactly the collapsed rows,
so the invites recompute from it under the same guard.
Verified by seeding a database at the old version with the two rows that used
to abort the migration, and booting: it reaches version 87, repairs the
collapsed row, and leaves the NULL and legacy rows untouched. Replaying the
statements changes zero rows.
The sqlite migration list is hand-maintained, so the file alone did nothing —
a fresh boot still reported 87 files and stopped at version 86. MySQL and
Postgres read their directories, so only this one needed the entry.
Caught by booting a local instance rather than by the suite: the client tests
pin `CURRENT_SCHEMA_VERSION`, which fails once the entry exists but says
nothing while the file is simply unregistered.
Verified against a seeded database at the old version. `John.Smith@iCloud.com`
went from `johnsmith@icloud.com` to `john.smith@icloud.com`, `a.b+tag@me.com`
from `ab@me.com` to `a.b@me.com` — subaddressing still folded — and a gmail row
was left alone.
`cleanEmail()` applied gmail's `dots_dont_matter` rule to icloud.com, me.com
and mac.com. Apple allocates the exact string: `j.smith@icloud.com` and
`jsmith@icloud.com` are two mailboxes with two owners. Collapsing them made one
canonical address stand for both.
What that canonical address feeds:
- `findEmailOwner` falls back to it, so an OIDC sign-in whose verified claim is
the dotted address resolves to the account holding the undotted one, and
`linkProviderToUser` has nothing left to object to -- the provider verified a
real mailbox and the account's address is confirmed.
- The unique index on COALESCE(clean_email, lower(email)) reads it, so the
rightful owner of a dotted Apple address is refused signup because the
undotted account already "owns" it.
- Pending share invites are stored and claimed on it.
Subaddressing is untouched: Apple does fold `+tag`, so that rule stays. Gmail
keeps both rules -- there the fold is correct.
A migration recomputes `clean_email` for those domains on all three engines,
since rows written before this still hold the collapsed value and the index
reads the stored one. Verified the SQL against the function rather than by
inspection: same five inputs through `cleanEmail()` and through sqlite give
byte-identical output. It can only split rows apart, never merge them, so it
cannot collide with the unique index.
Not changed: `findUserByEmail` still matches on the canonical address. The
cross-match it documents is correct where the provider really does fold --
`foo.bar+tag@gmail.com` is the same inbox as `foobar@gmail.com`. Requiring an
exact match would break that for no benefit here; the fix is for the data to
say what the provider does.
The in-flight claim's lease (2x the invoke timeout) is also a future score,
so polling for "held" could read it before the failed attempt's backoff
hold was written. Wait for deferAfterFailure to settle instead.