Password signup, save-account and a provider-verified OIDC sign-up take an
address a team seat holds unconfirmed; the seat's copy is cleared at the
write. A seat address that was confirmed still blocks.
- /login/otp and /login/recovery-code share a per-account bucket charged on
wrong codes; an otp-login token completes one sign-in.
- /send-pass-recovery-email adds an IP bucket, a per-account cap and captcha,
and reuses the outstanding recovery token instead of rotating it.
- Reauth tokens carry the rejection reason; /signup re-attaches a temp
account only from an expired session.
* fix(events): resolve relative fs: paths like fs operations do
* docs(events): use bare relative paths in examples, test both relative forms
* docs(events): clarify that relative fs: subjects come back resolved
---------
Co-authored-by: Daniel Salazar <daniel.salazar@puter.com>
* fix: a picker popup hands its opener a token once the user has picked
A picker opens because the site asked, not because the user agreed to
anything, so the hand-off now waits for the answer the popup was opened
to get. The exchange still runs at boot — the app row and host_app_uid
are what the dialog is opened against — but the token travels with
fileOpenPicked, directoryPicked or fileSaved. Cancelling sends nothing.
deliversTokenAtBoot is the one gate every boot-time hand-off reads, so
the plain exchange, the cross-origin-isolated /login/set path, temp-user
creation and manual signup all follow the same rule. A failed exchange
still reports and closes, having nothing to defer.
* fix: the review's findings on the picker hand-off
Double-clicking a file answers the opener from openItem.js, not from the
UIWindow button, so that pick returned the item and no token. The
deliverer is a global now and every path that answers an opener calls
it: both UIWindow handlers, openItem, and the save flow.
A failed exchange posted `{success: false, token: null}`, and the SDK's
handler sets that as readily as a real token, so a picker that failed
wiped whatever the site already held. The failure now goes only where a
boot-time token would have.
* fix: settle the url file-access gate from a grant held on a folder above
The gate asks `/auth/check-permissions` with `app_uid`, which resolved
through the permission scan. `fs:` access inherits down a tree and only
ACL walks that, so a grant on a folder never answered for a file inside
it and the gate asked again, writing another per-file row on each Allow.
The `app_uid` branch now routes `fs:<uid>:<mode>` through
`ACLService.check` as the app-under-user actor, which is the same
authority the operation itself will face. Everything else keeps the
scan, including `puter.perms.holds`, which sends no `app_uid`.
* fix: the review's findings on the consent gate
A permission that is not a string reached `startsWith` and faulted the
whole request, where it used to answer false inside the catch; both
parsers now refuse a non-string, which also covers the other callers
that read a request's list.
The path spelling of a permission took the scan while the uid spelling
took ACL, so one file could answer two ways. Both spellings resolve to
their entry and take the same route.
The checks run together rather than one after another: the caller gives
up after five seconds and prompts, and a deep path is several reads.
* fix: pin the token sender, scope the grant flag, and stop two stragglers
Five of the SDK auth-plumbing items.
**The global `puter.token` handler pinned origin but not source.** Its three
siblings all pin both, and each says why: the GUI origin is shared by every
frame on that domain, so origin alone identifies a domain, not a sender.
`signIn` pins its popup and binds `msg_id`, the consent dialog pins its popup,
and the reauth wait pins the embedder. This one pinned the embedder only when
`env === 'app'`. Since `authenticateWithPuter` now hands off to `signIn`, which
adopts the token itself, the embedder is the only sender this handler is left
serving, so it pins in every environment.
**`is_grant_user_app_permission` was a request-wide flag held across an await.**
It tells the `app-root-dir` rewriter that a row is being written, and while it
is set that rewriter resolves to a real `fs:` permission instead of
`PERMISSION_FOR_NOTHING_IN_PARTICULAR`. Anything evaluating an `app-root-dir:`
permission beside the grant -- the grant route runs one under `Promise.all`
with fs work -- read the same flag. It now lives in a scope derived for that
call, so what the write needs to say about itself is not also said to its
neighbours. `runInDerivedContext` is new: `runWithContext` starts an empty
store, which would drop the `actor` the rewriter reads.
**`fs.stat` and `fs.readdir` deduped on parameters alone**, so two identities
in one realm could share a result. They now key on origin and token as
`os/user.js` does, which documents the reason.
**`signIn`'s relay poll never stopped.** Closing the popup settles the promise;
the loop kept asking `/login/wait` every second for the life of the page.
**`triggerReauth`'s five-minute timeout was never cleared**, so a successful
reauth left a timer to fire into a settled promise.
The concurrency test parks the rewrite and reads the flag while it is held --
my first attempt read before the flag was even set and passed either way.
* fix: scope the fs cache per identity, and retarget two paths that never ran
The remaining SDK auth-plumbing items.
**`puter._cache` outlived the identity that filled it.** `_clearAuthToken`
dropped the token and the stored session but left the cache and `whoami` as
they were, and its keys were raw paths -- `item:/alice/Documents` and a
`~`-relative key alike, identical whoever asked. A GUI user switch without a
reload could be served the previous user's stat or readdir. Keys now carry the
origin and token they were read as, built in one place so the writers, the
readers, the invalidations and the warm-up probes cannot drift apart, and
signing out drops the cache and `whoami` with the token.
**The driver permission prompt could not fire.** Both paths waited for a 200
carrying `{success:false, error:{code:'permission_denied'}}`. The driver route
answers `success: true` or throws, so a refusal arrives as a 403 -- the shape
the retry now reads. Its two tests synthesised the old envelope, so they passed
against a response the API does not send; they use the real one now, and the
prompt goes from never called to called once.
This is a behaviour change: a driver call refused for permissions will now
prompt, which is what the path was written to do. Happy to delete it instead if
the prompt is not wanted.
The half-written copy of the same idea in `utils.js` is gone -- it prompted for
one hard-coded interface and its retry was a `// todo`.
**The verification-gate single-flight was keyed on nothing**, so a phone-gated
request in flight beside an email-gated one took the email answer, replayed,
failed again and surfaced raw without ever asking. Keyed per gate: same gate
still shares one dialog, a different gate gets its own.
* test: ask the cache the same way the code does
The API suites built `item:<path>` keys by hand, so they stopped finding
anything once the real keys carried the identity they were read as. They go
through `fsCacheKey` now, which is the point: a test that spells the key itself
passes until the moment the spelling changes, and then fails for a reason that
has nothing to do with the behaviour it is guarding.
These run against a live server, so a plain `vitest run` does not reach them --
they were green locally and red in CI.
* fix: the review's findings on this branch
Eight, and three of them were regressions this branch introduced.
**Pinning the token sender to `globalThis.parent` in every environment broke
the case it was meant to protect.** On a top-level page `parent` is the page
itself, and the token arrives from a popup, so the handler stopped accepting
anything. The GUI posts `puter.token` to `window.opener` for the file pickers
too, not only for sign-in, and this handler is their only consumer -- a site
calling `showOpenFilePicker()` would have signed the user in, dropped the token
and run the following read with none. It now pins the embedder when framed and
a window this SDK opened when not; `UI.js` already tracked its picker popups
for the same reason, so that set is shared rather than duplicated.
**The GUI still spelled a cache key by hand.** `updateSubdomainsForItems` built
`item:<path>`, which no longer matches what the SDK writes, so its `get` could
not hit and its `set` wrote an orphan nothing reads and nothing evicts -- a
published website's badge quietly stopped appearing on a cached render. It goes
through `fsCacheKey` now, and that is the last hand-spelled key in the tree.
**Retargeting the driver permission prompt to 403 was wrong and is reverted to
the deletion the ticket asked for.** The route refuses with `forbidden`, not
`permission_denied`, so the retarget still missed it; meanwhile a driver that
raises `permission_denied` for its own reasons would have opened a prompt that
had nothing to do with the refusal. And the prompt asks for
`driver:<iface>:<method>` while the route checks `service:<driver>:ii:<iface>`,
so granting it changes nothing. The path cannot work as written.
The rest were pre-existing or latent:
`cacheWhoami_` wrote its answer without checking the token it was issued under
was still current, so a reauth landing mid-flight could repopulate `whoami`
with the previous identity. `_dropIdentityCaches` also left `whoamiCache_`,
which other call sites await.
`setAuthToken` now drops the identity caches when the token actually changes:
keys carry the token and the store is persisted with no TTL, so a rotation that
did not go through `_clearAuthToken` orphaned a whole key space with nothing to
reclaim it.
The verification-gate entry was registered after the body could already have
deleted it -- a synchronous throw stranded a resolved "not verified" in the map
for the life of the page. It registers first and clears on a microtask.
`signIn`'s poll is bounded as well as settle-checked: on a cross-origin
isolated page opened from a gesture, nothing watches the popup, so closing it
never settled the promise and the loop ran on exactly as before.
`runInDerivedContext` says in its docblock that callee writes do not survive;
only the one flag needs the scope, and a future caller should not expect to
hand a value back through it.
* fix: the review's findings on the SDK auth plumbing
**The relay bound stopped asking without saying so.** The deadline left the
`signIn` promise pending and its listener attached. On an isolated page opened
from a gesture the relay is the only thing that can settle it -- nothing watches
the popup there -- so a user who took longer than the deadline was left with a
promise that never resolved, and `authenticateWithPuter` queues every later
implicit-auth call behind it. It rejects now.
**Pinning the token source left `signIn`'s own popup untracked.** Only the
picker popups registered, so on a top-level page the global handler bailed on a
sign-in it had opened itself: `puter.onAuth` never fired, and
`puterAuthState.authGranted` was never set. `openAuthPopup` is the one place
auth popups are opened, so they register there.
**A closed popup is ours too.** Rejecting one that had already closed loses a
message from a popup that posts and then closes itself, which is why `UI.js`
tolerates the same case. Kept, with the set bounded instead.
**The gate key carried the factors**, so one route naming them and another not
opened two dialogs for the same gate. It keys on the gate, as `ctx.done` does.
**`fsCacheKey` left `~` unresolved**, so a writer and an invalidation spelling
the same place differently missed each other -- the drift the one key builder
was meant to end.
The dead `opts.permission` went with the path that used to read it.
Four workerd-runner failures are left: `ai > chat stream adapter supports
cancellation`, `kv > a GUI boot key written by another client reads fresh` and
two cache-warmer ones. All four fail the same way on `origin/main`, and the node
and browser runners pass.
* chore: one-line the comments added here
* chore: reflow the comments the removed retry cause left ragged
The 2FA spinner showed "verifying", the files selection bar showed a
"done" button, and two login and signup errors showed their key names,
because i18n() returns the key itself when en.js has no entry for it.
The login error now uses the existing login_email_username_required key,
which already has translations. A test checks that every literal i18n()
key in the GUI exists in en.js.
The in-flight claim's lease (2x the invoke timeout) is also a future score, so
polling heldForMs > 0 could pass before the failed attempt's hold was written:
the backoff test read 600 instead of 2000, and the 404/hang/429/suspended-owner
tests ended while their attempt was still in flight, letting it land in the
next test. Wait on the subscription's failure counter instead, which only moves
after the hold or gap marker is written.
Durable subscribe, list, fetch and unsubscribe now apply the same scoped
token check as the handler, worker and kv-handle routes:
`isAccessTokenActor(actor) && !isAccountContext(actor)`. An access token
an app issued (effectiveApp = the app) is treated the same as one the
account issued: subscribe answers 403 `events_durable_requires_account`,
list and fetch return an empty page, unsubscribe answers 404.
`rowInActorScope` returns false for these tokens, which covers durable
unsubscribe and ack. The session verbs that also use it already refuse
access tokens at the socket handshake.
Full-access tokens, app tokens and sessions are unchanged.
The duplicate-username check in the signup placeholder claim, change-username
and save-account runs before the user-row write, so a concurrent write can
take the name in between and the unique index rejects ours. Each path now
re-reads the holder from the primary on a unique violation and returns the
400 its pre-check returns for a taken name, instead of a 500. In all three,
the home-directory rename, rename budget charge and username-changed event
run after the user-row write, so nothing has happened yet when it fails.
The boot-key batch now starts only when puter.env is 'gui'. In every other
environment a get() of one of those names reads that key on its own, like
any other key, instead of fetching all 14 and serving them for 4 s.
- signup: a username taken by a concurrent signup between the duplicate
check and the insert now returns the pre-check's 400 instead of
surfacing the unique-key error as a 500. The holder is read from the
primary.
- fs move: when the UPDATE trips the (parent_id, name) unique key after
the collision probe missed the occupant, re-read it from the primary and
resolve it the same way the probe would have (overwrite, dedupe, or the
409 item_with_same_name_exists), then retry the move once.
- server health: a check that hits the per-check timeout reports how long
the event loop was blocked during it, like the database-liveness latency
failure already does.
- flush follows LastEvaluatedKey until the namespace is empty and rejects
when a batch delete fails; cached reads for every key a page returned are
still invalidated and the flushed event still fires.
- decr validates the amount map before negating it, so it accepts exactly
what incr accepts.
- del checks the key like every other store method (oversized key is a 400).
- list offset emulation treats a missing continuation key as end of data
instead of restarting from the first key.
- Array get maps results through a Map instead of a linear find per key.
- The driver gates includeTotal on credit like value listings.
- New driver caps in drivers/kv/limits.ts: 1,000 keys per array get, 1,000
items per batchPut (both rejected past the cap), list limit lowered to
1,000. Published in rate-limits-and-quotas.md and the set/list pages.
- puter.kv.list() no longer throws synchronously for invalid options; it
returns a rejected promise, or for stream: true an iterator whose first
next() rejects. Error codes unchanged.
- Writes through puter.kv (set, del, incr, decr, add, remove, update,
expire, expireAt, flush) stop the boot-key batch from serving the keys
they touch, so the next get() reads the store. Untouched boot keys are
still served from the batch.
- Client-side key/value size checks measure UTF-8 bytes, and values are
measured as their JSON encoding, matching the store. Same codes and
messages.
- stat: only path-addressed results are cached; uid-addressed stats no
longer share one cache key.
- readdir: overloads declare the page result for recursive + limit/offset
and for the positional page/stream forms.
- move: only a not-found destination lookup means "new path"; other
lookup failures reject. A uid destination skips the lookup.
- getAbsolutePathForApp always resolves paths. read/copy/move/delete use
getAbsolutePathOrUidForApp, which keeps sending UID-shaped strings as uids.
- Docs for read, copy, move, delete and readdir.
Two findings from review, both of which contradict what the last commit
claimed.
**The backfill could merge keys, not only split them.** I asserted it "can only
split rows apart, so it cannot collide". Not true, and reproducible: a row
predating the `clean_email` column has a NULL there, so its index key is
`lower(email)` -- filling it in folds `+tag`, and `john+news@me.com` lands on a
key `john@me.com` already owns. A legacy row holding an un-lowercased
`clean_email` collides the same way once lowercased. Either throws, and the two
engines that re-run every file on every boot would then refuse to start until
someone operated on the database by hand.
Both updates now fire only where the stored value is exactly what the old rule
produced and differs from the correct one. That makes "splits only" true: the
sole change is dots coming back. NULLs and legacy values are left alone.
The same predicate makes them write-idempotent. Postgres has no no-op-update
suppression, so without it every boot rewrote a tuple per Apple row forever.
**The pending-invite residue misdelivers rather than just going missing.**
`claimPendingShares` looks the row up by the canonical address, so whoever
confirms `jsmith@icloud.com` claims the invites addressed to
`j.smith@icloud.com` -- someone else's file, handed over on confirmation. I had
written this off in the PR as unfindable invites that "expire on their own";
`share` has no expiry column and the lookup has no time predicate, so they are
claimable indefinitely.
Fixed here rather than later: the typed address is kept as `invitedAddress`
whenever it differed from the canonical, which is exactly the collapsed rows,
so the invites recompute from it under the same guard.
Verified by seeding a database at the old version with the two rows that used
to abort the migration, and booting: it reaches version 87, repairs the
collapsed row, and leaves the NULL and legacy rows untouched. Replaying the
statements changes zero rows.
The sqlite migration list is hand-maintained, so the file alone did nothing —
a fresh boot still reported 87 files and stopped at version 86. MySQL and
Postgres read their directories, so only this one needed the entry.
Caught by booting a local instance rather than by the suite: the client tests
pin `CURRENT_SCHEMA_VERSION`, which fails once the entry exists but says
nothing while the file is simply unregistered.
Verified against a seeded database at the old version. `John.Smith@iCloud.com`
went from `johnsmith@icloud.com` to `john.smith@icloud.com`, `a.b+tag@me.com`
from `ab@me.com` to `a.b@me.com` — subaddressing still folded — and a gmail row
was left alone.
`cleanEmail()` applied gmail's `dots_dont_matter` rule to icloud.com, me.com
and mac.com. Apple allocates the exact string: `j.smith@icloud.com` and
`jsmith@icloud.com` are two mailboxes with two owners. Collapsing them made one
canonical address stand for both.
What that canonical address feeds:
- `findEmailOwner` falls back to it, so an OIDC sign-in whose verified claim is
the dotted address resolves to the account holding the undotted one, and
`linkProviderToUser` has nothing left to object to -- the provider verified a
real mailbox and the account's address is confirmed.
- The unique index on COALESCE(clean_email, lower(email)) reads it, so the
rightful owner of a dotted Apple address is refused signup because the
undotted account already "owns" it.
- Pending share invites are stored and claimed on it.
Subaddressing is untouched: Apple does fold `+tag`, so that rule stays. Gmail
keeps both rules -- there the fold is correct.
A migration recomputes `clean_email` for those domains on all three engines,
since rows written before this still hold the collapsed value and the index
reads the stored one. Verified the SQL against the function rather than by
inspection: same five inputs through `cleanEmail()` and through sqlite give
byte-identical output. It can only split rows apart, never merge them, so it
cannot collide with the unique index.
Not changed: `findUserByEmail` still matches on the canonical address. The
cross-match it documents is correct where the provider really does fold --
`foo.bar+tag@gmail.com` is the same inbox as `foobar@gmail.com`. Requiring an
exact match would break that for no benefit here; the fix is for the data to
say what the provider does.
The in-flight claim's lease (2x the invoke timeout) is also a future score,
so polling for "held" could read it before the failed attempt's backoff
hold was written. Wait for deferAfterFailure to settle instead.