This is what makes 0.9 a minor rather than a patch: a few responses change
shape, one protocol moves, and an approval changes whose authority an agent
acts with.
An approved action runs as the person who approved it. When an agent asks to
do something and gets a yes, the call runs with the authority of the person the
agent works for — what they can reach, it can reach, and nothing more. Only that
person can approve, on every surface, including a colleague typing "approve" in
a shared channel. Anyone else is refused with 403 APPROVAL_NOT_OWNED.
A revoked connection says so. When a provider reports that an OAuth grant
has been withdrawn, or an expired token has nothing to refresh with, the call
fails with 409 CONNECTOR_REAUTH_REQUIRED, naming the account to reconnect. It
used to be a 502, which read as an outage worth retrying. A provider that is
actually unreachable is still a 502 OAUTH_FLOW_ERROR.
Spending limits are a percentage.GET /usage/organizations/{organization_id}/limits no longer returns
limit_usd or remaining_usd; each window carries used_percent, as
/usage/me/limits already did. used_usd and reserved_usd stay. What a plan
includes can be retuned, and a dollar allowance invited planning against a
number Lemma does not promise.
Agent Host speaks protocol 3. The host's HTTP device API — the long poll,
event batches, harness publishing, pairing and self-revoke — is replaced by one
WebSocket link, including the two colon aliases 0.8.0 kept for hosts already in
the field. A 0.8 Desktop paired to a hosted server is not stranded: it shows
"Needs updating", keeps its pairing, and updating it is all it takes. Anything
else calling the old routes gets 410 AGENT_HOST_UPGRADE_REQUIRED.
No CLI version is refused./health reports latest_cli_version, a
response to an older lemma carries an X-Lemma-Latest-CLI header, and the CLI
turns that into a one-line hint at most once a day. LEMMA_UPDATE_CHECK=0 turns
it off.
Smaller things you may meet. Conversations list by last activity rather than by
when they began, each carries last_activity_at, and the list takes search.
App releases and function revisions page with limit and page_token. The
SNOOZE toolset is gone: wait_for replaces the snooze tool and is always
available. An existing agent or an imported pod bundle that names SNOOZE
simply drops it, but creating or updating an agent with it is refused as an
unknown toolset, so remove it from those requests. dm_conversation_reset_after_hours
leaves the surface response and is ignored if sent. lemma chat exits non-zero
when the run fails on the server, so a script that trusted exit code 0 will now
see the failure. The datastore changes socket rejects a bad session with a
4401 close, which the current SDKs and lemma datastore watch answer by
refreshing once and reconnecting. A deployment that sells plans can refuse a new
pod or member with POD_LIMIT_REACHED or MEMBER_LIMIT_REACHED; the
open-source build has no plans. A shared file link lasts 24 hours by default and
up to seven days, up from three hours and one day.
If you use lemma-sdk for Python or TypeScript, or the lemma CLI, upgrade the
package; the generated clients carry every field above.
If you host Lemma yourself: E2B_ALLOW_PUBLIC_TRAFFIC is gone, and sandboxes
are always created closed to the internet. A deleted pod's datastore tables are
dropped by a daily cleanup 30 days after the pod goes;
DATASTORE_ORPHAN_SCHEMA_RETENTION_DAYS changes the window and 0 turns it
off. The chat inactivity reset is one deployment setting,
SURFACE_DM_CONVERSATION_RESET_AFTER_HOURS.
A new workspace
The workspace the September development updates described is now the app, and
Desktop serves it too. The operator app remains as Lemma Harness.
Resources open beside the conversation, on a divider you can drag, and as a
sheet over it on a phone. Apps take the whole view, because squeezed into half
of it they were cramped. History is ordered by what was answered last, pages
all the way back, and can be searched. A pod shows who has access to it, and
adding someone by email invites them to the organization if they are not in it
yet, landing them in the pod when they accept. Installing from GitHub walks
through choose, review and install, with the plan in plain words. Billing and
organization settings live in the workspace, and a per-seat plan shows what the
whole team would pay rather than one seat.
"Teammate" now names the kind of thing you hire and nothing else. An agent that
is acting is called by its name, and humans are people.
Signing in, and arriving
Sign-in asks for your address first and answers with the one door it opens: a
password, a provider button, or a code already on its way. An account made on
WhatsApp is passwordless, and the web used to offer it a password, Google and a
reset email that never arrived. It gets its code now.
A new account is asked once for a name and, where the deployment can verify
one, a phone proved by sending a single WhatsApp message, so no unproved number
is trusted. Someone who messages on WhatsApp before they have an account can
make one there, and a pending pod invitation is accepted on the way rather than
producing a new organization. lemma auth login works against the new
workspace, and asks which account you are handing over before it does. A
session whose refresh keeps failing stops, rather than refreshing every second
and a half for ever.
A browser you can watch, and take over
The sandbox has a browser the agent drives and you can see, now streamed over
VNC: a click lands where you clicked, and copy and paste work both ways. You can
take the wheel, pop it out into its own window, or go full screen. When a site
needs a person to sign in, the agent asks in the conversation and you answer
there. The login is the browser's own, so it survives a suspend and a
reconnect, and you can make it forget a site. An agent can record what it did,
and a page a bot check blocked is told apart from a real one. The file explorer
is a tree, rooted where the files actually are, and reaches files over eight
megabytes.
On hosted Lemma, an update no longer wipes a workspace. Lemma's own code was
baked into the sandbox image, so a release that touched it replaced the image
and the disk that came with it. That code is now installed into the running
sandbox, and the image changes only when the machine underneath does.
Agents that wait, stop, and say what failed
wait_for is the one way an agent waits — on a time, a process in its sandbox,
or a sub-agent — and it costs nothing while it waits, where before it was
sleep in a shell or a poll that gave up after about half a minute. A run that
goes on very long, or keeps failing, pauses and asks whether to carry on through
the same approval that reaches Slack, Teams, Telegram and email. That is a
backstop set well above ordinary work, not a schedule.
A model that thinks before it answers is told how much room it has, so a run no
longer dies after a dozen tool calls with nothing to show. A tool that failed is
reported as a failure, where it was often read as a success, and a pause for a
person no longer counts as one.
Widgets are files the agent writes and then displays, so fixing one is editing
a few lines rather than retyping all of it. A widget can hand you your next
question, and fits what it shows. The pod's own agent knows its name, what it
may do, and its schedules, apps and channels, and reaches pod data through its
tools rather than shelling out. On a plan with a spending limit, that agent
could not be priced because of how it searched its tools, and so could not run.
It can.
Connectors and channels
Any file argument on any connector takes a pod file,
{"pod_path": "/me/q3.pdf"}, read with your access and handed to the provider
in the form it wants. Native Gmail can send a message and create a draft.
Shopify connects with OAuth and asks for the store's subdomain, instead of a
pasted token that expired and could not be refreshed. Eight Composio toolkits
that could only ever fail on connect now say they need the organization's own
OAuth app. A connector your organization had not set up yet no longer fails to
connect from the workspace.
A Slack app two people connected answers as one bot. A threaded reply is
answered only in a thread the agent is in, so two colleagues talking under
somebody else's message no longer start a run each. WhatsApp numbers come from
a pool, each answering with its own credentials. A WhatsApp conversation keeps
its history: the channel had been trimming it to the last forty messages, tool
chatter included, before the agent ever saw it. A private chat stays one
conversation when the number or surface it arrives on changes, and when a
message reaches nobody, the reason is recorded instead of lost.
Desktop
Desktop's settings live in the workspace, under This Mac. A coding agent there —
Claude Code, Codex or OpenCode — can be talked to while it works: a message is
steered into the running turn where the agent supports it, and delivered next
where it does not. It can be allowed to run commands on the Mac itself,
confined, and you can pick its model and effort. Server setup configures the
local server's model, email and keys, with a button that tests each. Sharing an
install asks natively before anyone else is let in, and walls their sandboxes
off from the Mac.
The in-app updater had refused every install over a data check that nothing
ever recorded. Only a Postgres major-version change stops an update now, an
update downloads only the parts of the runtime that changed, and quitting takes
seconds rather than seventeen.
Under the floor
The API drops a step that took about 36 seconds on every boot, because pod
schemas are readable from the moment they are created. The backend stops
growing in steps it never gave back: a statement cache kept a new entry for
every size of bulk write. A worker lane that died no longer leaves the worker
reporting healthy while a Redis stream grows until memory runs out, and streams
have a hard ceiling. One pod's bulk import can no longer hold up every other
pod's record triggers, and twelve places that held the event loop or a database
connection across slow work no longer do.
Organization editors can invite, re-role and remove people, up to their own
level and never an owner. An editor could make themselves admin of a pod they
had no access to, and cannot now. A personal agent is no longer named to
everyone on the organization page.
These changes are merged into main after v0.8.0. This is a development update, not a new stable version. Desktop builds are available through the September 24 nightly.
A workspace centered on your teammates
The new frontend brings the teammate workspace into the platform repository. Conversations, apps, files, and teammate profiles share one interface. The operator application remains available as Lemma Harness.
The landing experience now includes distinct working examples for launch production, sales, customer onboarding, and research. The examples show different jobs and the tools that go with them.
See what desktop is doing
First-start workspace downloads show progress when the size is known. The file explorer explains that the workspace is being prepared instead of presenting startup as a failure.
Codex tool cards show clearer commands, output, and failure details. Host agents are directed to the sandbox browser visible inside Lemma for browser work.
More predictable updates
Desktop update checks distinguish a PostgreSQL major-version change from an ordinary update. Nightly packaging and update-feed fixes improve the installation path.
A development update covering changes merged into main after v0.8.0. Availability depends on the build or deployment you use.
Keep channel conversations in the right place
Slack bot routing identifies the surface that owns the conversation. Threaded replies are handled when the thread is one the teammate participates in, rather than treating every thread as a request.
When a message cannot reach its destination, delivery failures provide a clearer explanation.
Take over the browser when needed
The browser pane improves human takeover and paste support. Sign-in and reconnection fixes help keep the browser usable as work passes between a person and an agent.
Manage the organization from the workspace
Billing controls and organization settings are available from the frontend. Per-seat pricing accounts for the team purchasing the plan.
A development update covering changes merged into main after v0.8.0. These are not release notes for a new stable version.
Widgets you can come back to
Widgets can live as files the agent can reopen and edit. Interactive widgets can hand a follow-up question back to the conversation, keeping the next step with the work that prompted it.
Rendering fixes improve how widget content and ordinary text are distinguished.
See the browser your teammate uses
Sandbox work gains a browser surface, file access, and saved browser logins. Browser recordings make it possible for an agent to capture its work for later review.
Better context for the job
Teammates receive clearer instructions about their identity, workspace, and available resources. Connector and schedule fixes address failures that prevented tools and scheduled work from running correctly.
This is the change that makes 0.8 a minor rather than a patch. Three
inconsistencies were spread across the API, each of them harmless on its own and
each of them something a client had to learn twice.
{org_id} becomes {organization_id} in the fourteen identity and agent
operations that still spelled it the short way; usage and connectors already
said organization_id, and there was no reason for the split beyond the order
the modules were written in. GET /pods/organization/{organization_id} was the
only org-scoped collection not nested under its organization, and is now
GET /organizations/{organization_id}/pods. And three routes named an action
with a colon while twenty named it with a slash, so the slash wins on count.
None of the operation ids move, so anything written against the SDKs by name is
untouched — both SDK layers call the generated clients positionally, and a path
rename stops at the generated signature. Two of the colon routes,
/agent-host/pairings:complete and /agent-host/events:append, keep an alias:
the Rust agent host hard-codes them and an installed host keeps calling what it
shipped with, so removing them outright would strand every paired computer
already in the field.
If you call the HTTP API directly, those are the three edits. If you use
lemma-sdk for Python or TypeScript, or the lemma CLI, upgrade the package
and there is nothing else to do.
Lemma runs on Windows
Lemma had never been run on Windows. It nearly booted after four blockers, and
then kept not-quite-working in ways that only running it could find: a path is
not an identity there, so a wrapper-launched agent outlived every run it
belonged to; an installer cannot replace a file that is open; every sandbox
start began the download the previous start was still doing; and an upgrade was
a choice between the old release and nothing.
Those are fixed, along with the shell paths that could each stop the app, and
the nightly can now build a Windows host pack — it could not, because of an API
nobody calls.
GitHub, Slack and Gmail are connectors Lemma speaks natively
GitHub is a real GitHub App now. Lemma acts as the App against an installation,
so a schedule keeps working after its author leaves and a triggered run has an
identity with nobody present; and as the person for the sandbox's git and gh
and for pod publishing, because work in someone's checkout should carry their
name. GitHub events wake a pod. "GitHub is connected" used to be true and
useless — authorizing is not installing — and now says which of the two it
means. A published pod no longer carries its author's installation with it.
Slack and Gmail stop going through a vendored package and become native OpenAPI
connectors, which meant teaching the shared machinery the three things they
need: form-encoded bodies, failure reported inside a 200, and cursor
pagination. Slack answering {"ok": false} with HTTP 200 had been reaching
agents as a successful result, so an agent read a failure as data.
An MCP server connected with an API key came up with no tools; it comes up with
its tools, described in the server's own words, on a session that survives.
A run says what it costs, while it runs
Model work that had been running for free is metered per request, budgets are
enforced as the run proceeds rather than after it, and the usage screens say
what a cost is made of and how much of a plan is left. A refusal now names the
budget it hit, which is how an agent that could not look at an image explained
itself instead of failing the whole run over one picture.
Apps and functions can be rolled back
Both keep a bounded version history, and a rollback is a safe operation rather
than a redeploy of something you hope you still have. An app's version history
shipped in an earlier version with no way to open it; there is a way now.
Lem is a real agent
The pod's assistant had no agents row — a conversation with a null agent id
was the assistant — so it could not be scheduled, could not be a trigger's
target, and every surface that wanted to name it faked it differently. It has
one row, one identity and one name everywhere, and a schedule can wake it with
an instruction for the occasion rather than an agent whose entire purpose is one
sentence.
Self-hosting on a server
There were four ways to run Lemma and none of them was "on a server".
deploy/compose/ is that missing path: api and worker as separate containers, a
one-shot migrate that stops a deploy rather than crash-looping against a schema
it cannot use, and images pinned by digest from the release manifest — so a
compose deployment of a version runs the same bytes as a desktop install of it.
Desktop stops freezing, and stops being unreadable
Five locks were held across work that waits, and the tray froze for eleven
seconds. The splash parsed 356 KB before it could draw anything. A killed
install left 2.2 GB behind. A guest asked a system with no init to power itself
off. Those are gone, along with a sandbox that could quietly fill the guest's
disk, and a settings pane that asked a dead daemon for ever.
The agent host was 16,678 lines in ten files. Nothing under desktop/ is over
600 lines now, daemon.rs is not 3,261 of them, and the twenty guards that had
been reading the wrong file read the right one.
Under the floor
The composition root is empty and then gone: modules talk through contracts and
events, and imports through the root fell from 195 to 61 on the way. The
standard the code is held to is written down, with an id per rule and the check
that enforces it, and the gates were changed to match what it says. A unit test
can no longer reach production. A default a test cannot replace is a build
failure. Both DMG pipelines verify the same things, after each had been checking
something the other thought mattered.
A forged From: could make a mailbox write to a stranger, and a channel could
make an agent answer nobody. Neither can now.
0.7.1 shipped an auto-updater and a signing key. What it did not have was a
feed a pre-release build could follow, or a version it could compare — every
nightly called itself the same number, so the updater asked whether 0.7.1 was
newer than 0.7.1 and was always told no. The update path only ever ran on
release day, which is the worst day to discover it.
Nightlies now follow a feed of their own at an address that does not move, and
carry an ordered version, so the update they offer is the one after the one you
have. That means the path a real release takes is exercised every night rather
than once a version.
Two things behind it were worth fixing on their own. The version a build stamps
lives in three places, and stamping only two of them produced an install that
met its own runtime and refused it — so the build now fails if those numbers
disagree, naming both, instead of signing and publishing something that cannot
start. And every installed runtime used to be kept forever: 1.7 GiB per version,
in Application Support, discovered from a full disk. An install now keeps the
release it is running and the one before it, which is enough to step back
without downloading it all again.
An agent run can take as long as the work does
A local agent working steadily on a real task was stopped at minute fifty:
Agent Host run deadline elapsed. The number came from the credential the run
was dispatched with, which expired in an hour and which nothing renewed. Something
renews it now, and has for a while — but the ceiling stayed behind, and every run
since had been cut short for a reason that no longer applied.
Runs may now last four hours. A host that crashes, sleeps or loses its network
is still failed within about two minutes, because that was never the deadline's
job: the run lease notices, and it is ninety seconds long.
An Agent Host that outlived the app that started it used to hold its workspace
lock forever, so every later launch found local agents unreachable and said only
"Fetch is aborted". A leftover is now reclaimed before the next one starts.
Conversations you can keep talking to
A running conversation accepts a follow-up message instead of making you wait
for the turn to end. A stopped run is one you can carry on from. An interrupted
run resumes from the queue rather than waiting for a sweep, and one busy
conversation no longer freezes the rest of the pod.
Around that: sub-agents appear as chips with tabs that name themselves,
approvals show the one that is actually pending and look approved the moment you
approve them, and a message typed past a card is treated as a message rather
than as that card's answer.
Documents, files and apps
A .docx renders as the document it is instead of a parser error. A shared file
renders the way the workspace renders it, rather than redirecting past it. An
image somebody sends is one the agent can open. A published app can be added to
a home screen, and an app an agent has just built opens instead of reporting
itself unavailable.
A link inside an app that points at one of your pod files used to end on a raw
JSON error in a window with no way back. It now lands on a page that says what
happened and offers to open the file in your workspace.
Pods, exports and surfaces
Pods can be renamed, and conversations can be named, archived and pinned.
Pod export stopped leaving files behind. It had been reading a preview of the
file tree — capped at three entries per directory — and treating it as the
pod's contents, so a migration of 266 files exported 104 of them and reported
success. Skills were dropped from every bundle. Both are fixed, the caps that
remain say what they skipped, and a bundle this version writes can be imported
by it.
Surfaces now share one delivery seam, so email stops being the exception, and
each agent can have its own Slack bot instead of the second one being
unreachable.
Faster, and quieter about it
Cold start is about seven seconds shorter: the window sizes itself correctly,
the Continue screen stops being in the way, and the backend spends less time
importing things it does not need yet.
Failures that nobody was watching now reach Slack, a backend whose worker has
died reports itself unready instead of healthy, and the scenario suite stopped
reporting its own noise as failures.
Installing the 0.7.0 nightly and watching it fail surfaced two ways the app
could lose a user's data outright: Postgres 18 moved where it stores its
cluster and the volume mount never followed, so a fresh install wrote its
database into a container's throwaway volume instead of the one people are
told holds their data; and a secrets file that went missing — not just
corrupted — was treated as a first run instead of the permanent loss it is.
Both are fixed, and the auto-updater has a signing key for the first time, so
a release can finally be followed by another one without a manual reinstall.
That sits on top of a full adversarial review of the app, its CI, and its
recovery paths, which found and closed 10 blockers, 26 serious issues, and
40+ polish items — the app now gets itself unstuck instead of leaving a dead
end, and a quit that used to hang now closes. Pod apps and pod functions,
which worked everywhere except Desktop, now carry a session and an address
like they do on the web.
Underneath, the Agent Host — the bridge to Codex, Claude Code, and other
ACP-compatible coding agents — reattaches to a run you walked away from,
retries a failed run without dying on two missing methods, and stops
rejecting a run for naming the harness revision the host itself just
published.
Agents remember, and you choose what they run on
Agents can now hold onto things: /memory for facts a whole pod should
share, /me for what's private to you, both built on the pod-file model that
already backs everything else an agent reads. Memory is a capability you
grant like any other, not something every agent silently has — and it
collapses what an agent can do from twelve switches down to five.
Choosing a model is a popover now, not a window that stops the app — the chip
in the composer opens with the default, five recent picks by number key, and
the full catalog one row away. Models themselves moved into Pod settings,
next to Access and Automation, because every question people ask of the
catalog — what a chat runs on, why an agent can't reach a laptop — starts
from inside a pod. And an agent's reasoning is recorded as reasoning: it can
no longer be shown, or returned, as its answer.
Messaging gets more precise
An agent can now say where a message goes. message_user takes a channel —
email, Slack, Teams, Telegram, WhatsApp — and either delivers there or
refuses, never silently swaps it for another. A pod's assistant gets its
inbound mailbox the moment the pod exists, not the first time it happens to
send something. A second Slack workspace or Teams tenant no longer goes
permanently unreachable, WhatsApp gets progress messages on runs that can't
stream, and a voice note is transcribed once, in the language it was actually
spoken.
Security and correctness sweep
Every open CodeQL finding is closed, ruff format is now enforced across all
883 first-party Python files, and the scenario suite runs at the SSRF
strictness a real deployment uses instead of a permissive stand-in. Closing
the last 19 gaps in the scenario suite's coverage found and fixed 17 real
bugs on the way, including one that let a pod's sole owner make it
permanently unadministrable.
Fixes
A widget in a conversation renders at its own height instead of being cut
off at 480px, and opens in its own tab instead of overwriting another
Retrying a failed run works again on every surface that offers it — the
API, Telegram's /retry, and the retry button
A message sent while an agent is working now reaches the run that's
already going, instead of starting a second one
Files sent from a conversation arrive as an attachment, not a link only a
Lemma account can open
The pod settings tab formerly called "Access" is now "Members"
Agents stream their replies into Slack as they think, App Home carries setup,
and a pod can be configured without leaving the workspace.
Content and documentation
Docs, the blog and the changelog now render from MDX with a shared component
vocabulary, syntax highlighting that follows your appearance, and structured
data on every public page.
Fixes
Share links no longer reflect provider errors back to the page
Callback pages share one design across every provider