$ Latch v2: Self-Heal Local-Exec and Driving Grok Bot From Grok Bot
Latch v1 gave the agent hands. Version 2 adds bounded recovery for that path: a LaunchAgent that relaunches Grok Bot when the process is down or the heartbeat is stale, an optional outside heal request for disconnects the Mac cannot see, and a bounded way for Latch to operate Grok Bot's own UI when structured tools run out. Relaunch proven on self-heal kit 1.4.2; current kit 1.5.0 adds observe-only logging for silent cloud drops.

When I first wrote about Latch, the problem was interaction.
An agent could be perfectly capable of reasoning about what needed to happen next and still get stuck because the next action lived on my Mac. There was a button to click, a dialog to inspect, something in Terminal or Chrome, or some other piece of graphical state that did not have a useful API.
Latch became the small layer that gave the agent hands.
The second version exposes a different problem: giving the agent hands is not enough if the path to those hands can disappear.
And, once you accept that the Mac UI itself is a legitimate bounded interface, another slightly strange thing follows: sometimes the application Latch needs to operate is Grok Bot itself.
So Latch v2 has two additions that look different on the surface but are really extensions of the same boundary:
- relaunch Grok Bot when its process is down or its heartbeat is stale, or when an outside beacon requests it;
- let the agent use the local UI as a controlled fallback when structured interfaces run out.
I think this makes the role of Latch clearer.
Latch is becoming the Mac-side reliability and interaction layer beneath the agent — not another agent inside the harness.
Latch v1 gave the agent hands. V2 makes the path to those hands more recoverable.
The failure that the agent cannot fix
The first new capability came from a fairly quiet failure.
The Mac was awake.
The network was there.
Applications were still running.
But Grok Bot no longer had its local-exec connection to the machine.
ListMachines reported the Mac as disconnected, which meant commands against that machine were effectively finished until the connection returned.
This is different from the laptop sleeping or losing Wi-Fi. The machine can be healthy while the local bridge is not.
The obvious fix is also obvious to a human:
close Grok Bot, reopen it, wait for the machine to reconnect.
The problem is that the agent cannot necessarily do that.
Once local-exec is gone, asking the same agent to run a command on the Mac to restore local-exec is circular.
It is roughly:
Cloud agent
|
| local-exec
v
Grok Bot.app
|
v
MacIf the middle link disappears, anything above it can understand the failure but cannot use that path to repair the failure.
This is not really a prompting problem.
It is a supervision problem.
And the recovery mechanism has to sit outside the failure domain it is responsible for repairing.
Put the watchdog below the agent
On macOS, there is already a boring and well-understood place for this kind of responsibility: launchd.
So Latch v2 adds a LaunchAgent that runs independently of the Grok Bot session and checks whether the local-exec side still looks healthy.
The important distinction is not that Latch has a shell script now.
The important distinction is that the script does not need the agent connection in order to run.
At a high level, the installation looks like this:
| Piece | Responsibility |
|---|---|
~/Library/Application Support/Latch/bin/grok-bot-local-exec-heal.sh | Detect the unhealthy local state and perform recovery |
~/Library/LaunchAgents/com.latch.grok-bot-local-exec-heal.plist | Run the watchdog periodically and at login |
~/Library/Logs/GrokBotLocalExecHeal.log | Human-readable recovery log |
~/Library/Logs/GrokBotLocalExecHeal-last.json | Small machine-readable recovery receipt |
~/Library/Application Support/Latch/grok-bot-local-exec-heal.disable | Explicit soft-off switch |
~/Library/Application Support/Latch/grok-bot-local-exec-heal.request | One-shot “relaunch now” from anyone who can still reach the Mac |
The self-heal kit is open source, and its README carries the install details and knobs. I don’t think the interesting part of the design is whether a heartbeat timeout is 150 or 180 seconds.
The interesting part is the contract.
Latch treats the bridge as unhealthy when one of a small set of local signals says it is no longer viable: the expected process is gone, the heartbeat has gone stale or stopped moving, or startup reached a non-ready outcome. Someone who can still reach the Mac can also ask for a relaunch by leaving the request file.
The LaunchAgent checks periodically. When it decides recovery is warranted, it tries to relaunch Grok Bot gently rather than immediately killing it.
The sequence is deliberately uninteresting:
detect unhealthy state
|
v
request normal quit
|
v
wait
|
v
reopen Grok Bot in the background
|
v
wait for local readiness
|
v
record healed, or say why not
|
v
cool down before another recoveryThere is a fallback if the process refuses to exit, but the normal case is intended to be uneventful.
The readiness step matters more than it looks. A process that exists is not a connection that works. The kit records healed only when Grok Bot is running again and its heartbeat is fresh, or its startup reports ready, within a bounded wait. Otherwise the receipt says heal_incomplete or heal_failed, with a hint about what to escalate.
Relaunched is not the same as recovered.
That is what I want from this layer.
Recovery infrastructure should be less clever than the system it supervises.
What it does not recover
It is important to keep the boundary honest.
The watchdog does not wake a sleeping Mac.
It does not open a closed lid.
It does not restore a dead network connection.
It does not sign the operator back into an account.
It does not bypass macOS security boundaries.
It handles a narrower local case: the Mac is available, but the expected Grok Bot process is gone or the heartbeat is stale. A configured beacon also lets an outside box ask the Mac for the same relaunch.
Outside of that, Latch fails closed.
I would rather have an explicit edge where the human has to return than describe every unavailable capability as something automation should somehow solve.
There is another tradeoff worth saying directly.
If I intentionally press Cmd+Q, the watchdog cannot read my mind and know whether I meant “quit forever” or “the app died.”
Unless the disable file is present, Grok Bot can therefore come back roughly a minute later.
That is the cost of choosing unattended recovery.
The fix is not another heuristic. The fix is a visible off switch.
Predictable automation is better than automation that tries too hard to guess intent.
The disconnect the Mac cannot see
One failure sits awkwardly between what the watchdog fixes and what it leaves alone.
Grok Bot’s cloud side knows whether a machine is connected. ListMachines reports it as connected. The Mac never gets that flag. Nothing on disk and nothing on localhost mirrors it in a way a LaunchAgent could use.
So the cloud can decide the Mac is disconnected while, on the Mac, the process is alive and the heartbeat keeps moving. To the watchdog, that machine looks healthy. By every signal it has, it is.
The receipt is honest about this. GrokBotLocalExecHeal-last.json carries a field called cloudConnectObservable, and it is always false.
It would be easy to “fix” this by tightening the thresholds until the watchdog relaunches more often. That would only turn a missing signal into guesswork, and a relaunch on a guess is just a new way to interrupt the person using the machine.
The signal has to come from the side that can see it.
An outside signal, with the Mac still owning the fix
So the kit has an optional worker-beacon/ mode.
The agent that actually sees the disconnect can POST a heal_request to a small Cloudflare Worker the operator deploys. The Worker is a heal-only inbox. It accepts one kind of request in one strict shape. Requests are short-lived and single-use, and the Worker rate-limits them per machine.
The Mac polls that inbox on its normal tick. Nothing calls in to the Mac. When a request is waiting, the LaunchAgent runs the same quit-and-open heal it runs for a local request file, records reason=beacon_request, and skips the normal cooldown. The readiness gate still applies.
agent sees connected=false
|
| POST heal_request
v
operator's Worker (heal-only inbox)
^
| outbound poll, once per tick
|
LaunchAgent on the Mac
|
v
same quit, reopen, and readiness gateThe Worker applies the limits. The Mac still owns the recovery.
A few boundaries hold it together:
- It is off unless the operator configures it. Without a beacon, the kit is exactly the local watchdog above.
- The off switch still wins. While the disable file is present, the heal is skipped and Grok Bot is left alone.
- One beacon relaunch per outage. A second request within the hour after a beacon relaunch is not honored. That is the point where a human should look.
- The Worker can only ask for a heal. It cannot run anything else on the Mac, and its tokens live outside the repository.
Kit 1.4.2 also makes the installer keep an operator’s beacon settings across a reinstall, so updating the kit does not quietly switch the outside signal off.
What has been proven, and what has not
On October 5, at about 11:23 AM PT, the live end-to-end test of that path passed. The kit’s notes call it Prove C. The request came from a remote box, and nothing touched the request file on the Mac. Grok Bot relaunched, the pid changed, and the receipt came back healed and ready.
The negative cases passed too: a malformed request, an expired one, a replay, and a second request inside the rate limit were all refused.
A separate live check covered the off switch. With the disable file present, the heal was skipped and the pid did not change. With the file removed, the agent went back to its normal checks.
That proves the outside signal can reach the existing heal path.
It does not prove that Latch can detect every outage by itself.
On October 7, Prove D confirmed that gap: the cloud-connection helper exited while the main app and heartbeat still looked healthy, so kit 1.4.2 kept reporting healthy.
Kit 1.5.0, now released, detects and logs that silent cloud drop in observe-only mode; it does not heal it.
And the limits from before do not move. Latch still cannot wake a sleeping Mac, open a lid, restore a lost network connection, or infer ListMachines.connected locally. With 1.4.2, an agent that can see the disconnect still has to send the request.
The agent also needs a recovery protocol
Putting a watchdog on the Mac only solves half of the problem.
The agents using the machine also need to know what a disconnection means.
Without that shared behavior, it is easy to turn one dead connection into a pile of retries.
The Latch-side protocol is intentionally small:
ListMachines
|
| connected=false
v
soft-park Mac-dependent work
|
v
wait for the local watchdog, then re-poll
|
| still connected=false
v
send one heal_request, if a beacon is set up
|
v
re-poll ListMachines
|
| connected=true
v
read GrokBotLocalExecHeal-last.json
|
v
resume from verified stateIf the machine does not come back after the recovery window, then ask the human.
Do not invent a remote wake capability.
Do not keep throwing commands at a transport that has already told you it is unavailable.
Do not send another heal request because the first one did not work. One per outage, then escalate.
Do not assume that a reconnect proves what happened. Read the receipt.
This distinction between the skill and the supervisor has become useful to me.
A skill can describe what an agent should do around a failure.
The LaunchAgent can actually remain present when the agent path disappears.
One describes behavior.
The other enforces a piece of local reality.
That fits closely with how I have been thinking about harness-driven development. The harness is not only the model or the prompt. It includes the operational pieces that make the model’s actions dependable.
Long-running agents make this much easier to see.
Then the remaining interface is Grok Bot itself
The second half of Latch v2 sounds more recursive than it really is.
Latch can now drive Grok Bot.app.
So, yes, there are situations where Grok Bot can effectively use Latch to operate Grok Bot.
That is a funny sentence.
But I don’t think the recursion is the interesting idea.
The interesting part is what happens when the structured interfaces run out.
Most agent systems have some hierarchy of preferred interaction:
API / structured tool
↓
CLI / shell
↓
application-specific automation
↓
computer UII would nearly always prefer the top of that list.
Structured tools are easier to reason about, easier to verify, and much less ambiguous than moving a cursor across a screen.
But the existence of a better interface does not mean every action has one.
There are still product surfaces where the last useful move is graphical.
Grok Bot itself contains some of them.
An Allow card may need a click.
An Assist flow may need to be advanced.
A draft may need its Send action.
A share or publish surface may need to be opened.
A setting may exist only in the desktop interface.
These are not all the same kind of action, and Latch should not pretend they are.
But they share something important: the interface is the UI.
When that happens, the UI can become another bounded tool surface.
The UI is a tool, not permission to do anything
This is where I think agent-first development becomes relevant.
Designing for agents does not mean eliminating graphical interfaces or pretending an API exists for everything.
It means being explicit about what capabilities an agent has, how it can discover them, and where the boundaries are.
Latch treats computer use the same way.
Grok-Bot-on-Grok-Bot does not mean:
the agent is allowed to click whatever gets in its way.
It means a registered Mac can expose a constrained interaction capability with rules around it.
Those rules remain intentionally strict:
- mouse-first interaction;
- no HID keyboard injection;
- clipboard plus human paste where typing is required;
- macOS Accessibility remains an explicit permission boundary;
- only registered Macs participate;
- the local Mac lease is always released after the interaction;
- unexpected states fail closed.
Those rules came from Latch v1 and I do not want v2 to weaken them just because the UI happens to belong to Grok Bot.
There is a major difference between performing a gesture and making a judgment.
Computer Use can click an Allow card when the authorization to do so is already part of the task.
That does not imply that Latch should manufacture authorization where none was given.
It can advance an Assist form whose intended values are already known.
That does not mean it should invent missing information.
It can move through a mechanical share flow.
That does not turn every consequential confirmation into something an agent should silently approve.
The UI is another interface.
The human boundary still exists above it.
Why mouse-first still matters
There is a temptation, once you have local computer access, to turn the Mac into a general input-injection surface.
Latch deliberately does not do that.
For graphical work, the mouse is easier to constrain and easier to observe.
For text, the preferred path remains clipboard-based rather than pretending to be an invisible hardware keyboard.
That is partly a safety decision, but it also makes the automation easier to reason about.
If Latch puts text on the clipboard and a human paste action completes it, the transition is explicit.
If an agent can synthesize arbitrary keyboard events at arbitrary moments, the boundary between application automation and control of the entire user session gets much fuzzier.
I already learned in the first version of Latch that macOS Accessibility deserves respect.
V2 does not reinterpret that lesson.
It builds on it.
Both features are actually the same architectural move
Self-heal and Grok-Bot-on-Grok-Bot look like two unrelated release notes.
One is a LaunchAgent.
One is computer use.
But both exist because the agent itself should not own every layer it depends on.
The self-heal path says:
If the agent loses its path to the Mac, something below that path must be able to attempt a bounded relaunch when it has a reliable signal.
The UI path says:
If the agent reaches the end of its structured interfaces, something below the reasoning layer can provide a bounded physical interaction surface.
In both cases, Latch sits closer to the machine than the agent does.
That gives the architecture a clearer shape:
┌──────────────────────────────────────┐
│ Agent / Harness │
│ reasoning, planning, skills, tools │
└──────────────────┬───────────────────┘
│
│ structured intent
▼
┌──────────────────────────────────────┐
│ Latch │
│ │
│ reliability interaction │
│ ─────────── ──────────── │
│ self-heal mouse/UI │
│ receipts clipboard │
│ outside signal lease control │
│ │
└──────────────────┬───────────────────┘
│
▼
┌──────────────────────────────────────┐
│ macOS │
│ Grok Bot · Chrome · Terminal · UI │
└──────────────────────────────────────┘That is a more useful mental model for Latch than “a bot that can click things.”
It is the local reliability and interaction layer underneath the agent.
The release is intentionally small
This is Latch v2: relaunch proven on self-heal kit 1.4.2, with 1.5.0 now detecting and logging silent cloud drops without healing them.
The version includes:
- a Mac LaunchAgent that relaunches on a down process or stale heartbeat, with a readiness gate;
- an optional outside signal through
worker-beacon/, for the disconnects the Mac cannot see; - bounded Grok Bot UI operation through Latch’s existing computer-use discipline.
The exact timers, plist, health checks, install flow, and operational details belong in the kit’s README rather than turning this post into a deployment guide.
Latch itself is public as a Grok Bot template:

Latch by Kobi
Drives a real Mac safely for Grok Bot: mouse-first UI automation, Accessibility without lockouts, and a hard hand-back when you need the machine. For anyone who wants Mac UI playbooks other bots can hand off to.
There is also an observable product gap underneath the watchdog.
Ideally, Grok Bot’s own local-exec lifecycle becomes reliable enough that an external watchdog is unnecessary. And if Grok Bot wrote its own connection state somewhere on the Mac, no bot would have to declare the machine down. If you hit the same class of failure, the useful long-term thing to do is also send it upstream through Help → Send Feedback.
I do not particularly want Latch to own process resurrection forever.
That would be missing the point.
Infrastructure can be temporary and still be worth building
One thing I increasingly like about building around fast-moving agent systems is that some infrastructure is allowed to be temporary.
There is a tendency in software to treat deletion as evidence that the original work was wasted.
I think the opposite can be true here.
If Grok Bot eventually absorbs reliable local-exec recovery, the self-heal LaunchAgent should disappear.
If the Mac can ever read its own connection state, the beacon should go first.
If richer structured interfaces replace one of the UI interactions, Latch should stop clicking that UI.
If macOS or the agent harness grows a safer primitive, use the safer primitive.
The goal is not to preserve Latch’s implementation.
The goal is to preserve the operating capability.
That is also why I would rather ship a small watchdog now than turn “the agent sometimes disconnects” into a permanent ritual where a human checks the Dock every few hours.
The current primitive is appropriate for the current gap.
When the gap moves, the primitive can move with it.
Latch v1 started with a fairly practical question: how do I let an agent safely operate the Mac when APIs stop?
V2 adds another one: what can attempt to restore that capability when the agent itself cannot reach it?
The answer, at least for now, is to put a deliberately boring layer underneath the clever one.
And once that layer exists, even Grok Bot’s own UI becomes just another bounded surface it can mediate.
That is the practice I care about more than the slogan:
put reliability below the thing that can fail, keep interaction narrower than authority, and delete the workaround when the platform finally makes it unnecessary.