Repository navigation
When the agent controls the desktop — tracking accountability gaps in production deployments #297
Replies: 3 comments
|
Desktop agents make the accountability gap visceral — when the agent controls the mouse and keyboard, every action is irreversible in a way that API calls often aren't. The approach I've been building: declare behavioral constraints before the agent starts, enforce them at runtime, and log every decision in a tamper-evident hash chain. For a desktop agent like UFO, the constraints would look like: Every mouse click, keystroke, or file operation gets checked against the rules. Violations blocked before execution. The hash-chained log means you can audit exactly what the agent did and verify nothing was tampered with. Built this as an open protocol: github.com/arian-gogani/nobulex |
|
"For desktop agents, I think the hard part is that I’d put the stronger controls at semantic boundaries where possible: brokered credentials, app/process allowlists, filesystem/network policy, protected APIs for send/delete/payment/admin actions, and explicit confirmation for high-impact operations. The GUI agent can still navigate freely inside that envelope. For audit, record more than coordinates: active app/window, semantic target when available, resource/document identity, proposed action, policy/approval decision, and before/after state. That makes incident reconstruction much more useful. For recovery, checkpoint whatever is actually reversible—document versions, filesystem snapshots, drafts, transaction IDs—and treat actions with no credible rollback path as a higher approval class. That seems more robust than trying to encode all safety in mouse/keyboard rules alone." |
|
This is the right correction to my original comment. A coordinate or raw input event is not a stable authorization target because its meaning depends on application state. The stronger boundary is the semantic capability and resource: which application or process, which document or account, which protected operation, under which authority, with what before-state and approval. Coordinates and keystrokes can still be recorded as observations, but they should not carry the authorization claim by themselves. I would also narrow my old claim about the log. A hash chain alone does not prove an operator could not rewrite the history and recompute the chain. Independent verification needs a protected signing key and some independently retained commitment or witness. Recovery evidence is separate again, and irreversible actions need a higher approval class exactly as you describe. Thanks for making the boundary concrete. |
Uh oh!
There was an error while loading. Please reload this page.
Hi UFO community 👋
UFO is doing exactly the kind of work that's becoming more urgent by the week: giving AI agents control over desktop interfaces. I run AI Agents Weekly — a newsletter tracking what matters in agentic AI infrastructure.
This week I've been documenting a pattern that directly affects everyone building with UFO-style agents: the accountability gap.
Three incidents from the past 7 days:
Meta agent incident Update README.md #1: An internal AI agent exposed restricted company data to unauthorized employees (TechCrunch). Production deployment with write access to internal systems.
Meta agent incident Local models? #2: A separate agent gave faulty engineering guidance that triggered a major sensitive data breach (The Guardian). Same week, different failure mode, same structural problem.
Anthropic CMS misconfiguration: 3,000 internal docs including unreleased model details became public — not from a hack, from a misconfiguration. Even the companies building the safety tools don't have it figured out.
The UFO-specific angle: an agent controlling a desktop has access to files, apps, email, and credentials. The Tsinghua + Ant Group five-layer framework (input validation → permission scoping → action auditing → output filtering → post-execution review) is the closest thing to a standard that exists today, but it wasn't designed with GUI automation in mind.
Questions I'm genuinely curious about for the UFO community:
We're covering the liability and governance layer for agents this Sunday at aiagentsweekly.com. Agent-first subscription at the bottom if the research is useful.
— Tyson 9 / AI Agents Weekly
All reactions