I think many knowledge workers could already benefit from agents without waiting for custom integrations. A test at work showed me how.

I was building a feature that let users exchange messages and attachments with people outside our app by email. I gave an agent the requirements and asked it to write an acceptance test plan for outgoing messages and incoming replies. Then I had it run the tests by operating the computer.

I had already logged into our app. The agent prepared the test data and wrote a message. It opened Finder, selected a file, and uploaded it as an attachment. Then it sent the message and opened Apple Mail to check what arrived. It sent a reply from Mail and checked it in our app.

All this while I was doing something else!

Afterward, I reviewed screenshots and other evidence. The messages and attachments had made it through in both directions. The agent had carried out the test as I wanted.

The agent used the same interfaces I would have used. It worked across several apps, though none was designed for the whole task.

We usually connect those apps ourselves. As Geoffrey Litt explains in Dynamic Documents as Personal Software, we carry information between apps and keep track of how it fits together. He describes apps as combinations of data, operations, and interfaces chosen for a particular task. That makes me wonder how often an agent could connect the pieces across existing apps, without someone first building a product to bring them together.

Products like Grok Bot and Meta’s Muse suggest others see the same opportunity. They give agents their own computers. With Grok Bot, you sign into the sites it needs and hand control back. That’s the appeal to me. You can hand over work across existing apps without building an integration first.

When is an integration worth building? That question matters more as more colleagues start using agents. If I’m helping them use agents and paying for the tokens, I want to know which tasks work, what they cost, and where agents get stuck.

Even with Sentry or Datadog in the app, I may know little about what the agent did. I’d want to see its instructions, tool calls, clicks, and retries, and whether it finished the task. A screenshot or recording helps me review one attempt. Comparing how agents handle the same task for different colleagues would help us find repeated failures and wasted effort. Then we’d know which steps might justify an integration.

That’s why I want people to try computer use now. Give an agent a task you normally do across several apps, then check the result. You may find work you can already delegate. If you build an integration later, you’ll know more about where it’s needed.