Skip to content
← Radar
33

Show Xa11y Cross-platform Desktop Automation Via

8 signals · 1 source

Evidence

  • Show HN: Xa11y – cross-platform desktop automation via accessibility trees — I built xa11y to simplify cross-platform desktop automation.<p>Originally I was interested in using computer use agents to automate testing of desktop applications. The tools I tried gave agents the ability to click a mouse and fed them screenshots, but the agents struggled to use the mouse accurately. A couple projects like OmniParser use specialized screenshot reading models and element labeling instead of pixel coordinates for interactions. But even when they work, they're slow and expensive to run since they rely on models.<p>I then looked at accessibility APIs as a way to read the screen. Initial experiments worked and I found that accessibility APIs are a well-known tool for desktop automation. But when I wrote tests for a cross-platform desktop app, I didn't find a library that worked across Mac, Windows, and Linux. Additionally, my initial tests were flaky because elements take time to appear and get new IDs as the UI re-renders.<p>I built xa11y around two ideas to close these gaps: 1) create an accessibility abstraction that works across all platforms and 2) emulate Playwright's auto-waiting, selectors, and locator patterns which make web automation robust.<p>The hardest part was finding the right abstraction which faithfully represents the platform-specific APIs in a common interface. One example: on Windows, the accessibility tree for an entire app can be read in one API call, but on Linux each element attribute requires a separate call. The library originally fetched the entire accessibility tree for each query, then walked it to return results, but on Linux for large apps, reading all the data could take 10 seconds or more. As a result, the library evaluates the filter as it walks the tree (instead of after) and only returns the matching elements (not subtrees).<p>The library is written in Rust, has Python and JS bindings, and is MIT licensed. Blog post with more details: <a href="https://crowecawcaw.github.io/general/2026/05/30/accessibility-for-computer-use.html" rel="nofollow">https://crowecawcaw.github.io/general/2026/05/30/accessibili...</a><p>Any feedback welcome!

    HACKER_NEWS

  • Show HN: Natively – comms for agents built by agents — i have more agents than i have real humans who report to me now. as a serial founder/builder/domain hoarder, i have 5+ projects ongoing at all times and in the AI era, i can spin up a team of agents to work on them simultaneously.<p>so at this point, across my businesses, i have some flavor of a gas city/company-wide ai harness for each along with instinct personally and it seems every SaaS app has an agent too. oh and my wife, who works 3 full time jobs (mother of 4, corp job, real estate agent) also has agents.<p>BUT NONE OF THEM COMMUNICATE WITH EACH OTHER, at least not well. they started by using my emails and after i woke up one day to 500+ unread emails (i'm an inbox zero kinda guy) i realized it was all in a single thread where my agents were going back and forth.<p>so i asked my agents to collaborate together on some better way to communicate with each other that isn't adapting something intended for humans but instead is built for and by agents. my one contribution was the domain, natively.io =)<p>some forward looking stuff though i did give to the initiative: - the future will likely be agents communicating with agents on behalf of humans, businesses, and heck agents themselves - direct messages, group messages, and announcement/follower style - build for agents, not humans - fully open source MIT license<p>give this URL to your agents and they can start communicating in the agentic world: <a href="http://natively.io/BOOTSTRAP.html" rel="nofollow">http://natively.io/BOOTSTRAP.html</a>

    HACKER_NEWS

  • Show HN: Charter – Operate production-safe agents that run on your own infra — Even though everyone is talking about AI agents doing everything for them at work, many companies still aren't using autonomous agents to automate real operational work internally. Usually it's because of the complexity of bootstrapping an agent from scratch that’s production-safe (won’t spend all your company’s budget, burn through compute costs because it runs too often, or call a tool that breaks a customer). Even with today's agent-building frameworks, running agents reliably in production often means stitching together LangChain, Temporal, observability tools, approval systems, and custom guardrail logic. And that’s just for running an agent safely: there also has to be a management system so someone can observe and operate all their agents (upgrade it, roll it back, stop it). We wanted to build the end-to-end internal agent infrastructure that every company has to build today to use production-safe agents internally as an open-source platform.<p>The important thing is that getting an agent running safely is quick and simple: someone can first define their agent tying it to a version, and their lifecycle and runtime policies in YAML. Next they register their agents on workers in YAML and run those workers wherever they want. This means that if one AKS cluster has permissions for the tools needed for agents X and Y and another has permissions for agents A and B, they would each register those agents on their respective clusters. Finally they can use the CLI to apply and run the agent, and then manage it with approvals, responding to questions its asking, upgrade it/roll it back/pause it (via either CLI or a UI/portal running on localhost that are communicating with the control plane).<p>Some of the agent lifecycle management can be automated by the system: metric-based lifecycle policies make it so a user can say “if this new agent version I’m trying out (with a different prompt) fails 10 times, roll it back to the old version” or “if the new agent version is spending over a certain amount in the last 5 runs block it from being run anymore till I debug and resume it”. This is useful for teams that want to run many agents without needing someone on standby to intervene if something goes wrong on any one of them.<p>The control plane is self-hostable or can be run via a managed cloud offering, and the self-hostable option combined with someone doing local inference is actually a completely air-gapped system.<p>Give it a try: <a href="https://github.com/boundflow/charter" rel="nofollow">https://github.com/boundflow/charter</a> and let us know what you think.

    HACKER_NEWS

🔒 5 more evidence quotes with Pro

See every signal, who said it, and where — across all sources.

Unlock with Pro