Network operations teams are increasingly being asked to hand over real, actionable decisions to autonomous artificial intelligence agents. However, trusting an agent enough to let it act independently within a live production environment presents an entirely different challenge compared to the initial process of building it. While engineers have spent decades perfecting how to develop software, applying those same safety margins to AI agents operating critical digital infrastructure remains a frontier industry challenge.
Software engineering solved a remarkably similar trust problem years ago through established practices like version control, rigorous code review, and staged rollouts. Now, network infrastructure firm Selector is applying that exact same operational discipline to AI agents in network operations with the introduction of a new development and runtime environment called Foundry.
Selector develops a comprehensive NetOps platform that was updated earlier this year to provide a correlated, unified view across distributed corporate infrastructure, spanning branch offices, colocation facilities, on-premises data centers, and public cloud environments. With the addition of Foundry, network operations teams gain the capability to build, test, version, and govern their own custom AI agents directly inside the company’s existing NetOps platform, bridging the gap between visibility and automated remediation.
Nitin Kumar, co-founder and CEO of Selector, pointed out that while traditional observability tools have grown increasingly sophisticated at diagnosing network anomalies, they have historically stopped short of taking corrective action.
"We told you what was wrong with the network, we told you how to find the data, we did all of that," Kumar said in an interview with Network World. "But what to do, what to fix, how to do the fixing, that was still manual or still built into their playbooks. We believe that’s the next level of simplification we can bring in."
How the platform works
To achieve this level of operational trust, Foundry treats an AI agent in the exact same manner an enterprise software engineering team would treat any other critical piece of code, carefully managing every stage from initial framework choice all the way through to final production rollout.
Under the hood, Foundry agents run on Pydantic AI. Kumar noted that Selector deliberately avoided heavier, more complex agent frameworks such as CrewAI in favor of a simpler and more deterministic approach. In the Foundry environment, agents are defined purely as configuration, functioning similarly to infrastructure as code, which then drives a common orchestrator with domain-specific code operating underneath it.
When it comes to review and rollback procedures, customers commit their agents directly to their own internal Git repositories. This means any changes must go through the customer’s existing, trusted pull-request workflow and peer-review processes. Before an agent is ever permitted to reach a live production environment, it undergoes rigorous back-testing where it is replayed against the customer’s historical incident data and systematically compared against the recorded outcome of each past event. If an agent fails this promotion phase or misbehaves during testing, the change rolls back in a single step, preventing faulty logic from ever touching the network.
Guardrails are another foundational pillar of the platform. Kumar described two distinct kinds of limits designed to tightly constrain agent behavior. The first mechanism caps operational costs and token usage, limiting the exact number of model calls an agent is permitted to make before the system flags it as broken or unresponsive. The second mechanism constrains the logical conclusions an agent is allowed to draw during an investigation.
"If you’re reporting an AWS outage, you cannot blame a GCP. You cannot hallucinate and do that," Kumar emphasized, highlighting the critical need for factual accuracy in automated network management.
Regarding the ultimate level of autonomy, Kumar expressed a pragmatic view. While full automation remains the long-term goal for many enterprise technology leaders, he does not expect complete hands-off management to become the immediate reality across the industry.
"The goal is to be completely non-human in the loop, but I don’t think that’s going to be reality," Kumar noted. "You will get to a state where things will just happen automatically, but it will be a very asymptotic conversion."
The workflows agents handle
Foundry agents are purposefully built around complex infrastructure support workflows that have traditionally consumed significant time from human engineers. Handling a typical network incident today generally means working through several sequential steps by hand, waiting patiently for each diagnostic test to finish, and subsequently closing or opening an incident ticket depending on the final outcome. That entire sequence is work that Kumar described as entirely manual until now.
"If something happens, you need to be able to check if the underlying provider has maintenance going on," Kumar explained. "If it’s not maintenance, you need to be able to file a ticket for the provider."
When a major cloud outage strikes, an enterprise agent must first determine whether the reported problem is real—verifying, for instance, whether a third-party cloud provider such as Amazon Web Services is actually experiencing a regional outage or if the enterprise’s local connection is operating normally.
"This ability to troubleshoot an outage and what could be done, that’s the most common workflow that we expect," Kumar said.
In practice, when an outage occurs within a supported environment, a Foundry agent activates entirely on its own accord. It immediately issues a structured plan of execution, followed by continuous, periodic status updates to operations dashboards while the outage remains ongoing, keeping human stakeholders informed without requiring manual intervention during the critical early moments of an incident.
What’s next
Looking ahead, Kumar’s roadmap focuses heavily on the broader concept of agent interoperability. He observed that Selector will not be the only platform where autonomous agents run, arguing that the broader tech industry urgently needs standardized communication protocols, such as Model Context Protocol (MCP) and agent-to-agent (A2A) standards, before agents built on entirely different vendor platforms can successfully communicate and collaborate with each other.
At the same time, Selector is pursuing active expansions into monitoring modern neocloud infrastructure, building directly upon a cloud product the company recently launched to compete with established observability vendors such as Datadog. Kumar stated that existing data center monitoring tooling is simply not sufficient as high-performance AI infrastructure continues to scale rapidly, adding that he expects to share concrete customer deployment results within the next six months.
What ultimately differentiates Selector from established rivals in the NetOps space, according to Kumar, is the fundamental shift toward allowing customers to build and customize their own agents rather than forcing them to rely on a fixed, rigid set of capabilities provided out-of-the-box by the vendor.
"We called it Foundry for a specific reason: You can create things on your own because your data is yours, your workflows are yours," Kumar concluded. "We can only guess what your workflows are, what you want to be doing, and we don’t want to curtail your innovation. We don’t want to artificially bound what you want to do."